Three-dimensional data encoding method, three-dimensional data decoding method, three-dimensional data encoding device, and three-dimensional data decoding device

By calculating the attribute information of three-dimensional points, predicting residuals, performing binarization and arithmetic encoding, the problem of difficult to reduce the amount of three-dimensional data encoding in the prior art is solved, and efficient coding and transmission are achieved.

CN120047548APending Publication Date: 2025-05-27PANASONIC INTELLECTUAL PROPERTY CORP OF AMERICA
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510115738.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2018-06-06
Filing Date
2019-05-30
Publication Date
2025-05-27

AI Technical Summary

Technical Problem

The prior art is difficult to effectively reduce the encoding amount in three-dimensional data encoding, resulting in inefficient data transmission and storage.

Method used

By calculating the predicted value of the attribute information of three-dimensional points, calculating the prediction residual, binarizing and arithmetic encoding, the prefix and suffix parts are encoded using different contexts.

Benefits of technology

It realizes effective reduction of the amount of three-dimensional data encoding, improves encoding efficiency, and reduces the need for data transmission and storage.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120047548A_ABST
    Figure CN120047548A_ABST
Patent Text Reader

Abstract

Provided are a three-dimensional data encoding method, a three-dimensional data decoding method, a three-dimensional data encoding device, and a three-dimensional data decoding device. A three-dimensional data encoding method for encoding a three-dimensional point having attribute information, in which a prediction value of the attribute information of the three-dimensional point is calculated, a prediction residual, which is the difference between the attribute information of the three-dimensional point and the prediction value, is calculated, and the prediction residual is calculated. In the present invention, a prediction residual is generated by performing binarization on the prediction residual to generate binary data, and the binary data is arithmetically encoded, the binary data including a prefix part and a suffix part, and in the arithmetic encoding, different contexts are used for the prefix part and the suffix part.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] This application is a divisional application of the patent application with the Chinese Patent Application No. 201980037127.1 (International Application No. PCT / JP2019 / 021636) and the invention title of "3D Data Encoding Method, 3D Data Decoding Method, 3D Data Encoding Apparatus, and 3D Data Decoding Apparatus", which was filed on May 30, 2019. Technical Field

[0002] The present disclosure relates to a 3D data encoding method, a 3D data decoding method, a 3D data encoding apparatus, and a 3D data decoding apparatus. Background Art

[0003] In large fields such as computer vision, map information, monitoring, infrastructure inspection, or video distribution for automobiles or robots to work autonomously, devices or services that make flexible use of 3D data will be popularized in the future. 3D data is obtained by various methods such as distance sensors such as rangefinders, stereo cameras, or combinations of multiple monocular cameras.

[0004] As a representation method of 3D data, there is a representation method called point cloud, which represents the shape of a 3D structure by a point group in a 3D space. The position and color of the point group are stored in the point cloud. Although it is expected that the point cloud will become the mainstream as a representation method of 3D data, the data volume of the point group is very large. Therefore, in the storage or transmission of 3D data, like 2D moving images (as an example, MPEG-4 AVC or HEVC standardized by MPEG), data volume compression needs to be performed by encoding.

[0005] In addition, for the compression of point clouds, a part is supported by publicly available libraries (PointCloud Library) that perform point cloud correlation processing.

[0006] In addition, there is a well-known technology that uses 3D map data to retrieve facilities around a vehicle and display them (for example, refer to Patent Document 1).

[0007] Prior Art Documents

[0008] Patent Documents

[0009] Patent Document 1 International Publication No. 2014 / 020663 Summary of the Invention

[0010] Problems to be Solved by the Invention

[0011] It is desired to reduce the encoding amount in the encoding of 3D data.

[0012] An object of the present disclosure is to provide a three-dimensional data encoding method, a three-dimensional data decoding method, a three-dimensional data encoding apparatus, or a three-dimensional data decoding apparatus capable of reducing the amount of encoding.

[0013] Means for solving the problem

[0014] A three-dimensional data encoding method according to one aspect of the present disclosure is a three-dimensional data encoding method for encoding three-dimensional points having attribute information, wherein a predicted value of the attribute information of the three-dimensional points is calculated, a difference between the attribute information of the three-dimensional points and the predicted value, that is, a prediction residual, is calculated, binary data is generated by binarizing the prediction residual, arithmetic coding is performed on the binary data, the binary data includes a prefix part, that is, a prefix part, and a suffix part, that is, a suffix part, and in the arithmetic coding, different contexts are used for the prefix part and the suffix part.

[0015] A three-dimensional data decoding method according to one aspect of the present disclosure is a three-dimensional data decoding method for decoding three-dimensional points having attribute information, wherein a predicted value of the attribute information of the three-dimensional points is calculated, binary data is generated by performing arithmetic decoding on encoded data included in a bitstream, a prediction residual is generated by multi-valuing the binary data, and a decoded value of the attribute information of the three-dimensional points is calculated by adding the predicted value and the prediction residual, the binary data includes a prefix part, that is, a prefix part, and a suffix part, that is, a suffix part, and in the arithmetic decoding, different contexts are used for the prefix part and the suffix part.

[0016] A three-dimensional data encoding apparatus according to one aspect of the present disclosure is a three-dimensional data encoding apparatus for encoding three-dimensional points having attribute information, and includes a processor and a memory, the processor uses the memory to calculate a predicted value of the attribute information of the three-dimensional points, calculate a difference between the attribute information of the three-dimensional points and the predicted value, that is, a prediction residual, generate binary data by binarizing the prediction residual, perform arithmetic coding on the binary data, the binary data includes a prefix part, that is, a prefix part, and a suffix part, that is, a suffix part, and in the arithmetic coding, different contexts are used for the prefix part and the suffix part.

[0017] A three-dimensional data decoding device according to an aspect of the present disclosure is a three-dimensional data decoding device that decodes three-dimensional points having attribute information. The device includes a processor and a memory. The processor uses the memory to calculate a predicted value of the attribute information of the three-dimensional points, generates binary data by performing arithmetic decoding on the encoded data included in the bitstream, generates a prediction residual by multi-valuing the binary data, and calculates a decoded value of the attribute information of the three-dimensional points by adding the predicted value and the prediction residual. The binary data includes a prefix part and a suffix part. In the arithmetic decoding, different contexts are used for the prefix part and the suffix part.

[0018] A three-dimensional data encoding method according to an aspect of the present disclosure is a three-dimensional data encoding method that encodes three-dimensional points having attribute information. The method calculates a predicted value of the attribute information of the three-dimensional points, calculates a difference between the attribute information of the three-dimensional points and the predicted value, i.e., a prediction residual, generates binary data by binarizing the prediction residual, and performs arithmetic encoding on the binary data.

[0019] A three-dimensional data decoding method according to an aspect of the present disclosure is a three-dimensional data decoding method that decodes three-dimensional points having attribute information. The method calculates a predicted value of the attribute information of the three-dimensional points, generates binary data by performing arithmetic decoding on the encoded data included in the bitstream, generates a prediction residual by multi-valuing the binary data, and calculates a decoded value of the attribute information of the three-dimensional points by adding the predicted value and the prediction residual.

[0020] Advantageous Effects of the Invention

[0021] The present disclosure can provide a three-dimensional data encoding method, a three-dimensional data decoding method, a three-dimensional data encoding device, or a three-dimensional data decoding device capable of reducing the encoding amount. BRIEF DESCRIPTION OF THE DRAWINGS

[0022] Figure 1 Shows the configuration of the encoded three-dimensional data of Embodiment 1.

[0023] Figure 2 Shows an example of the prediction structure between SPCs belonging to the lowest layer of GOS in Embodiment 1.

[0024] Figure 3 Shows an example of the inter-layer prediction structure of Embodiment 1.

[0025] Figure 4 Shows an example of the encoding order of GOS in Embodiment 1.

[0026] Figure 5Shows an example of the encoding order of GOS in Embodiment 1.

[0027] Figure 6 Is a block diagram of the three-dimensional data encoding device in Embodiment 1.

[0028] Figure 7 Is a flowchart of the encoding process in Embodiment 1.

[0029] Figure 8 Is a block diagram of the three-dimensional data decoding device in Embodiment 1.

[0030] Figure 9 Is a flowchart of the decoding process in Embodiment 1.

[0031] Figure 10 Shows an example of the meta information in Embodiment 1.

[0032] Figure 11 Shows a configuration example of SWLD in Embodiment 2.

[0033] Figure 12 Shows an operation example of the server and the client in Embodiment 2.

[0034] Figure 13 Shows an operation example of the server and the client in Embodiment 2.

[0035] Figure 14 Shows an operation example of the server and the client in Embodiment 2.

[0036] Figure 15 Shows an operation example of the server and the client in Embodiment 2.

[0037] Figure 16 Is a block diagram of the three-dimensional data encoding device in Embodiment 2.

[0038] Figure 17 Is a flowchart of the encoding process in Embodiment 2.

[0039] Figure 18 Is a block diagram of the three-dimensional data decoding device in Embodiment 2.

[0040] Figure 19 Is a flowchart of the decoding process in Embodiment 2.

[0041] Figure 20 Shows a configuration example of WLD in Embodiment 2.

[0042] Figure 21 Shows an example of the octree structure of WLD in Embodiment 2.

[0043] Figure 22Shows a configuration example of the SWLD of Embodiment 2.

[0044] Figure 23 Shows an example of the octree structure of the SWLD of Embodiment 2.

[0045] Figure 24 Is a block diagram of the three-dimensional data production device of Embodiment 3.

[0046] Figure 25 Is a block diagram of the three-dimensional data transmission device of Embodiment 3.

[0047] Figure 26 Is a block diagram of the three-dimensional information processing device of Embodiment 4.

[0048] Figure 27 Is a block diagram of the three-dimensional data production device of Embodiment 5.

[0049] Figure 28 Shows the configuration of the system of Embodiment 6.

[0050] Figure 29 Is a block diagram of the client device of Embodiment 6.

[0051] Figure 30 Is a block diagram of the server of Embodiment 6.

[0052] Figure 31 Is a flowchart of the three-dimensional data production process performed by the client device of Embodiment 6.

[0053] Figure 32 Is a flowchart of the sensor information transmission process performed by the client device of Embodiment 6.

[0054] Figure 33 Is a flowchart of the three-dimensional data production process performed by the server of Embodiment 6.

[0055] Figure 34 Is a flowchart of the three-dimensional map transmission process performed by the server of Embodiment 6.

[0056] Figure 35 Shows the configuration of a modified example of the system of Embodiment 6.

[0057] Figure 36 Shows the configuration of the server and client device of Embodiment 6.

[0058] Figure 37 Is a block diagram of the three-dimensional data encoding device of Embodiment 7.

[0059] Figure 38 Shows an example of the prediction residual of Embodiment 7.

[0060] Figure 39 Shows an example of the volume of Embodiment 7.

[0061] Figure 40 Shows an example of the octree representation of the volume of Embodiment 7.

[0062] Figure 41 Shows an example of the bit string of the volume of Embodiment 7.

[0063] Figure 42 Shows an example of the octree representation of the volume of Embodiment 7.

[0064] Figure 43 Shows an example of the volume of Embodiment 7.

[0065] Figure 44 Is a diagram for explaining the intra prediction processing of Embodiment 7.

[0066] Figure 45 Is a diagram for explaining the rotation and translation processing of Embodiment 7.

[0067] Figure 46 Shows an example of the syntax of the RT application flag and RT information of Embodiment 7.

[0068] Figure 47 Is a diagram for explaining the inter prediction processing of Embodiment 7.

[0069] Figure 48 Is a block diagram of the three-dimensional data decoding device of Embodiment 7.

[0070] Figure 49 Is a flowchart of the three-dimensional data encoding process performed by the three-dimensional data encoding device of Embodiment 7.

[0071] Figure 50 Is a flowchart of the three-dimensional data decoding process performed by the three-dimensional data decoding device of Embodiment 7.

[0072] Figure 51 Is a diagram showing the reference relationship in the octree structure of Embodiment 8.

[0073] Figure 52 Is a diagram showing the reference relationship in the spatial region of Embodiment 8.

[0074] Figure 53 Is a diagram showing an example of adjacent reference nodes of Embodiment 8.

[0075] Figure 54 Is a diagram showing the relationship between the parent node and the node of Embodiment 8.

[0076] Figure 55It is a diagram showing an example of occupancy rate encoding for the parent node of Embodiment 8.

[0077] Figure 56 It is a block diagram showing a three-dimensional data encoding device of Embodiment 8.

[0078] Figure 57 It is a block diagram showing a three-dimensional data decoding device of Embodiment 8.

[0079] Figure 58 It is a flowchart showing a three-dimensional data encoding process of Embodiment 8.

[0080] Figure 59 It is a flowchart showing a three-dimensional data decoding process of Embodiment 8.

[0081] Figure 60 It is a diagram showing an example of switching of the encoding table of Embodiment 8.

[0082] Figure 61 It is a diagram showing the reference relationship in the spatial region of Variant 1 of Embodiment 8.

[0083] Figure 62 It is a diagram showing a syntactic example of the header information of Variant 1 of Embodiment 8.

[0084] Figure 63 It is a diagram showing a syntactic example of the header information of Variant 1 of Embodiment 8.

[0085] Figure 64 It is a diagram showing an example of adjacent reference nodes of Variant 2 of Embodiment 8.

[0086] Figure 65 It is a diagram showing examples of object nodes and adjacent nodes of Variant 2 of Embodiment 8.

[0087] Figure 66 It is a diagram showing the reference relationship in the octree structure of Variant 3 of Embodiment 8.

[0088] Figure 67 It is a diagram showing the reference relationship in the spatial region of Variant 3 of Embodiment 8.

[0089] Figure 68 It is a diagram showing an example of three-dimensional points of Embodiment 9.

[0090] Figure 69 It is a diagram showing an example of LoD setting of Embodiment 9.

[0091] Figure 70 It is a diagram showing an example of a threshold value used in LoD setting of Embodiment 9.

[0092] Figure 71It is a diagram showing an example of the attribute information used in the predicted value of Embodiment 9.

[0093] Figure 72 It is a diagram showing an example of the exponential Golomb code of Embodiment 9.

[0094] Figure 73 It is a diagram showing the processing for the exponential Golomb code of Embodiment 9.

[0095] Figure 74 It is a diagram showing a syntactic example of the attribute header of Embodiment 9.

[0096] Figure 75 It is a diagram showing a syntactic example of the attribute data of Embodiment 9.

[0097] Figure 76 It is a flowchart of the three-dimensional data encoding process of Embodiment 9.

[0098] Figure 77 It is a flowchart of the attribute information encoding process of Embodiment 9.

[0099] Figure 78 It is a diagram showing the processing for the exponential Golomb code of Embodiment 9.

[0100] Figure 79 It is a diagram showing an example of the inverse table representing the relationship between the remaining code and its value of Embodiment 9.

[0101] Figure 80 It is a flowchart of the three-dimensional data decoding process of Embodiment 9.

[0102] Figure 81 It is a flowchart of the attribute information decoding process of Embodiment 9.

[0103] Figure 82 It is a block diagram of the three-dimensional data encoding device of Embodiment 9.

[0104] Figure 83 It is a block diagram of the three-dimensional data decoding device of Embodiment 9.

[0105] Figure 84 It is a flowchart of the three-dimensional data encoding process of Embodiment 9.

[0106] Figure 85 It is a flowchart of the three-dimensional data decoding process of Embodiment 9. Detailed implementation mode

[0107] A three-dimensional data encoding method according to one aspect of the present disclosure is a three-dimensional data encoding method for encoding three-dimensional points with attribute information, calculating a predicted value of the attribute information of the three-dimensional points, calculating a difference between the attribute information of the three-dimensional points and the predicted value, i.e., a prediction residual, generating binary data by binarizing the prediction residual, and performing arithmetic coding on the binary data.

[0108] Thus, the three-dimensional data encoding method calculates the prediction residual of the attribute information, and then binarizes and arithmetically encodes the prediction residual, thereby being able to reduce the encoding amount of the encoded data of the attribute information.

[0109] For example, it may also be that in the arithmetic coding, different coding tables are used for each bit of the binary data.

[0110] Thus, the three-dimensional data encoding method can improve the encoding efficiency.

[0111] For example, it may also be that in the arithmetic coding, the lower the bit position of the binary data, the larger the number of coding tables used.

[0112] For example, it may also be that in the arithmetic coding, according to the value of the upper bit of the target bit included in the binary data, a coding table used in the arithmetic coding of the target bit is selected.

[0113] Thus, the three-dimensional data encoding method can select a coding table according to the value of the upper bit, and thus can improve the encoding efficiency.

[0114] For example, it may also be that in the binarization, when the prediction residual is less than a threshold, the binary data is generated by binarizing the prediction residual with a fixed number of bits, and when the prediction residual is equal to or greater than the threshold, the binary data including a first coding representing the threshold with the fixed number of bits and a second coding obtained by binarizing the value obtained by subtracting the threshold from the prediction residual with exponential Golomb is generated, and in the arithmetic coding, different arithmetic coding methods are used for the first coding and the second coding.

[0115] Thus, the three-dimensional data encoding method can, for example, arithmetically encode the first coding and the second coding by arithmetic coding methods respectively suitable for the first coding and the second coding, and thus can improve the encoding efficiency.

[0116] For example, it may also be that the three-dimensional data encoding method further quantizes the prediction residual, and in the binarization, the quantized prediction residual is binarized, and the threshold is changed according to the quantization scale in the quantization.

[0117] Accordingly, the three-dimensional data encoding method can use an appropriate threshold corresponding to the quantization scale, and thus can improve the encoding efficiency.

[0118] For example, it may also be that the second encoding includes a prefix part and a suffix part, and in the arithmetic encoding, different encoding tables are used for the prefix part and the suffix part.

[0119] Accordingly, the three-dimensional data encoding method can improve the encoding efficiency.

[0120] A three-dimensional data decoding method according to one aspect of the present disclosure is a three-dimensional data decoding method for decoding three-dimensional points having attribute information, calculating a predicted value of the attribute information of the three-dimensional points, generating binary data by performing arithmetic decoding on the encoded data included in the bitstream, generating a prediction residual by multi-valuing the binary data, and calculating a decoded value of the attribute information of the three-dimensional points by adding the predicted value and the prediction residual.

[0121] Accordingly, the three-dimensional data decoding method can calculate the prediction residual of the attribute information, and thus can appropriately decode the bitstream of the attribute information generated by binarizing and arithmetic encoding the prediction residual.

[0122] For example, it may also be that in the arithmetic decoding, different encoding tables are used for each bit of the binary data.

[0123] Accordingly, the three-dimensional data decoding method can appropriately decode the bitstream with improved encoding efficiency.

[0124] For example, it may also be that in the arithmetic decoding, the larger the lower bits of the binary data, the larger the number of encoding tables used.

[0125] For example, it may also be that in the arithmetic decoding, according to the value of the upper bit of the target bit included in the binary data, an encoding table used in the arithmetic decoding of the target bit is selected.

[0126] Accordingly, the three-dimensional data decoding method can appropriately decode the bitstream with improved encoding efficiency.

[0127] For example, it may also be that in the multi-valuing, the first value is generated by multi-valuing the first encoding of the fixed number of bits included in the binary data. When the first value is less than the threshold, the first value is determined as the prediction residual. When the first value is equal to or greater than the threshold, the second value is generated by multi-valuing the exponential Golomb code, i.e., the second encoding, included in the binary data. The prediction residual is generated by adding the first value and the second value. In the arithmetic decoding, different arithmetic decoding methods are used for the first encoding and the second encoding.

[0128] Thereby, the three-dimensional data decoding method can appropriately decode a bitstream with improved encoding efficiency.

[0129] For example, it may also be that the three-dimensional data decoding method further inverse-quantizes the prediction residual. In the addition, the predicted value is added to the prediction residual after inverse-quantization, and the threshold is changed according to the quantization scale in the inverse quantization.

[0130] Thereby, the three-dimensional data decoding method can appropriately decode a bitstream with improved encoding efficiency.

[0131] For example, it may also be that the second encoding includes a prefix part and a suffix part, and different coding tables are used for the prefix part and the suffix part in the arithmetic decoding.

[0132] In addition, a three-dimensional data encoding device according to one aspect of the present disclosure is a three-dimensional data encoding device that encodes three-dimensional points having attribute information, and includes a processor and a memory. The processor uses the memory to calculate a predicted value of the attribute information of the three-dimensional points, calculates a difference between the attribute information of the three-dimensional points and the predicted value, i.e., a prediction residual, generates binary data by binarizing the prediction residual, and performs arithmetic coding on the binary data.

[0133] Thereby, the three-dimensional data encoding device calculates the prediction residual of the attribute information, and then binarizes and arithmetically encodes the prediction residual, so that the encoding amount of the encoded data of the attribute information can be reduced.

[0134] In addition, a three-dimensional data decoding device according to one aspect of the present disclosure is a three-dimensional data decoding device that decodes three-dimensional points having attribute information, and includes a processor and a memory. The processor uses the memory to calculate a predicted value of the attribute information of the three-dimensional points, generates binary data by arithmetically decoding the encoded data included in the bitstream, generates a prediction residual by multi-valuing the binary data, and calculates a decoded value of the attribute information of the three-dimensional points by adding the predicted value and the prediction residual.

[0135] Accordingly, the three-dimensional data decoding apparatus can calculate the prediction residual of the attribute information, and further, appropriately decode the bit stream of the attribute information generated by binarizing and arithmetic-coding the prediction residual.

[0136] In addition, these general or specific forms can be implemented by a system, a method, an integrated circuit, a computer program, or a recording medium such as a computer-readable CD-ROM, and can be implemented by any combination of a system, a method, an integrated circuit, a computer program, and a recording medium.

[0137] Hereinafter, embodiments will be specifically described with reference to the drawings. In addition, all the embodiments to be described below are specific examples showing the present disclosure. The numerical values, shapes, materials, constituent elements, arrangement positions and connection forms of the constituent elements, steps, the order of steps, etc. shown in the following embodiments are all examples, and the gist thereof is not to limit the present disclosure. And, among the constituent elements of the following embodiments, the constituent elements not described in the technical solution showing the uppermost concept are described as optional constituent elements.

[0138] (Embodiment 1)

[0139] First, the data structure of the encoded three-dimensional data (hereinafter also referred to as encoded data) related to the present embodiment will be described. Figure 1 The configuration of the encoded three-dimensional data related to the present embodiment is shown.

[0140] In the present embodiment, the three-dimensional space is divided into spaces (SPCs) corresponding to pictures in the encoding of moving images, and the three-dimensional data is encoded in units of space. The space is further divided into volumes (VLMs) corresponding to macroblocks or the like in moving image encoding, and prediction and transformation are performed in units of VLM. A volume includes a plurality of voxels (VXLs) which are the smallest units corresponding to position coordinates. In addition, prediction means, similar to the prediction performed in two-dimensional images, referring to other processing units, generating predicted three-dimensional data similar to the processing unit of the object to be processed, and encoding the difference between the predicted three-dimensional data and the processing unit of the object to be processed. And this prediction includes not only spatial prediction referring to other prediction units at the same time, but also temporal prediction referring to prediction units at different times.

[0141] For example, when a three-dimensional data encoding apparatus (hereinafter also referred to as an encoding apparatus) encodes a three-dimensional space represented by point cloud data or the like, it encodes each point of the point cloud or a plurality of points included in a voxel together according to the size of the voxel. If the voxel is subdivided, the three-dimensional shape of the point cloud can be represented with high precision, and if the size of the voxel is increased, the three-dimensional shape of the point cloud can be represented roughly.

[0142] In addition, although the following description is given by taking the case where the three-dimensional data is point cloud as an example, the three-dimensional data is not limited to point cloud and can be any form of three-dimensional data.

[0143] Moreover, a voxel of a hierarchical structure can be used. In this case, in the n-th level, it is possible to sequentially show whether there are sampling points in the levels below the (n - 1)-th level (the lower layer of the n-th level). For example, when only decoding the n-th level, if there are sampling points in the levels below the (n - 1)-th level, it can be regarded that there are sampling points at the center of the voxel of the n-th level for decoding.

[0144] Moreover, the encoding device obtains point group data through a distance sensor, a stereo camera, a monocular camera, a gyroscope, an inertial sensor, or the like.

[0145] Regarding space, similar to the encoding of moving images, it is at least classified into any one of the following three prediction structures, and these three prediction structures are: an intra-frame space (I-SPC) that can be decoded independently, a predictive space (P-SPC) that can only be unidirectionally referenced, and a bidirectional space (B-SPC) that can be bidirectionally referenced. Moreover, space has two types of time information, namely, a decoding time and a display time.

[0146] Moreover, as Figure 1 shown, as a processing unit including multiple spaces, there is a GOS (Group Of Space) as a random access unit. Moreover, as a processing unit including multiple GOSs, there is a world space (WLD).

[0147] The spatial region occupied by the world space is associated with the absolute position on the earth through GPS or latitude and longitude information, etc. This position information is stored as meta information. In addition, the meta information can be included in the encoded data or transmitted separately from the encoded data.

[0148] Moreover, within a GOS, all SPCs can be three-dimensionally adjacent, or there can be an SPC that is not three-dimensionally adjacent to other SPCs.

[0149] In addition, hereinafter, the processes such as encoding, decoding, or referencing corresponding to the three-dimensional data included in processing units such as GOS, SPC, or VLM are also simply referred to as encoding, decoding, or referencing the processing unit. Moreover, the three-dimensional data included in the processing unit includes at least one group of a spatial position such as three-dimensional coordinates and characteristic values such as color information.

[0150] Next, the prediction structure of SPCs in a GOS will be described. Multiple SPCs within the same GOS, or multiple VLMs within the same SPC, although they occupy different spaces from each other, have the same time information (decoding time and display time).

[0151] Also, within a GOS, the SPC that is the first in the decoding order is an I-SPC. Also, there are two types of GOSs in the GOS: a closed GOS and an open GOS. A closed GOS is a GOS in which all SPCs within the GOS can be decoded when starting decoding from the first I-SPC. In an open GOS, within the GOS, a part of the SPCs whose display time is earlier than that of the first I-SPC refer to different GOSs and can only be decoded in that GOS.

[0152] In addition, in encoded data such as map information, there are cases where the WLD is decoded in the direction opposite to the encoding order. If there is a dependency between GOSs, it is difficult to perform reverse playback. Therefore, in such cases, a closed GOS is basically adopted.

[0153] Also, the GOS has a layer structure in the height direction, and encoding or decoding is sequentially performed starting from the SPCs of the bottom layer.

[0154] Figure 2 An example of the prediction structure between SPCs belonging to the bottommost layer of the GOS is shown. Figure 3 An example of the inter-layer prediction structure is shown.

[0155] There is one or more I-SPCs within the GOS. Although there are objects such as people, animals, cars, bicycles, traffic lights, or buildings that serve as land marks in the three-dimensional space, it is particularly effective when encoding small-sized objects as I-SPCs. For example, when a three-dimensional data decoding device (hereinafter also referred to as a decoding device) decodes the GOS with a low processing amount or at high speed, only the I-SPCs within the GOS are decoded.

[0156] Also, the encoding device can switch the encoding interval or occurrence frequency of the I-SPCs according to the density of the objects within the WLD.

[0157] Also, in Figure 3 In the configuration shown, the encoding device or the decoding device sequentially performs encoding or decoding for multiple layers starting from the lower layer (layer 1). Accordingly, for example, for an automatically moving vehicle or the like, the priority of the data near the ground with a large amount of information can be increased.

[0158] In addition, in the encoded data used in a drone or the like, within the GOS, encoding or decoding can be sequentially performed starting from the SPCs of the upper layer in the height direction.

[0159] Also, the encoding device or the decoding device may encode or decode multiple layers in such a way that the decoding device can roughly grasp the GOS and gradually increase the resolution. For example, the encoding device or the decoding device may encode or decode in the order of layer 3, 8, 1, 9...

[0160] Next, the corresponding methods for static objects and dynamic objects will be described.

[0161] In a three-dimensional space, there are static objects or scenes such as buildings or roads (collectively referred to as static objects hereafter), and dynamic objects such as vehicles or people (referred to as dynamic objects hereafter). Detection of objects can be performed separately by extracting feature points from data of point clouds or images captured by a stereo camera, etc. Here, an example of an encoding method for dynamic objects will be described.

[0162] The first method is a method of encoding without distinguishing between static objects and dynamic objects. The second method is a method of distinguishing between static objects and dynamic objects by identification information.

[0163] For example, GOS is used as an identification unit. In this case, the GOS including the SPC constituting the static object and the GOS including the SPC constituting the dynamic object are distinguished within the encoded data or by identification information stored separately from the encoded data.

[0164] Alternatively, SPC is used as an identification unit. In this case, only the SPC including the VLM constituting the static object and the SPC including the VLM constituting the dynamic object are distinguished by the above-mentioned identification information.

[0165] Alternatively, VLM or VXL can be used as an identification unit. In this case, the VLM or VXL including the static object and the VLM or VXL including the dynamic object are distinguished by the above-mentioned identification information.

[0166] Also, the encoding device may encode the dynamic object as one or more VLMs or SPCs, and encode the VLM or SPC including the static object and the SPC including the dynamic object as different GOSs. And when the size of the GOS becomes variable according to the size of the dynamic object, the encoding device stores the size of the GOS as meta information separately.

[0167] Also, the encoding device encodes the static object and the dynamic object independently of each other, and for the world space composed of the static object, the dynamic object can be overlapped. At this time, the dynamic object is composed of one or more SPCs, and each SPC corresponds to one or more SPCs of the static object overlapping the SPC. In addition, the dynamic object may not be represented by SPCs, but may be represented by one or more VLMs or VXLs.

[0168] Furthermore, the encoding device can encode static objects and dynamic objects as different streams from each other.

[0169] Furthermore, the encoding device can also generate a GOS including one or more SPCs constituting a dynamic object. Moreover, the encoding device can set the GOS (GOS_M) including the dynamic object and the GOS of the static object corresponding to the spatial region of GOS_M to have the same size (occupy the same spatial region). In this way, overlapping processing can be performed in units of GOS.

[0170] The P-SPC or B-SPC constituting the dynamic object can also refer to the SPCs included in different encoded GOSs. When the position of the dynamic object changes over time and the same dynamic object is encoded as GOSs at different times, cross-GOS reference is effective from the viewpoint of compression ratio.

[0171] Furthermore, the above first method and second method can also be switched according to the use of the encoded data. For example, when encoding three-dimensional data to be applied as a map, since separation from the dynamic object is desired, the encoding device adopts the second method. In addition, when the encoding device encodes three-dimensional data of an event such as a concert or a sports event, if separation of the dynamic object is not required, the first method is adopted.

[0172] Furthermore, the decoding time and display time of the GOS or SPC can be stored in the encoded data or as meta-information. And the time information of the static objects can all be the same. At this time, the actual decoding time and display time can be determined by the decoding device. Or, as the decoding time, different values can be assigned to each GOS or SPC, and as the display time, the same value can be assigned to all. Moreover, as shown in the decoder mode in dynamic image encoding such as HRD (Hypothetical Reference Decoder) of HEVC, the decoder has a buffer of a specified size, and as long as the bitstream is read at a specified bit rate according to the decoding time, a model that will not be damaged and can be guaranteed to be decoded can be imported.

[0173] Next, the configuration of GOSs in the world space will be described. The coordinates of the three-dimensional space in the world space are represented by three mutually orthogonal coordinate axes (x-axis, y-axis, z-axis). By setting a specified rule in the encoding order of GOSs, GOSs adjacent in space can be encoded continuously in the encoded data. For example, in Figure 4In the example shown, continuous encoding is performed on the GOS in the xz plane. After the encoding of all GOSs in one xz plane is completed, the value of the y-axis is updated. That is, as encoding progresses, the world space extends in the y-axis direction. Also, the index number of the GOS is set to the encoding order.

[0174] Here, the three-dimensional space of the world space corresponds one-to-one with the GPS or geographical absolute coordinates such as latitude and longitude. Alternatively, the three-dimensional space can be represented by the relative position with respect to a preset reference position. The directions of the x-axis, y-axis, and z-axis of the three-dimensional space are represented as direction vectors determined based on latitude and longitude, etc., and this direction vector is stored together with the encoded data as meta information.

[0175] Also, the size of the GOS is set to be fixed, and the encoding device stores this size as meta information. Also, the size of the GOS can be switched, for example, according to whether it is in the city, indoors, or outdoors, etc. That is, the size of the GOS can be switched according to the quantity or nature of the object having the value as information. Alternatively, the encoding device can appropriately switch the size of the GOS or the interval of the I-SPC within the GOS in the same world space according to the density of the object, etc. For example, the encoding device sets the size of the GOS to be smaller and the interval of the I-SPC within the GOS to be shorter when the density of the object is higher.

[0176] In Figure 5 the example, in the area from the 3rd to the 10th GOS, since the density of the object is high, in order to achieve fine-grained random access, the GOS is subdivided. Also, the 7th to 10th GOSs are respectively located behind the 3rd to 6th GOSs.

[0177] Next, the configuration and operation flow of the three-dimensional data encoding device according to the present embodiment will be described. Figure 6 is a block diagram of the three-dimensional data encoding device 100 according to the present embodiment. Figure 7 is a flowchart showing an operation example of the three-dimensional data encoding device 100.

[0178] Figure 6 The three-dimensional data encoding device 100 shown generates encoded three-dimensional data 112 by encoding three-dimensional data 111. This three-dimensional data encoding device 100 includes: an acquisition unit 101, an encoding area determination unit 102, a division unit 103, and an encoding unit 104.

[0179] As Figure 7 shown, first, the acquisition unit 101 acquires three-dimensional data 111 as point cloud data (S101).

[0180] Next, the encoding region determination unit 102 determines the region to be encoded from the spatial region corresponding to the acquired point group data (S102). For example, the encoding region determination unit 102 determines the spatial region around the position of the user or the vehicle as the region to be encoded.

[0181] Next, the segmentation unit 103 segments the point group data included in the region to be encoded into respective processing units. Here, the processing units are the above-mentioned GOS and SPC, etc. And the region to be encoded corresponds to the above-mentioned world space, for example. Specifically, the segmentation unit 103 segments the point group data into processing units according to the size of the GOS set in advance, the presence or absence or size of the dynamic object (S103). And the segmentation unit 103 determines the start position of the SPC that becomes the beginning in the encoding order in each GOS.

[0182] Next, the encoding unit 104 generates the encoded three-dimensional data 112 by sequentially encoding a plurality of SPCs within each GOS (S104).

[0183] In addition, here, after the region to be encoded is segmented into GOS and SPC, an example of encoding each GOS is shown, but the order of processing is not limited to the above. For example, after determining the composition of one GOS, the GOS can be encoded, and then the order of determining the composition of the GOS, etc. can be followed.

[0184] In this way, the three-dimensional data encoding device 100 generates the encoded three-dimensional data 112 by encoding the three-dimensional data 111. Specifically, the three-dimensional data encoding device 100 segments the three-dimensional data into random access units, that is, segments it into first processing units (GOS) corresponding to three-dimensional coordinates respectively, segments the first processing units (GOS) into a plurality of second processing units (SPC), and segments the second processing units (SPC) into a plurality of third processing units (VLM). And the third processing unit (VLM) includes one or more voxels (VXL), and the voxel (VXL) is the smallest unit corresponding to the position information.

[0185] Next, the three-dimensional data encoding device 100 generates the encoded three-dimensional data 112 by encoding each of the plurality of first processing units (GOS). Specifically, the three-dimensional data encoding device 100 encodes each of the plurality of second processing units (SPC) in each first processing unit (GOS). And the three-dimensional data encoding device 100 encodes each of the plurality of third processing units (VLM) in each second processing unit (SPC).

[0186] For example, when the first processing unit (GOS) of the processing object is a closed GOS, the three-dimensional data encoding device 100 encodes the second processing unit (SPC) of the processing object included in the first processing unit (GOS) of the processing object with reference to other second processing units (SPCs) included in the first processing unit (GOS) of the processing object. That is, the three-dimensional data encoding device 100 does not refer to the second processing units (SPCs) included in the first processing unit (GOS) different from the first processing unit (GOS) of the processing object.

[0187] Moreover, when the first processing unit (GOS) of the processing object is an open GOS, the three-dimensional data encoding device 100 encodes the second processing unit (SPC) of the processing object included in the first processing unit (GOS) of the processing object with reference to other second processing units (SPCs) included in the first processing unit (GOS) of the processing object or the second processing units (SPCs) included in the first processing unit (GOS) different from the first processing unit (GOS) of the processing object.

[0188] In addition, the three-dimensional data encoding device 100 selects one of the first type (I-SPC) of other second processing units (SPCs), the second type (P-SPC) of one other second processing unit (SPC), and the third type of two other second processing units (SPCs) as the type of the second processing unit (SPC) of the processing object, and encodes the second processing unit (SPC) of the processing object according to the selected type.

[0189] Next, the configuration and operation process of the three-dimensional data decoding device according to the present embodiment will be described. Figure 8 It is a block diagram of the three-dimensional data decoding device 200 according to the present embodiment. Figure 9 It is a flowchart showing an operation example of the three-dimensional data decoding device 200.

[0190] Figure 8 The shown three-dimensional data decoding device 200 generates decoded three-dimensional data 212 by decoding the encoded three-dimensional data 211. Here, the encoded three-dimensional data 211 is, for example, the encoded three-dimensional data 112 generated by the three-dimensional data encoding device 100. The three-dimensional data decoding device 200 includes: an acquisition unit 201, a decoding start GOS determination unit 202, a decoding SPC determination unit 203, and a decoding unit 204.

[0191] First, acquisition unit 201 acquires encoded three-dimensional data 211 (S201). Next, decoding start GOS determination unit 202 determines the GOS to be decoded (S202). Specifically, decoding start GOS determination unit 202 refers to the meta-information within the encoded three-dimensional data 211 or stored separately from the encoded three-dimensional data, and determines the GOS including the spatial position, object, or SPC corresponding to the time at which decoding starts as the GOS to be decoded.

[0192] Next, decoding SPC determination unit 203 determines the type (I, P, B) of the SPC to be decoded within the GOS (S203). For example, decoding SPC determination unit 203 determines (1) whether to decode only I-SPC, (2) whether to decode I-SPC and P-SPC, (3) whether to decode all types. Additionally, in cases where the type of SPC to be decoded, such as decoding all SPCs, is pre-specified, this step may not be performed.

[0193] Next, decoding unit 204 acquires the SPC that is the start in the decoding order (same as the encoding order) within the GOS, the address position where it starts within the encoded three-dimensional data 211, acquires the encoded data of the start SPC from this address position, and decodes each SPC in sequence starting from this start SPC (S204). And the above address position is stored in meta-information or the like.

[0194] In this way, three-dimensional data decoding device 200 decodes decoded three-dimensional data 212. Specifically, three-dimensional data decoding device 200 generates decoded three-dimensional data 212 of the first processing unit (GOS) as a random access unit by decoding each of the encoded three-dimensional data 211 of the first processing unit (GOS) corresponding to the three-dimensional coordinates respectively. More specifically, three-dimensional data decoding device 200 decodes each of the multiple second processing units (SPCs) within each first processing unit (GOS). And three-dimensional data decoding device 200 decodes each of the multiple third processing units (VLM) within each second processing unit (SPC).

[0195] The following explains the meta-information for random access. This meta-information is generated by three-dimensional data encoding device 100 and is included in encoded three-dimensional data 112 (211).

[0196] In the random access of conventional two-dimensional moving images, decoding starts from the first frame of the random access unit near the specified time. However, in the world space, random access is envisioned not only for time but also for (coordinates or objects, etc.).

[0197] Therefore, in order to at least achieve random access to the three elements of coordinates, objects, and time, a table is prepared in which the indexes of each element are corresponding to the GOS indexes. Moreover, the GOS index is corresponding to the address of the I-SPC that is the start of the GOS. Figure 10 An example of the table included in the meta information is shown. In addition, it is not necessary to use Figure 10 all the tables shown. At least one table can be used.

[0198] Hereinafter, as an example, random access starting from coordinates will be described. When accessing the coordinates (x2, y2, z2), first referring to the coordinate-GOS table, it can be known that the location with the coordinates (x2, y2, z2) is included in the second GOS. Then, referring to the GOS address table, since it can be known that the address of the I-SPC at the start of the second GOS is addr(2), the decoding unit 204 obtains data from this address and starts decoding.

[0199] In addition, the address can be an address in the logical format or a physical address of an HDD or a memory. Also, information for determining a file segment can be used instead of the address. For example, a file segment is a unit obtained by segmenting one or more GOSs, etc.

[0200] Moreover, in the case where an object spans multiple GOSs, the GOSs to which the multiple objects belong can also be shown in the object GOS table. If the multiple GOSs are closed GOSs, the encoding device and the decoding device can perform encoding or decoding in parallel. In addition, if the multiple GOSs are open GOSs, by referring to each other among the multiple GOSs, the compression efficiency can be further improved.

[0201] Examples of objects include people, animals, cars, bicycles, traffic lights, or buildings that are land marks. For example, when the three-dimensional data encoding device 100 encodes in the world space, characteristic points unique to the object are extracted from a three-dimensional point cloud or the like, the object is detected based on the characteristic points, and the detected object can be set as a random access point.

[0202] In this way, the three-dimensional data encoding device 100 generates the first information, which shows a plurality of first processing units (GOSs) and the three-dimensional coordinates corresponding to each of the plurality of first processing units (GOSs). And encoding the three-dimensional data 112(211) includes this first information. And the first information further shows at least one of the object, time, and data storage destination corresponding to each of the plurality of first processing units (GOSs).

[0203] The three-dimensional data decoding device 200 obtains the first information from the encoded three-dimensional data 211, uses the first information to determine the encoded three-dimensional data 211 of the first processing unit corresponding to the specified three-dimensional coordinates, object, or time, and decodes the encoded three-dimensional data 211.

[0204] Examples of other meta-information are described below. In addition to the meta-information for random access, the three-dimensional data encoding device 100 can also generate and store the following meta-information. Also, the three-dimensional data decoding device 200 can use this meta-information during decoding.

[0205] In the case of using three-dimensional data as map information, etc., a profile is specified according to the use, and the information indicating the profile can be included in the meta-information. For example, a profile for urban areas or suburbs is specified, or a profile for flying objects is specified, and the maximum or minimum size of the world space, SPC, or VLM is defined respectively. For example, in the profile for urban areas, more detailed information is required than in the suburbs, so the minimum size of the VLM is set smaller.

[0206] The meta-information can also include a tag value indicating the type of the object. This tag value corresponds to the VLM, SPC, or GOS that constitutes the object. The tag value can be set according to the type of the object. For example, the tag value "0" represents "person", the tag value "1" represents "car", and the tag value "2" represents "traffic signal". Or, in the case where the type of the object is difficult to determine or does not need to be determined, a tag value indicating properties such as size, or whether it is a dynamic object or a static object can also be used.

[0207] Also, the meta-information can include information indicating the range of the spatial region occupied by the world space.

[0208] Also, the meta-information can store the size of the SPC or VXL as the header information shared by the entire stream of encoded data or multiple SPCs such as SPCs within the GOS.

[0209] Also, the meta-information can include identification information such as a distance sensor or a camera used in the generation of the point cloud, or information indicating the position accuracy of the point group within the point cloud.

[0210] Also, the meta-information can include information indicating whether the world space is composed only of static objects or contains dynamic objects.

[0211] A modification example of the present embodiment is described below.

[0212] An encoding device or a decoding device can encode or decode two or more SPCs or GOSs that are different from each other in parallel. The GOSs encoded or decoded in parallel can be determined based on meta information indicating the spatial positions of the GOSs, etc.

[0213] In a case where three-dimensional data is used as a spatial map when a vehicle or a flying object moves, or in a case where such a spatial map is generated, etc., the encoding device or the decoding device can encode or decode GOSs or SPCs included in a space determined based on GPS, path information, or magnification ratio, etc.

[0214] Moreover, the decoding device can also start decoding sequentially from the space close to its own position or travel path. The encoding device or the decoding device can also perform encoding or decoding by making the priority of the space far from its own position or travel path lower than that of the close space. Here, reducing the priority means reducing the processing order, reducing the resolution (post-filtering), or reducing the image quality (improving the encoding efficiency. For example, increasing the quantization step size), etc.

[0215] Moreover, when the decoding device decodes the encoded data hierarchically encoded in a space, it can also decode only the lower hierarchy.

[0216] Moreover, the decoding device can also start decoding from the lower hierarchy first according to the magnification ratio or use of the map.

[0217] Moreover, in applications such as self-position estimation or object recognition performed during the automatic driving of an automobile or a robot, the encoding device or the decoding device can also reduce the resolution of the area outside the area within a specified height from the road surface (the area to be recognized) to perform encoding or decoding.

[0218] Moreover, the encoding device can also encode the point clouds representing the spatial shapes of the indoor and outdoor spaces independently. For example, by separating the GOS representing the indoor (indoor GOS) from the GOS representing the outdoor (outdoor GOS), the decoding device can select the GOS to be decoded according to the viewpoint position when using the encoded data.

[0219] Moreover, the encoding device can make the indoor GOS and the outdoor GOS with close coordinates adjacent in the encoding stream to perform encoding. For example, the encoding device corresponds the identifiers of the two, and stores the information indicating the identifiers corresponding to each other in the encoding stream or in the meta information stored separately. Accordingly, the decoding device can refer to the information in the meta information to identify the indoor GOS and the outdoor GOS with close coordinates.

[0220] Also, the encoding device can also switch the size of the GOS or SPC between the indoor GOS and the outdoor GOS. For example, the encoding device sets the size of the GOS to be smaller indoors than outdoors. Also, the encoding device can also change the accuracy when extracting feature points from the point cloud or the accuracy of object detection, etc., between the indoor GOS and the outdoor GOS.

[0221] Also, the encoding device can attach information for the decoding device to distinguish and display dynamic objects from static objects to the encoded data. Accordingly, the decoding device can combine and represent dynamic objects with red frames or explanatory text, etc. In addition, the decoding device can also represent only with a red frame or explanatory text instead of the dynamic object. And the decoding device can represent more detailed object categories. For example, a car can use a red frame and a person can use a yellow frame.

[0222] Also, the encoding device or the decoding device can determine whether to encode or decode dynamic objects and static objects as different SPCs or GOSs according to the appearance frequency of the dynamic objects, or the ratio of static objects to dynamic objects, etc. For example, when the appearance frequency or ratio of the dynamic objects exceeds the threshold, the SPC or GOS in which the dynamic objects and static objects are mixed is allowed, and when the appearance frequency or ratio of the dynamic objects does not exceed the threshold, the SPC or GOS in which the dynamic objects and static objects are mixed is not allowed.

[0223] When the dynamic object is detected not from the point cloud but from the two-dimensional image information of the camera, the encoding device can separately obtain information (such as a frame or text) for identifying the detection result and the object position, and encode these as part of the three-dimensional encoded data. In this case, the decoding device overlays and displays auxiliary information (frame or text) representing the dynamic object on the decoding result of the static object.

[0224] Also, the encoding device can change the density of the VXL or VLM according to the complexity of the shape of the static object, etc. For example, the more complex the shape of the static object, the denser the encoding device sets the VXL or VLM. Moreover, the encoding device can determine the quantization step, etc., when quantifying the spatial position or color information according to the density of the VXL or VLM. For example, the denser the VXL or VLM, the smaller the encoding device sets the quantization step.

[0225] As described above, the encoding device or the decoding device according to the present embodiment performs spatial encoding or decoding in a spatial unit having coordinate information.

[0226] Also, the encoding device and the decoding device perform encoding or decoding in volume units within the space. The volume includes voxels, which are the smallest units corresponding to the position information.

[0227] Further, the encoding device and the decoding device perform encoding or decoding by establishing a correspondence table that associates each element of spatial information including coordinates, objects, and time, etc. with a GOP, or a correspondence table between the elements, to establish a correspondence between any elements. Further, the decoding device determines coordinates using the value of the selected element, and determines a volume, voxel, or space based on the coordinates, and decodes the space including the volume or voxel, or the determined space.

[0228] Further, the encoding device determines a volume, voxel, or space that can be selected by an element through feature point extraction or object recognition, and encodes it as a volume, voxel, or space that can be randomly accessed.

[0229] The space is divided into three types, namely: I-SPC that can be encoded or decoded by the space alone, P-SPC that is encoded or decoded with reference to any one processed space, and B-SPC that is encoded or decoded with reference to any two processed spaces.

[0230] One or more volumes correspond to static objects or dynamic objects. The space containing static objects and the space containing dynamic objects are encoded or decoded as different GOSs from each other. That is, the SPC containing static objects and the SPC containing dynamic objects are assigned to different GOSs.

[0231] Dynamic objects are encoded or decoded for each object and correspond to one or more spaces containing only static objects. That is, multiple dynamic objects are encoded separately, and the encoded data of the multiple dynamic objects obtained corresponds to the SPC containing only static objects.

[0232] The encoding device and the decoding device perform encoding or decoding by increasing the priority of I-SPC within the GOS. For example, the encoding device performs encoding in a manner that reduces the degradation of I-SPC (after decoding, the original three-dimensional data can be reproduced more faithfully). Further, the decoding device decodes only I-SPC, for example.

[0233] The encoding device may change the frequency of using I-SPC according to the density or numerical value (quantity) of the objects in the world space to perform encoding. That is, the encoding device changes the frequency of selecting I-SPC according to the quantity or density of the objects included in the three-dimensional data. For example, the encoding device increases the usage frequency of the I space as the density of the objects in the world space is greater.

[0234] Further, the encoding device sets random access points in units of GOS, and stores information indicating the spatial region corresponding to the GOS in the header information.

[0235] The encoding device, for example, uses a default value as the spatial size of the GOS. Additionally, the encoding device can also change the size of the GOS according to the value (quantity) or density of the object or dynamic object. For example, when the object or dynamic object is denser or has a larger quantity, the encoding device sets the spatial size of the GOS to be smaller.

[0236] Moreover, the space or volume includes a group of feature points derived using information obtained from sensors such as depth sensors, gyroscopes, or cameras. The coordinates of the feature points are set as the center positions of the voxels. And through the subdivision of the voxels, high-precision position information can be achieved.

[0237] The group of feature points is derived using multiple pictures. The multiple pictures have at least the following two types of time information: actual time information, and the same time information in multiple pictures corresponding to the space (for example, the encoding time for rate control, etc.).

[0238] Moreover, encoding or decoding is performed in units of GOS including one or more spaces.

[0239] The encoding device and the decoding device predict the P space or B space within the GOS of the object to be processed by referring to the spaces within the GOS that have been processed.

[0240] Alternatively, the encoding device and the decoding device do not refer to different GOSs, and use the processed spaces within the GOS of the object to be processed to predict the P space or B space within the GOS of the object to be processed.

[0241] Moreover, the encoding device and the decoding device send or receive the encoded stream in units of a world space including one or more GOSs.

[0242] Moreover, the GOS has a layer structure in at least one direction within the world space, and the encoding device and the decoding device perform encoding or decoding starting from the lower layer. For example, the GOS that can be randomly accessed belongs to the lowest layer. The GOS belonging to the upper layer only refers to the GOSs belonging to the layers below the same layer. That is, the GOS is spatially divided in a predefined direction and includes multiple layers each having one or more SPCs. The encoding device and the decoding device perform encoding or decoding for each SPC by referring to the SPCs included in the same layer as or lower than the SPC.

[0243] Moreover, the encoding device and the decoding device continuously perform encoding or decoding on the GOSs within the unit of the world space including multiple GOSs. The encoding device and the decoding device write or read the information indicating the order (direction) of encoding or decoding as metadata. That is, the encoded data includes the information indicating the encoding order of multiple GOSs.

[0244] Furthermore, the encoding device and the decoding device perform encoding or decoding in parallel for two or more different spaces or GOSs.

[0245] Furthermore, the encoding device and the decoding device perform encoding or decoding on the spatial information (coordinates, size, etc.) of the space or GOS.

[0246] Furthermore, the encoding device and the decoding device perform encoding or decoding on the space or GOS included in a specific space determined according to external information such as GPS, path information, or magnification related to its own position or / and area size.

[0247] The encoding device or the decoding device performs encoding or decoding by making the priority of the space far from its own position lower than that of the space close to its own position.

[0248] The encoding device sets a direction in the world space according to magnification or use, and encodes the GOS having a layer structure in that direction. And the decoding device preferentially decodes from the lower layer for the GOS having a layer structure in one direction of the world space set according to magnification or use.

[0249] The encoding device changes the feature point extraction, object recognition accuracy, or spatial area size, etc. included in the indoor and outdoor spaces. However, the encoding device and the decoding device encode or decode the indoor GOS and the outdoor GOS with adjacent coordinates in the world space, and also encode or decode by corresponding these identifiers.

[0250] (Embodiment 2)

[0251] When using the encoded data of the point cloud for an actual device or service, in order to suppress the network bandwidth, it is desired to transmit and receive the required information according to the use. However, such a function does not exist in the existing encoding structure of three-dimensional data, so there is no corresponding encoding method.

[0252] What will be described in this embodiment is a three-dimensional data encoding method and a three-dimensional data encoding device for providing a function of transmitting and receiving the required information according to the use in the encoded data of the three-dimensional point cloud, and a three-dimensional data decoding method and a three-dimensional data decoding device for decoding the encoded data.

[0253] Define a voxel (VXL) having a certain amount or more of feature amount as a feature voxel (FVXL), and define the world space (WLD) composed of FVXL as a sparse world space (SWLD). Figure 11It shows a sparse world space and a configuration example of the world space. In SWLD, it includes: FGOS, which is a GOS composed of FVXL; FSPC, which is an SPC composed of FVXL; and FVLM, which is a VLM composed of FVXL. The data structures and prediction structures of FGOS, FSPC, and FVLM can be the same as those of GOS, SPC, and VLM.

[0254] The feature quantity refers to a feature quantity that represents the three-dimensional position information of VXL or the visible light information at the VXL position. In particular, more feature quantities can be detected at the corners and edges of three-dimensional objects, etc. Specifically, although the feature quantity is the three-dimensional feature quantity or the visible light feature quantity described below, as long as it is a feature quantity that represents the position, brightness, or color information of VXL, etc., it can be any feature quantity.

[0255] As the three-dimensional feature quantity, the SHOT feature quantity (Signature of Histograms of OrienTations), the PFH feature quantity (Point Feature Histograms), or the PPF feature quantity (Point Pair Feature) is adopted.

[0256] The SHOT feature quantity is obtained by dividing the periphery of VXL and calculating the inner product of the normal vector of the reference point and the divided area, and then performing histogramming. The SHOT feature quantity has the characteristics of high dimensionality and high feature expressiveness.

[0257] The PFH feature quantity is obtained by selecting multiple two-point groups near VXL, calculating the normal vector, etc. based on these two points, and then performing histogramming. Since the PFH feature quantity is a histogram feature, it is robust to a small amount of interference and has the characteristic of high feature expressiveness.

[0258] The PPF feature quantity is a feature quantity calculated using the normal vector, etc. according to two VXLs. In this PPF feature quantity, since all VXLs are used, it is robust to occlusion.

[0259] And, as the feature quantity of visible light, SIFT (Scale-Invariant Feature Transform), SURF (Speeded Up Robust Features), or HOG (Histogram of Oriented Gradients), etc. that adopt information such as the brightness gradient information of the image can be used.

[0260] The SWLD is generated by calculating the above-described feature amounts from each VXL of the WLD and extracting the FVXL. Here, the SWLD can be updated each time the WLD is updated, or it can be updated periodically after a certain period of time regardless of the update timing of the WLD.

[0261] The SWLD can be generated for each feature amount. For example, as shown by SWLD1 based on the SHOT feature amount and SWLD2 based on the SIFT feature amount, the SWLD can be generated separately for each feature amount and used according to the application. Also, the feature amounts of the calculated FVXLs can be stored as feature amount information in each FVXL.

[0262] Next, the method of using the sparse world space (SWLD) will be described. Since the SWLD only contains the feature voxels (FVXL), the data size is generally smaller compared to the WLD that includes all the VXLs.

[0263] In an application that uses feature amounts to achieve a certain purpose, by using the information of the SWLD instead of the WLD, it is possible to suppress the read time from the hard disk and also suppress the bandwidth and transmission time during network transmission. For example, as map information, the WLD and the SWLD are stored in the server in advance, and by switching the transmitted map information to the WLD or the SWLD according to the demand from the client, it is possible to suppress the network bandwidth and the transmission time. Specific examples are shown below.

[0264] Figure 12 And Figure 13 Shows an example of using the SWLD and the WLD. As Figure 12 shown, when the client 1 as a vehicle-mounted device needs map information for its own position determination, the client 1 sends a request (S301) for obtaining map data for its own position estimation to the server. The server sends the SWLD to the client 1 according to this acquisition request (S302). The client 1 uses the received SWLD to determine its own position (S303). At this time, the client 1 obtains the VXL information around the client 1 by various methods such as a distance sensor such as a rangefinder, a stereo camera, or a combination of multiple monocular cameras, and estimates its own position information based on the obtained VXL information and the SWLD. Here, the own position information includes the three-dimensional position information and the orientation of the client 1.

[0265] As Figure 13As shown, when the client 2, which is a vehicle-mounted device, needs map information for uses such as map rendering of a three-dimensional map or the like, the client 2 sends a request for obtaining map data for map rendering to the server (S311). The server sends the WLD to the client 2 according to this acquisition request (S312). The client 2 uses the received WLD for map rendering (S313). At this time, the client 2, for example, uses an image captured by its own visible light camera or the like and the WLD obtained from the server to create a conceptual image, and depicts the created image on a screen such as an in-vehicle navigation system.

[0266] As described above, the server sends the SWLD to the client in applications that mainly require the feature amounts of each VXL for self-position estimation, and sends the WLD to the client in cases where detailed VXL information is required, such as map rendering. Accordingly, map data can be efficiently transmitted and received.

[0267] In addition, the client can determine which of the SWLD and the WLD it needs and request the server to send the SWLD or the WLD. Also, the server can determine which of the SWLD or the WLD should be sent according to the status of the client or the network.

[0268] Next, a method for switching the transmission and reception of the sparse world space (SWLD) and the world space (WLD) will be described.

[0269] The reception of the WLD or the SWLD can be switched according to the network bandwidth. Figure 14 A working example in this case is shown. For example, when a low-speed network with a network bandwidth such as that in an LTE (Long Term Evolution) environment is used, when the client accesses the server via the low-speed network (S321), it obtains the SWLD as map information from the server (S322). Also, when a high-speed network with a surplus network bandwidth such as in a WiFi environment is used, the client accesses the server via the high-speed network (S323) and obtains the WLD from the server (S324). Accordingly, the client can obtain appropriate map information according to the network bandwidth of the client.

[0270] Specifically, the client receives the SWLD via LTE outdoors, and when entering an indoor area such as a facility, it obtains the WLD via WiFi. Accordingly, the client can obtain more detailed map information for the indoor area.

[0271] In this way, the client can request the WLD or SWLD from the server according to the frequency band of the network it uses. Alternatively, the client can send the information indicating the frequency band of the network it uses to the server, and the server sends appropriate data (WLD or SWLD) to the client according to this information. Or, the server can determine the network bandwidth of the client and send appropriate data (WLD or SWLD) to the client.

[0272] Moreover, the reception of the WLD or SWLD can be switched according to the moving speed. Figure 15 A working example in this case is shown. For example, when the client is moving at high speed (S331), the client receives the SWLD from the server (S332). In addition, when the client is moving at low speed (S333), the client receives the WLD from the server (S334). Accordingly, the client can both suppress the network bandwidth and obtain map information according to the speed. Specifically, when the client is driving on a highway, by receiving the SWLD with less data volume, the map information can be updated at an appropriate speed approximately. In addition, when the client is driving on an ordinary road, by receiving the WLD, more detailed map information can be obtained.

[0273] In this way, the client can request the WLD or SWLD from the server according to its own moving speed. Alternatively, the client can send the information indicating its own moving speed to the server, and the server sends appropriate data (WLD or SWLD) to the client according to this information. Or, the server can determine the moving speed of the client and send appropriate data (WLD or SWLD) to the client.

[0274] Also, it can be that the client first obtains the SWLD from the server and then obtains the WLD of the important areas therein. For example, when the client obtains map data, it first obtains the general map information with the SWLD, screens out the areas where features such as buildings, signs, or people appear more, and then obtains the WLD of the screened areas. Accordingly, the client can both suppress the amount of received data from the server and obtain the detailed information of the required areas.

[0275] Also, it can be that the server separately creates the SWLD for each object according to the WLD, and the client receives them separately according to the usage. Accordingly, the network bandwidth can be suppressed. For example, the server pre-identifies people or vehicles from the WLD and creates the SWLD of people and the SWLD of vehicles. When the client wants to obtain information about the people around, it receives the SWLD of people, and when it wants to obtain information about vehicles, it receives the SWLD of vehicles. And the types of such SWLD can be distinguished according to the information (marks or types, etc.) attached to the head.

[0276] Next, the configuration and operation process of the three-dimensional data encoding device (e.g., server) according to this embodiment will be described. Figure 16 is a block diagram of the three-dimensional data encoding device 400 according to this embodiment. Figure 17 is a flowchart of the three-dimensional data encoding process performed by the three-dimensional data encoding device 400.

[0277] Figure 16 The three-dimensional data encoding device 400 shown generates encoded three-dimensional data 413 and 414 as an encoded stream by encoding the input three-dimensional data 411. Here, the encoded three-dimensional data 413 is the encoded three-dimensional data corresponding to the WLD, and the encoded three-dimensional data 414 is the encoded three-dimensional data corresponding to the SWLD. The three-dimensional data encoding device 400 includes: an acquisition unit 401, an encoding region determination unit 402, an SWLD extraction unit 403, a WLD encoding unit 404, and an SWLD encoding unit 405.

[0278] As Figure 17 shown, first, the acquisition unit 401 acquires the input three-dimensional data 411 (S401) as point cloud data in a three-dimensional space.

[0279] Next, the encoding region determination unit 402 determines the spatial region to be encoded according to the spatial region where the point cloud data exists (S402).

[0280] Next, the SWLD extraction unit 403 defines the spatial region to be encoded as the WLD, calculates the feature amount according to each VXL included in the WLD. And the SWLD extraction unit 403 extracts the VXL whose feature amount is above a preset threshold, defines the extracted VXL as the FVXL, and generates the extracted three-dimensional data 412 by adding the FVXL to the SWLD (S403). That is, the extracted three-dimensional data 412 whose feature amount is above the threshold is extracted from the input three-dimensional data 411.

[0281] Next, the WLD encoding unit 404 generates the encoded three-dimensional data 413 corresponding to the WLD by encoding the input three-dimensional data 411 corresponding to the WLD (S404). At this time, the WLD encoding unit 404 attaches information for distinguishing that the encoded three-dimensional data 413 is a stream containing the WLD to the header of the encoded three-dimensional data 413.

[0282] And, the SWLD encoding unit 405 generates the encoded three-dimensional data 414 corresponding to the SWLD by encoding the extracted three-dimensional data 412 corresponding to the SWLD (S405). At this time, the SWLD encoding unit 405 attaches information for distinguishing that the encoded three-dimensional data 414 is a stream containing the SWLD to the header of the encoded three-dimensional data 414.

[0283] Also, the processing order of the process of generating the encoded three-dimensional data 413 and the process of generating the encoded three-dimensional data 414 may be opposite to the above. Also, part or all of the above processes may be executed in parallel.

[0284] Information assigned to the headers of the encoded three-dimensional data 413 and 414 is defined as a parameter such as "world_type", for example. When world_type = 0, it indicates that the stream contains WLD, and when world_type = 1, it indicates that the stream contains SWLD. When defining more other categories, the assigned value can be increased, such as world_type = 2. Also, a specific flag may be included in one of the encoded three-dimensional data 413 and 414. For example, the encoded three-dimensional data 414 may be assigned a flag indicating that the stream contains SWLD. In this case, the decoding device can determine whether the stream contains WLD or SWLD based on the presence or absence of the flag.

[0285] Also, the encoding method used by the WLD encoding unit 404 when encoding WLD may be different from the encoding method used by the SWLD encoding unit 405 when encoding SWLD.

[0286] For example, since SWLD data is decimated, the correlation with surrounding data may be lower compared to WLD. Therefore, in the encoding method for SWLD, inter-frame prediction among intra-frame prediction and inter-frame prediction is prioritized compared to the encoding method for WLD.

[0287] Also, it may be that the representation method of the three-dimensional position is different between the encoding method for SWLD and the encoding method for WLD. For example, it may be that the three-dimensional position of FVXL is represented by three-dimensional coordinates in FWLD, and the three-dimensional position is represented by an octree described later in WLD, and vice versa.

[0288] Also, the SWLD encoding unit 405 encodes in such a way that the data size of the encoded three-dimensional data 414 of SWLD is smaller than the data size of the encoded three-dimensional data 413 of WLD. As described above, for example, the correlation between data may decrease for SWLD compared to WLD. Accordingly, the encoding efficiency decreases, and the data size of the encoded three-dimensional data 414 may be larger than the data size of the encoded three-dimensional data 413 of WLD. Therefore, when the data size of the obtained encoded three-dimensional data 414 is larger than the data size of the encoded three-dimensional data 413 of WLD, the SWLD encoding unit 405 re-encodes to regenerate the encoded three-dimensional data 414 with a reduced data size.

[0289] For example, the SWLD extraction unit 403 regenerates the extracted three-dimensional data 412 with a reduced number of extracted feature points, and the SWLD encoding unit 405 encodes the extracted three-dimensional data 412. Alternatively, the quantization level in the SWLD encoding unit 405 can be made coarser. For example, in the octree structure described later, by rounding the data in the bottom layer, the quantization level can be made coarser.

[0290] Moreover, when the SWLD encoding unit 405 cannot make the data size of the encoded three-dimensional data 414 of SWLD smaller than the data size of the encoded three-dimensional data 413 of WLD, the SWLD encoding unit 405 may not generate the encoded three-dimensional data 414 of SWLD. Alternatively, the encoded three-dimensional data 413 of WLD can be copied to the encoded three-dimensional data 414 of SWLD. That is, the encoded three-dimensional data 413 of WLD can be directly used as the encoded three-dimensional data 414 of SWLD.

[0291] Next, the configuration and the working process of the three-dimensional data decoding device (such as a client) according to the present embodiment will be described. Figure 18 FIG. is a block diagram of the three-dimensional data decoding device 500 according to the present embodiment. Figure 19 FIG. is a flowchart of the three-dimensional data decoding process performed by the three-dimensional data decoding device 500.

[0292] Figure 18 The shown three-dimensional data decoding device 500 decodes the encoded three-dimensional data 511 to generate the decoded three-dimensional data 512 or 513. Here, the encoded three-dimensional data 511 is, for example, the encoded three-dimensional data 413 or 414 generated by the three-dimensional data encoding device 400.

[0293] The three-dimensional data decoding device 500 includes: an acquisition unit 501, a header analysis unit 502, a WLD decoding unit 503, and an SWLD decoding unit 504.

[0294] As shown in Figure 19 , first, the acquisition unit 501 acquires the encoded three-dimensional data 511 (S501). Next, the header analysis unit 502 analyzes the header of the encoded three-dimensional data 511 to determine whether the encoded three-dimensional data 511 includes a WLD stream or an SWLD stream (S502). For example, the determination is made by referring to the above-mentioned parameter of world_type.

[0295] When the encoded three-dimensional data 511 is a stream containing WLD (Yes in S503), the WLD decoding unit 503 decodes the encoded three-dimensional data 511 to generate decoded three-dimensional data 512 of WLD (S504). On the other hand, when the encoded three-dimensional data 511 is a stream containing SWLD (No in S503), the SWLD decoding unit 504 decodes the encoded three-dimensional data 511 to generate decoded three-dimensional data 513 of SWLD (S505).

[0296] Also, similar to the encoding device, the decoding method used by the WLD decoding unit 503 when decoding WLD and the decoding method used by the SWLD decoding unit 504 when decoding SWLD can be different. For example, in the decoding method for SWLD, inter-frame prediction in intra-frame prediction and inter-frame prediction can be given priority compared to the decoding method for WLD.

[0297] Also, in the decoding method for SWLD and the decoding method for WLD, the representation methods of three-dimensional positions can be different. For example, in SWLD, the three-dimensional position of FVXL can be represented by three-dimensional coordinates, and in WLD, the three-dimensional position can be represented by an octree described later, and vice versa.

[0298] Next, the octree representation as a method of representing three-dimensional positions will be described. The VXL data included in the three-dimensional data is converted into an octree structure and then encoded. Figure 20 An example of VXL of WLD is shown. Figure 21 Shows Figure 20 the octree structure of the WLD shown. In Figure 20 the example shown, there are three VXLs 1 to 3 as VXLs (hereinafter, valid VXLs) containing point groups. As Figure 21 shown, the octree structure is composed of nodes and leaf nodes. Each node has a maximum of 8 nodes or leaf nodes. Each leaf node has VXL information. Here, Figure 21 among the leaf nodes shown, leaf nodes 1, 2, and 3 respectively represent Figure 20 the VXL1, VXL2, and VXL3 shown.

[0299] Specifically, each node and leaf node correspond to a three-dimensional position. Node 1 corresponds to Figure 20 all the blocks shown. The block corresponding to Node 1 is divided into 8 blocks. Among the 8 blocks, the blocks including valid VXLs are set as nodes, and the other blocks are set as leaf nodes. The block corresponding to the node is further divided into 8 nodes or leaf nodes, and the number of times this process is repeated is the same as the number of levels in the tree structure. And all the blocks in the lowest layer are set as leaf nodes.

[0300] Also,Figure 22 shows an example of an SWLD generated from the Figure 20 shown WLD. Figure 20 The results of feature quantity extraction of the shown VXL1 and VXL2 are judged as FVXL1 and FVXL2 and added to the SWLD. In addition, since VXL3 is not judged as FVXL, it is not included in the SWLD. Figure 23 shows Figure 22 the octree structure of the shown SWLD. In Figure 23 the shown octree structure, Figure 21 the leaf node 3 corresponding to VXL3 shown is deleted. Accordingly, Figure 21 the node 3 shown has no valid VXL and is changed to a leaf node. In this way, generally speaking, the number of leaf nodes of the SWLD is smaller than that of the WLD, and the encoded three-dimensional data of the SWLD is also smaller than that of the WLD.

[0301] The following describes a modification example of the present embodiment.

[0302] For example, it may also be the case that when a client such as an in-vehicle device estimates its own position, receives the SWLD from the server, uses the SWLD to estimate its own position, and performs obstacle detection, various methods such as a distance sensor such as a rangefinder, a stereo camera, or a combination of multiple monocular cameras are used to perform obstacle detection based on the three-dimensional information of the surrounding area obtained by itself.

[0303] And generally speaking, it is difficult to include VXL data of a flat area in the SWLD. For this reason, the server maintains a downsampled world space (SubWLD) obtained by downsampling the WLD for detecting stationary obstacles, and can send the SWLD and the SubWLD to the client. Accordingly, both the network bandwidth can be suppressed and the client side can perform its own position estimation and obstacle detection.

[0304] And when the client quickly depicts three-dimensional map data, it may be convenient if the map information is in a grid structure. Then, the server can generate a grid based on the WLD and keep it in advance as a grid world space (MWLD). For example, when the client needs to perform rough three-dimensional depiction, it receives the MWLD, and when it needs to perform detailed three-dimensional depiction, it receives the WLD. Accordingly, the network bandwidth can be suppressed.

[0305] Also, although the server sets the VXLs with feature amounts above the threshold as FVXLs from each VXL, the FVXLs can also be calculated by different methods. For example, if the server determines that VXLs, VLMs, SPCs, or GOSs that make up signals or intersections are required for its own position estimation, driving assistance, or autonomous driving, etc., they can be included in the SWLD as FVXLs, FVLMs, FSPCs, or FGOSs. Also, the above determination can be made manually. In addition, FVXLs obtained by the above method can be added to the FVXLs, etc. set based on the feature amount. That is, the SWLD extraction unit 403 can further extract, from the input three-dimensional data 411, the data corresponding to the object having a predetermined attribute as the extracted three-dimensional data 412.

[0306] Also, different labels can be assigned to the situations required for these uses from the feature amount. The server can separately hold the FVXLs required for its own position estimation, driving assistance, or autonomous driving, such as signals or intersections, as the upper layer of the SWLD (for example, the lane world space).

[0307] Also, the server can attach attributes to the VXLs in the WLD in units of random access or prescribed units. The attributes include, for example, information indicating whether it is required or not required for its own position estimation, or information indicating whether it is important as traffic information such as a signal or an intersection. The attributes can also include the correspondence relationship with Features (intersections or roads, etc.) in the lane information (GDF: Geographic Data Files, etc.).

[0308] Also, as a method for updating the WLD or SWLD, the following method can be adopted.

[0309] Update information showing changes in people, construction, or street trees (facing the trajectory), etc. is loaded into the server as a point cloud or metadata. The server updates the WLD based on this loading, and then uses the updated WLD to update the SWLD.

[0310] Also, when the client detects a mismatch between the three-dimensional information generated by itself during its own position estimation and the three-dimensional information received from the server, the three-dimensional information generated by itself can be sent to the server together with an update notification. In this case, the server updates the SWLD using the WLD. If the SWLD is not updated, the server determines that the WLD itself is old.

[0311] Also, as the header information of the encoded stream, although information for distinguishing between WLD and SWLD is attached, for example, in a case where there are multiple world spaces such as a grid world space or a lane world space, information for distinguishing them can be attached to the header information. Also, in a case where there are multiple SWLDs with different feature amounts, information for distinguishing them separately can also be attached to the header information.

[0312] Also, although SWLD is composed of FVXLs, it can also include VXLs that are not judged as FVXLs. For example, SWLD can include adjacent VXLs used when calculating the feature amounts of FVXLs. Accordingly, even when no feature amount information is attached to each FVXL of SWLD, the client can calculate the feature amounts of FVXLs when receiving SWLD. Also, at this time, SWLD can include information for distinguishing whether each VXL is an FVXL or a VXL.

[0313] As described above, the three-dimensional data encoding device 400 extracts the extracted three-dimensional data 412 (second three-dimensional data) whose feature amount is equal to or greater than the threshold from the input three-dimensional data 411 (first three-dimensional data), and generates the encoded three-dimensional data 414 (first encoded three-dimensional data) by encoding the extracted three-dimensional data 412.

[0314] Accordingly, the three-dimensional data encoding device 400 generates the encoded three-dimensional data 414 obtained by encoding data whose feature amount is equal to or greater than the threshold. In this way, compared with the case of directly encoding the input three-dimensional data 411, the data amount can be reduced. Therefore, the three-dimensional data encoding device 400 can reduce the data amount during transmission.

[0315] Also, the three-dimensional data encoding device 400 further generates the encoded three-dimensional data 413 (second encoded three-dimensional data) by encoding the input three-dimensional data 411.

[0316] Accordingly, the three-dimensional data encoding device 400 can selectively transmit the encoded three-dimensional data 413 and the encoded three-dimensional data 414, for example, according to the usage purpose and the like.

[0317] Also, the extracted three-dimensional data 412 is encoded by the first encoding method, and the input three-dimensional data 411 is encoded by the second encoding method different from the first encoding method.

[0318] Accordingly, the three-dimensional data encoding device 400 can adopt appropriate encoding methods for the input three-dimensional data 411 and the extracted three-dimensional data 412 respectively.

[0319] Also, in the first encoding method, among intra prediction and inter prediction, inter prediction is prioritized compared with the second encoding method.

[0320] Accordingly, the three-dimensional data encoding device 400 can increase the priority of inter-frame prediction for the extracted three-dimensional data 412 where the correlation between adjacent data is likely to be low.

[0321] Moreover, in the first encoding method and the second encoding method, the representation methods of three-dimensional positions are different. For example, in the second encoding method, the three-dimensional position is represented by an octree, and in the first encoding method, the three-dimensional position is represented by three-dimensional coordinates.

[0322] Accordingly, the three-dimensional data encoding device 400 can adopt a more appropriate representation method of three-dimensional positions for three-dimensional data with different numbers of data (the number of VXL or FVXL).

[0323] Moreover, in at least one of the encoded three-dimensional data 413 and 414, there is an identifier indicating whether the encoded three-dimensional data is the encoded three-dimensional data obtained by encoding the input three-dimensional data 411 or the encoded three-dimensional data obtained by encoding a part of the input three-dimensional data 411. That is, this identifier indicates whether the encoded three-dimensional data is the encoded three-dimensional data 413 of WLD or the encoded three-dimensional data 414 of SWLD.

[0324] Accordingly, the decoding device can easily determine whether the acquired encoded three-dimensional data is the encoded three-dimensional data 413 or the encoded three-dimensional data 414.

[0325] Moreover, the three-dimensional data encoding device 400 encodes the extracted three-dimensional data 412 such that the data amount of the encoded three-dimensional data 414 is less than the data amount of the encoded three-dimensional data 413.

[0326] Accordingly, the three-dimensional data encoding device 400 can make the data amount of the encoded three-dimensional data 414 less than the data amount of the encoded three-dimensional data 413.

[0327] Moreover, the three-dimensional data encoding device 400 further extracts, as the extracted three-dimensional data 412, the data corresponding to an object having a predetermined attribute from the input three-dimensional data 411. For example, an object having a predetermined attribute refers to an object required in self-position estimation, driving assistance, or autonomous driving, etc., such as a signal or an intersection.

[0328] Accordingly, the three-dimensional data encoding device 400 can generate the encoded three-dimensional data 414 including the data required by the decoding device.

[0329] Moreover, the three-dimensional data encoding device 400 (server) further sends one of the encoded three-dimensional data 413 and 414 to the client according to the state of the client.

[0330] Accordingly, the three-dimensional data encoding device 400 can send appropriate data according to the state of the client.

[0331] Moreover, the state of the client includes the communication status of the client (such as network bandwidth) or the moving speed of the client.

[0332] Furthermore, the three-dimensional data encoding device 400 further sends one of the encoded three-dimensional data 413 and 414 to the client according to the request of the client.

[0333] Accordingly, the three-dimensional data encoding device 400 can send appropriate data according to the request of the client.

[0334] In addition, the three-dimensional data decoding device 500 according to the present embodiment decodes the encoded three-dimensional data 413 or 414 generated by the above three-dimensional data encoding device 400.

[0335] That is, the three-dimensional data decoding device 500 decodes the encoded three-dimensional data 414 obtained by encoding the extracted three-dimensional data 412 whose feature amount extracted from the input three-dimensional data 411 is above the threshold value by the first decoding method. And the three-dimensional data decoding device 500 decodes the encoded three-dimensional data 413 obtained by encoding the input three-dimensional data 411 by using a second decoding method different from the first decoding method.

[0336] Accordingly, the three-dimensional data decoding device 500 can selectively receive, for example, according to the usage purpose, etc., the encoded three-dimensional data 414 and the encoded three-dimensional data 413 obtained by encoding the data whose feature amount is above the threshold value. Accordingly, the three-dimensional data decoding device 500 can reduce the amount of data during transmission. Moreover, the three-dimensional data decoding device 500 can adopt appropriate decoding methods for the input three-dimensional data 411 and the extracted three-dimensional data 412 respectively.

[0337] In addition, in the first decoding method, among intra prediction and inter prediction, inter prediction is prioritized compared with the second decoding method.

[0338] Accordingly, the three-dimensional data decoding device 500 can increase the priority of inter prediction for the extracted three-dimensional data whose correlation between adjacent data is likely to become low.

[0339] In addition, in the first decoding method and the second decoding method, the representation methods of three-dimensional positions are different. For example, in the second decoding method, the three-dimensional position is represented by an octree, and in the first decoding method, the three-dimensional position is represented by three-dimensional coordinates.

[0340] Accordingly, the three-dimensional data decoding device 500 can adopt a more appropriate representation method of three-dimensional positions for three-dimensional data with different numbers of data (the number of VXL or FVXL).

[0341] Moreover, at least one of the encoded three-dimensional data 413 and 414 includes an identifier indicating whether the encoded three-dimensional data is obtained by encoding the input three-dimensional data 411 or by encoding a part of the input three-dimensional data 411. The three-dimensional data decoding device 500 identifies the encoded three-dimensional data 413 and 414 with reference to this identifier.

[0342] Accordingly, the three-dimensional data decoding device 500 can easily determine whether the obtained encoded three-dimensional data is the encoded three-dimensional data 413 or the encoded three-dimensional data 414.

[0343] Moreover, the three-dimensional data decoding device 500 further notifies the server of the state of the client (the three-dimensional data decoding device 500). The three-dimensional data decoding device 500 receives one of the encoded three-dimensional data 413 and 414 sent from the server according to the state of the client.

[0344] Accordingly, the three-dimensional data decoding device 500 can receive appropriate data according to the state of the client.

[0345] Moreover, the state of the client includes the communication status of the client (such as network bandwidth) or the moving speed of the client.

[0346] Moreover, the three-dimensional data decoding device 500 further requests one of the encoded three-dimensional data 413 and 414 from the server and receives one of the encoded three-dimensional data 413 and 414 sent from the server according to this request.

[0347] Accordingly, the three-dimensional data decoding device 500 can receive appropriate data corresponding to the use.

[0348] (Embodiment 3)

[0349] In this embodiment, a method for transmitting and receiving three-dimensional data between vehicles is described. For example, three-dimensional data is transmitted and received between the host vehicle and surrounding vehicles.

[0350] Figure 24 is a block diagram of the three-dimensional data production device 620 according to this embodiment. The three-dimensional data production device 620 is included in the host vehicle, for example, and produces denser third three-dimensional data 636 by synthesizing the received second three-dimensional data 635 and the first three-dimensional data 632 produced by the three-dimensional data production device 620.

[0351] The three-dimensional data production device 620 includes: a three-dimensional data production unit 621, a request range determination unit 622, a search unit 623, a reception unit 624, a decoding unit 625, and a synthesis unit 626.

[0352] First, the 3D data creation unit 621 creates the first 3D data 632 using the sensor information 631 detected by the sensors equipped on its own vehicle. Next, the request range determination unit 622 determines the request range, which refers to the 3D space range where the data in the created first 3D data 632 is insufficient.

[0353] Next, the search unit 623 searches for surrounding vehicles that hold the 3D data within the request range, and sends the request range information 633 indicating the request range to the surrounding vehicles determined through the search. Next, the receiving unit 624 receives the encoded 3D data 634 (S624), which is the encoded stream of the request range, from the surrounding vehicles. Additionally, the search unit 623 can send requests to all vehicles existing within the determined range without discrimination, and receive the encoded 3D data 634 from the responding parties. Also, the search unit 623 is not limited to vehicles, and can also send requests to objects such as traffic lights or signs, and receive the encoded 3D data 634 from such objects.

[0354] Next, the received encoded 3D data 634 is decoded by the decoding unit 625 to obtain the second 3D data 635. Next, the first 3D data 632 and the second 3D data 635 are synthesized by the synthesis unit 626 to create a denser third 3D data 636.

[0355] Next, the configuration and operation of the 3D data transmission device 640 according to the present embodiment will be described. Figure 25 It is a block diagram of the 3D data transmission device 640.

[0356] The 3D data transmission device 640, for example, is included among the above-mentioned surrounding vehicles, processes the fifth 3D data 652 created by the surrounding vehicles into the sixth 3D data 654 requested by its own vehicle, generates the encoded 3D data 634 by encoding the sixth 3D data 654, and sends the encoded 3D data 634 to its own vehicle.

[0357] The 3D data transmission device 640 includes: a 3D data creation unit 641, a receiving unit 642, an extraction unit 643, an encoding unit 644, and a sending unit 645.

[0358] First, the 3D data creation unit 641 creates the fifth 3D data 652 using the sensor information 651 detected by the sensors equipped on the surrounding vehicles. Next, the receiving unit 642 receives the request range information 633 sent from its own vehicle.

[0359] Next, the extraction unit 643 extracts the three-dimensional data of the requested range represented by the requested range information 633 from the fifth three-dimensional data 652, and processes the fifth three-dimensional data 652 into the sixth three-dimensional data 654. Next, the encoding unit 644 encodes the sixth three-dimensional data 654 to generate the encoded three-dimensional data 634 as an encoded stream. Then, the transmission unit 645 transmits the encoded three-dimensional data 634 to its own vehicle.

[0360] In addition, although an example in which the own vehicle is equipped with the three-dimensional data production device 620 and the surrounding vehicles are equipped with the three-dimensional data transmission device 640 has been described here, each vehicle may also have the functions of the three-dimensional data production device 620 and the three-dimensional data transmission device 640.

[0361] (Embodiment 4)

[0362] In the present embodiment, the operation related to abnormal conditions in the own position estimation based on the three-dimensional map will be described.

[0363] The use of autonomous movement of moving bodies such as the autonomous driving of motor vehicles, robots, or flying objects such as drones will expand in the future. As an example of a method for realizing such autonomous movement, there is a method in which the moving body estimates its own position in the three-dimensional map (own position estimation) and travels according to the map.

[0364] The own position estimation is realized by matching the three-dimensional map with the three-dimensional information around the own vehicle (hereinafter referred to as the own vehicle detection three-dimensional data) obtained by sensors such as a distance measuring instrument (LIDAR, etc.) or a stereo camera mounted on the own vehicle, and estimating the position of the own vehicle in the three-dimensional map.

[0365] As shown in the HD map proposed by HERE Corporation, etc., the three-dimensional map can include not only three-dimensional point clouds, but also two-dimensional map data such as the shapes of roads and intersections, or information that changes in real time such as congestion and accidents. The three-dimensional map is composed of multiple layers such as three-dimensional data, two-dimensional data, and metadata that changes in real time. The device can obtain only the required data, or can also refer to the required data.

[0366] The data of the point cloud can be the above-mentioned SWLD, or can include point group data that is not feature points. And the transmission and reception of the data of the point cloud are basically performed in one or more random access units.

[0367] As a method for matching a three-dimensional map with three-dimensional data detected by the vehicle itself, the following method can be adopted. For example, the device compares the shapes of point groups in the point clouds of each other, and determines the part with a high similarity between feature points as the same position. Further, when the three-dimensional map is composed of SWLD, the device compares the feature points constituting the SWLD with the three-dimensional feature points extracted from the three-dimensional data detected by the vehicle itself and performs matching.

[0368] Here, in order to estimate the vehicle's own position with high accuracy, the following (A) and (B) need to be satisfied. (A) It is already possible to obtain a three-dimensional map and three-dimensional data detected by the vehicle itself. (B) Their accuracies satisfy a predetermined standard. However, in the following abnormal situations, (A) or (B) cannot be satisfied.

[0369] (1) The three-dimensional map cannot be obtained through the communication path.

[0370] (2) There is no three-dimensional map, or the obtained three-dimensional map is damaged.

[0371] (3) The sensor of the vehicle itself malfunctions, or due to bad weather, the generation accuracy of the three-dimensional data detected by the vehicle itself is insufficient.

[0372] The operations for coping with these abnormal situations will be described below. Although the operations will be described below taking a vehicle as an example, the following methods can also be applied to all moving objects such as robots or drones that perform autonomous movement.

[0373] What will be described below is the configuration and operation of the three-dimensional information processing device according to the present embodiment for coping with abnormal situations in the three-dimensional map or the three-dimensional data detected by the vehicle itself. Figure 26 It is a block diagram showing a configuration example of the three-dimensional information processing device 700 according to the present embodiment.

[0374] The three-dimensional information processing device 700 is mounted on a moving object such as a motor vehicle, for example. As Figure 26 shown, the three-dimensional information processing device 700 includes: a three-dimensional map acquisition unit 701, a vehicle self-detection data acquisition unit 702, an abnormal situation determination unit 703, a coping operation determination unit 704, and an operation control unit 705.

[0375] In addition, the three-dimensional information processing device 700 may include a camera for obtaining a two-dimensional image, or may include a two-dimensional or one-dimensional sensor (not shown) such as a sensor using ultrasonic waves or lasers for one-dimensional data for detecting a structural object or a moving object around the vehicle itself. Further, the three-dimensional information processing device 700 may include a communication unit (not shown) for obtaining a three-dimensional map through a mobile communication network such as 4G or 5G, or vehicle-to-vehicle communication, or road-to-vehicle communication.

[0376] The three-dimensional map acquisition unit 701 acquires a three-dimensional map 711 near the driving route. For example, the three-dimensional map acquisition unit 701 acquires the three-dimensional map 711 through a mobile communication network, vehicle-to-vehicle communication, or road-to-vehicle communication.

[0377] Next, the own vehicle detection data acquisition unit 702 acquires own vehicle detection three-dimensional data 712 based on sensor information. For example, the own vehicle detection data acquisition unit 702 generates own vehicle detection three-dimensional data 712 based on the sensor information obtained by the sensors equipped on the own vehicle.

[0378] Next, the abnormal situation determination unit 703 detects an abnormal situation by performing a pre-determined inspection on at least one of the acquired three-dimensional map 711 and the own vehicle detection three-dimensional data 712. That is, the abnormal situation determination unit 703 determines whether at least one of the acquired three-dimensional map 711 and the own vehicle detection three-dimensional data 712 is abnormal.

[0379] When an abnormal situation is detected, the countermeasure work determination unit 704 determines the countermeasure work for the abnormal situation. Next, the work control unit 705 controls the work of each processing unit required in the implementation of the countermeasure work, such as the three-dimensional map acquisition unit 701.

[0380] In addition, when no abnormal situation is detected, the three-dimensional information processing device 700 ends the processing.

[0381] Furthermore, the three-dimensional information processing device 700 estimates the own position of the vehicle equipped with the three-dimensional information processing device 700 by using the three-dimensional map 711 and the own vehicle detection three-dimensional data 712. Next, the three-dimensional information processing device 700 uses the result of the own position estimation to make the vehicle perform autonomous driving.

[0382] Accordingly, the three-dimensional information processing device 700 acquires map data (three-dimensional map 711) including first three-dimensional position information via a channel. For example, the first three-dimensional position information is encoded in units of partial spaces having three-dimensional coordinate information, and the first three-dimensional position information includes a plurality of random access units. Each of the plurality of random access units is an aggregate of one or more partial spaces and can be independently decoded. For example, the first three-dimensional position information is data (SWLD) in which feature points where three-dimensional feature amounts are above a specified threshold are encoded.

[0383] Furthermore, the three-dimensional information processing device 700 generates second three-dimensional position information (own vehicle detection three-dimensional data 712) based on the information detected by the sensors. Next, the three-dimensional information processing device 700 determines whether the first three-dimensional position information or the second three-dimensional position information is abnormal by performing an abnormality determination process on the first three-dimensional position information or the second three-dimensional position information.

[0384] When the three-dimensional information processing device 700 determines that the first three-dimensional position information or the second three-dimensional position information is abnormal, it determines the response operation for the abnormality. Next, the three-dimensional information processing device 700 executes the control required for the implementation of the response operation.

[0385] Accordingly, the three-dimensional information processing device 700 can detect the abnormality of the first three-dimensional position information or the second three-dimensional position information and can perform the response operation.

[0386] (Embodiment 5)

[0387] In the present embodiment, a method for transmitting three-dimensional data to a following vehicle and the like will be described.

[0388] Figure 27 FIG. is a block diagram showing a configuration example of a three-dimensional data production device 810 according to the present embodiment. The three-dimensional data production device 810 is mounted on a vehicle, for example. The three-dimensional data production device 810 transmits and receives three-dimensional data to and from external traffic cloud monitoring, a preceding vehicle, or a following vehicle, and at the same time produces and accumulates the three-dimensional data.

[0389] The three-dimensional data production device 810 includes: a data reception unit 811, a communication unit 812, a reception control unit 813, a format conversion unit 814, a plurality of sensors 815, a three-dimensional data production unit 816, a three-dimensional data synthesis unit 817, a three-dimensional data storage unit 818, a communication unit 819, a transmission control unit 820, a format conversion unit 821, and a data transmission unit 822.

[0390] The data reception unit 811 receives three-dimensional data 831 from traffic cloud monitoring or a preceding vehicle. The three-dimensional data 831 includes, for example, point clouds, visible light images, depth information, sensor position information, or speed information that contain information on areas that cannot be detected by the sensors 815 of the own vehicle.

[0391] The communication unit 812 communicates with traffic cloud monitoring or a preceding vehicle and sends a data transmission request or the like to traffic cloud monitoring or a preceding vehicle.

[0392] The reception control unit 813 exchanges information such as corresponding formats with the communication partner via the communication unit 812 and establishes communication with the communication partner.

[0393] The format conversion unit 814 generates the three-dimensional data 832 by performing format conversion or the like on the three-dimensional data 831 received by the data reception unit 811. Further, when the three-dimensional data 831 is compressed or encoded, the format conversion unit 814 performs decompression or decoding processing.

[0394] The plurality of sensors 815 are a group of sensors such as LiDAR, visible light cameras, or infrared cameras that acquire information about the outside of the vehicle, and generate sensor information 833. For example, when the sensor 815 is a laser sensor such as LiDAR, the sensor information 833 is three-dimensional data such as point cloud (point group data). Additionally, the sensor 815 may not be plural.

[0395] The three-dimensional data creation unit 816 generates three-dimensional data 834 based on the sensor information 833. The three-dimensional data 834 includes, for example, information such as point cloud, visible light images, depth information, sensor position information, or speed information.

[0396] The three-dimensional data synthesis unit 817 synthesizes the three-dimensional data 832 created by traffic cloud monitoring or the like of the preceding vehicle into the three-dimensional data 834 created based on the sensor information 833 of its own vehicle, thereby enabling the construction of three-dimensional data 835 that includes the space in front of the preceding vehicle that cannot be detected by the sensors 815 of its own vehicle.

[0397] The three-dimensional data storage unit 818 stores the generated three-dimensional data 835 and the like.

[0398] The communication unit 819 communicates with traffic cloud monitoring or the following vehicle, and sends a data transmission request or the like to traffic cloud monitoring or the following vehicle.

[0399] The transmission control unit 820 exchanges information such as the corresponding format with the communication partner via the communication unit 819, and establishes communication with the communication partner. Further, the transmission control unit 820 determines the transmission area of the space of the three-dimensional data to be transmitted based on the three-dimensional data construction information of the three-dimensional data 832 generated by the three-dimensional data synthesis unit 817 and the data transmission request from the communication partner.

[0400] Specifically, the transmission control unit 820 determines the transmission area that includes the space in front of its own vehicle that cannot be detected by the sensors of the following vehicle in accordance with the data transmission request from traffic cloud monitoring or the following vehicle. Further, the transmission control unit 820 determines the transmission area by judging, based on the three-dimensional data construction information, whether there is an update to the space that can be transmitted or the space that has already been transmitted. For example, the transmission control unit 820 determines as the transmission area the area that is both specified by the data transmission request and where the corresponding three-dimensional data 835 exists. Further, the transmission control unit 820 notifies the format conversion unit 821 of the format corresponding to the communication partner and the transmission area.

[0401] The format conversion unit 821 generates the three-dimensional data 837 by converting the three-dimensional data 836 of the transmission area in the three-dimensional data 835 stored in the three-dimensional data storage unit 818 into a format corresponding to the receiving side. Additionally, the format conversion unit 821 can also compress or encode the three-dimensional data 837 to reduce the data volume.

[0402] The data transmission unit 822 transmits the three-dimensional data 837 to the traffic cloud monitoring or the following vehicle. This three-dimensional data 837 includes, for example, the point cloud in front of the host vehicle containing information on the area that is a blind spot for the following vehicle, visible light images, depth information, or sensor position information, etc.

[0403] In addition, although the format conversion units 814 and 821 are taken as examples for format conversion and the like here, format conversion may not be performed.

[0404] With this configuration, the three-dimensional data production device 810 obtains the three-dimensional data 831 of the area that cannot be detected by the sensors 815 of the host vehicle from the outside, and generates the three-dimensional data 835 by synthesizing the three-dimensional data 831 and the three-dimensional data 834 based on the sensor information 833 detected by the sensors 815 of the host vehicle. Accordingly, the three-dimensional data production device 810 can generate the three-dimensional data of the range that cannot be detected by the sensors 815 of the host vehicle.

[0405] Moreover, the three-dimensional data production device 810 can transmit the three-dimensional data of the space in front of the host vehicle that cannot be detected by the sensors of the following vehicle to the traffic cloud monitoring or the following vehicle, etc., in accordance with a data transmission request from the traffic cloud monitoring or the following vehicle.

[0406] (Embodiment 6)

[0407] In the example to be described in Embodiment 5, the client device such as a vehicle transmits the three-dimensional data to other vehicles or servers such as traffic cloud monitoring. In this embodiment, the client device transmits the sensor information obtained by the sensor to the server or other client devices.

[0408] First, the configuration of the system according to this embodiment will be described. Figure 28 The configuration of the three-dimensional map and the transceiver system of the sensor information according to this embodiment is shown. This system includes a server 901, and client devices 902A and 902B. Additionally, when not specifically distinguishing between the client devices 902A and 902B, they are also denoted as the client device 902.

[0409] The client device 902 is, for example, an in-vehicle device mounted on a moving body such as a vehicle. The server 901 is, for example, a traffic cloud monitoring device or the like that can communicate with a plurality of client devices 902.

[0410] The server 901 sends a three-dimensional map composed of point clouds to the client device 902. In addition, the composition of the three-dimensional map is not limited to point clouds and can also be represented by other three-dimensional data such as a grid structure.

[0411] The client device 902 sends sensor information obtained by the client device 902 to the server 901. The sensor information includes, for example, at least one of LiDAR acquired information, visible light image, infrared image, depth image, sensor position information, and speed information.

[0412] Regarding the data transmitted and received between the server 901 and the client device 902, it can be compressed when reducing data is desired, and can be not compressed when maintaining data accuracy is desired. When compressing the data, for example, an octree-based three-dimensional compression method can be adopted in the point cloud. Also, a two-dimensional image compression method can be adopted for visible light images, infrared images, and depth images. The two-dimensional image compression method is, for example, MPEG-4 AVC or HEVC standardized by MPEG.

[0413] Also, the server 901 sends the three-dimensional map managed by the server 901 to the client device 902 in accordance with a transmission request for the three-dimensional map from the client device 902. In addition, the server 901 may send the three-dimensional map without waiting for a transmission request for the three-dimensional map from the client device 902. For example, the server 901 may broadcast the three-dimensional map to one or more client devices 902 in a pre-specified space. Also, the server 901 may send a three-dimensional map suitable for the position of the client device 902 to the client device 902 that has received a transmission request once at regular intervals. Also, the server 901 may send the three-dimensional map to the client device 902 whenever the three-dimensional map managed by the server 901 is updated.

[0414] The client device 902 sends a transmission request for the three-dimensional map to the server 901. For example, when the client device 902 wants to estimate its own position while driving, the client device 902 sends a transmission request for the three-dimensional map to the server 901.

[0415] In addition, in the following situations, the client device 902 may also send a request to the server 901 to send a 3D map. When the 3D map held by the client device 902 is relatively old, the client device 902 may also send a request to the server 901 to send a 3D map. For example, when a certain period of time has passed since the client device 902 obtained the 3D map, the client device 902 may also send a request to the server 901 to send a 3D map.

[0416] It may also be that, before a certain moment when the client device 902 is about to leave the space shown in the 3D map held by the client device 902, the client device 902 sends a request to the server 901 to send a 3D map. For example, it may also be that when the client device 902 is within a pre-specified distance from the boundary of the space shown in the 3D map held by the client device 902, the client device 902 sends a request to the server 901 to send a 3D map. Moreover, when the movement path and movement speed of the client device 902 are known, the moment when the client device 902 leaves the space shown in the 3D map held by the client device 902 can be predicted based on the known movement path and movement speed.

[0417] When the error in the position comparison between the 3D data generated by the client device 902 based on sensor information and the 3D map is above a certain range, the client device 902 may send a request to the server 901 to send a 3D map.

[0418] The client device 902 sends the sensor information to the server 901 in accordance with the request to send the sensor information sent from the server 901. In addition, the client device 902 may also send the sensor information to the server 901 without waiting for the request to send the sensor information from the server 901. For example, when the client device 902 has received a request to send the sensor information from the server 901 once, it may regularly send the sensor information to the server 901 within a certain period. It may also be that when the error in the position comparison between the 3D data generated by the client device 902 based on sensor information and the 3D map obtained from the server 901 is above a certain range, the client device 902 determines that there is a possibility that the 3D map around the client device 902 has changed, and sends this judgment result together with the sensor information to the server 901.

[0419] Server 901 sends a request to client device 902 to send sensor information. For example, server 901 receives the location information of client device 902 such as GPS from client device 902. When server 901 determines, based on the location information of client device 902, that client device 902 is approaching a space with less information in the 3D map managed by server 901, in order to regenerate the 3D map, server 901 sends a request to client device 902 to send sensor information. Also, server 901 may send a request to send sensor information when it wants to update the 3D map, when it wants to confirm road conditions such as during snow accumulation or disasters, or when it wants to confirm traffic jams or accident situations.

[0420] Also, client device 902 may set the amount of sensor information to be sent to server 901 according to the communication state or frequency band at the time of receiving the request to send sensor information from server 901. Setting the amount of sensor information to be sent to server 901, for example, means increasing or decreasing the data itself or selecting an appropriate compression method.

[0421] Figure 29 It is a block diagram showing a configuration example of client device 902. Client device 902 receives a 3D map composed of point clouds etc. from server 901, and estimates its own position based on 3D data created from the sensor information of client device 902. And client device 902 sends the acquired sensor information to server 901.

[0422] Client device 902 includes: data receiving unit 1011, communication unit 1012, reception control unit 1013, format conversion unit 1014, multiple sensors 1015, 3D data creation unit 1016, 3D image processing unit 1017, 3D data storage unit 1018, format conversion unit 1019, communication unit 1020, transmission control unit 1021, and data transmission unit 1022.

[0423] Data receiving unit 1011 receives 3D map 1031 from server 901. 3D map 1031 is data including point clouds such as WLD or SWLD. 3D map 1031 may include either compressed data or uncompressed data.

[0424] Communication unit 1012 communicates with server 901 and sends a data transmission request (e.g., a request to send a 3D map) etc. to server 901.

[0425] Reception control unit 1013 exchanges information such as corresponding formats with the communication partner via communication unit 1012 and establishes communication with the communication partner.

[0426] The format conversion unit 1014 generates a 3D map 1032 by performing format conversion and the like on the 3D map 1031 received by the data reception unit 1011. Further, when the 3D map 1031 is compressed or encoded, the format conversion unit 1014 performs decompression or decoding processing. In addition, when the 3D map 1031 is uncompressed data, the format conversion unit 1014 does not perform decompression or decoding processing.

[0427] The multiple sensors 1015 are a group of sensors mounted on the client device 902 such as a LiDAR, a visible light camera, an infrared camera, or a depth sensor, which are used to obtain information about the outside of the vehicle, and generate sensor information 1033. For example, when the sensor 1015 is a laser sensor such as a LiDAR, the sensor information 1033 is 3D data such as point cloud (point group data). In addition, the number of sensors 1015 may not be multiple.

[0428] The 3D data creation unit 1016 creates 3D data 1034 around its own vehicle based on the sensor information 1033. For example, the 3D data creation unit 1016 uses the information obtained by the LiDAR and the visible light image obtained by the visible light camera to create point cloud data with color information around its own vehicle.

[0429] The 3D image processing unit 1017 performs its own position estimation processing and the like of its own vehicle by using the received 3D map 1032 such as point cloud and the 3D data 1034 around its own vehicle generated based on the sensor information 1033. Alternatively, the 3D image processing unit 1017 may synthesize the 3D map 1032 and the 3D data 1034 to create 3D data 1035 around its own vehicle, and use the created 3D data 1035 to perform its own position estimation processing.

[0430] The 3D data storage unit 1018 stores the 3D map 1032, the 3D data 1034, the 3D data 1035, and the like.

[0431] The format conversion unit 1019 generates sensor information 1037 by converting the sensor information 1033 into a format corresponding to the receiving side. In addition, the format conversion unit 1019 may reduce the data amount by compressing or encoding the sensor information 1037. Further, when format conversion is not required, the format conversion unit 1019 may omit the processing. Moreover, the format conversion unit 1019 may control the data amount of the data to be transmitted according to the specified transmission range.

[0432] The communication unit 1020 communicates with the server 901 and receives a data transmission request (a transmission request for sensor information) and the like from the server 901.

[0433] The transmission control unit 1021 exchanges information such as corresponding formats with the communication partner via the communication unit 1020 to establish communication.

[0434] The data transmission unit 1022 transmits the sensor information 1037 to the server 901. The sensor information 1037 includes, for example, information obtained by LiDAR, a luminance image (visible light image) obtained by a visible light camera, an infrared image obtained by an infrared camera, a depth image obtained by a depth sensor, sensor position information, and speed information, etc., which are obtained by multiple sensors 1015.

[0435] Next, the configuration of the server 901 will be described. Figure 30 It is a block diagram showing a configuration example of the server 901. The server 901 receives the sensor information sent from the client device 902, and creates three-dimensional data based on the received sensor information. The server 901 updates the three-dimensional map managed by the server 901 using the created three-dimensional data. And the server 901 sends the updated three-dimensional map to the client device 902 according to the transmission request of the three-dimensional map from the client device 902.

[0436] The server 901 includes: a data reception unit 1111, a communication unit 1112, a reception control unit 1113, a format conversion unit 1114, a three-dimensional data creation unit 1116, a three-dimensional data synthesis unit 1117, a three-dimensional data storage unit 1118, a format conversion unit 1119, a communication unit 1120, a transmission control unit 1121, and a data transmission unit 1122.

[0437] The data reception unit 1111 receives the sensor information 1037 from the client device 902. The sensor information 1037 includes, for example, information obtained by LiDAR, a luminance image (visible light image) obtained by a visible light camera, an infrared image obtained by an infrared camera, a depth image obtained by a depth sensor, sensor position information, and speed information, etc.

[0438] The communication unit 1112 communicates with the client device 902 and sends a data transmission request (for example, a transmission request for sensor information) etc. to the client device 902.

[0439] The reception control unit 1113 exchanges information such as corresponding formats with the communication partner via the communication unit 1112 to establish communication.

[0440] When the received sensor information 1037 is compressed or encoded, the format conversion unit 1114 generates the sensor information 1132 by performing decompression or decoding processing. Additionally, when the sensor information 1037 is uncompressed data, the format conversion unit 1114 does not perform decompression or decoding processing.

[0441] The three-dimensional data creation unit 1116 creates three-dimensional data 1134 of the surroundings of the client device 902 based on the sensor information 1132. For example, the three-dimensional data creation unit 1116 uses the information obtained by LiDAR and the visible light image obtained by the visible light camera to create point cloud data with color information of the surroundings of the client device 902.

[0442] The three-dimensional data synthesis unit 1117 synthesizes the three-dimensional data 1134 created based on the sensor information 1132 with the three-dimensional map 1135 managed by the server 901, thereby updating the three-dimensional map 1135.

[0443] The three-dimensional data storage unit 1118 stores the three-dimensional map 1135 and the like.

[0444] The format conversion unit 1119 generates the three-dimensional map 1031 by converting the three-dimensional map 1135 into a format corresponding to the receiving side. Additionally, the format conversion unit 1119 can also reduce the data volume by compressing or encoding the three-dimensional map 1135. And when format conversion is not required, the format conversion unit 1119 can also omit the processing. And the format conversion unit 1119 can control the data volume to be sent according to the specified sending range.

[0445] The communication unit 1120 communicates with the client device 902 and receives a data sending request (a sending request for the three-dimensional map) and the like from the client device 902.

[0446] The transmission control unit 1121 exchanges information such as the corresponding format with the communication partner via the communication unit 1120, thereby establishing communication.

[0447] The data sending unit 1122 sends the three-dimensional map 1031 to the client device 902. The three-dimensional map 1031 is data including point clouds such as WLD or SWLD. Either compressed data or uncompressed data may be included in the three-dimensional map 1031.

[0448] Next, the workflow of the client device 902 will be described. Figure 31 It is a flowchart showing the work when the client device 902 obtains the three-dimensional map.

[0449] First, the client device 902 requests the server 901 to send a three-dimensional map (point cloud, etc.) (S1001). At this time, the client device 902 also sends the position information of the client device 902 obtained through GPS, etc. Accordingly, the server 901 can be requested to send a three-dimensional map related to this position information.

[0450] Next, the client device 902 receives the three-dimensional map from the server 901 (S1002). If the received three-dimensional map is compressed data, the client device 902 decodes the received three-dimensional map to generate an uncompressed three-dimensional map (S1003).

[0451] Next, the client device 902 creates three-dimensional data 1034 around the client device 902 based on the sensor information 1033 obtained from multiple sensors 1015 (S1004). Next, the client device 902 estimates its own position of the client device 902 using the three-dimensional map 1032 received from the server 901 and the three-dimensional data 1034 created based on the sensor information 1033 (S1005).

[0452] Figure 32 This is a flowchart showing the operation when the client device 902 sends sensor information. First, the client device 902 receives a request to send sensor information from the server 901 (S1011). The client device 902 that has received the send request sends the sensor information 1037 to the server 901 (S1012). In addition, when the sensor information 1033 includes multiple pieces of information obtained through multiple sensors 1015, the client device 902 compresses each piece of information in a compression method suitable for each piece of information to generate the sensor information 1037.

[0453] Next, the operation process of the server 901 will be described. Figure 33 This is a flowchart showing the operation when the server 901 obtains sensor information. First, the server 901 requests the client device 902 to send sensor information (S1021). Next, the server 901 receives the sensor information 1037 sent from the client device 902 in accordance with this request (S1022). Next, the server 901 creates three-dimensional data 1134 using the received sensor information 1037 (S1023). Next, the server 901 reflects the created three-dimensional data 1134 in the three-dimensional map 1135 (S1024).

[0454] Figure 34It is a flowchart showing the operations when the server 901 sends a 3D map. First, the server 901 receives a 3D map sending request from the client device 902 (S1031). The server 901 that has received the 3D map sending request sends the 3D map 1031 to the client device 902 (S1032). At this time, the server 901 can extract the 3D map in the vicinity corresponding to the location information of the client device 902 and send the extracted 3D map. And it can be that the server 901 compresses the 3D map composed of point clouds, for example, using a compression method such as an octree, and sends the compressed 3D map.

[0455] Hereinafter, a modified example of the present embodiment will be described.

[0456] The server 901 uses the sensor information 1037 received from the client device 902 to create 3D data 1134 near the location of the client device 902. Next, the server 901 matches the created 3D data 1134 with the 3D map 1135 of the same area managed by the server 901 and calculates the difference between the 3D data 1134 and the 3D map 1135. When the difference is equal to or greater than a predetermined threshold, the server 901 determines that some abnormality has occurred around the client device 902. For example, when the ground surface subsides due to natural disasters such as earthquakes, a large difference may be considered to occur between the 3D map 1135 managed by the server 901 and the 3D data 1134 created based on the sensor information 1037.

[0457] The sensor information 1037 may also include at least one of the type of the sensor, the performance of the sensor, and the model of the sensor. Also, it may be that a category ID corresponding to the performance of the sensor is attached to the sensor information 1037. For example, when the sensor information 1037 is information obtained by LiDAR, it is possible to consider assigning an identifier according to the performance of the sensor. For example, category 1 is assigned to a sensor capable of obtaining information with an accuracy of several millimeters, category 2 is assigned to a sensor capable of obtaining information with an accuracy of several centimeters, and category 3 is assigned to a sensor capable of obtaining information with an accuracy of several meters. Also, the server 901 may estimate the performance information of the sensor from the model of the client device 902. For example, when the client device 902 is mounted on a vehicle, the server 901 may determine the specification information of the sensor based on the model of the vehicle. In this case, the server 901 may obtain the information of the vehicle model in advance or include this information in the sensor information. Also, it may be that the server 901 uses the obtained sensor information 1037 to switch the degree of correction for the three-dimensional data 1134 created using the sensor information 1037. For example, when the sensor performance is high accuracy (category 1), the server 901 does not perform correction on the three-dimensional data 1134. When the sensor performance is low accuracy (category 3), the server 901 applies correction suitable for the accuracy of the sensor to the three-dimensional data 1134. For example, the server 901 increases the degree (intensity) of correction as the accuracy of the sensor is lower.

[0458] The server 901 may also send a request to send sensor information to a plurality of client devices 902 existing in a certain space at the same time. When the server 901 receives a plurality of sensor information from the plurality of client devices 902, it is not necessary to use all the sensor information for the creation of the three-dimensional data 1134. For example, the sensor information to be used may be selected according to the performance of the sensor. For example, when updating the three-dimensional map 1135, the server 901 may select high-accuracy sensor information (category 1) from the received plurality of sensor information and use the selected sensor information to create the three-dimensional data 1134.

[0459] The server 901 is not limited to servers such as traffic cloud monitoring, and may also be other client devices (in-vehicle). Figure 35 The system configuration in this case is shown.

[0460] For example, the client device 902C sends a request to the nearby client device 902A to send sensor information, and obtains the sensor information from the client device 902A. Then, the client device 902C uses the obtained sensor information of the client device 902A to create three-dimensional data and updates the three-dimensional map of the client device 902C. In this way, the client device 902C can utilize the performance of the client device 902C to generate a three-dimensional map of the space that can be obtained from the client device 902A. For example, this may occur when the performance of the client device 902C is high.

[0461] Moreover, in this case, the client device 902A that provided the sensor information is given the right to obtain the highly accurate three-dimensional map generated by the client device 902C. The client device 902A receives the highly accurate three-dimensional map from the client device 902C according to this right.

[0462] It can also be that the client device 902C sends a request to send sensor information to a plurality of nearby client devices 902 (client device 902A and client device 902B). When the sensor of the client device 902A or the client device 902B has high performance, the client device 902C can use the sensor information obtained through this high-performance sensor to create three-dimensional data.

[0463] Figure 36 It is a block diagram showing the functional configurations of the server 901 and the client device 902. The server 901 includes, for example: a three-dimensional map compression / decoding processing unit 1201 that compresses and decodes a three-dimensional map, and a sensor information compression / decoding processing unit 1202 that compresses and decodes sensor information.

[0464] The client device 902 includes: a three-dimensional map decoding processing unit 1211 and a sensor information compression processing unit 1212. The three-dimensional map decoding processing unit 1211 receives the encoded data of the compressed three-dimensional map, decodes the encoded data, and obtains the three-dimensional map. The sensor information compression processing unit 1212 does not compress the three-dimensional data created from the obtained sensor information, but compresses the sensor information itself and sends the encoded data of the compressed sensor information to the server 901. According to this configuration, the client device 902 can keep the processing unit (device or LSI) for decoding the three-dimensional map (point cloud, etc.) inside, without having to keep the processing unit for compressing the three-dimensional data of the three-dimensional map (point cloud, etc.) inside. In this way, the cost and power consumption of the client device 902 can be suppressed.

[0465] As described above, the client device 902 according to the present embodiment is mounted on a moving body, and creates three-dimensional data 1034 of the periphery of the moving body based on sensor information 1033 obtained by a sensor 1015 mounted on the moving body and showing the peripheral situation of the moving body. The client device 902 estimates the own position of the moving body by using the created three-dimensional data 1034. The client device 902 transmits the obtained sensor information 1033 to the server 901 or other moving bodies 902.

[0466] Accordingly, the client device 902 transmits the sensor information 1033 to the server 901 or the like. In this way, compared with the case of transmitting three-dimensional data, there is a possibility of reducing the amount of data to be transmitted. And since there is no need to perform processing such as compression or encoding of three-dimensional data in the client device 902, the amount of processing of the client device 902 can be reduced. Therefore, the client device 902 can achieve a reduction in the amount of transmitted data or a simplification of the device configuration.

[0467] In addition, the client device 902 further sends a transmission request for a three-dimensional map to the server 901, and receives a three-dimensional map 1031 from the server 901. The client device 902 estimates its own position by using the three-dimensional data 1034 and the three-dimensional map 1032 in the estimation of its own position.

[0468] Moreover, the sensor information 1033 includes at least one of information obtained by a laser sensor, a luminance image (visible light image), an infrared image, a depth image, the position information of the sensor, and the speed information of the sensor.

[0469] And the sensor information 1033 includes information showing the performance of the sensor.

[0470] In addition, the client device 902 encodes or compresses the sensor information 1033, and transmits the encoded or compressed sensor information 1037 to the server 901 or other moving bodies 902 when transmitting the sensor information. Accordingly, the client device 902 can reduce the amount of data transmitted.

[0471] For example, the client device 902 includes a processor and a memory, and the processor uses the memory to perform the above processing.

[0472] In addition, the server 901 according to the present embodiment can communicate with the client device 902 mounted on a moving body, and receives sensor information 1037 obtained by a sensor 1015 mounted on the moving body and showing the peripheral situation of the moving body. The server 901 creates three-dimensional data 1134 of the periphery of the moving body based on the received sensor information 1037.

[0473] Accordingly, the server 901 uses the sensor information 1037 sent from the client device 902 to create three-dimensional data 1134. In this way, compared with the case where the client device 902 sends three-dimensional data, there is a possibility of reducing the amount of data to be sent. Also, since it is not necessary to perform processing such as compression or encoding of the three-dimensional data on the client device 902, the processing amount of the client device 902 can be reduced. In this way, the server 901 can achieve a reduction in the amount of data transmitted or a simplification of the device configuration.

[0474] Furthermore, the server 901 further sends a transmission request for the sensor information to the client device 902.

[0475] Furthermore, the server 901 further uses the created three-dimensional data 1134 to update the three-dimensional map 1135, and sends the three-dimensional map 1135 to the client device 902 in accordance with a transmission request for the three-dimensional map 1135 from the client device 902.

[0476] Moreover, the sensor information 1037 includes at least one of information obtained by a laser sensor, a luminance image (visible light image), an infrared image, a depth image, the position information of the sensor, and the speed information of the sensor.

[0477] Moreover, the sensor information 1037 includes information indicating the performance of the sensor.

[0478] Furthermore, the server 901 corrects the three-dimensional data according to the performance of the sensor. Accordingly, this three-dimensional data creation method can improve the quality of the three-dimensional data.

[0479] Moreover, in receiving the sensor information, the server 901 receives a plurality of sensor information 1037 from a plurality of client devices 902, and selects the sensor information 1037 to be used in creating the three-dimensional data 1134 according to a plurality of information indicating the performance of the sensor included in the plurality of sensor information 1037. Accordingly, the server 901 can improve the quality of the three-dimensional data 1134.

[0480] Moreover, the server 901 decodes or decompresses the received sensor information 1037, and creates the three-dimensional data 1134 according to the decoded or decompressed sensor information 1132. Accordingly, the server 901 can reduce the amount of data transmitted.

[0481] For example, the server 901 includes a processor and a memory, and the processor uses the memory to perform the above processing.

[0482] (Embodiment 7)

[0483] In this embodiment, a method for encoding and decoding three-dimensional data using inter-frame prediction processing will be described.

[0484] Figure 37 FIG. 4 is a block diagram of a three-dimensional data encoding apparatus 1300 according to this embodiment. The three-dimensional data encoding apparatus 1300 generates an encoded bitstream (hereinafter also simply referred to as a bitstream) as an encoded signal by encoding three-dimensional data. As Figure 37 shown, the three-dimensional data encoding apparatus 1300 includes: a segmentation unit 1301, a subtraction unit 1302, a transformation unit 1303, a quantization unit 1304, an inverse quantization unit 1305, an inverse transformation unit 1306, an addition unit 1307, a reference volume memory 1308, an intra-frame prediction unit 1309, a reference space memory 1310, an inter-frame prediction unit 1311, a prediction control unit 1312, and an entropy encoding unit 1313.

[0485] The segmentation unit 1301 divides each space (SPC) included in the three-dimensional data into a plurality of volumes (VLM) as encoding units. Further, the segmentation unit 1301 performs octree representation (octree conversion) on the voxels within each volume. In addition, the segmentation unit 1301 may make the space and the volume the same size and perform octree representation on the space. Further, the segmentation unit 1301 may attach information (such as depth information) required for octree conversion to the head of the bitstream.

[0486] The subtraction unit 1302 calculates the difference between the volume (encoding target volume) output from the segmentation unit 1301 and the predicted volume generated by intra-frame prediction or inter-frame prediction described later, and outputs the calculated difference as a prediction residual to the transformation unit 1303. Figure 38 FIG. 5 shows an example of calculating the prediction residual. In addition, the bit strings of the encoding target volume and the predicted volume shown here are, for example, position information indicating the positions of three-dimensional points (for example, point clouds) included in the volume.

[0487] Hereinafter, octree representation and the voxel scanning order will be described. After the volume is transformed into an octree structure (octree conversion), it is encoded. The octree structure is composed of nodes and leaf nodes. Each node has eight nodes or leaf nodes, and each leaf node has voxel (VXL) information. Figure 39 FIG. 6 shows a configuration example of a volume including a plurality of voxels. Figure 40 FIG. 7 shows Figure 39 an example of transforming the volume shown in FIG. 6 into an octree structure. Here, Figure 40 among the leaf nodes shown in FIG. 7, leaf nodes 1, 2, and 3 respectively represent Figure 39 the voxels VXL1, VXL2, and VXL3 shown in FIG. 6, and represent VXLs (hereinafter referred to as valid VXLs) including a point group.

[0488] An octree is represented by a binary sequence of 0s and 1s, for example. For example, when a node or a valid VXL is set to the value 1 and the rest are set to the value 0, the binary sequence shown is assigned to each node and leaf node. Then, the binary sequence is scanned in a breadth-first or depth-first scan order. For example, when scanned in breadth-first order, the binary sequence shown in A is obtained. When scanned in depth-first order, the binary sequence shown in B is obtained. The binary sequence obtained by this scan is encoded by entropy coding, thereby reducing the amount of information. Figure 40 Next, the depth information in the octree representation is described. The depth in the octree representation is used for controlling up to which granularity the point cloud information contained in the volume is retained. If the depth is set large, the point cloud information can be reproduced at a finer level, but the amount of data for representing the nodes and leaf nodes increases. On the contrary, if the depth is set small, although the amount of data can be reduced, point cloud information at multiple different positions and with different colors will be regarded as the same position and the same color, so the information originally possessed by the point cloud information will be lost. Figure 41 For example, Figure 41 shows an example of representing the octree with depth = 2 shown in

[0489] as an octree with depth = 1.

[0490] The octree shown in Figure 42 has less data volume than the octree shown in Figure 40 . That is, Figure 42 the octree shown in Figure 40 has fewer bits after binary serialization than the octree shown in Figure 42 . Here, Figure 42 the leaf node 1 and leaf node 2 shown in Figure 40 become represented by the leaf node 1 shown in Figure 41 . That is, the information that the leaf node 1 and leaf node 2 shown in Figure 40 are at different positions is lost.

[0491] Figure 43 shows the volume corresponding to the octree shown in Figure 42 . Figure 39 The VXL1 and VXL2 shown in Figure 43 correspond to the VXL12 shown in Figure 39 . In this case, the three-dimensional data encoding device 1300 generates Figure 43The color information of VXL12 shown. For example, the three-dimensional data encoding device 1300 calculates the average value, median value, weighted average value, etc. of the color information of VXL1 and VXL2 as the color information of VXL12. In this way, the three-dimensional data encoding device 1300 can control the reduction of the data amount by changing the depth of the octree.

[0492] The three-dimensional data encoding device 1300 can also set the depth information of the octree using any one of the world space unit, space unit, and volume unit. And at this time, the three-dimensional data encoding device 1300 can also attach the depth information to the header information of the world space, the header information of the space, or the header information of the volume. Also, the same value can be used as the depth information in all world spaces, spaces, and volumes at different times. In this case, the three-dimensional data encoding device 1300 can also attach the depth information to the header information for managing the world space for all times.

[0493] When the voxel contains color information, the transformation unit 1303 applies a frequency transformation such as an orthogonal transformation to the prediction residual of the color information of the voxel in the volume. For example, the transformation unit 1303 scans the prediction residual in a certain scan order to create a one-dimensional arrangement. After that, the transformation unit 1303 transforms the created one-dimensional arrangement into the frequency domain by applying a one-dimensional orthogonal transformation. Accordingly, when the values of the prediction residuals in the volume are close, the values of the frequency components in the low-frequency band become larger, and the values of the frequency components in the high-frequency band become smaller. Therefore, the quantization unit 1304 can more effectively reduce the encoding amount.

[0494] And the transformation unit 1303 can also use an orthogonal transformation of two dimensions or more instead of using a one-dimensional orthogonal transformation. For example, the transformation unit 1303 maps the prediction residual into a two-dimensional arrangement in a certain scan order and applies a two-dimensional orthogonal transformation to the obtained two-dimensional arrangement. And the transformation unit 1303 can also select the orthogonal transformation method to be used from multiple orthogonal transformation methods. In this case, the three-dimensional data encoding device 1300 attaches information indicating which orthogonal transformation method is used to the bitstream. And it can be that the transformation unit 1303 selects the orthogonal transformation method to be used from multiple orthogonal transformation methods with different dimensions. In this case, the three-dimensional data encoding device 1300 attaches information indicating which dimension of the orthogonal transformation method is used to the bitstream.

[0495] For example, the transform unit 1303 aligns the scan order of the prediction residual with the scan order (such as breadth-first or depth-first) in the octree within the volume. Accordingly, since there is no need to attach information indicating the scan order of the prediction residual to the bitstream, the overhead can be reduced. Also, the transform unit 1303 may apply a scan order different from the scan order of the octree. In this case, the 3D data encoding device 1300 attaches information indicating the scan order of the prediction residual to the bitstream. Accordingly, the 3D data encoding device 1300 can efficiently encode the prediction residual. Also, it may be that the 3D data encoding device 1300 attaches information (such as a flag) indicating whether the scan order of the octree is applied to the bitstream, and in the case where the scan order of the octree is not applied, attaches information indicating the scan order of the prediction residual to the bitstream.

[0496] The transform unit 1303 can transform not only the prediction residual of color information but also other attribute information that the voxels have. For example, it may be that the transform unit 1303 transforms and encodes information such as reflectance obtained when acquiring point clouds through LiDAR or the like.

[0497] When the space does not have attribute information such as color information, the transform unit 1303 can skip the processing. Also, the 3D data encoding device 1300 can attach information (a flag) indicating whether to skip the processing of the transform unit 1303 to the bitstream.

[0498] The quantization unit 1304 quantizes the frequency components of the prediction residual generated by the transform unit 1303 using quantization control parameters to generate quantization coefficients. Thereby, the amount of information is reduced. The generated quantization coefficients are output to the entropy encoding unit 1313. The quantization unit 1304 can control the quantization control parameters in world space units, space units, or volume units. At this time, the 3D data encoding device 1300 attaches the quantization control parameters to respective header information and the like. Also, the quantization unit 1304 can change the weights for quantization control according to the frequency components of each prediction residual. For example, the quantization unit 1304 can perform fine quantization on low-frequency components and rough quantization on high-frequency components. In this case, the 3D data encoding device 1300 can attach a parameter indicating the weights of the respective frequency components to the header.

[0499] When the space does not have attribute information such as color information, the quantization unit 1304 can skip the processing. Also, the 3D data encoding device 1300 can attach information (a flag) indicating whether to skip the processing of the quantization unit 1304 to the bitstream.

[0500] The inverse quantization unit 1305 uses the quantization control parameter to perform inverse quantization on the quantization coefficients generated by the quantization unit 1304. Accordingly, the inverse quantization coefficients of the prediction residual are generated, and the generated inverse quantization coefficients are output to the inverse transform unit 1306.

[0501] The inverse transform unit 1306 applies an inverse transform to the inverse quantization coefficients generated in the inverse quantization unit 1305, thereby generating the prediction residual after the inverse transform is applied. Since the prediction residual after the inverse transform is applied is the prediction residual generated after quantization, it may not be exactly the same as the prediction residual output by the transform unit 1303.

[0502] The addition unit 1307 adds the prediction volume generated by the inverse transform unit 1306 after the inverse transform is applied and the prediction volume used in the generation of the prediction residual before quantization and generated by intra-frame prediction or inter-frame prediction described later to generate a reconstructed volume. The reconstructed volume is stored in the reference volume memory 1308 or the reference space memory 1310.

[0503] The intra-frame prediction unit 1309 generates a prediction volume of the volume to be encoded using the attribute information of the adjacent volumes stored in the reference volume memory 1308. The attribute information includes voxel color information or reflectivity. The intra-frame prediction unit 1309 generates a predicted value of the color information or reflectivity of the volume to be encoded.

[0504] Figure 44 It is a diagram for explaining the operation of the intra-frame prediction unit 1309. For example, Figure 44 As shown, the intra-frame prediction unit 1309 generates a prediction volume of the volume to be encoded (volume idx = 3) based on the adjacent volume (volume idx = 0). Here, the volume idx is identifier information attached to the volumes in the space, and different values are assigned to each volume. The order of assignment of the volume idx may be the same as the encoding order or different from the encoding order. For example, as Figure 44 The predicted value of the color information of the volume to be encoded shown, the intra-frame prediction unit 1309 uses the average value of the color information of the voxels included in the adjacent volume with volume idx = 0 as the adjacent volume. In this case, by subtracting the predicted value of the color information from the color information of each voxel included in the volume to be encoded, a prediction residual is generated. The processing after the transform unit 1303 is performed on the prediction residual. And, in this case, the three-dimensional data encoding device 1300 attaches the adjacent volume information and the prediction mode information to the bitstream. Here, the adjacent volume information is information showing the adjacent volume used in the prediction, for example, showing the volume idx of the adjacent volume used in the prediction. And, the prediction mode information shows the mode used in the generation of the prediction volume. The mode is, for example, the average value mode that generates a predicted value based on the average value of the voxels in the adjacent volume, or the median value mode that generates a predicted value based on the median value of the voxels in the adjacent volume, etc.

[0505] The intra prediction unit 1309 can also generate a prediction volume based on multiple adjacent volumes. For example, in the configuration shown in Figure 44 , the intra prediction unit 1309 generates a prediction volume 0 based on the volume with volume idx = 0, and generates a prediction volume 1 based on the volume with volume idx = 1. Then, the intra prediction unit 1309 generates the average of the prediction volume 0 and the prediction volume 1 as the final prediction volume. In this case, the three-dimensional data encoding device 1300 can also attach the multiple volume idxs of the multiple volumes used in the generation of the prediction volume to the bitstream.

[0506] Figure 45 The inter prediction process according to this embodiment is shown in terms of the mode. The inter prediction unit 1311 performs encoding (inter prediction) on the space (SPC) at a certain time T_Cur using the encoded spaces at different times T_LX. In this case, the inter prediction unit 1311 applies rotation and translation processing to the encoded spaces at different times T_LX to perform the encoding process.

[0507] Moreover, the three-dimensional data encoding device 1300 attaches RT information related to the rotation and translation processing of the spaces at different times T_LX to the bitstream. The different times T_LX are, for example, the time T_L0 before the certain time T_Cur. At this time, the three-dimensional data encoding device 1300 can also attach the RT information RT_L0 related to the rotation and translation processing of the space at time T_L0 to the bitstream.

[0508] Alternatively, the different times T_LX are, for example, the time T_L1 after the certain time T_Cur. At this time, the three-dimensional data encoding device 1300 can attach the RT information RT_L1 related to the rotation and translation processing of the space at time T_L1 to the bitstream.

[0509] Alternatively, the inter prediction unit 1311 performs encoding (bi-prediction) by referring to the spaces at both different times T_L0 and time T_L1. In this case, the three-dimensional data encoding device 1300 can attach both the RT information RT_L0 and RT_L1 related to the rotation and translation applied to the spaces respectively to the bitstream.

[0510] In addition, although T_L0 is set as the time before T_Cur and T_L1 is set as the time after T_Cur above, it is not limited thereto. For example, both T_L0 and T_L1 can be the times before T_Cur. Or, both T_L0 and T_L1 can be the times after T_Cur.

[0511] Alternatively, when the three-dimensional data encoding device 1300 performs encoding with reference to spaces at multiple different times, the RT information related to the rotation and translation applied to each space is appended to the bitstream. For example, the three-dimensional data encoding device 1300 manages the multiple encoded spaces to be referred to through two reference lists (L0 list and L1 list). When the first reference space in the L0 list is set as L0R0, the second reference space in the L0 list is set as L0R1, the first reference space in the L1 list is set as L1R0, and the second reference space in the L1 list is set as L1R1, the three-dimensional data encoding device 1300 appends the RT information RT_L0R0 of L0R0, the RT information RT_L0R1 of L0R1, the RT information RT_L1R0 of L1R0, and the RT information RT_L1R1 of L1R1 to the bitstream. For example, the three-dimensional data encoding device 1300 appends this RT information to the head of the bitstream.

[0512] Alternatively, when the three-dimensional data encoding device 1300 performs encoding with reference to reference spaces at multiple different times, it determines whether rotation and translation are applied for each reference space. At this time, the three-dimensional data encoding device 1300 may append information (such as an RT application flag) indicating whether rotation and translation are applied for each reference space to the header information of the bitstream. For example, the three-dimensional data encoding device 1300 calculates the RT information and the ICP error value for each reference space to be referred to according to the encoding target space by using the ICP (Interactive Closest Point) algorithm. When the ICP error value is equal to or less than a predetermined fixed value, the three-dimensional data encoding device 1300 determines that rotation and translation are not required and sets the RT application flag to OFF (invalid). In addition, when the ICP error value is greater than the above-mentioned fixed value, the three-dimensional data encoding device 1300 sets the RT application flag to ON (valid) and appends the RT information to the bitstream.

[0513] Figure 46 A syntax example of appending the RT information and the RT application flag to the head is shown. In addition, the number of bits allocated to each syntax can be determined according to the range that the syntax can take. For example, when the number of reference spaces included in the reference list L0 is 8, 3 bits can be allocated to MaxRefSpc_l0. The number of allocated bits can be changed according to the values that each syntax can take, or the number of allocated bits can be fixed regardless of the values that can be taken. When the number of allocated bits is fixed, the three-dimensional data encoding device 1300 can append this fixed number of bits to other header information.

[0514] Here, Figure 46The shown MaxRefSpc_l0 indicates the number of reference spaces included in the reference list L0. RT_flag_l0[i] is the RT application flag for the reference space i in the reference list L0. When RT_flag_l0[i] is 1, rotation and translation are applied to the reference space i. When RT_flag_l0[i] is 0, rotation and translation are not applied to the reference space i.

[0515] R_l0[i] and T_l0[i] are the RT information of the reference space i in the reference list L0. R_l0[i] is the rotation information of the reference space i in the reference list L0. The rotation information indicates the content of the applied rotation process, such as a rotation matrix or a quaternion, etc. T_l0[i] is the translation information of the reference space i in the reference list L0. The translation information indicates the content of the applied translation process, such as a translation vector, etc.

[0516] MaxRefSpc_l1 indicates the number of reference spaces included in the reference list L1. RT_flag_l1[i] is the RT application flag for the reference space i in the reference list L1. When RT_flag_l1[i] is 1, rotation and translation are applied to the reference space i. When RT_flag_l1[i] is 0, rotation and translation are not applied to the reference space i.

[0517] R_l1[i] and T_l1[i] are the RT information of the reference space i in the reference list L1. R_l1[i] is the rotation information of the reference space i in the reference list L1. The rotation information indicates the content of the applied rotation process, such as a rotation matrix or a quaternion, etc. T_l1[i] is the translation information of the reference space i in the reference list L1. The translation information indicates the content of the applied translation process, such as a translation vector, etc.

[0518] The inter-frame prediction unit 1311 generates a predicted volume of the coding target volume by using the information of the encoded reference spaces stored in the reference space memory 1310. As described above, before generating the predicted volume of the coding target volume, the inter-frame prediction unit 1311 uses the ICP (Interactive Closest Point) algorithm in the coding target space and the reference space to obtain the RT information in order to make the positional relationship between the coding target space and the entire reference space closer. Then, the inter-frame prediction unit 1311 applies rotation and translation processing to the reference space by using the obtained RT information, thereby obtaining the reference space B. After that, the inter-frame prediction unit 1311 generates a predicted volume of the coding target volume in the coding target space by using the information in the reference space B. Here, the three-dimensional data coding device 1300 attaches the RT information used to obtain the reference space B to the header information of the coding target space, etc.

[0519] In this way, the inter-frame prediction unit 1311 applies rotation and translation processing to the reference space, so that after making the positional relationship between the coding target space and the overall reference space closer, it uses the information of the reference space to generate a prediction volume. In this way, the accuracy of the prediction volume can be improved. Moreover, since the prediction residual can be suppressed, the coding amount can be reduced. In addition, although an example of performing ICP using the coding target space and the reference space is shown here, it is not limited thereto. For example, in order to reduce the processing amount, the inter-frame prediction unit 1311 can also perform ICP using at least one of the coding target space with the voxel or point cloud number extracted and the reference space with the voxel or point cloud number extracted, so as to obtain the RT information.

[0520] Moreover, when the ICP error value obtained from the result of ICP is smaller than a preset first threshold value, that is, for example, when the positional relationship between the coding target space and the reference space is close, the inter-frame prediction unit 1311 can determine that rotation and translation processing are not required and does not perform rotation and translation. In this case, the three-dimensional data coding device 1300 can refrain from attaching the RT information to the bitstream, thereby being able to suppress the overhead.

[0521] Moreover, when the ICP error value is larger than a preset second threshold value, the inter-frame prediction unit 1311 determines that the shape change in the space is large, and can apply intra-frame prediction to all volumes of the coding target space. Hereinafter, the space to which intra-frame prediction is applied is referred to as the intra-frame space. Moreover, the second threshold value is a value larger than the above-mentioned first threshold value. And it is not limited to ICP. As long as it is a method for obtaining the RT information from two voxel sets or two point cloud sets, any method can be applied.

[0522] Moreover, when attribute information such as shape or color is included in the three-dimensional data, as the prediction volume of the coding target volume in the coding target space, the inter-frame prediction unit 1311, for example, searches for the volume in the reference space that is closest to the shape or color attribute information of the coding target volume. And this reference space is, for example, the reference space after the above-mentioned rotation and translation processing. The inter-frame prediction unit 1311 generates a prediction volume based on the volume (reference volume) obtained through the search. Figure 47 It is a diagram for explaining the generation operation of the prediction volume. The inter-frame prediction unit 1311 is for Figure 47In the case where the volume of the coding target (volume idx = 0) shown is coded using inter-frame prediction, while sequentially scanning the reference volumes in the reference space, the volume with the smallest difference, i.e., the prediction residual, between the coding target volume and the reference volumes is searched for. The inter-frame prediction unit 1311 selects the volume with the smallest prediction residual as the prediction volume. The prediction residual between the coding target volume and the prediction volume is coded by the processing after the transform unit 1303. Here, the prediction residual refers to the difference between the attribute information of the coding target volume and the attribute information of the prediction volume. Further, the three-dimensional data coding device 1300 attaches the volume idx of the reference volume in the reference space that is used as the prediction volume to the head of the bit stream.

[0523] In Figure 47 In the example shown, the reference volume with volume idx = 4 in the reference space L0R0 is selected as the prediction volume of the coding target volume. Then, the prediction residual between the coding target volume and the reference volume and the reference volume idx = 4 are coded and attached to the bit stream.

[0524] In addition, although an example of generating a prediction volume of attribute information has been described here, the same processing can also be performed for the prediction volume of position information.

[0525] The prediction control unit 1312 controls which of intra-frame prediction and inter-frame prediction is used to code the coding target volume. Here, the mode including intra-frame prediction and inter-frame prediction is referred to as the prediction mode. For example, the prediction control unit 1312 calculates the prediction residual in the case where the coding target volume is predicted by intra-frame prediction and the prediction residual in the case where it is predicted by inter-frame prediction as evaluation values, and selects the prediction mode with the smaller evaluation value. Alternatively, the prediction control unit 1312 may perform orthogonal transformation, quantization, and entropy coding on the prediction residual of intra-frame prediction and the prediction residual of inter-frame prediction, respectively, to calculate the actual coding amount, and use the calculated coding amount as the evaluation value to select the prediction mode. Further, overhead information (such as reference volume idx information) other than the prediction residual may be added to the evaluation value. Also, the prediction control unit 1312 may normally select intra-frame prediction even when the coding target space is predetermined to be coded in the intra-frame space.

[0526] The entropy coding unit 1313 generates a coded signal (coded bit stream) by performing variable-length coding on the input from the quantization unit 1304, i.e., the quantization coefficients. Specifically, the entropy coding unit 1313, for example, binarizes the quantization coefficients and performs arithmetic coding on the obtained binary signal.

[0527] Next, a three-dimensional data decoding device that decodes the coded signal generated by the three-dimensional data coding device 1300 will be described. Figure 48FIG. 0 is a block diagram of a three-dimensional data decoding apparatus 1400 according to the present embodiment. The three-dimensional data decoding apparatus 1400 includes an entropy decoding unit 1401, an inverse quantization unit 1402, an inverse transform unit 1403, an addition unit 1404, a reference volume memory 1405, an intra prediction unit 1406, a reference space memory 1407, an inter prediction unit 1408, and a prediction control unit 1409.

[0528] The entropy decoding unit 1401 performs variable length decoding on the encoded signal (encoded bit stream). For example, the entropy decoding unit 1401 performs arithmetic decoding on the encoded signal to generate a binary signal, and generates quantization coefficients based on the generated binary signal.

[0529] The inverse quantization unit 1402 performs inverse quantization on the quantization coefficients input from the entropy decoding unit 1401 using quantization parameters attached to the bit stream or the like, thereby generating inverse quantization coefficients.

[0530] The inverse transform unit 1403 performs an inverse transform on the inverse quantization coefficients input from the inverse quantization unit 1402, thereby generating a prediction residual. For example, the inverse transform unit 1403 performs an inverse orthogonal transform on the inverse quantization coefficients based on the information attached to the bit stream, thereby generating a prediction residual.

[0531] The addition unit 1404 adds the prediction residual generated by the inverse transform unit 1403 and the prediction volume generated by intra prediction or inter prediction to generate a reconstructed volume. The reconstructed volume is output as decoded three-dimensional data, and is stored in the reference volume memory 1405 or the reference space memory 1407.

[0532] The intra prediction unit 1406 generates a prediction volume by intra prediction using the reference volume in the reference volume memory 1405 and the information attached to the bit stream. Specifically, the intra prediction unit 1406 obtains prediction mode information and adjacent volume information (e.g., volume idx) attached to the bit stream, and generates a prediction volume in the mode indicated by the prediction mode information using the adjacent volume indicated by the adjacent volume information. In addition, the details of these processes are the same as those of the intra prediction unit 1309 described above, except that the information attached to the bit stream is used.

[0533] The inter-frame prediction unit 1408 generates a prediction volume through inter-frame prediction by using the reference space in the reference space memory 1407 and the information appended to the bitstream. Specifically, the inter-frame prediction unit 1408 applies rotation and translation processing to the reference space by using the RT information of each reference space appended to the bitstream, and generates a prediction volume by using the reference space after the application. In addition, when the RT application flag for each reference space exists in the bitstream, the inter-frame prediction unit 1408 applies rotation and translation processing to the reference space according to the RT application flag. In addition, regarding the details of the above processing, except for using the information appended to the bitstream, it is the same as the processing of the above inter-frame prediction unit 1311.

[0534] Regarding whether to decode the volume to be decoded by intra-frame prediction or inter-frame prediction, it will be controlled by the prediction control unit 1409. For example, the prediction control unit 1409 selects intra-frame prediction or inter-frame prediction according to the information appended to the bitstream and indicating the prediction mode to be used. In addition, when it is predetermined that the decoding target space is decoded as an intra-frame space, the prediction control unit 1409 may usually select intra-frame prediction.

[0535] A modification example of the present embodiment will be described below. In the present embodiment, although rotation and translation are applied in units of space as an example, rotation and translation may also be applied in smaller units. For example, the three-dimensional data encoding device 1300 may divide the space into sub-spaces and apply rotation and translation in units of sub-spaces. In this case, the three-dimensional data encoding device 1300 generates RT information for each sub-space and appends the generated RT information to the head of the bitstream. And, the three-dimensional data encoding device 1300 may apply rotation and translation in units of volume as the encoding unit. In this case, the three-dimensional data encoding device 1300 generates RT information in units of encoding volume and appends the generated RT information to the head of the bitstream. Moreover, the above can be combined. That is, the three-dimensional data encoding device 1300 may apply rotation and translation in a large unit first, and then apply rotation and translation in a smaller unit. For example, it may be that the three-dimensional data encoding device 1300 applies rotation and translation in units of space, and applies different rotations and translations to each of the multiple volumes included in the obtained space.

[0536] Also, although in this embodiment, rotation and translation are applied to the reference space as an example, it is not limited thereto. For example, it may be that the three-dimensional data encoding device 1300 applies a scaling process to change the size of the three-dimensional data. Also, the three-dimensional data encoding device 1300 may apply any one or two of rotation, translation, and scaling. Also, as described above, when the processing is applied in different units in multiple stages, the types of processing applied in each unit may be different. For example, it may be that rotation and translation are applied in the space unit, and translation is applied in the volume unit.

[0537] In addition, regarding these modification examples, the three-dimensional data decoding device 1400 can be similarly applied.

[0538] As described above, the three-dimensional data encoding device 1300 according to this embodiment performs the following processing. Figure 48 It is a flowchart of the inter-frame prediction process performed by the three-dimensional data encoding device 1300.

[0539] First, the three-dimensional data encoding device 1300 uses the position information of the three-dimensional points included in the object three-dimensional data (for example, the encoding object space) and the reference three-dimensional data (for example, the reference space) at different times to generate prediction position information (for example, the prediction volume) (S1301). Specifically, the three-dimensional data encoding device 1300 generates prediction position information by applying rotation and translation processing to the position information of the three-dimensional points included in the reference three-dimensional data.

[0540] In addition, the three-dimensional data encoding device 1300 performs rotation and translation processing in the first unit (for example, space), and generates prediction position information in the second unit (for example, volume) that is finer than the first unit. For example, it may be that the three-dimensional data encoding device 1300 searches for the volume in the multiple volumes included in the reference space after rotation and translation processing, where the difference between the encoded object volume included in the encoding object space and the position information is the smallest, and uses the obtained volume as the prediction volume. In addition, the three-dimensional data encoding device 1300 may also perform rotation and translation processing and generation of prediction position information in the same unit.

[0541] And it may be that the three-dimensional data encoding device 1300 applies the first rotation and translation processing to the position information of the three-dimensional points included in the reference three-dimensional data in the first unit (for example, space), and applies the second rotation and translation processing to the position information of the three-dimensional points obtained through the first rotation and translation processing in the second unit (for example, volume) that is finer than the first unit, thereby generating prediction position information.

[0542] Here, the position information of the three-dimensional points and the prediction position information are as Figure 41As shown, it is represented in an octree structure. For example, the position information of three-dimensional points and the predicted position information are represented in a width-first scanning order in terms of depth and width in the octree structure. Alternatively, the position information of three-dimensional points and the predicted position information are represented in a depth-first scanning order in terms of depth and width in the octree structure.

[0543] And, as Figure 46 shown, the three-dimensional data encoding device 1300 encodes an RT application flag indicating whether rotation and translation processing are applied to the position information of three-dimensional points included in the reference three-dimensional data. That is, the three-dimensional data encoding device 1300 generates an encoded signal (encoded bitstream) including the RT application flag. And, the three-dimensional data encoding device 1300 encodes RT information indicating the content of rotation and translation processing. That is, the three-dimensional data encoding device 1300 generates an encoded signal (encoded bitstream) including the RT information. Additionally, it can be that the three-dimensional data encoding device 1300 encodes the RT information when the RT application flag indicates that rotation and translation processing are applied, and does not encode the RT information when the RT application flag indicates that rotation and translation processing are not applied.

[0544] And, the three-dimensional data includes, for example, the position information of three-dimensional points and the attribute information (such as color information) of each three-dimensional point. The three-dimensional data encoding device 1300 generates predicted attribute information (S1302) using the attribute information of the three-dimensional points included in the reference three-dimensional data.

[0545] Next, the three-dimensional data encoding device 1300 encodes the position information of the three-dimensional points included in the target three-dimensional data using the predicted position information. For example, as Figure 38 shown, the three-dimensional data encoding device 1300 calculates the difference, i.e., the differential position information, between the position information of the three-dimensional points included in the target three-dimensional data and the predicted position information (S1303).

[0546] And, the three-dimensional data encoding device 1300 encodes the attribute information of the three-dimensional points included in the target three-dimensional data using the predicted attribute information. For example, the three-dimensional data encoding device 1300 calculates the difference, i.e., the differential attribute information, between the attribute information of the three-dimensional points included in the target three-dimensional data and the predicted attribute information (S1304). Next, the three-dimensional data encoding device 1300 performs transformation and quantization on the calculated differential attribute information (S1305).

[0547] Finally, the three-dimensional data encoding device 1300 encodes the differential position information and the quantized differential attribute information (e.g., entropy encoding) (S1306). That is, the three-dimensional data encoding device 1300 generates an encoded signal (encoded bitstream) including the differential position information and the differential attribute information.

[0548] In addition, when the attribute information is not included in the three-dimensional data, the three-dimensional data encoding apparatus 1300 may also not perform steps S1302, S1304, and S1305. Moreover, the three-dimensional data encoding apparatus 1300 may also perform only one of the encoding of the position information of the three-dimensional points and the encoding of the attribute information of the three-dimensional points.

[0549] Moreover, Figure 49 The order of the processes shown is merely an example and is not limited thereto. For example, since the processes for the position information (S1301, S1303) and the processes for the attribute information (S1302, S1304, S1305) are independent of each other, they may be executed in any order, or a part of them may be processed in parallel.

[0550] As described above, in the present embodiment, the three-dimensional data encoding apparatus 1300 generates predicted position information by using the position information of the three-dimensional points included in the target three-dimensional data and the reference three-dimensional data at different times, and encodes the differential position information, which is the difference between the position information of the three-dimensional points included in the target three-dimensional data and the predicted position information. Accordingly, since the data amount of the encoded signal can be reduced, the encoding efficiency can be improved.

[0551] Moreover, in the present embodiment, the three-dimensional data encoding apparatus 1300 generates predicted attribute information by using the attribute information of the three-dimensional points included in the reference three-dimensional data, and encodes the differential attribute information, which is the difference between the attribute information of the three-dimensional points included in the target three-dimensional data and the predicted attribute information. Accordingly, since the data amount of the encoded signal can be reduced, the encoding efficiency can be improved.

[0552] For example, the three-dimensional data encoding apparatus 1300 includes a processor and a memory, and the processor uses the memory to perform the above-described processes.

[0553] Figure 48 is a flowchart of the inter-frame prediction process performed by the three-dimensional data decoding apparatus 1400.

[0554] First, the three-dimensional data decoding apparatus 1400 decodes (e.g., entropy decodes) the differential position information and the differential attribute information according to the encoded signal (encoded bitstream) (S1401).

[0555] Further, the three-dimensional data decoding device 1400 decodes an RT application flag indicating whether rotation and translation processing are applicable to the position information of the three-dimensional points included in the reference three-dimensional data, based on the encoded signal. Further, the three-dimensional data decoding device 1400 decodes RT information indicating the details of the rotation and translation processing. In addition, when the RT application flag indicates that rotation and translation processing are applicable, the three-dimensional data decoding device 1400 decodes the RT information, and when the RT application flag indicates that rotation and translation processing are not applicable, the three-dimensional data decoding device 1400 does not decode the RT information.

[0556] Next, the three-dimensional data decoding device 1400 performs inverse quantization and inverse transformation on the decoded differential attribute information (S1402).

[0557] Next, the three-dimensional data decoding device 1400 generates predicted position information (e.g., predicted volume) by using the position information of the three-dimensional points included in the object three-dimensional data (e.g., decoding target space) and the reference three-dimensional data (e.g., reference space) at different times (S1403). Specifically, the three-dimensional data decoding device 1400 generates predicted position information by applying rotation and translation processing to the position information of the three-dimensional points included in the reference three-dimensional data.

[0558] More specifically, when the RT application flag indicates that rotation and translation processing are applicable, the three-dimensional data decoding device 1400 applies rotation and translation processing to the position information of the three-dimensional points included in the reference three-dimensional data indicated by the RT information. When the RT application flag indicates that rotation and translation processing are not applicable, the three-dimensional data decoding device 1400 does not apply rotation and translation processing to the position information of the three-dimensional points included in the reference three-dimensional data.

[0559] In addition, the three-dimensional data decoding device 1400 can perform rotation and translation processing in a first unit (e.g., space), and can generate predicted position information in a second unit (e.g., volume) finer than the first unit. In addition, the three-dimensional data decoding device 1400 can also perform rotation and translation processing and generation of predicted position information in the same unit.

[0560] It may be that the three-dimensional data decoding device 1400 applies first rotation and translation processing to the position information of the three-dimensional points included in the reference three-dimensional data in a first unit (e.g., space), and applies second rotation and translation processing to the position information of the three-dimensional points obtained by the first rotation and translation processing in a second unit (e.g., volume) finer than the first unit, thereby generating predicted position information.

[0561] Here, the position information of the three-dimensional points and the predicted position information are, for example Figure 41As shown, it is represented in an octree structure. For example, the position information and predicted position information of three-dimensional points are represented in a scan order that gives priority to the width over the depth in the octree structure. Alternatively, the position information and predicted position information of three-dimensional points are represented in a scan order that gives priority to the depth over the width in the octree structure.

[0562] The three-dimensional data decoding device 1400 generates predicted attribute information by using the attribute information of the three-dimensional points included in the reference three-dimensional data (S1404).

[0563] Next, the three-dimensional data decoding device 1400 decodes the encoded position information included in the encoded signal by using the predicted position information, thereby restoring the position information of the three-dimensional points included in the target three-dimensional data. Here, the encoded position information is, for example, differential position information, and the three-dimensional data decoding device 1400 restores the position information of the three-dimensional points included in the target three-dimensional data by adding the differential position information and the predicted position information (S1405).

[0564] Furthermore, the three-dimensional data decoding device 1400 decodes the encoded attribute information included in the encoded signal by using the predicted attribute information, thereby restoring the attribute information of the three-dimensional points included in the target three-dimensional data. Here, the encoded attribute information is, for example, differential attribute information, and the three-dimensional data decoding device 1400 restores the attribute information of the three-dimensional points included in the target three-dimensional data by adding the differential attribute information and the predicted attribute information (S1406).

[0565] Alternatively, in the case where the attribute information is not included in the three-dimensional data, the three-dimensional data decoding device 1400 may not execute steps S1402, S1404, and S1406. Also, the three-dimensional data decoding device 1400 may perform only one of the decoding of the position information of the three-dimensional points and the decoding of the attribute information of the three-dimensional points.

[0566] And Figure 50 The order of the processing shown is an example and is not limited thereto. For example, since the processing for the position information (S1403, S1405) and the processing for the attribute information (S1402, S1404, S1406) are independent of each other, they can be performed in any order, and a part of them can also be processed in parallel.

[0567] (Embodiment 8)

[0568] A control method for the reference during encoding of occupancy encoding in the present embodiment will be described. In addition, hereinafter, the operation of the three-dimensional data encoding device will be mainly described, but the same processing can also be performed in the three-dimensional data decoding device.

[0569] Figure 51 and Figure 52 is a diagram showing the reference relationship involved in this embodiment, Figure 51 is a diagram showing the reference relationship on an octree structure, Figure 52 is a diagram showing the reference relationship on a spatial region.

[0570] In this embodiment, when the three-dimensional data encoding device encodes the encoding information of a node of an encoding object (hereinafter referred to as an object node), it refers to the encoding information of each node in the parent node to which the object node belongs. However, it does not refer to the encoding information of each node in other nodes (hereinafter referred to as parent adjacent nodes) at the same level as the parent node. That is, the three-dimensional data encoding device sets the parent adjacent nodes as non-referable or prohibits reference.

[0571] In addition, the three-dimensional data encoding device may also permit reference to the encoding information in the parent node to which the parent node belongs (hereinafter referred to as the grandparent node). That is, the three-dimensional data encoding device may also refer to the encoding information of the parent node and the grandparent node to which the object node belongs to encode the encoding information of the object node.

[0572] Here, the encoding information is, for example, occupancy encoding. When the three-dimensional data encoding device encodes the occupancy encoding of an object node, it refers to the information indicating whether each node in the parent node to which the object node belongs contains a point cloud (hereinafter referred to as occupancy information). In other words, when the three-dimensional data encoding device encodes the occupancy encoding of an object node, it refers to the occupancy encoding of the parent node. On the other hand, the three-dimensional data encoding device does not refer to the occupancy information of each node in the parent adjacent nodes. That is, the three-dimensional data encoding device does not refer to the occupancy encoding of the parent adjacent nodes. In addition, the three-dimensional data encoding device may also refer to the occupancy information of each node in the grandparent node. That is, the three-dimensional data encoding device may also refer to the occupancy information of the parent node and the parent adjacent nodes.

[0573] For example, when the three-dimensional data encoding device encodes the occupancy encoding of an object node, it uses the occupancy encoding of the parent node or the grandparent node to which the object node belongs to switch the encoding table used when performing entropy encoding on the occupancy encoding of the object node. In addition, the details will be described later. At this time, the three-dimensional data encoding device may also not refer to the occupancy encoding of the parent adjacent nodes. Thus, the three-dimensional data encoding device can appropriately switch the encoding table according to the information of the occupancy encoding of the parent node or the grandparent node when encoding the occupancy encoding of the object node, so that the encoding efficiency can be improved. In addition, since the three-dimensional data encoding device does not refer to the parent adjacent nodes, it can suppress the confirmation process of the information of the parent adjacent nodes and the memory capacity for storing this process. In addition, it becomes easy to scan and encode the occupancy encoding of each node of the octree in depth-first order.

[0574] Hereinafter, an example of switching the code table using the occupancy rate coding of the parent node will be described. Figure 53 It is a diagram showing an example of an object node and adjacent reference nodes. Figure 54 It is a diagram showing the relationship between the parent node and the nodes. Figure 55 It is a diagram showing an example of the occupancy rate coding of the parent node. Here, the adjacent reference node refers to the node that is referred to during the coding of the object node among the nodes that are spatially adjacent to the object node. In Figure 53 In the example shown, the adjacent nodes are the nodes belonging to the same layer as the object node. In addition, as the reference adjacent nodes, node X adjacent to the object block in the x direction, node Y adjacent in the y direction, and node Z adjacent in the z direction are used. That is, one adjacent block is set as the reference adjacent block in each of the x, y, and z directions.

[0575] In addition, Figure 54 The node numbers shown in Figure 55 are an example, and the relationship between the node numbers and the positions of the nodes is not limited to this. In addition, in Figure 55 node 0 is assigned to the lower bits and node 7 is assigned to the upper bits, but the assignment can also be made in the reverse order. In addition, each node can be assigned to any bit.

[0576] The three-dimensional data encoding device determines the code table for entropy encoding the occupancy rate coding of the object node by the following formula, for example.

[0577] CodingTable = (FlagX << 2) + (FlagY << 1) + (FlagZ)

[0578] Here, CodingTable represents the code table for the occupancy rate coding of the object node, and represents any one of the values 0 to 7. FlagX is the occupancy information of the adjacent node X, and represents 1 if the adjacent node X contains (occupies) a point group, and represents 0 if not. FlagY is the occupancy information of the adjacent node Y, and represents 1 if the adjacent node Y contains (occupies) a point group, and represents 0 if not. FlagZ is the occupancy information of the adjacent node Z, and represents 1 if the adjacent node Z contains (occupies) a point group, and represents 0 if not.

[0579] In addition, since the information indicating whether the adjacent node is occupied is included in the occupancy rate coding of the parent node, the three-dimensional data encoding device can also use the value shown in the occupancy rate coding of the parent node to select the code table.

[0580] From the above, it can be seen that the three-dimensional data encoding device switches the code table using the information indicating whether the adjacent nodes of the object node contain a point group, thereby being able to improve the encoding efficiency.

[0581] In addition, as Figure 53 shown, the three-dimensional data encoding device can also switch adjacent reference nodes according to the spatial positions of the object nodes within the parent node. That is, the three-dimensional data encoding device can also switch the adjacent nodes for reference among the multiple adjacent nodes according to the spatial position within the parent node of the object node.

[0582] Next, a structural example of the three-dimensional data encoding device and the three-dimensional data decoding device will be described. Figure 56 is a block diagram of the three-dimensional data encoding device 2100 according to the present embodiment. Figure 56 The three-dimensional data encoding device 2100 shown includes an octree generation unit 2101, a geometric information calculation unit 2102, a coding table selection unit 2103, and an entropy encoding unit 2104.

[0583] The octree generation unit 2101 generates, for example, an octree based on the input three-dimensional points (point cloud), and generates occupancy rate encodings for each node included in the octree. The geometric information calculation unit 2102 obtains occupancy information indicating whether an adjacent reference node of an object node is occupied. For example, the geometric information calculation unit 2102 obtains the occupancy information of the adjacent reference node from the occupancy rate encoding of the parent node to which the object node belongs. In addition, as Figure 53 shown, the geometric information calculation unit 2102 can also switch adjacent reference nodes according to the position within the parent node of the object node. In addition, the geometric information calculation unit 2102 does not refer to the occupancy information of each node within the parent adjacent node.

[0584] The coding table selection unit 2103 uses the occupancy information of the adjacent reference node calculated by the geometric information calculation unit 2102 to select a coding table used in the entropy encoding of the occupancy rate encoding of the object node. The entropy encoding unit 2104 uses the selected coding table to generate a bitstream by performing entropy encoding on the occupancy rate encoding. In addition, the entropy encoding unit 2104 can also attach information indicating the selected coding table to the bitstream.

[0585] Figure 57 is a block diagram of the three-dimensional data decoding device 2110 according to the present embodiment. Figure 57 The three-dimensional data decoding device 2110 shown includes an octree generation unit 2111, a geometric information calculation unit 2112, a coding table selection unit 2113, and an entropy decoding unit 2114.

[0586] The octree generation unit 2111 generates an octree for a certain space (node) using the header information of the bitstream and the like. The octree generation unit 2111 generates a large space (root node) using, for example, the sizes in the x-axis, y-axis, and z-axis directions of a certain space attached to the header information, and generates an octree by dividing this space into two in the x-axis, y-axis, and z-axis directions respectively to generate eight small spaces A (nodes A0 to A7). In addition, nodes A0 to A7 are sequentially set as the target nodes.

[0587] The geometric information calculation unit 2112 obtains occupancy information indicating whether adjacent reference nodes of the target node are occupied. For example, the geometric information calculation unit 2112 obtains the occupancy information of adjacent reference nodes from the occupancy rate encoding of the parent node to which the target node belongs. In addition, as Figure 53 shown, the geometric information calculation unit 2112 can also switch adjacent reference nodes according to the position within the parent node of the target node. In addition, the geometric information calculation unit 2112 does not refer to the occupancy information of each node within the parent adjacent node.

[0588] The encoding table selection unit 2113 selects an encoding table (decoding table) used in the entropy decoding of the occupancy rate encoding of the target node using the occupancy information of the adjacent reference nodes calculated by the geometric information calculation unit 2112. The entropy decoding unit 2114 generates three-dimensional points by performing entropy decoding on the occupancy rate encoding using the selected encoding table. In addition, the encoding table selection unit 2113 decodes and obtains the information of the selected encoding table attached to the bitstream, and the entropy decoding unit 2114 can also use the encoding table indicated by the obtained information.

[0589] Each bit of the occupancy rate encoding (8 bits) included in the bitstream indicates whether a point cloud is included in each of the eight small spaces A (nodes A0 to node A7) respectively. In addition, further, the three-dimensional data decoding device divides the small space node A0 into eight small spaces B (nodes B0 to B7) and generates an octree, decodes the occupancy rate encoding, and obtains information indicating whether a point cloud is included in each node of the small space B. In this way, the three-dimensional data decoding device decodes the occupancy rate encoding of each node while generating an octree from a large space to a small space.

[0590] Hereinafter, the processing flows of the three-dimensional data encoding device and the three-dimensional data decoding device will be described. Figure 58 is a flowchart of the three-dimensional data encoding process in the three-dimensional data encoding device. First, the three-dimensional data encoding device determines (defines) a space (target node) that includes part or all of the input three-dimensional point cloud (S2101). Next, the three-dimensional data encoding device divides the target node into eight to generate eight small spaces (nodes) (S2102). Next, the three-dimensional data encoding device generates the occupancy rate encoding of the target node according to whether a point cloud is included in each node (S2103).

[0591] Next, the three-dimensional data encoding device calculates (obtains) the occupancy information of the adjacent reference nodes of the object node from the occupancy rate encoding of the parent node of the object node (S2104). Next, the three-dimensional data encoding device selects the coding table used in the entropy encoding based on the determined occupancy information of the adjacent reference nodes of the object node (S2105). Next, the three-dimensional data encoding device performs entropy encoding on the occupancy rate encoding of the object node using the selected coding table (S2106).

[0592] In addition, the three-dimensional data encoding device repeatedly performs the process of dividing each node into 8 parts and encoding the occupancy rate encoding of each node until the node cannot be divided (S2107). That is, the processes of steps S2102 to S2106 are recursively repeated.

[0593] Figure 59 It is a flowchart of the three-dimensional data decoding method in the three-dimensional data decoding device. First, the three-dimensional data decoding device determines (defines) the space (object node) to be decoded using the header information of the bitstream (S2111). Next, the three-dimensional data decoding device divides the object node into 8 parts to generate 8 small spaces (nodes) (S2112). Next, the three-dimensional data decoding device calculates (obtains) the occupancy information of the adjacent reference nodes of the object node from the occupancy rate encoding of the parent node of the object node (S2113).

[0594] Next, the three-dimensional data decoding device selects the coding table used in the entropy decoding based on the occupancy information of the adjacent reference nodes (S2114). Next, the three-dimensional data decoding device performs entropy decoding on the occupancy rate encoding of the object node using the selected coding table (S2115).

[0595] In addition, the three-dimensional data decoding device repeatedly performs the process of dividing each node into 8 parts and decoding the occupancy rate encoding of each node until the node cannot be divided (S2116). That is, the processes of steps S2112 to S2115 are recursively repeated.

[0596] Next, an example of the switching of the coding table will be described. Figure 60 It is a diagram showing an example of the switching of the coding table. For example, like the coding table 0 shown in Figure 60 , the same context model can also be applied to multiple occupancy rate encodings. In addition, different context models can also be assigned to each occupancy rate encoding. Thus, the context model can be assigned according to the occurrence probability of the occupancy rate encoding, so the encoding efficiency can be improved. In addition, a context model that updates the probability table according to the occurrence frequency of the occupancy rate encoding can also be used. In addition, a context model that fixes the probability table can also be used.

[0597] Hereinafter, Modification Example 1 of the present embodiment will be described. Figure 61This is a diagram showing the reference relationship in this modified example. In the above-described embodiment, the three-dimensional data encoding device does not encode with reference to the occupancy rate of the parent adjacent node, but it is also possible to switch whether to encode with reference to the occupancy rate of the parent adjacent node according to specific conditions.

[0598] For example, when the three-dimensional data encoding device encodes while performing a breadth-first scan of the octree, it encodes the occupancy rate encoding of the target node with reference to the occupancy information of the nodes within the parent adjacent node. On the other hand, when the three-dimensional data encoding device encodes while performing a depth-first scan of the octree, it prohibits referring to the occupancy information of the nodes within the parent adjacent node. In this way, according to the scan order (encoding order) of the nodes of the octree, the nodes that can be referred to are appropriately switched, thereby enabling an improvement in encoding efficiency and a suppression of the processing load.

[0599] In addition, the three-dimensional data encoding device can attach information such as whether to encode the octree in breadth-first or depth-first to the header of the bitstream. Figure 62 This is a diagram showing a syntactic example of the header information in this case. Figure 62 The shown octree_scan_order is encoding order information (encoding order flag) indicating the encoding order of the octree. For example, when octree_scan_order is 0, it indicates breadth-first, and when it is 1, it indicates depth-first. Thus, the three-dimensional data decoding device can know whether the bitstream is encoded in breadth-first or depth-first by referring to octree_scan_order, and can appropriately decode the bitstream.

[0600] In addition, the three-dimensional data encoding device can also attach information indicating whether to prohibit referring to the parent adjacent node to the header information of the bitstream. Figure 63 This is a diagram showing a syntactic example of the header information in this case. limit_refer_flag is prohibition switching information (prohibition switching flag) indicating whether to prohibit referring to the parent adjacent node. For example, when limit_refer_flag is 1, it indicates prohibiting referring to the parent adjacent node, and when it is 0, it indicates no reference restriction (permission to refer to the parent adjacent node).

[0601] That is, the three-dimensional data encoding device determines whether to prohibit referring to the parent adjacent node, and based on the result of the above determination, switches whether to prohibit or permit referring to the parent adjacent node. In addition, the three-dimensional data encoding device generates a bitstream including the prohibition switching information, which is the result of the above determination and indicates whether to prohibit referring to the parent adjacent node.

[0602] In addition, the three-dimensional data decoding device obtains the prohibition switching information indicating whether to prohibit referring to the parent adjacent node from the bitstream, and based on the prohibition switching information, switches whether to prohibit or permit referring to the parent adjacent node.

[0603] Thus, the three-dimensional data encoding device can control the reference to the parent adjacent node and generate a bitstream. In addition, the three-dimensional data decoding device can obtain information indicating whether the reference to the parent adjacent node is prohibited from the header of the bitstream.

[0604] In addition, in the present embodiment, as an example of the encoding process for prohibiting the reference to the parent adjacent node, the occupancy encoding process is described as an example, but it is not necessarily limited thereto. For example, the same method can also be applied when encoding other information of the nodes of the octree. For example, when encoding other attribute information such as the color, normal vector, or reflectance attached to the node, the method of the present embodiment can also be applied. In addition, the same method can also be applied when encoding the encoding table or the predicted value.

[0605] Next, a second modification example of the present embodiment will be described. In the above description, as Figure 53 shown, an example of using three reference adjacent nodes is shown, but four or more reference adjacent nodes can also be used. Figure 64 FIG. is a diagram showing an example of an object node and reference adjacent nodes.

[0606] For example, the three-dimensional data encoding device calculates, by the following formula, the encoding table for entropy encoding the occupancy encoding of the object node shown in Figure 64 FIG.

[0607] CodingTable = (FlagX0 << 3) + (FlagX1 << 2) + (FlagY << 1) + (FlagZ)

[0608] Here, CodingTable represents the encoding table for the occupancy encoding of the object node and represents any one of values 0 to 15. FlagXN is the occupancy information of the adjacent node XN (N = 0...1), and represents 1 if the adjacent node XN contains (occupies) the point group, and represents 0 if not. FlagY is the occupancy information of the adjacent node Y, and represents 1 if the adjacent node Y contains (occupies) the point group, and represents 0 if not. FlagZ is the occupancy information of the adjacent node Z, and represents 1 if the adjacent node Z contains (occupies) the point group, and represents 0 if not.

[0609] At this time, if the adjacent node is, for example, Figure 64 the adjacent node X0 as shown, which cannot be referenced (reference prohibited), the three-dimensional data encoding device can also use a fixed value such as 1 (occupied) or 0 (not occupied) as a substitute value.

[0610] Figure 65 FIG. is a diagram showing an example of an object node and adjacent nodes. As Figure 65As shown, in the case where adjacent nodes cannot be referred to (reference prohibited), the occupancy rate encoding of the grandparent node of the object node can also be referred to, and the occupancy information of the adjacent nodes can be calculated. For example, instead of Figure 65 the adjacent node X0 shown, the three-dimensional data encoding device can also use the occupancy information of the adjacent node G0 to calculate FlagX0 in the above formula, and use the calculated FlagX0 to determine the value of the encoding table. In addition, Figure 65 the adjacent node G0 shown is an adjacent node that can determine whether it is occupied by the occupancy rate encoding of the grandparent node. The adjacent node X1 is an adjacent node that can determine whether it is occupied by the occupancy rate encoding of the parent node.

[0611] Hereinafter, Modification Example 3 of the present embodiment will be described. Figure 66 and Figure 67 are diagrams showing the reference relationship related to this modification example, Figure 66 is a diagram showing the reference relationship on an octree structure, Figure 67 is a diagram showing the reference relationship on a spatial region.

[0612] In this modification example, when the three-dimensional data encoding device encodes the encoding information of the node to be encoded (hereinafter referred to as the object node 2), it refers to the encoding information of each node in the parent node to which the object node 2 belongs. That is, the three-dimensional data encoding device permits referring to the information (such as occupancy information) of the child nodes of the first node whose parent node is the same as the parent node of the object node among the plurality of adjacent nodes. For example, when the three-dimensional data encoding device encodes Figure 66 the occupancy rate encoding of the object node 2 shown, it refers to the nodes existing in the parent node to which the object node 2 belongs. For example, Figure 66 the occupancy rate encoding of the object node shown. As Figure 67 shown, Figure 66 the occupancy rate encoding of the object node shown indicates whether each node in the object node adjacent to the object node 2, for example, is occupied. Therefore, the three-dimensional data encoding device can switch the encoding table of the occupancy rate encoding of the object node 2 according to the finer shape of the object node, and thus can improve the encoding efficiency.

[0613] The three-dimensional data encoding device can also calculate the encoding table for entropy encoding the occupancy rate encoding of the object node 2 by, for example, the following formula.

[0614] CodingTable = (FlagX1 << 5) + (FlagX2 << 4) + (FlagX3 << 3) + (FlagX4 << 2) + (FlagY << 1) + (FlagZ)

[0615] Here, CodingTable represents a coding table for encoding the occupancy rate of object node 2, representing any one of values 0 to 63. FlagXN is the occupancy information of adjacent node XN (N = 1...4). If the adjacent node XN contains (occupies) a point group, it represents 1; otherwise, it represents 0. FlagY is the occupancy information of adjacent node Y. If the adjacent node Y contains (occupies) a point group, it represents 1; otherwise, it represents 0. FlagZ is the occupancy information of adjacent node Y. If the adjacent node Z contains (occupies) a point group, it represents 1; otherwise, it represents 0.

[0616] In addition, the three-dimensional data encoding device can also change the calculation method of the coding table according to the node position of object node 2 within the parent node.

[0617] In addition, when it is not prohibited to refer to the parent adjacent node, the three-dimensional data encoding device can refer to the coding information of each node within the parent adjacent node. For example, when it is not prohibited to refer to the parent adjacent node, it is permitted to refer to the information (such as occupancy information) of the child nodes of the third node whose parent node is different from that of the object node. For example, in Figure 65 the example shown, the three-dimensional data encoding device refers to the occupancy rate encoding of adjacent node X0 whose parent node is different from that of the object node, and obtains the occupancy information of the child nodes of adjacent node X0. The three-dimensional data encoding device switches the coding table used in the entropy coding of the occupancy rate encoding of the object node based on the obtained occupancy information of the child nodes of adjacent node X0.

[0618] As described above, the three-dimensional data encoding device according to this embodiment encodes the information (such as occupancy rate encoding) of the object nodes included in the N-ary tree structure (where N is an integer greater than or equal to 2) of the multiple three-dimensional points included in the three-dimensional data. As Figure 51 and Figure 52 shown, in the above encoding, the three-dimensional data encoding device permits referring to the information (such as occupancy information) of the first node whose parent node is the same as that of the object node among the multiple adjacent nodes that are spatially adjacent to the object node, and prohibits referring to the information (such as occupancy information) of the second node whose parent node is different from that of the object node. In other words, in the above encoding, the three-dimensional data encoding device permits referring to the information of the parent node (such as occupancy rate encoding), and prohibits referring to the information (such as occupancy rate encoding) of other nodes (parent adjacent nodes) at the same layer as the parent node.

[0619] Accordingly, the three-dimensional data encoding device can improve the encoding efficiency by referring to the information of the first node among a plurality of adjacent nodes that are spatially adjacent to the object node and whose parent node is the same as the parent node of the object node. In addition, since the three-dimensional data encoding device does not refer to the information of the second node among the plurality of adjacent nodes whose parent node is different from the parent node of the object node, the processing amount can be reduced. In this way, the three-dimensional data encoding device can improve the encoding efficiency and reduce the processing amount.

[0620] For example, the three-dimensional data encoding device further determines whether to prohibit referring to the information of the second node, and in the above encoding, based on the result of the above determination, switches whether to prohibit or permit referring to the information of the second node. The three-dimensional data encoding device further generates a bitstream including prohibition switching information (for example, Figure 63 the limit_refer_flag shown), where the prohibition switching information is the result of the above determination and indicates whether to prohibit referring to the information of the second node.

[0621] Accordingly, the three-dimensional data encoding device can switch whether to prohibit referring to the information of the second node. In addition, the three-dimensional data decoding device can appropriately perform decoding processing using the prohibition switching information.

[0622] For example, the information of the object node is information indicating whether there are three-dimensional points for each of the child nodes belonging to the object node (for example, occupancy encoding), the information of the first node is information indicating whether there are three-dimensional points at the first node (the occupancy information of the first node), and the information of the second node is information indicating whether there are three-dimensional points at the second node (the occupancy information of the second node).

[0623] For example, in the above encoding, the three-dimensional data encoding device selects an encoding table based on whether there are three-dimensional points at the first node, and uses the selected encoding table to perform entropy encoding on the information of the object node (for example, occupancy encoding).

[0624] For example, as Figure 66 and Figure 67 shown, in the above encoding, the three-dimensional data encoding device permits referring to the information (for example, occupancy information) of the child nodes of the first node among the plurality of adjacent nodes.

[0625] Accordingly, since the three-dimensional data encoding device can refer to more detailed information of adjacent nodes, the encoding efficiency can be improved.

[0626] For example, as Figure 53 shown, in the above encoding, the three-dimensional data encoding device switches the adjacent nodes to be referred to among the plurality of adjacent nodes according to the spatial position within the parent node of the object node.

[0627] Accordingly, based on the spatial position within the parent node of the object node, the three-dimensional data encoding device can refer to appropriate adjacent nodes.

[0628] For example, the three-dimensional data encoding device includes a processor and a memory, and the processor uses the memory to perform the above-described processing.

[0629] In addition, the three-dimensional data decoding device according to the present embodiment decodes information (such as occupancy encoding) of an object node included in an N-ary tree structure (where N is an integer of 2 or more) of a plurality of three-dimensional points included in the three-dimensional data. As Figure 51 and Figure 52 shown, in the above decoding, the three-dimensional data decoding device permits referring to information (such as occupancy information) of a first node whose parent node is the same as the parent node of the object node among a plurality of adjacent nodes that are spatially adjacent to the object node, and prohibits referring to information (such as occupancy information) of a second node whose parent node is different from the parent node of the object node. In other words, in the above decoding, the three-dimensional data decoding device permits referring to information of the parent node (such as occupancy encoding), and prohibits referring to information (such as occupancy encoding) of other nodes (parent adjacent nodes) at the same level as the parent node.

[0630] Accordingly, by referring to information of the first node whose parent node is the same as the parent node of the object node among a plurality of adjacent nodes that are spatially adjacent to the object node, the three-dimensional data decoding device can improve the encoding efficiency. In addition, since the three-dimensional data decoding device does not refer to information of the second node whose parent node is different from the parent node of the object node among the plurality of adjacent nodes, the processing amount can be reduced. In this way, the three-dimensional data decoding device can improve the encoding efficiency and reduce the processing amount.

[0631] For example, the three-dimensional data decoding device further obtains prohibition switching information (such as, Figure 63 shown as limit_refer_flag) indicating whether to prohibit referring to information of the second node from the bitstream, and in the above decoding, based on the prohibition switching information, switches whether to prohibit or permit referring to information of the second node.

[0632] Accordingly, the three-dimensional data decoding device can appropriately perform decoding processing using the prohibition switching information.

[0633] For example, the information of the object node is information (such as occupancy encoding) indicating whether there are three-dimensional points for each of the child nodes belonging to the object node, the information of the first node is information (the occupancy information of the first node) indicating whether there are three-dimensional points in the first node, and the information of the second node is information (the occupancy information of the second node) indicating whether there are three-dimensional points in the second node.

[0634] For example, in the above decoding, the three-dimensional data decoding device selects a coding table based on the presence or absence of three-dimensional points at the first node, and uses the selected coding table to perform entropy decoding on the information of the object node (e.g., occupancy coding).

[0635] For example, as Figure 66 and Figure 67 shown, in the above decoding, the three-dimensional data decoding device permits referring to the information (e.g., occupancy information) of the child nodes of the first node among multiple adjacent nodes.

[0636] Thus, the three-dimensional data decoding device can refer to more detailed information of adjacent nodes, and therefore can improve the coding efficiency.

[0637] For example, as Figure 53 shown, in the above decoding, the three-dimensional data decoding device switches the adjacent nodes to be referred to among multiple adjacent nodes according to the spatial position within the parent node of the object node.

[0638] Thus, the three-dimensional data decoding device can refer to appropriate adjacent nodes according to the spatial position within the parent node of the object node.

[0639] For example, the three-dimensional data decoding device includes a processor and a memory, and the processor uses the memory to perform the above processing.

[0640] (Embodiment 9)

[0641] The information of the three-dimensional point cloud includes position information (geometry) and attribute information (attribute). The position information includes coordinates (x coordinate, y coordinate, z coordinate) based on a certain point. In the case of encoding the position information, instead of directly encoding the coordinates of each three-dimensional point, a method is used to reduce the encoding amount by representing the position of each three-dimensional point using an octree representation and encoding the information of the octree.

[0642] On the other hand, the attribute information includes information such as color information (RGB, YUV, etc.), reflectivity, and normal vector representing each three-dimensional point. For example, the three-dimensional data encoding device can encode the attribute information using a coding method different from the position information.

[0643] In this embodiment, a method for encoding attribute information will be described. In addition, in this embodiment, integer values are used as the values of the attribute information for description. For example, when each color component of color information RGB or YUV has an 8-bit precision, each color component takes an integer value from 0 to 255. When the value of the reflectance has a 10-bit precision, the value of the reflectance takes an integer value from 0 to 1023. In addition, when the bit precision of the attribute information is a fractional precision, the three-dimensional data encoding device may also round the value to an integer value after multiplying the value by a scaling value so that the value of the attribute information becomes an integer value. In addition, the three-dimensional data encoding device may also attach the scaling value to the head of the bitstream or the like.

[0644] As a method for encoding the attribute information of a three-dimensional point, consider calculating a predicted value of the attribute information of the three-dimensional point and encoding the difference (prediction residual) between the original value of the attribute information and the predicted value. For example, when the value of the attribute information of the three-dimensional point p is Ap and the predicted value is Pp, the three-dimensional data encoding device encodes the absolute value of the difference Diffp = |Ap - Pp|. In this case, if the predicted value Pp can be generated with high precision, the value of the absolute difference Diffp becomes smaller. Therefore, for example, by using a coding table with a smaller number of bits generated for smaller values to perform entropy coding on the absolute difference Diffp, the coding amount can be reduced.

[0645] As a method for generating a predicted value of the attribute information, consider using the attribute information of other three-dimensional points, that is, reference three-dimensional points, located around the object three-dimensional point to be encoded. Here, the reference three-dimensional point refers to a three-dimensional point within a pre-specified distance range from the object three-dimensional point. For example, when there are an object three-dimensional point p = (x1, y1, z1) and a three-dimensional point q = (x2, y2, z2), the three-dimensional data encoding device calculates the Euclidean distance d(p, q) between the three-dimensional point p and the three-dimensional point q shown in (Equation A1).

[0646]

Equation 1

[0647]

[0648] When the Euclidean distance d(p, q) between three-dimensional data is less than a predetermined threshold THd, it is determined that the position of the three-dimensional point q is close to the position of the target three-dimensional point p, and it is determined that the value of the attribute information of the three-dimensional point q is used in the generation of the predicted value of the attribute information of the target three-dimensional point p. In addition, the distance calculation method can also be other methods, such as the Mahalanobis distance, etc. In addition, the three-dimensional data encoding device can also determine not to use three-dimensional points outside a pre-specified distance range from the target three-dimensional point for prediction processing. For example, when there is a three-dimensional point r and the distance d(p, r) between the target three-dimensional point p and the three-dimensional point r is equal to or greater than the threshold THd, the three-dimensional data encoding device can also determine not to use the three-dimensional point r for prediction. In addition, the three-dimensional data encoding device can also attach information indicating the threshold THd to the head of the bitstream, etc.

[0649] Figure 68 FIG. is an example showing three-dimensional points. In this example, the distance d(p, q) between the target three-dimensional point p and the three-dimensional point q is less than the threshold THd. Therefore, the three-dimensional data encoding device determines the three-dimensional point q as the reference three-dimensional point of the target three-dimensional point p, and determines to use the value of the attribute information Aq of the three-dimensional point q in the generation of the predicted value Pp of the attribute information Ap of the target three-dimensional point p.

[0650] On the other hand, the distance d(p, r) between the target three-dimensional point p and the three-dimensional point r is equal to or greater than the threshold THd. Therefore, the three-dimensional data encoding device determines that the three-dimensional point r is not the reference three-dimensional point of the target three-dimensional point p, and determines not to use the value of the attribute information Ar of the three-dimensional point r in the generation of the predicted value Pp of the attribute information Ap of the target three-dimensional point p.

[0651] In addition, when encoding the attribute information of the three-dimensional point using the predicted value, the three-dimensional data encoding device uses the three-dimensional point whose attribute information has been encoded and decoded as the reference three-dimensional point. Similarly, when decoding the attribute information of the target three-dimensional point of the decoding object using the predicted value, the three-dimensional data decoding device uses the three-dimensional point whose attribute information has been decoded as the reference three-dimensional point. Thereby, the same predicted value can be generated during encoding and decoding, so that the bitstream of the three-dimensional point generated by encoding can be correctly decoded on the decoding side.

[0652] In addition, when encoding the attribute information of the three-dimensional point, consider classifying each three-dimensional point into multiple levels using the position information of the three-dimensional point and then encoding. Here, each classified level is called LoD (Level of Detail). Use Figure 69 To explain the generation method of LoD.

[0653] First, the 3D data encoding device selects an initial point a0 and assigns it to LoD0. Next, the 3D data encoding device extracts points a1 whose distance from point a0 is greater than the threshold Thres_LoD[0] of LoD0 and assigns them to LoD0. Next, the 3D data encoding device extracts points a2 whose distance from point a1 is greater than the threshold Thres_LoD[0] of LoD0 and assigns them to LoD0. In this way, the 3D data encoding device constructs LoD0 such that the distance between each point within LoD0 is greater than the threshold Thres_LoD[0].

[0654] Next, the 3D data encoding device selects a point b0 that has not been assigned an LoD and assigns it to LoD1. Next, the 3D data encoding device extracts points b1 whose distance from point b0 is greater than the threshold Thres_LoD[1] of LoD1 and that have not been assigned an LoD and assigns them to LoD1. Next, the 3D data encoding device extracts points b2 whose distance from point b1 is greater than the threshold Thres_LoD[1] of LoD1 and that have not been assigned an LoD and assigns them to LoD1. In this way, the 3D data encoding device constructs LoD1 such that the distance between each point within LoD1 is greater than the threshold Thres_LoD[1].

[0655] Next, the 3D data encoding device selects a point c0 that has not been assigned an LoD and assigns it to LoD2. Next, the 3D data encoding device extracts points c1 whose distance from point c0 is greater than the threshold Thres_LoD[2] of LoD2 and that have not been assigned an LoD and assigns them to LoD2. Next, the 3D data encoding device extracts points c2 whose distance from point c1 is greater than the threshold Thres_LoD[2] of LoD2 and that have not been assigned an LoD and assigns them to LoD2. In this way, the 3D data encoding device constructs LoD2 such that the distance between each point within LoD2 is greater than the threshold Thres_LoD[2]. For example, as Figure 70 shown, the thresholds Thres_LoD[0], Thres_LoD[1], and Thres_LoD[2] of each LoD are set.

[0656] In addition, the 3D data encoding device may also attach information indicating the thresholds of each LoD to the header of the bitstream or the like. For example, in the case of the example shown in Figure 70 the 3D data encoding device may also attach the thresholds Thres_LoD[0], Thres_LoD[1], and Thres_LoD[2] to the header.

[0657] In addition, the 3D data encoding device may also assign all 3D points that have not been assigned an LoD to the bottom layer of the LoD. In this case, the 3D data encoding device can reduce the encoding amount of the header by not attaching the threshold of the bottom layer of the LoD to the header. For example, in Figure 70In the case of the example shown, the three-dimensional data encoding device attaches the thresholds Thres_LoD[0] and Thres_LoD[1] to the header and does not attach Thres_LoD[2] to the header. In this case, the three-dimensional data decoding device can also estimate the value of Thres_LoD[2] as 0. Additionally, the three-dimensional data encoding device can also attach the number of levels of LoD to the header. Thereby, the three-dimensional data decoding device can use the number of levels of LoD to determine the lowest-level LoD.

[0658] In addition, as Figure 70 shown, the threshold values of each layer of LoD are set to be larger for the upper layers, so that the upper layers (the layers closer to LoD0) are sparser point groups with a greater distance between three-dimensional points, and the lower layers are denser point groups with a smaller distance between three-dimensional points. In addition, in Figure 70 the example shown, LoD0 is the uppermost layer.

[0659] In addition, the method for selecting the initial three-dimensional points when setting each LoD can also depend on the encoding order during position information encoding. For example, the three-dimensional data encoding device selects the three-dimensional point that was first encoded during position information encoding as the initial point a0 of LoD0, and based on the initial point a0, selects points a1 and a2 to form LoD0. Moreover, the three-dimensional data encoding device can also select the three-dimensional point with the earliest encoded position information among the three-dimensional points that do not belong to LoD0 as the initial point b0 of LoD1. That is, the three-dimensional data encoding device can also select the three-dimensional point with the earliest encoded position information among the three-dimensional points that do not belong to the upper layers (LoD0 to LoDn - 1) of LoDn as the initial point n0 of LoDn. Thereby, by using the same initial point selection method during decoding, the three-dimensional data decoding device can construct the same LoD as during encoding, and thus can appropriately decode the bitstream. Specifically, the three-dimensional data decoding device selects the three-dimensional point with the earliest decoded position information among the three-dimensional points that do not belong to the upper layers of LoDn as the initial point n0 of LoDn.

[0660] Next, a method for predicting the attribute information of three-dimensional points using LoD will be described. For example, in the case of sequentially encoding starting from the three-dimensional points included in LoD0, the three-dimensional data encoding device generates the target three-dimensional points included in LoD1 using the encoded (hereinafter, also simply referred to as "encoded") attribute information included in LoD0 and LoD1. In this way, the three-dimensional data encoding device generates a predicted value of the attribute information of the three-dimensional points included in LoDn using the encoded attribute information included in LoDn’ (n’ <= n). That is, the three-dimensional data encoding device does not use the attribute information of the three-dimensional points included in the lower layer of LoDn in the calculation of the predicted value of the attribute information of the three-dimensional points included in LoDn.

[0661] For example, the three-dimensional data encoding device generates a predicted value of the attribute information of the three-dimensional points by calculating the average value of the attribute values of N or fewer three-dimensional points among the encoded three-dimensional points around the target three-dimensional points of the encoding object. In addition, the three-dimensional data encoding device may attach the value of N to the head of the bitstream or the like. Additionally, the three-dimensional data encoding device may also change the value of N for each three-dimensional point and attach the value of N to each three-dimensional point. Thereby, an appropriate N can be selected for each three-dimensional point, so the accuracy of the predicted value can be improved. Therefore, the prediction residual can be reduced. In addition, the three-dimensional data encoding device may also attach the value of N to the head of the bitstream and fix the value of N within the bitstream. Thereby, it is not necessary to encode or decode the value of N for each three-dimensional point, so the processing amount can be reduced. Additionally, the three-dimensional data encoding device may also encode the value of N separately for each LoD. Thereby, by selecting an appropriate N for each LoD, the encoding efficiency can be improved.

[0662] Alternatively, the three-dimensional data encoding device may also calculate the predicted value of the attribute information of the three-dimensional points by the weighted average of the attribute information of the surrounding N encoded three-dimensional points. For example, the three-dimensional data encoding device uses the distance information between the target three-dimensional point and each of the surrounding N three-dimensional points to calculate the weight.

[0663] When the three-dimensional data encoding device encodes the value of N separately for each LoD, for example, the value of N is set larger for the upper layer of LoD and smaller for the lower layer. In the upper layer of LoD, the distance between the three-dimensional points belonging to this layer is far, so it is possible to improve the prediction accuracy by setting a large value of N and selecting multiple surrounding three-dimensional points for averaging. In addition, since in the lower layer of LoD, the distance between the three-dimensional points belonging to this layer is close, it is possible to perform efficient prediction while suppressing the processing amount of averaging by setting a small value of N.

[0664] Figure 71It is a diagram showing an example of the attribute information used in the predicted values. As described above, the predicted value of the point P included in LoDn is generated using the encoded surrounding points P' included in LoDN’ (N’ <= N). Here, the surrounding points P' are selected based on the distance from the point P. For example, the attribute information of points a0, a1, a2, b0, and b1 is used to generate Figure 71 the predicted value of the attribute information of the point b2 shown.

[0665] According to the value of N described above, the selected surrounding points change. For example, when N = 5, a0, a1, a2, b0, and b1 are selected as the surrounding points of the point b2. When N = 4, points a0, a1, a2, and b1 are selected based on the distance information.

[0666] The predicted value is calculated by weighted average depending on the distance. For example, in Figure 71 the example shown, the predicted value a2p of the point a2 is calculated by the weighted average of the attribute information of the points a0 and a1 as shown in (Equation A2) and (Equation A3). Additionally, A i is the value of the attribute information of the point ai.

[0667]

Equation 2

[0668]

[0669] Additionally, the predicted value b2p of the point b2 is calculated by the weighted average of the attribute information of the points a0, a1, a2, b0, and b1 as shown in (Equation A4) to (Equation A6). Additionally, B i is the value of the attribute information of the point bi.

[0670]

Equation 3

[0671]

[0672] Furthermore, the three-dimensional data encoding device can also calculate the difference value (prediction residual) between the value of the attribute information of the three-dimensional point and the predicted value generated from the surrounding points, and quantize the calculated prediction residual. For example, the three-dimensional data encoding device quantizes by dividing the prediction residual by the quantization scale (also called the quantization step). In this case, the smaller the quantization scale, the smaller the error (quantization error) that may be generated due to quantization. On the contrary, the larger the quantization scale, the larger the quantization error.

[0673] In addition, the three-dimensional data encoding device may also change the quantization scale used for each LoD. For example, the higher the layer of the three-dimensional data encoding device, the smaller the quantization scale, and the lower the layer, the larger the quantization scale. Since the value of the attribute information of the three-dimensional points belonging to the upper layer may be used as the predicted value of the attribute information of the three-dimensional points belonging to the lower layer, the quantization scale of the upper layer can be reduced to suppress the quantization error generated in the upper layer. By improving the accuracy of the predicted value, the encoding efficiency can be improved. In addition, the three-dimensional data encoding device may also attach the quantization scale used for each LoD to the header or the like. Thus, the three-dimensional data decoding device can correctly decode the quantization scale, and thus can appropriately decode the bitstream.

[0674] In addition, the three-dimensional data encoding device may also transform the signed integer value (signed quantization value) that is the quantized prediction residual into an unsigned integer value (unsigned quantization value). Thus, when entropy encoding the prediction residual, there is no need to consider the generation of negative integers. In addition, the three-dimensional data encoding device does not necessarily need to transform the signed integer value into an unsigned integer value. For example, the sign bit may also be entropy encoded separately.

[0675] The prediction residual is calculated by subtracting the predicted value from the original value. For example, as shown in (Equation A7), the prediction residual a2r of point a2 is calculated by subtracting the predicted value a2p of point a2 from the value A of the attribute information of point a2. 2 As shown in (Equation A8), the prediction residual b2r of point b2 is calculated by subtracting the predicted value b2p of point b2 from the value B of the attribute information of point b2. 2

[0676] a2r = A 2 - a2p...(Equation A7)

[0677] b2r = B 2 - b2p...(Equation A8)

[0678] In addition, the prediction residual is quantized by dividing by QS (Quantization Step). For example, the quantization value a2q of point a2 is calculated by (Equation A9). The quantization value b2q of point b2 is calculated by (Equation A10). Here, QS_LoD0 is the QS for LoD0, and QS_LoD1 is the QS for LoD1. That is, QS can be changed according to LoD.

[0679] a2q = a2r / QS_LoD0...(Equation A9)

[0680] b2q = b2r / QS_LoD1...(Equation A10)

[0681] In addition, as described below, the three-dimensional data encoding device converts the signed integer value, which is the quantization value described above, into an unsigned integer value. When the signed integer value a2q is less than 0, the three-dimensional data encoding device sets the unsigned integer value a2u to -1 - (2 × a2q). When the signed integer value a2q is 0 or greater, the three-dimensional data encoding device sets the unsigned integer value a2u to 2 × a2q.

[0682] Similarly, when the signed integer value b2q is less than 0, the three-dimensional data encoding device sets the unsigned integer value b2u to -1 - (2 × b2q). When the signed integer value b2q is 0 or greater, the three-dimensional data encoding device sets the unsigned integer value b2u to 2 × b2q.

[0683] In addition, the three-dimensional data encoding device can also perform entropy encoding on the quantized prediction residual (unsigned integer value). For example, binary arithmetic coding can also be applied after binarizing the unsigned integer value.

[0684] In addition, in this case, the three-dimensional data encoding device can also switch the binarization method according to the value of the prediction residual. For example, when the prediction residual pu is less than the threshold R_TH, the three-dimensional data encoding device binarizes the prediction residual pu with a fixed number of bits required to represent the threshold R_TH. In addition, when the prediction residual pu is equal to or greater than the threshold R_TH, the three-dimensional data encoding device uses, for example, Exponential-Golomb to binarize the binarized data of the threshold R_TH and the value (pu - R_TH).

[0685] For example, when the threshold R_TH is 63 and the prediction residual pu is less than 63, the three-dimensional data encoding device binarizes the prediction residual pu with 6 bits. In addition, when the prediction residual pu is 63 or greater, the three-dimensional data encoding device uses Exponential-Golomb to binarize the binarized data of the threshold R_TH (111111) and (pu - 63), and then performs arithmetic coding.

[0686] In a more specific example, when the prediction residual pu is 32, the three-dimensional data encoding device generates 6-bit binary data (100000) and performs arithmetic coding on this bit string. In addition, when the prediction residual pu is 66, the three-dimensional data encoding device generates the binary data (111111) representing the threshold R_TH using Exponential-Golomb and the bit string (00100) of the value 3 (66 - 63), and performs arithmetic coding on this bit string (111111 + 00100).

[0687] In this way, the three-dimensional data encoding device switches the binarization method according to the magnitude of the prediction residual, so that encoding can be performed while suppressing a sharp increase in the number of binarization bits in the case where the prediction residual becomes large. In addition, the three-dimensional data encoding device may also attach the threshold value R_TH to the head of the bitstream or the like.

[0688] For example, in the case of encoding at a high bit rate, that is, in the case of a small quantization scale, the quantization error becomes small and the prediction accuracy becomes high. As a result, the prediction residual may not become large. Therefore, in this case, the three-dimensional data encoding device sets the threshold value R_TH to be large. Thereby, the possibility of encoding the binarized data of the threshold value R_TH becomes low, and the encoding efficiency is improved. On the contrary, in the case of encoding at a low bit rate, that is, in the case of a large quantization scale, the quantization error becomes large and the prediction accuracy deteriorates. As a result, it is possible that the prediction residual becomes large. Therefore, in this case, the three-dimensional data encoding device sets the threshold value R_TH to be small. Thereby, a sharp increase in the bit length of the binarized data can be prevented.

[0689] In addition, the three-dimensional data encoding device may also switch the threshold value R_TH for each LoD and attach the threshold value R_TH of each LoD to the head or the like. That is, the three-dimensional data encoding device may also switch the binarization method for each LoD. For example, in the upper layer, since the distance between three-dimensional points is far, the prediction accuracy deteriorates, and as a result, it is possible that the prediction residual becomes large. Therefore, the three-dimensional data encoding device sets the threshold value R_TH to be small for the upper layer to prevent a sharp increase in the bit length of the binarized data. In addition, in the lower layer, since the distance between three-dimensional points is close, the prediction accuracy becomes high, and as a result, it is possible that the prediction residual becomes small. Therefore, the three-dimensional data encoding device improves the encoding efficiency by setting the threshold value R_TH to be large for the layer.

[0690] Figure 72 It is a diagram showing an example of the exponential Golomb code and shows the relationship between the value before binarization (multi-value) and the bits (encoded) after binarization. In addition, it is also possible to Figure 72 reverse the 0 and 1 shown.

[0691] In addition, the three-dimensional data encoding device applies arithmetic coding to the binarized data of the prediction residual. Thereby, the encoding efficiency can be improved. In addition, when applying arithmetic coding, in the binarized data, in the part binarized with n bits, that is, the n-bit code, and the part binarized using the exponential Golomb, that is, the remaining code, the tendency of the occurrence probability of 0 and 1 of each bit may be different. Therefore, the three-dimensional data encoding device may also switch the application method of arithmetic coding by the n-bit code and the remaining code.

[0692] For example, for n-bit encoding, the three-dimensional data encoding device performs arithmetic encoding on each bit using a different encoding table (probability table). At this time, the three-dimensional data encoding device can also change the number of encoding tables used for each bit. For example, the three-dimensional data encoding device uses 1 encoding table to perform arithmetic encoding on the leading bit b0 of the n-bit encoding. In addition, the three-dimensional data encoding device uses 2 encoding tables for the next bit b1. Furthermore, the three-dimensional data encoding device switches the encoding table used in the arithmetic encoding of bit b1 according to the value (0 or 1) of b0. Similarly, the three-dimensional data encoding device uses 4 encoding tables for the next bit b2. In addition, the three-dimensional data encoding device switches the encoding table used in the arithmetic encoding of bit b2 according to the values (0 to 3) of b0 and b1.

[0693] In this way, when the three-dimensional data encoding device performs arithmetic encoding on each bit bn-1 of the n-bit encoding, it uses 2 n-1 encoding tables. In addition, the three-dimensional data encoding device switches the encoding table used according to the values (occurrence patterns) of the bits before bn-1. Thereby, the three-dimensional data encoding device can use an appropriate encoding table for each bit, and thus can improve the encoding efficiency.

[0694] In addition, the three-dimensional data encoding device can also reduce the number of encoding tables used for each bit. For example, when performing arithmetic encoding on each bit bn-1, the three-dimensional data encoding device can also switch 2 m encoding tables according to the values (occurrence patterns) of the previous m bits (m < n-1) of bn-1. Thereby, the number of encoding tables used for each bit can be suppressed, and the encoding efficiency can be improved. In addition, the three-dimensional data encoding device can also update the occurrence probabilities of 0 and 1 in each encoding table according to the values of the actually generated binarized data. In addition, the three-dimensional data encoding device can also fix the occurrence probabilities of 0 and 1 in the encoding tables of a part of the bits. Thereby, the number of times of updating the occurrence probability can be suppressed, and thus the processing amount can be reduced.

[0695] For example, when the n-bit encoding is b0b1b2…bn-1, the encoding table used for b0 is 1 (CTb0). The encoding tables used for b1 are 2 (CTb10, CTb11). In addition, the encoding table used is switched according to the value (0 to 1) of b0. The encoding tables used for b2 are 4 (CTb20, CTb21, CTb22, CTb23). Furthermore, the encoding table used is switched according to the values (0 to 3) of b0 and b1. The encoding tables used for bn-1 are 2 n-1 (CTbn0, CTbn1, …, CTbn(2 n-1 -1)). In addition, the encoding table used is switched according to the values (0 to 2 n-1 -1) of b0b1…bn-2.

[0696] In addition, the three-dimensional data encoding device may also apply arithmetic coding of m-ary with values set from 0 to 2 n -1 to the n-bit coding without binarization (m = 2 n ). In addition, when the three-dimensional data encoding device performs arithmetic coding on the n-bit coding using m-ary, the three-dimensional data decoding device may also restore the n-bit coding by arithmetic decoding of m-ary.

[0697] Figure 73 is a diagram for explaining the processing when, for example, the residual coding is an exponential Golomb code. As Figure 73 shown, the part where binarization is performed using exponential Golomb, that is, the residual coding, includes a prefix part and a suffix part. For example, the three-dimensional data encoding device switches the coding tables in the prefix part and the suffix part. That is, the three-dimensional data encoding device performs arithmetic coding on each bit included in the prefix part using the coding table for the prefix, and performs arithmetic coding on each bit included in the suffix part using the coding table for the suffix.

[0698] In addition, the three-dimensional data encoding device may also update the occurrence probabilities of 0 and 1 in each coding table according to the value of the actually generated binarized data. Or, the three-dimensional data encoding device may fix the occurrence probabilities of 0 and 1 in a certain coding table. Thereby, the number of times of updating the occurrence probability can be suppressed, and thus the processing amount can be reduced. For example, the three-dimensional data encoding device may update the occurrence probability for the prefix part and fix the occurrence probability for the suffix part.

[0699] In addition, the three-dimensional data encoding device decodes the quantized prediction residual through inverse quantization and reconstruction, and uses the decoded prediction residual, that is, the decoded value, for prediction after the three-dimensional point to be encoded. Specifically, the three-dimensional data encoding device calculates the inverse quantization value by multiplying the quantized prediction residual (quantized value) by the quantization scale, and obtains the decoded value (reconstructed value) by adding the inverse quantization value and the predicted value.

[0700] For example, the inverse quantization value a2iq of point a2 is calculated using the quantization value a2q of point a2 through (Equation A11). The inverse quantization value b2iq of point b2 is calculated using the quantization value b2q of point b2 through (Equation A12). Here, QS_LoD0 is the QS for LoD0, and QS_LoD1 is the QS for LoD1. That is, the QS can be changed according to LoD.

[0701] a2iq = a2q × QS_LoD0 … (Equation A11)

[0702] b2iq = b2q × QS_LoD1 … (Equation A12)

[0703] For example, as shown in (Equation A13), the decoded value a2rec of point a2 is calculated by adding the inverse quantization value a2iq of point a2 to the predicted value a2p of point a2. As shown in (Equation A14), the decoded value b2rec of point b2 is calculated by adding the inverse quantization value b2iq of point b2 to the predicted value b2p of point b2.

[0704] a2rec = a2iq + a2p … (Equation A13)

[0705] b2rec = b2iq + b2p … (Equation A14)

[0706] Next, a syntax example of the bitstream according to this embodiment will be described. Figure 74 FIG. is a syntax example of the attribute header of this embodiment. The attribute header is the header information of the attribute information. As Figure 74 shown, the attribute header includes the number of levels information (NumLoD), the three-dimensional point number information (NumOfPoint[i]), the level threshold (Thres_Lod[i]), the surrounding point number information (NumNeighorPoint[i]), the prediction threshold (THd[i]), the quantization scale (QS[i]), and the binarization threshold (R_TH[i]).

[0707] The number of levels information (NumLoD) represents the number of levels of the LoD used.

[0708] The three-dimensional point number information (NumOfPoint[i]) represents the number of three-dimensional points belonging to level i. In addition, the three-dimensional data encoding device may also attach the three-dimensional point total number information (AllNumOfPoint) indicating the total number of three-dimensional points to another header. In this case, the three-dimensional data encoding device may not attach NumOfPoint[NumLoD - 1] indicating the number of three-dimensional points belonging to the bottom layer to the header. In this case, the three-dimensional data decoding device can calculate NumOfPoint[NumLoD - 1] by (Equation A15). Thereby, the encoding amount of the header can be reduced.

[0709]

Equation 4

[0710]

[0711] The hierarchical threshold (Thres_Lod[i]) is used to set the threshold for level i. The three-dimensional data encoding device and the three-dimensional data decoding device form LoDi such that the distance between each point within LoDi is greater than the threshold Thres_LoD[i]. Additionally, the three-dimensional data encoding device may not attach the value of Thres_Lod[NumLoD - 1] (the bottommost level) to the header. In this case, the three-dimensional data decoding device estimates the value of Thres_Lod[NumLoD - 1] as 0. Thereby, the encoding amount of the header can be reduced.

[0712] The number of surrounding points information (NumNeighorPoint[i]) represents the upper limit value of the number of surrounding points used in the generation of the predicted value of the three-dimensional points belonging to level i. When the number of surrounding points M is less than NumNeighorPoint[i] (M < NumNeighorPoint[i]), the three-dimensional data encoding device may also use M surrounding points to calculate the predicted value. Additionally, when it is not necessary to separate the value of NumNeighorPoint[i] in each LoD, the three-dimensional data encoding device may also attach one number of surrounding points information (NumNeighorPoint) used in all LoDs to the header.

[0713] The prediction threshold (THd[i]) represents the upper limit value of the distance between the surrounding three-dimensional points used in the prediction of the object three-dimensional points to be encoded or decoded at level i and the object three-dimensional points. The three-dimensional data encoding device and the three-dimensional data decoding device do not use the three-dimensional points whose distance from the object three-dimensional points is farther than THd[i] for prediction. Additionally, when it is not necessary to separate the value of THd[i] in each LoD, the three-dimensional data encoding device may also attach one prediction threshold (THd) used in all LoDs to the header.

[0714] The quantization scale (QS[i]) represents the quantization scale used in the quantization and inverse quantization at level i.

[0715] The binarization threshold (R_TH[i]) is the threshold for switching the binarization method of the prediction residual of the three-dimensional points belonging to level i. For example, when the prediction residual is less than the threshold R_TH, the three-dimensional data encoding device binarizes the prediction residual pu with a fixed number of bits, and when the prediction residual is greater than or equal to the threshold R_TH, it uses exponential Golomb to binarize the binarized data of the threshold R_TH and the value of (pu - R_TH). Additionally, when it is not necessary to switch the value of R_TH[i] in each LoD, the three-dimensional data encoding device may also attach one binarization threshold (R_TH) used in all LoDs to the header.

[0716] In addition, R_TH[i] can also be the maximum value represented by n bits. For example, in 6 bits, R_TH is 63, and in 8 bits, R_TH is 255. Additionally, instead of encoding the maximum value represented by n bits as the binarization threshold, the three-dimensional data encoding device can encode the number of bits. For example, when R_TH[i]=63, the three-dimensional data encoding device can append the value 6 to the header, and when R_TH[i]=255, the three-dimensional data encoding device can append the value 8 to the header. Additionally, the three-dimensional data encoding device can also define the minimum number of bits (minimum bit number) representing R_TH[i] and append the relative bit number with respect to the minimum value to the header. For example, when R_TH[i]=63 and the minimum bit number is 6, the three-dimensional data encoding device can append the value 0 to the header, and when R_TH[i]=255 and the minimum bit number is 6, the three-dimensional data encoding device can append the value 2 to the header.

[0717] In addition, the three-dimensional data encoding device can also perform entropy encoding on at least one of NumLoD, Thres_Lod[i], NumNeighborPoint[i], THd[i], QS[i], and R_TH[i] and append it to the header. For example, the three-dimensional data encoding device can also perform arithmetic encoding by binarizing each value. Additionally, in order to suppress the processing amount, the three-dimensional data encoding device can also encode each value with a fixed length.

[0718] In addition, the three-dimensional data encoding device may not append at least one of NumLoD, Thres_Lod[i], NumNeighborPoint[i], THd[i], QS[i], and R_TH[i] to the header. For example, the value of at least one of them can also be specified by a profile or level such as a standard. Thus, the bit amount of the header can be reduced.

[0719] Figure 75 is a diagram showing a syntactic example of the attribute data of the present embodiment. This attribute data includes encoded data of attribute information of a plurality of three-dimensional points. As Figure 75 shown, the attribute data includes an n-bit code and a remaining code.

[0720] The n-bit code is encoded data of the prediction residual of the value of the attribute information or a part thereof. The bit length of the n-bit code depends on the value of R_TH[i]. For example, when the value shown by R_TH[i] is 63, the n-bit code is 6 bits, and when the value shown by R_TH[i] is 255, the n-bit code is 8 bits.

[0721] The remaining code is the coded data after exponential Golomb coding in the coded data of the prediction residual of the value of the attribute information. When the n-bit code is the same as R_TH[i], the remaining code is coded or decoded. In addition, the three-dimensional data decoding device decodes the prediction residual by adding the value of the n-bit code and the value of the remaining code. Furthermore, when the n-bit code is not the same value as R_TH[i], the remaining code may not be coded or decoded.

[0722] Hereinafter, the process flow in the three-dimensional data coding device will be described. Figure 76 It is a flowchart of the three-dimensional data coding process performed by the three-dimensional data coding device.

[0723] First, the three-dimensional data coding device codes the position information (geometry) (S3001). For example, the three-dimensional data coding is performed using octree representation.

[0724] After the three-dimensional data coding device codes the position information, when the position of the three-dimensional point changes due to quantization or the like, the three-dimensional data coding device reassigns the attribute information of the original three-dimensional point to the changed three-dimensional point (S3002). For example, the three-dimensional data coding device performs the reassignment by interpolating the value of the attribute information according to the amount of change in the position. For example, the three-dimensional data coding device detects N three-dimensional points before the change that are close to the changed three-dimensional position, and performs weighted averaging on the values of the attribute information of the N three-dimensional points. For example, in the weighted averaging, the three-dimensional data coding device determines the weight based on the distance from the changed three-dimensional position to each of the N three-dimensional points. Then, the three-dimensional data coding device determines the value obtained by the weighted averaging as the value of the attribute information of the changed three-dimensional point. In addition, when two or more three-dimensional points change to the same three-dimensional position due to quantization or the like, the three-dimensional data coding device may also assign the average value of the attribute information of the two or more three-dimensional points before the change as the value of the attribute information of the changed three-dimensional point.

[0725] Next, the three-dimensional data coding device codes the reassigned attribute information (Attribute) (S3003). For example, when coding multiple types of attribute information, the three-dimensional data coding device may also code the multiple types of attribute information in sequence. For example, when coding color and reflectance as the attribute information, the three-dimensional data coding device may also generate a bitstream in which the coding result of the reflectance is appended after the coding result of the color. In addition, the order of the multiple coding results of the attribute information appended to the bitstream is not limited to this order and can be any order.

[0726] In addition, the 3D data encoding device may also attach information indicating the start position of the encoded data of each attribute information in the bitstream to the header or the like. Thus, the 3D data decoding device can selectively decode the attribute information that needs to be decoded, and therefore can omit the decoding process of the attribute information that does not need to be decoded. Therefore, the processing amount of the 3D data decoding device can be reduced. In addition, the 3D data encoding device may also encode multiple types of attribute information in parallel and merge the encoding results into one bitstream. Thus, the 3D data encoding device can encode multiple types of attribute information at high speed.

[0727] Figure 77 is a flowchart of the attribute information encoding process (S3003). First, the 3D data encoding device sets the LoD (S3011). That is, the 3D data encoding device assigns each 3D point to any one of multiple LoDs.

[0728] Next, the 3D data encoding device starts a loop for each LoD (S3012). That is, the 3D data encoding device repeatedly performs the processes of steps S3013 to S3021 for each LoD.

[0729] Next, the 3D data encoding device starts a loop for each 3D point (S3013). That is, the 3D data encoding device repeatedly performs the processes of steps S3014 to S3020 for each 3D point.

[0730] First, the 3D data encoding device searches for a plurality of surrounding points, which are 3D points existing around the object 3D point used in the calculation of the predicted value of the object 3D point to be processed (S3014). Next, the 3D data encoding device calculates the weighted average of the values of the attribute information of the plurality of surrounding points and sets the obtained value as the predicted value P (S3015). Next, the 3D data encoding device calculates the difference between the attribute information of the object 3D point and the predicted value, that is, the prediction residual (S3016). Next, the 3D data encoding device calculates the quantization value by quantizing the prediction residual (S3017). Next, the 3D data encoding device performs arithmetic coding on the quantization value (S3018).

[0731] In addition, the 3D data encoding device calculates the inverse quantization value by inverse quantizing the quantization value (S3019). Next, the 3D data encoding device generates a decoded value by adding the inverse quantization value to the predicted value (S3020). Next, the 3D data encoding device ends the loop for each 3D point (S3021). In addition, the 3D data encoding device ends the loop for each LoD (S3022).

[0732] Hereinafter, the 3D data decoding process in the 3D data decoding device that decodes the bitstream generated by the above 3D data encoding device will be described.

[0733] The three-dimensional data decoding device generates decoded binary data by performing arithmetic decoding on the binary data of the attribute information in the bitstream generated by the three-dimensional data encoding device in the same manner as the three-dimensional data encoding device. In addition, in the three-dimensional data encoding device, when switching the application method of arithmetic coding between the part binarized with n bits (n-bit coding) and the part binarized with exponential Golomb (residual coding), the three-dimensional data decoding device performs decoding accordingly when applying arithmetic decoding.

[0734] For example, in the arithmetic decoding method of n-bit coding, the three-dimensional data decoding device performs arithmetic decoding on each bit using a different coding table (decoding table). At this time, the three-dimensional data decoding device can also change the number of coding tables used for each bit. For example, for the leading bit b0 of n-bit coding, 1 coding table is used for arithmetic decoding. In addition, the three-dimensional data decoding device uses 2 coding tables for the next bit b1. Furthermore, the three-dimensional data decoding device switches the coding table used in the arithmetic decoding of bit b1 according to the value of b0 (0 or 1). Similarly, the three-dimensional data decoding device further uses 4 coding tables for the next bit b2. In addition, the three-dimensional data decoding device switches the coding table used in the arithmetic decoding of bit b2 according to the values of b0 and b1 (0 to 3).

[0735] In this way, when performing arithmetic decoding on each bit bn-1 of n-bit coding, the three-dimensional data decoding device uses 2 n-1 coding tables. In addition, the three-dimensional data decoding device switches the coding table used according to the values of the bits before bn-1 (occurrence pattern). Thereby, the three-dimensional data decoding device can use an appropriate coding table for each bit to appropriately decode the bitstream with improved coding efficiency.

[0736] In addition, the three-dimensional data decoding device can also reduce the number of coding tables used for each bit. For example, when performing arithmetic decoding on each bit bn-1, the three-dimensional data decoding device can switch 2 m coding tables according to the values of the previous m bits (m < n - 1) of bn-1 (occurrence pattern). Thereby, the three-dimensional data decoding device can appropriately decode the bitstream with improved coding efficiency while suppressing the number of coding tables used for each bit. In addition, the three-dimensional data decoding device can also update the occurrence probabilities of 0 and 1 in each coding table according to the value of the actually generated binary data. In addition, the three-dimensional data decoding device can also fix the occurrence probabilities of 0 and 1 in the coding tables of a part of the bits. Thereby, the number of times of updating the occurrence probability can be suppressed, and thus the processing amount can be reduced.

[0737] For example, when n bits are encoded as b0b1b2…bn-1, there is 1 encoding table (CTb0) for b0. There are 2 encoding tables (CTb10, CTb11) for b1. In addition, the encoding table is switched according to the value of b0 (0 to 1). There are 4 encoding tables (CTb20, CTb21, CTb22, CTb23) for b2. In addition, the encoding table is switched according to the values of b0 and b1 (0 to 3). There are 2 n-1 encoding tables (CTbn0, CTbn1, …, CTbn(2 n-1 -1)) for bn-1. Also, the encoding table is switched according to the value of b0b1…bn-2 (0 to 2 n-1 -1).

[0738] For example, Figure 78 is a diagram for explaining the processing in the case where the residual code is an exponential Golomb code. As Figure 78 shown, the three-dimensional data encoding device uses exponential Golomb to binarize and encode a part (residual code), which includes a prefix part and a suffix part. For example, the three-dimensional data decoding device switches the encoding table between the prefix part and the suffix part. That is, the three-dimensional data decoding device uses the encoding table for the prefix to perform arithmetic decoding on each bit included in the prefix part, and uses the encoding table for the suffix to perform arithmetic decoding on each bit included in the suffix part.

[0739] In addition, the three-dimensional data decoding device can also update the occurrence probabilities of 0 and 1 in each encoding table according to the value of the binarized data generated during decoding. Or, the three-dimensional data decoding device can fix the occurrence probabilities of 0 and 1 in a certain encoding table. Thereby, the number of times of updating the occurrence probability can be suppressed, and thus the processing amount can be reduced. For example, the three-dimensional data decoding device can update the occurrence probability for the prefix part and fix the occurrence probability for the suffix part.

[0740] Also, the three-dimensional data decoding device multi-valued the binarized data of the predicted residual obtained by arithmetic decoding to match the encoding method used in the three-dimensional data encoding device, thereby decoding the quantized predicted residual (unsigned integer value). The three-dimensional data decoding device first calculates the value of the n-bit code decoded by performing arithmetic decoding on the binarized data encoded with n bits. Then, the three-dimensional data decoding device compares the value of the n-bit code with the value of R_TH.

[0741] When the value of the n-bit encoded data is the same as the value of R_TH, the three-dimensional data decoding device determines that there are bits encoded by exponential Golomb next, and performs arithmetic decoding on the binarized data encoded by exponential Golomb, that is, the remaining encoding. Then, the three-dimensional data decoding device calculates the value of the remaining encoding using a reverse table indicating the relationship between the remaining encoding and the value thereof according to the decoded remaining encoding. Figure 79 It is a diagram showing an example of a reverse table indicating the relationship between the remaining encoding and the value thereof. Next, the three-dimensional data decoding device obtains the quantized prediction residual after multi-valued quantization by adding the value of the obtained remaining encoding to R_TH.

[0742] On the other hand, when the value of the n-bit encoded data is not the same as the value of R_TH (the value is less than R_TH), the three-dimensional data decoding device directly determines the value of the n-bit encoded data as the quantized prediction residual after multi-valued quantization. Thus, the three-dimensional data decoding device can appropriately decode the bitstream generated by switching the binarization method according to the value of the prediction residual in the three-dimensional data encoding device.

[0743] In addition, when the threshold R_TH is attached to the head of the bitstream or the like, the three-dimensional data decoding device may also decode the value of the threshold R_TH from the head and use the decoded value of the threshold R_TH to switch the decoding method. Further, when the threshold R_TH is attached to the head or the like for each LoD, the three-dimensional data decoding device switches the decoding method for each LoD using the decoded threshold R_TH.

[0744] For example, when the threshold R_TH is 63 and the decoded value of the n-bit encoded data is 63, the three-dimensional data decoding device decodes the remaining encoding using exponential Golomb to obtain the value of the remaining encoding. For example, in Figure 79 the example shown, the remaining encoding is 00100, and 3 is obtained as the value of the remaining encoding. Then, the three-dimensional data decoding device obtains the value of the prediction residual 66 by adding the value 63 of the threshold R_TH and the value 3 of the remaining encoding.

[0745] In addition, when the decoded value of the n-bit encoded data is 32, the three-dimensional data decoding device sets the value of the n-bit encoded data 32 as the value of the prediction residual.

[0746] In addition, the 3D data decoding device transforms the decoded quantized prediction residual from an unsigned integer value into a signed integer value through, for example, a process opposite to that in the 3D data encoding device. Thereby, when the 3D data decoding device performs entropy encoding on the prediction residual, it can appropriately decode the bitstream generated without considering the generation of negative integers. In addition, the 3D data decoding device does not necessarily need to transform the unsigned integer value into a signed integer value. For example, when decoding a bitstream generated by separately performing entropy encoding on sign bits, it can also decode the sign bits.

[0747] The 3D data decoding device decodes the quantized prediction residual transformed into a signed integer value through inverse quantization and reconstruction, thereby generating a decoded value. In addition, the 3D data decoding device uses the generated decoded value for prediction after the 3D points to be decoded. Specifically, the 3D data decoding device calculates an inverse quantization value by multiplying the quantized prediction residual by the decoded quantization scale, and obtains a decoded value by adding the inverse quantization value and the predicted value.

[0748] The decoded unsigned integer value (unsigned quantization value) is transformed into a signed integer value through the following process. When the least significant bit (LSB) of the decoded unsigned integer value a2u is 1, the 3D data decoding device sets the signed integer value a2q to -((a2u + 1) >> 1). When the LSB of the unsigned integer value a2u is not 1, the 3D data decoding device sets the signed integer value a2q to (a2u >> 1).

[0749] Similarly, when the LSB of the decoded unsigned integer value b2u is 1, the 3D data decoding device sets the signed integer value b2q to -((b2u + 1) >> 1). When the LSB of the unsigned integer value n2u is not 1, the 3D data decoding device sets the signed integer value b2q to (b2u >> 1).

[0750] In addition, the details of the inverse quantization and reconstruction processes performed by the 3D data decoding device are the same as those in the 3D data encoding device.

[0751] Hereinafter, the flow of the process in the 3D data decoding device will be described. Figure 80 It is a flowchart of the 3D data decoding process performed by the 3D data decoding device. First, the 3D data decoding device decodes position information (geometry) from the bitstream (S3031). For example, the 3D data decoding device performs decoding using an octree representation.

[0752] Next, the three-dimensional data decoding device decodes the attribute information (Attribute) from the bitstream (S3032). For example, in the case of decoding multiple pieces of attribute information, the three-dimensional data decoding device can also decode the multiple pieces of attribute information sequentially. For example, in the case of decoding color and reflectance as attribute information, the three-dimensional data decoding device decodes the encoding result of the color and the encoding result of the reflectance in the order attached to the bitstream. For example, in the bitstream, if the encoding result of the reflectance is attached after the encoding result of the color, the three-dimensional data decoding device decodes the encoding result of the co...

Claims

1. A three-dimensional data encoding method, which is a three-dimensional data encoding method for encoding three-dimensional points with attribute information, wherein, calculate the predicted value of the attribute information of the three-dimensional point, calculate the difference between the attribute information of the three-dimensional point and the predicted value, i.e., the prediction residual, generate binary data by binarizing the prediction residual, perform arithmetic coding on the binary data, the binary data includes a prefix part, i.e., the prefix part, and a suffix part, i.e., the suffix part, in the arithmetic coding, different contexts are used for the prefix part and the suffix part.

2. The three-dimensional data encoding method according to claim 1, wherein, use the position information of the three-dimensional points to assign each of the three-dimensional points to a plurality of LoDs, i.e., levels of detail, for a certain level of LoD, use the encoded attribute information included in the LoD of the upper level of this level to generate the predicted value of the attribute information of the three-dimensional points included in the LoD of this level.

3. The three-dimensional data encoding method according to claim 1, wherein, in the arithmetic coding, different contexts are used for each bit of the binary data.

4. The three-dimensional data encoding method according to claim 3, wherein, in the arithmetic coding, the lower the bit position of the binary data, the more contexts are used.

5. The three-dimensional data encoding method according to any one of claims 1 to 4, wherein, in the arithmetic coding, according to the value of the upper bit of the target bit included in the binary data, select the context used in the arithmetic coding of the target bit.

6. The three-dimensional data encoding method according to any one of claims 1 to 4, wherein, in the binarization, when the prediction residual is less than the threshold, generate the binary data by binarizing the prediction residual with a fixed number of bits, when the prediction residual is greater than or equal to the threshold, generate the binary data including a first code representing the fixed number of bits of the threshold and a second code obtained by binarizing the value obtained by subtracting the threshold from the prediction residual using exponential Golomb, in the arithmetic coding, different arithmetic coding methods are used for the first code and the second code.

7. The three-dimensional data encoding method according to claim 6, wherein, the three-dimensional data encoding method further quantize the prediction residual, in the binarization, binarize the quantized prediction residual, the threshold is changed according to the quantization scale in the quantization.

8. A three-dimensional data decoding method, which is a three-dimensional data decoding method for decoding three-dimensional points with attribute information, wherein, calculate the predicted value of the attribute information of the three-dimensional point, generate binary data by performing arithmetic decoding on the encoded data included in the bitstream, generate a prediction residual by multi-valuing the binary data, calculate the decoded value of the attribute information of the three-dimensional point by adding the predicted value and the prediction residual, the binary data includes a prefix part, i.e., the prefix part, and a suffix part, i.e., the suffix part, In the arithmetic decoding, different contexts are used for the prefix part and the suffix part.

9. The three-dimensional data decoding method according to claim 8, wherein, using the position information of the three-dimensional points, each of the three-dimensional points is assigned to a plurality of LoDs (Levels of Detail); for a certain level of LoD, a predicted value of the attribute information of the three-dimensional points included in that level of LoD is generated using the encoded attribute information included in the LoD of the upper level of that level.

10. The three-dimensional data decoding method according to claim 8, wherein, in the arithmetic decoding, different contexts are used for each bit of the binary data.

11. The three-dimensional data decoding method according to claim 10, wherein, in the arithmetic decoding, the lower the bit position of the binary data, the larger the number of contexts used.

12. The three-dimensional data decoding method according to any one of claims 8 to 11, wherein, in the arithmetic decoding, according to the value of the upper bit of the target bit included in the binary data, a context used in the arithmetic decoding of the target bit is selected.

13. The three-dimensional data decoding method according to any one of claims 8 to 11, wherein, in the multi-valuing, a first value is generated by multi-valuing a first encoding of a fixed number of bits included in the binary data, when the first value is less than a threshold, the first value is determined as the prediction residual, when the first value is equal to or greater than the threshold, a second value is generated by multi-valuing an exponential Golomb code (i.e., a second encoding) included in the binary data, and the first value and the second value are added together to generate the prediction residual, in the arithmetic decoding, different arithmetic decoding methods are used for the first encoding and the second encoding.

14. The three-dimensional data decoding method according to claim 13, wherein, the three-dimensional data decoding method further inverse-quantizes the prediction residual, in the addition, the predicted value and the inverse-quantized prediction residual are added together, the threshold is changed according to the quantization scale in the inverse quantization.

15. A three-dimensional data encoding device is a three-dimensional data encoding device that encodes three-dimensional points with attribute information, wherein, it includes: a processor; and a memory, the processor uses the memory, calculates a predicted value of the attribute information of the three-dimensional points, calculates a difference between the attribute information of the three-dimensional points and the predicted value, i.e., a prediction residual, generates binary data by binarizing the prediction residual, arithmetically encodes the binary data, the binary data includes a prefix part (i.e., prefix part) and a suffix part (i.e., suffix part), in the arithmetic encoding, different contexts are used for the prefix part and the suffix part.

16. A three-dimensional data decoding device is a three-dimensional data decoding device that decodes three-dimensional points with attribute information, wherein, it includes: a processor; and a memory, the processor uses the memory, calculates a predicted value of the attribute information of the three-dimensional points, Arithmetic decoding is performed on the encoded data included in the bitstream to generate binary data, The prediction residual is generated by multi-valuing the binary data, The decoded value of the attribute information of the three-dimensional point is calculated by adding the predicted value and the prediction residual, The binary data includes a prefix part, i.e., the prefix part, and a suffix part, i.e., the suffix part, In the arithmetic decoding, different contexts are used for the prefix part and the suffix part.

Citation Information

Patent Citations

  • Map display device

    WO2014020663A1