Three-dimensional data encoding method, three-dimensional data decoding method, three-dimensional data encoding device, and three-dimensional data decoding device

By selecting the appropriate encoding table, entropy encoding of the N fork tree structure of the three-dimensional points in the three-dimensional data is solved, and effective compression and efficient decoding of the data volume are achieved.

CN111727460BActive Publication Date: 2025-05-06PANASONIC INTELLECTUAL PROPERTY CORP OF AMERICA
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN201980013461.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2018-01-26
Filing Date
2019-01-24
Publication Date
2025-05-06
Estimated Expiration
2039-01-24

AI Technical Summary

Technical Problem

In the prior art, the encoding efficiency of three-dimensional data is low, and it is difficult to effectively compress large-scale three-dimensional point cloud data.

Method used

Using the encoding table selected from multiple encoding tables, the bit string of N fork-tree structure of multiple three-dimensional points contained in the three-dimensional data is entropy-encoding, and the entropy encoding of the bit string is performed using context information.

Benefits of technology

Through this method, the encoding efficiency of three-dimensional data can be significantly improved, the amount of data can be reduced, while maintaining efficient decoding performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN111727460B_ABST
    Figure CN111727460B_ABST
Patent Text Reader

Abstract

The three-dimensional data encoding method uses a coding table selected from multiple coding tables to entropy encode a bit string of an N-ary tree structure (N is an integer greater than 2) representing multiple three-dimensional points contained in the three-dimensional data; the bit string contains N bits of information for each node in the N-ary tree structure; the N bits of information contain N 1-bit information indicating whether there is a three-dimensional point in each of the N child nodes of the corresponding node; in each of the multiple coding tables, a context is set for each bit of the N-bit information; in entropy coding, each bit of the N-bit information is entropy encoded using the context set for the bit in the selected coding table.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to a three-dimensional data encoding method, a three-dimensional data decoding method, a three-dimensional data encoding device, and a three-dimensional data decoding device. Background Art

[0002] In the future, devices and services that make use of 3D data will become more common in large fields such as computer vision, map information, monitoring, infrastructure inspection, or image distribution, which are used for autonomous operation of cars or robots. 3D data is obtained by various methods such as distance sensors such as rangefinders, stereo cameras, or a combination of multiple single-lens reflex cameras.

[0003] As a method of expressing three-dimensional data, there is a method called point cloud, which expresses the shape of a three-dimensional structure through a point group in a three-dimensional space (for example, refer to non-patent document 1). The position and color of the point group are stored in the point cloud. Although point cloud is expected to become the mainstream method of expressing three-dimensional data, the amount of point group data is very large. Therefore, in the accumulation or transmission of three-dimensional data, as with two-dimensional dynamic images (as an example, there are MPEG-4AVC or HEVC standardized by MPEG), it is necessary to compress the data volume through encoding.

[0004] Furthermore, compression of point clouds is partially supported by a public library (PointCloud Library) that performs point cloud association processing.

[0005] Furthermore, there is a known technique for searching for facilities around a vehicle using three-dimensional map data and displaying the facilities (for example, refer to Patent Document 1).

[0006] Prior art literature

[0007] Patent Literature

[0008] Patent Document 1 International Publication No. 2014 / 020663 Summary of the invention

[0009] Problem that the invention aims to solve

[0010] It is desirable to improve coding efficiency in encoding three-dimensional data.

[0011] An object of the present disclosure is to provide a three-dimensional data encoding method, a three-dimensional data decoding method, a three-dimensional data encoding device or a three-dimensional data decoding device that can improve encoding efficiency.

[0012] Means used to solve problems

[0013] A three-dimensional data encoding method in one form of the present disclosure uses a coding table selected from multiple coding tables to entropy encode a bit string of an N (N is an integer greater than 2) fork tree structure representing multiple three-dimensional points contained in three-dimensional data; the bit string contains N bits of information for each node in the N-fork tree structure; the N bits of information contain N 1-bit information indicating whether a three-dimensional point exists in each of the N child nodes of the corresponding node; in each of the multiple coding tables, a context is set for each bit of the N bits of information; in the entropy encoding, each bit of the N bits of information is entropy encoded using the context set for the bit in the selected coding table.

[0014] A three-dimensional data decoding method in one form of the present disclosure uses a coding table selected from multiple coding tables to entropy decode a bit string of an N (N is an integer greater than 2) fork tree structure representing multiple three-dimensional points contained in three-dimensional data; the bit string contains N bits of information for each node in the N-fork tree structure; the N bits of information contain N 1-bit information indicating whether a three-dimensional point exists in each of the N child nodes of the corresponding node; in each of the multiple coding tables, a context is set for each bit of the N bits of information; in the entropy decoding, each bit of the N bits of information is entropy decoded using the context set for the bit in the selected coding table.

[0015] Effects of the Invention

[0016] The present disclosure can provide a three-dimensional data encoding method, a three-dimensional data decoding method, a three-dimensional data encoding device, or a three-dimensional data decoding device that can improve encoding efficiency. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] Figure 1 The structure of the encoded three-dimensional data involved in Embodiment 1 is shown.

[0018] Figure 2 An example of a prediction structure between SPCs belonging to the lowest layer of the GOS involved in the first embodiment is shown.

[0019] Figure 3 An example of an inter-layer prediction structure according to the first embodiment is shown.

[0020] Figure 4 An example of the coding order of the GOS according to the first embodiment is shown.

[0021] Figure 5 An example of the coding order of the GOS according to the first embodiment is shown.

[0022] Figure 6This is a block diagram of the three-dimensional data encoding device involved in embodiment 1.

[0023] Figure 7 This is a flowchart of the encoding process involved in Implementation 1.

[0024] Figure 8 This is a block diagram of the three-dimensional data decoding device involved in Embodiment 1.

[0025] Fig. 9 This is a flowchart of the decoding process involved in Implementation 1.

[0026] Fig.10 An example of meta-information according to the first embodiment is shown.

[0027] Fig.11 A configuration example of a SWLD according to the second embodiment is shown.

[0028] Fig.12 An operation example of the server and the client according to the second embodiment is shown.

[0029] Fig.13 An operation example of the server and the client according to the second embodiment is shown.

[0030] Fig.14 An operation example of the server and the client according to the second embodiment is shown.

[0031] Fig.15 An operation example of the server and the client according to the second embodiment is shown.

[0032] Fig.16 This is a block diagram of a three-dimensional data encoding device according to the second embodiment.

[0033] Fig.17 This is a flowchart of the encoding process involved in Implementation Method 2.

[0034] Fig.18 This is a block diagram of a three-dimensional data decoding device according to the second embodiment.

[0035] Fig.19 This is a flowchart of the decoding process involved in Implementation Method 2.

[0036] Fig. 20 A configuration example of a WLD according to the second embodiment is shown.

[0037] Fig.21 An example of the octree structure of the WLD according to the second embodiment is shown.

[0038] Fig. 22 A configuration example of a SWLD according to the second embodiment is shown.

[0039] Fig.23 An example of the octree structure of the SWLD according to the second embodiment is shown.

[0040] Fig.24 This is a block diagram of a three-dimensional data creation device according to the third embodiment.

[0041] Fig.25 This is a block diagram of a three-dimensional data transmitting device according to Embodiment 3.

[0042] Fig.26 This is a block diagram of a three-dimensional information processing device according to a fourth embodiment.

[0043] Fig. 27 This is a block diagram of a three-dimensional data production device according to the fifth embodiment.

[0044] Fig.28 The configuration of the system according to the sixth embodiment is shown.

[0045] Fig.29 This is a block diagram of a client device according to the sixth embodiment.

[0046] Fig.30 This is a block diagram of the server involved in Implementation Example 6.

[0047] Fig.31 This is a flowchart of the three-dimensional data creation process performed by the client device involved in the sixth embodiment.

[0048] Fig.32 This is a flowchart of the sensor information transmission process performed by the client device according to the sixth embodiment.

[0049] Fig.33 This is a flowchart of the three-dimensional data creation process performed by the server involved in the sixth embodiment.

[0050] Fig.34 This is a flowchart of the three-dimensional map transmission processing performed by the server involved in the sixth embodiment.

[0051] Fig.35 The configuration of a modified example of the system according to the sixth embodiment is shown.

[0052] Fig.36 The configuration of the server and client device according to the sixth embodiment is shown.

[0053] Fig.37 This is a block diagram of a three-dimensional data encoding device involved in embodiment 7.

[0054] Fig.38An example of the prediction residual involved in Embodiment 7 is shown.

[0055] Fig.39 An example of the volume involved in Embodiment 7 is shown.

[0056] Fig.40 An example of octree representation of volume according to the seventh embodiment is shown.

[0057] Fig.41 An example of a bit string of volume involved in Implementation Example 7 is shown.

[0058] Fig.42 An example of octree representation of volume according to the seventh embodiment is shown.

[0059] Fig.43 An example of the volume involved in Embodiment 7 is shown.

[0060] Fig.44 This is a diagram for explaining the intra-frame prediction processing involved in Implementation Example 7.

[0061] Fig.45 This is a diagram used to illustrate the rotation and translation processing involved in embodiment 7.

[0062] Fig.46 An example of the syntax of the RT application flag and RT information involved in the seventh embodiment is shown.

[0063] Fig.47 This is a diagram used to illustrate the inter-frame prediction processing involved in embodiment 7.

[0064] Fig.48 This is a block diagram of a three-dimensional data decoding device according to the seventh embodiment.

[0065] Fig.49 This is a flowchart of a three-dimensional data encoding process performed by the three-dimensional data encoding device according to the seventh embodiment.

[0066] Fig.50 This is a flowchart of a three-dimensional data decoding process performed by the three-dimensional data decoding device according to the seventh embodiment.

[0067] Fig.51 The configuration of the distribution system according to the eighth embodiment is shown.

[0068] Fig.52 An example of the structure of a bit stream encoding a three-dimensional map according to the eighth embodiment is shown.

[0069] Fig.53 This is a diagram used to illustrate the improvement effect of coding efficiency involved in implementation mode 8.

[0070] Fig.54 This is a flowchart of the processing performed by the server involved in Implementation Example 8.

[0071] Fig.55 This is a flowchart of the processing performed by the client involved in Implementation Example 8.

[0072] Fig.56 A syntax example of a submap according to the eighth embodiment is shown.

[0073] Fig.57 The switching process of the coding type involved in Implementation Example 8 is shown in a schematic manner.

[0074] Fig.58 A syntax example of a submap according to the eighth embodiment is shown.

[0075] Fig.59 This is a flowchart of the three-dimensional data encoding process involved in the eighth embodiment.

[0076] Fig.60 This is a flowchart of the three-dimensional data decoding process involved in the eighth embodiment.

[0077] Fig.61 The operation of a modified example of the coding type switching process involved in Implementation 8 is shown in a schematic manner.

[0078] Fig.62 The operation of a modified example of the coding type switching process involved in Implementation 8 is shown in a schematic manner.

[0079] Fig.63 The operation of a modified example of the coding type switching process involved in Implementation 8 is shown in a schematic manner.

[0080] Fig.64 The operation of a modified example of the calculation process of the difference value involved in the eighth embodiment is schematically shown.

[0081] Fig.65 The operation of a modified example of the calculation process of the difference value involved in the eighth embodiment is schematically shown.

[0082] Fig.66 The operation of a modified example of the calculation process of the difference value involved in the eighth embodiment is schematically shown.

[0083] Fig.67 The operation of a modified example of the calculation process of the difference value involved in the eighth embodiment is schematically shown.

[0084] Fig.68 A syntax example of volume involved in Implementation Example 8 is shown.

[0085] Fig.69 This is a diagram showing an example of an important area involved in Implementation Example 9.

[0086] Fig.70 This is a diagram showing an example of an occupancy code involved in Implementation Example 9.

[0087] Fig.71 This is a diagram showing an example of a quadtree structure involved in Implementation Example 9.

[0088] Fig.72 This is a diagram showing an example of occupancy codes and position codes involved in Implementation Example 9.

[0089] Fig.73 This is a diagram showing an example of three-dimensional points obtained by LiDAR according to the ninth embodiment.

[0090] Fig.74 This is a diagram showing an example of an octree structure according to the ninth embodiment.

[0091] Fig.75 This is a diagram showing an example of mixed coding involved in Implementation Example 9.

[0092] Fig.76 This is a diagram for illustrating a method for switching between position coding and occupancy coding according to the ninth embodiment.

[0093] Fig.77 This is a diagram showing an example of a position-coded bit stream according to the ninth embodiment.

[0094] Fig.78 This is a diagram showing an example of a mixed coded bit stream involved in Implementation Example 9.

[0095] Fig.79 This is a diagram showing the tree structure of the occupancy code of the important three-dimensional points involved in Implementation Example 9.

[0096] Fig.80 This is a diagram showing the tree structure of the occupancy code of the non-important three-dimensional points involved in Implementation Example 9.

[0097] Fig.81 This is a diagram showing an example of a mixed coded bit stream involved in Implementation Example 9.

[0098] Fig.82 This is a diagram showing an example of a bit stream including coding mode information involved in Implementation Example 9.

[0099] Fig.83 This is a diagram showing a syntactic example involved in Implementation Method 9.

[0100] Fig.84This is a flowchart of the encoding process involved in Implementation Example 9.

[0101] Fig.85 This is a flowchart of the node encoding processing involved in implementation mode 9.

[0102] Fig.86 This is a flowchart of the decoding process involved in Implementation 9.

[0103] Fig.87 This is a flowchart of the node decoding process involved in Implementation Method 9.

[0104] Fig.88 This is a diagram showing an example of a tree structure related to implementation example 10.

[0105] Fig.89 This is a diagram showing an example of the number of valid leaf nodes that each branch has according to the tenth embodiment.

[0106] Fig.90 This is a diagram showing an application example of the encoding method related to implementation mode 10.

[0107] Fig.91 This is a diagram showing an example of a dense branch area related to embodiment 10.

[0108] Fig.92 This is a diagram showing an example of a dense three-dimensional point group related to embodiment 10.

[0109] Fig.93 This is a diagram showing an example of a sparse three-dimensional point group related to embodiment 10.

[0110] Fig.94 This is a flowchart of the encoding process related to implementation mode 10.

[0111] Fig.95 This is a flowchart of the decoding process related to implementation mode 10.

[0112] Fig.96 This is a flowchart of the encoding process related to implementation mode 10.

[0113] Fig.97 This is a flowchart of the decoding process related to implementation mode 10.

[0114] Fig.98 This is a flowchart of the encoding process related to implementation mode 10.

[0115] Fig.99 This is a flowchart of the decoding process related to implementation mode 10.

[0116] Fig.100 This is a flowchart showing the three-dimensional point separation process according to the tenth embodiment.

[0117] Fig.101 This is a diagram showing a syntax example related to implementation mode 10.

[0118] Fig.102 This is a diagram showing an example of dense branching related to implementation example 10.

[0119] Fig.103 This is a diagram showing an example of sparse branching related to implementation example 10.

[0120] Fig.104 This is a flowchart of the encoding processing of a variation example of implementation mode 10.

[0121] Fig.105 This is a flowchart of the decoding process of a variation of implementation mode 10.

[0122] Fig.106 This is a flowchart of the three-dimensional point separation process according to a variation of the tenth embodiment.

[0123] Fig.107 This is a diagram showing a syntactic example of a variation of implementation example 10.

[0124] Fig.108 This is a flowchart of the encoding process related to implementation mode 10.

[0125] Fig.109 This is a flowchart of the decoding process related to implementation mode 10.

[0126] Fig.110 This is a diagram showing an example of a tree structure related to implementation example 11.

[0127] Fig.111 This is a diagram showing an example of an occupancy code related to implementation example 11.

[0128] Fig.112 This is a diagram schematically showing the operation of the three-dimensional data encoding device related to embodiment 11.

[0129] Fig.113 This is a diagram showing an example of geometric information related to implementation example 11.

[0130] Fig.114 This is a diagram showing an example of selecting a coding table using geometric information related to implementation example 11.

[0131] Fig.115 This is a diagram showing an example of selecting a coding table using structural information related to implementation example 11.

[0132] Fig.116 This is a diagram showing an example of selecting a coding table for usage attribute information related to implementation example 11.

[0133] Fig.117 This is a diagram showing an example of selecting a coding table for usage attribute information related to implementation example 11.

[0134] Fig.118 This is a diagram showing a structural example of a bit stream related to implementation example 11.

[0135] Fig.119 This is a diagram showing an example of a coding table related to implementation example 11.

[0136] Fig.120 This is a diagram showing an example of a coding table related to implementation example 11.

[0137] Fig.121 This is a diagram showing a structural example of a bit stream related to implementation example 11.

[0138] Fig.122 This is a diagram showing an example of a coding table related to implementation example 11.

[0139] Fig.123 This is a diagram showing an example of a coding table related to implementation example 11.

[0140] Fig.124 This is a diagram showing an example of the bit number of the occupancy code related to implementation example 11.

[0141] Fig.125 This is a flowchart of the encoding process using geometric information related to implementation mode 11.

[0142] Fig.126 This is a flowchart of the decoding process using geometric information related to implementation mode 11.

[0143] Fig.127 This is a flowchart of the encoding process of using structural information related to implementation mode 11.

[0144] Fig.128 This is a flowchart of the decoding process of the usage structure information related to implementation mode 11.

[0145] Fig.129 This is a flowchart of the encoding process of usage attribute information related to implementation mode 11.

[0146] Fig.130 This is a flowchart of the decoding process of usage attribute information related to implementation mode 11.

[0147] Fig.131 This is a flowchart of the coding table selection process using geometric information related to implementation mode 11.

[0148] Fig.132This is a flowchart of the coding table selection process using structural information related to implementation mode 11.

[0149] Fig.133 This is a flowchart of the coding table selection process for using attribute information related to implementation mode 11.

[0150] Fig.134 This is a block diagram of a three-dimensional data encoding device related to embodiment 11.

[0151] Fig.135 This is a block diagram of a three-dimensional data decoding device related to embodiment 11. DETAILED DESCRIPTION

[0152] A three-dimensional data encoding method in one form of the present disclosure uses a coding table selected from multiple coding tables to entropy encode a bit string of an N (N is an integer greater than 2) fork tree structure representing multiple three-dimensional points contained in three-dimensional data; the bit string contains N bits of information for each node in the N-fork tree structure; the N bits of information contain N 1-bit information indicating whether a three-dimensional point exists in each of the N child nodes of the corresponding node; in each of the multiple coding tables, a context is set for each bit of the N bits of information; in the entropy encoding, each bit of the N bits of information is entropy encoded using the context set for the bit in the selected coding table.

[0153] Therefore, the three-dimensional data encoding method can improve encoding efficiency by switching the context for each bit.

[0154] For example, in the entropy coding, the coding table to be used may be selected from the plurality of coding tables based on whether or not a three-dimensional point exists in each of a plurality of adjacent nodes adjacent to the target node.

[0155] Therefore, the three-dimensional data encoding method can improve encoding efficiency by switching the encoding table based on whether there is a three-dimensional point in an adjacent node.

[0156] For example, in the entropy coding, the coding table can also be selected based on the configuration style representing the configuration positions of adjacent nodes where three-dimensional points exist among the multiple adjacent nodes; and for the configuration styles among the configuration styles that become the same configuration style through rotation, the same coding table can be selected.

[0157] Therefore, the three-dimensional data encoding method can suppress the increase of the encoding table.

[0158] For example, in the entropy coding, a coding table to be used may be selected from the plurality of coding tables based on the layer to which the target node belongs.

[0159] Therefore, the three-dimensional data encoding method can improve encoding efficiency by switching the encoding table based on the layer to which the object node belongs.

[0160] For example, in the entropy coding, a coding table to be used may be selected from the plurality of coding tables based on the normal vector of the object node.

[0161] Therefore, the three-dimensional data encoding method can improve the encoding efficiency by switching the encoding table based on the normal vector.

[0162] In addition, regarding a three-dimensional data decoding method in one form of the present disclosure, a coding table selected from multiple coding tables is used to entropy decode a bit string of an N (N is an integer greater than 2) fork tree structure representing multiple three-dimensional points contained in the three-dimensional data; the bit string contains N bits of information for each node in the N-fork tree structure; the N bits of information contain N 1-bit information indicating whether there is a three-dimensional point in each of the N child nodes of the corresponding node; in each of the multiple coding tables, a context is set for each bit of the N bits of information; in the entropy decoding, each bit of the N bits of information is entropy decoded using the context set for the bit in the selected coding table.

[0163] Therefore, the three-dimensional data decoding method can improve the encoding efficiency by switching the context for each bit.

[0164] For example, in the entropy decoding, the coding table to be used may be selected from the plurality of coding tables based on whether or not a three-dimensional point exists in each of a plurality of adjacent nodes adjacent to the target node.

[0165] Therefore, the three-dimensional data decoding method can improve the encoding efficiency by switching the encoding table based on whether there is a three-dimensional point in the adjacent node.

[0166] For example, in the entropy decoding, the coding table can also be selected based on the configuration pattern representing the configuration positions of adjacent nodes where three-dimensional points exist among the multiple adjacent nodes; and the same coding table can be selected for the configuration patterns among the configuration patterns that become the same configuration patterns through rotation.

[0167] Therefore, the three-dimensional data decoding method can suppress the increase of the coding table.

[0168] For example, in the entropy decoding, a coding table to be used may be selected from the plurality of coding tables based on the layer to which the target node belongs.

[0169] Therefore, the three-dimensional data decoding method can improve the encoding efficiency by switching the encoding table based on the layer to which the object node belongs.

[0170] For example, in the entropy decoding, a coding table to be used may be selected from the plurality of coding tables based on a normal vector of a node of the object.

[0171] Therefore, the three-dimensional data decoding method can improve the encoding efficiency by switching the encoding table based on the normal vector.

[0172] In addition, a three-dimensional data encoding device in one form related to the present disclosure includes a processor and a memory; the processor uses the memory to perform the following processing: using a coding table selected from a plurality of coding tables, entropy encoding a bit string of an N (N is an integer greater than 2) fork tree structure representing a plurality of three-dimensional points contained in the three-dimensional data; the bit string includes N bits of information for each node in the N-fork tree structure; the N bits of information include N 1-bit information indicating whether a three-dimensional point exists in each of the N child nodes of the corresponding node; in each of the plurality of coding tables, a context is set for each bit of the N bits of information; in the entropy encoding, entropy encoding is performed on each bit of the N bits of information using the context set for the bit in the selected coding table.

[0173] Therefore, the three-dimensional data encoding device can improve encoding efficiency by switching the context for each bit.

[0174] In addition, a three-dimensional data decoding device according to one form of the present disclosure includes a processor and a memory; the processor uses the memory to perform the following processing: entropy decoding is performed on a bit string of an N-ary tree structure (N is an integer greater than 2) representing multiple three-dimensional points contained in the three-dimensional data using a coding table selected from multiple coding tables; the bit string contains N bits of information for each node in the N-ary tree structure; the N bits of information contain N 1-bit information indicating whether there is a three-dimensional point in each of the N child nodes of the corresponding node; in each of the multiple coding tables, a context is set for each bit of the N-bit information; in the entropy decoding, each bit of the N-bit information is entropy decoded using the context set for the bit in the selected coding table.

[0175] Therefore, the three-dimensional data decoding device can improve encoding efficiency by switching the context for each bit.

[0176] In addition, these general or specific forms can be implemented by systems, methods, integrated circuits, computer programs, or computer-readable recording media such as CD-ROMs, and can be implemented by any combination of systems, methods, integrated circuits, computer programs, and recording media.

[0177] The following detailed description of the implementation mode is given with reference to the accompanying drawings. In addition, the implementation modes to be described below are all specific examples of the present disclosure. The numerical values, shapes, materials, constituent elements, configuration positions of constituent elements, connection forms, steps, order of steps, etc. shown in the following implementation modes are all examples, and the main purpose is not to limit the present disclosure. Furthermore, the constituent elements of the following implementation modes that are not recorded in the technical solution showing the highest concept are described as arbitrary constituent elements.

[0178] (Implementation Method 1)

[0179] First, the data structure of encoded three-dimensional data (hereinafter also referred to as encoded data) according to the present embodiment will be described. Figure 1 The structure of the encoded three-dimensional data involved in this embodiment is shown.

[0180] In this embodiment, the three-dimensional space is divided into a space (SPC) equivalent to a picture in the encoding of a dynamic image, and the three-dimensional data is encoded in units of space. The space is further divided into volumes (VLM) equivalent to macroblocks in dynamic image encoding, and prediction and conversion are performed in units of VLM. The volume includes a plurality of voxels (VXL), which are the smallest units corresponding to position coordinates. In addition, prediction means that, similar to the prediction performed in a two-dimensional image, predicted three-dimensional data similar to the processing unit of the processing object is generated with reference to other processing units, and the difference between the predicted three-dimensional data and the processing unit of the processing object is encoded. Furthermore, the prediction includes not only spatial prediction with reference to other prediction units at the same time, but also temporal prediction with reference to prediction units at different times.

[0181] For example, when encoding a three-dimensional space represented by point group data such as a point cloud, a three-dimensional data encoding device (hereinafter also referred to as an encoding device) encodes each point of the point group or a plurality of points contained in a voxel according to the size of the voxel. If the voxel is subdivided, the three-dimensional shape of the point group can be expressed with high accuracy, and if the size of the voxel is increased, the three-dimensional shape of the point group can be roughly expressed.

[0182] In addition, although the following description takes the case where the three-dimensional data is a point cloud as an example, the three-dimensional data is not limited to the point cloud, and can also be three-dimensional data in any form.

[0183] Furthermore, voxels of a hierarchical structure may be used. In this case, in the n-order hierarchy, whether or not a sampling point exists in the n-1-order or lower hierarchy (the lower hierarchy of the n-order hierarchy) may be sequentially indicated. For example, when decoding only the n-order hierarchy, if a sampling point exists in the n-1-order or lower hierarchy, decoding may be performed by assuming that a sampling point exists at the center of a voxel in the n-order hierarchy.

[0184] Furthermore, the encoding device obtains point group data through a distance sensor, a stereo camera, a monocular camera, a gyroscope, or an inertial sensor.

[0185] As with the encoding of moving images, the space is classified into at least one of the following three prediction structures: an intra-frame space (I-SPC) that can be decoded independently, a prediction space (P-SPC) that can only be referenced unidirectionally, and a bidirectional space (B-SPC) that can be referenced bidirectionally. In addition, the space has two types of time information: decoding time and display time.

[0186] And, if Figure 1 As shown, as a processing unit including a plurality of spaces, there is a GOS (Group Of Space) which is a random access unit. Also, as a processing unit including a plurality of GOS, there is a world space (WLD).

[0187] The spatial area occupied by the world space is associated with an absolute position on the earth through GPS or latitude and longitude information. This position information is stored as meta information. In addition, the meta information can be included in the encoded data or transmitted separately from the encoded data.

[0188] Furthermore, within the GOS, all SPCs may be three-dimensionally adjacent, or there may be SPCs that are not three-dimensionally adjacent to other SPCs.

[0189] In addition, the encoding, decoding or referencing of the three-dimensional data included in the processing unit such as GOS, SPC or VLM is also simply referred to as encoding, decoding or referencing the processing unit. The three-dimensional data included in the processing unit includes at least one set of spatial positions such as three-dimensional coordinates and characteristic values ​​such as color information.

[0190] Next, the prediction structure of the SPC in the GOS will be described. Although multiple SPCs in the same GOS or multiple VLMs in the same SPC occupy different spaces, they have the same time information (decoding time and display time).

[0191] Furthermore, in the GOS, the first SPC in the decoding order is the I-SPC. Furthermore, there are two types of GOS, the closed GOS and the open GOS. The closed GOS is a GOS that can decode all the SPCs in the GOS when decoding starts from the first I-SPC. In the open GOS, in the GOS, some SPCs earlier than the display time of the first I-SPC refer to different GOSs and can only be decoded in the GOS.

[0192] In addition, in the case of coded data such as map information, the WLD may be decoded in the reverse direction of the coding order. If there is a dependency between GOS, it is difficult to reproduce the data in the reverse direction. Therefore, in this case, a closed GOS is basically used.

[0193] Furthermore, the GOS has a layer structure in the height direction, and encoding or decoding is performed sequentially starting from the SPC of the bottom layer.

[0194] Figure 2 An example of a prediction structure between SPCs belonging to the lowest layer of the GOS is shown. Figure 3 An example of an inter-layer prediction structure is shown.

[0195] There are more than one I-SPC in the GOS. Although there are objects such as people, animals, cars, bicycles, traffic lights, or buildings that serve as land landmarks in the three-dimensional space, it is particularly effective to encode small-sized objects as I-SPCs. For example, when a three-dimensional data decoding device (hereinafter also referred to as a decoding device) decodes the GOS at a low processing amount or high speed, it only decodes the I-SPC in the GOS.

[0196] Furthermore, the encoding device may switch the encoding interval or the occurrence frequency of the I-SPC according to the density of the objects in the WLD.

[0197] And, in Figure 3 In the configuration shown, the encoding device or decoding device encodes or decodes the plurality of layers sequentially from the lower layer (layer 1). This allows, for example, autonomous vehicles to prioritize data near the ground with a large amount of information.

[0198] In addition, in the coded data used by a drone or the like, coding or decoding may be performed sequentially starting from the SPC of the upper layer in the height direction within the GOS.

[0199] Furthermore, the encoding device or decoding device may encode or decode multiple layers in such a way that the decoding device roughly grasps the GOS and can gradually increase the resolution. For example, the encoding device or decoding device may encode or decode in the order of layers 3, 8, 1, 9, ...

[0200] Next, the corresponding method of the static object and the dynamic object is described.

[0201] In three-dimensional space, there are static objects or scenes such as buildings and roads (hereinafter collectively referred to as static objects), and dynamic objects such as vehicles and people (hereinafter referred to as dynamic objects). Object detection can be performed by extracting feature points from point cloud data or images captured by stereo cameras. Here, an example of a method for encoding dynamic objects is described.

[0202] The first method is a method of encoding without distinguishing between static objects and dynamic objects. The second method is a method of distinguishing between static objects and dynamic objects using identification information.

[0203] For example, GOS is used as the identification unit. In this case, GOS including SPCs constituting static objects and GOS including SPCs constituting dynamic objects are distinguished within the coded data or by identification information stored separately from the coded data.

[0204] Alternatively, SPC is used as the identification unit. In this case, the SPC including only the VLM constituting the static object and the SPC including the VLM constituting the dynamic object are distinguished by the above-mentioned identification information.

[0205] Alternatively, VLM or VXL may be used as the identification unit. In this case, the VLM or VXL including the static object and the VLM or VXL including the dynamic object are distinguished by the above-mentioned identification information.

[0206] Furthermore, the encoding device may encode the dynamic object as one or more VLMs or SPCs, and encode the VLM or SPC including the static object and the SPC including the dynamic object as different GOSs. Furthermore, when the size of the GOS becomes variable according to the size of the dynamic object, the encoding device may store the size of the GOS separately as meta-information.

[0207] Furthermore, the encoding device encodes the static object and the dynamic object independently of each other, and the dynamic object can be overlapped with respect to the world space composed of the static object. In this case, the dynamic object is composed of one or more SPCs, and each SPC corresponds to one or more SPCs constituting the static object overlapped with the SPC. In addition, the dynamic object may not be represented by an SPC, but may be represented by one or more VLMs or VXLs.

[0208] Furthermore, the encoding device may encode static objects and dynamic objects as different streams.

[0209] Furthermore, the encoding device may generate a GOS including one or more SPCs constituting a dynamic object. Furthermore, the encoding device may set the GOS (GOS_M) including the dynamic object and the GOS of the static object corresponding to the spatial region of the GOS_M to be of the same size (occupying the same spatial region). In this way, overlapping processing can be performed in units of GOS.

[0210] The P-SPC or B-SPC constituting the dynamic object may also refer to the SPC included in the encoded different GOS. When the position of the dynamic object changes over time and the same dynamic object is encoded as the GOS at different times, cross-GOS reference is effective from the perspective of compression rate.

[0211] Furthermore, the first method and the second method may be switched according to the purpose of the encoded data. For example, when the encoded three-dimensional data is used as a map, it is desirable to separate the three-dimensional data from the dynamic objects, so the encoding device uses the second method. In addition, when the encoding device encodes the three-dimensional data of an event such as a concert or sports, if it is not necessary to separate the dynamic objects, the first method is used.

[0212] Furthermore, the decoding time and display time of GOS or SPC can be stored in the encoded data or as meta-information. Furthermore, the time information of static objects can all be the same. At this time, the actual decoding time and display time can be determined by the decoding device. Alternatively, different values ​​can be assigned to each GOS or SPC as the decoding time, and the same value can be assigned to all the display times. Moreover, as shown in the decoder mode in dynamic image encoding such as HEVC's HRD (Hypothetical Reference Decoder), the decoder has a buffer of a specified size. As long as the bit stream is read at a specified bit rate according to the decoding time, a model that will not be destroyed and is guaranteed to be decodable can be imported.

[0213] Next, the configuration of the GOS in the world space is described. The coordinates of the three-dimensional space in the world space are represented by three mutually orthogonal coordinate axes (x-axis, y-axis, and z-axis). By setting a prescribed rule in the encoding order of the GOS, spatially adjacent GOS can be encoded continuously in the encoded data. For example, Figure 4 In the example shown, the GOS in the xz plane is continuously encoded. After the encoding of all GOS in an xz plane is completed, the value of the y axis is updated. That is, as the encoding continues, the world space extends in the y axis direction. And the index number of the GOS is set as the encoding order.

[0214] Here, the three-dimensional space of the world space corresponds one-to-one to absolute geographical coordinates such as GPS, latitude and longitude. Alternatively, the three-dimensional space can be represented by a relative position relative to a preset reference position. The directions of the x-axis, y-axis, and z-axis of the three-dimensional space are represented as direction vectors determined based on latitude and longitude, etc., and the direction vectors are stored together with the encoded data as meta information.

[0215] Furthermore, the size of the GOS is set to be fixed, and the encoding device stores the size as meta-information. Furthermore, the size of the GOS can be switched, for example, depending on whether it is in the city, indoors, or outdoors. That is, the size of the GOS can be switched according to the amount or nature of objects that have value as information. Alternatively, the encoding device can appropriately switch the size of the GOS or the interval of the I-SPC in the GOS according to the density of the object, etc. in the same world space. For example, the encoding device sets the size of the GOS to be smaller and the interval of the I-SPC in the GOS to be shorter when the density of the object is higher.

[0216] exist Figure 5 In the example, in the area from the 3rd to the 10th GOS, since the density of objects is high, the GOS is subdivided to achieve random access of fine granularity. And, the 7th to the 10th GOS exist on the back of the 3rd to the 6th GOS, respectively.

[0217] Next, the configuration and operation flow of the three-dimensional data encoding device according to this embodiment will be described. Figure 6 It is a block diagram of the three-dimensional data encoding device 100 involved in this embodiment. Figure 7 : is a flowchart showing an operation example of the three-dimensional data encoding device 100 .

[0218] Figure 6 The three-dimensional data encoding device 100 shown generates encoded three-dimensional data 112 by encoding three-dimensional data 111. The three-dimensional data encoding device 100 includes an acquisition unit 101, an encoding region determination unit 102, a division unit 103, and an encoding unit 104.

[0219] like Figure 7 As shown, first, the acquisition unit 101 acquires three-dimensional data 111 as point group data ( S101 ).

[0220] Next, the coding region determination unit 102 determines a coding target region from the spatial region corresponding to the obtained point cloud data (S102). For example, the coding region determination unit 102 determines a spatial region around the position as the coding target region according to the position of the user or vehicle.

[0221] Next, the division unit 103 divides the point group data contained in the area of ​​the encoding object into each processing unit. Here, the processing unit is the above-mentioned GOS and SPC, etc. And, the area of ​​the encoding object corresponds to the above-mentioned world space, for example. Specifically, the division unit 103 divides the point group data into processing units according to the size of the pre-set GOS, the presence or size of the dynamic object (S103). And, the division unit 103 determines the starting position of the SPC that becomes the beginning in the encoding order in each GOS.

[0222] Next, the encoding unit 104 generates the encoded three-dimensional data 112 by sequentially encoding a plurality of SPCs in each GOS ( S104 ).

[0223] In addition, here, after the area to be coded is divided into GOS and SPC, although an example of coding each GOS is shown, the order of processing is not limited to the above. For example, after the composition of a GOS is determined, the GOS may be coded, and then the composition of the GOS may be determined.

[0224] In this way, the three-dimensional data encoding device 100 generates the encoded three-dimensional data 112 by encoding the three-dimensional data 111. Specifically, the three-dimensional data encoding device 100 divides the three-dimensional data into random access units, that is, into first processing units (GOS) corresponding to three-dimensional coordinates, divides the first processing unit (GOS) into a plurality of second processing units (SPC), and divides the second processing unit (SPC) into a plurality of third processing units (VLM). In addition, the third processing unit (VLM) includes more than one voxel (VXL), and the voxel (VXL) is the smallest unit corresponding to the position information.

[0225] Next, the three-dimensional data encoding device 100 generates the encoded three-dimensional data 112 by encoding each of the plurality of first processing units (GOS). Specifically, the three-dimensional data encoding device 100 encodes each of the plurality of second processing units (SPC) in each first processing unit (GOS). Furthermore, the three-dimensional data encoding device 100 encodes each of the plurality of third processing units (VLM) in each second processing unit (SPC).

[0226] For example, when the first processing unit (GOS) of the processing object is a closed GOS, the three-dimensional data encoding device 100 performs encoding with reference to other second processing units (SPCs) included in the first processing unit (GOS) of the processing object for the second processing unit (SPC) of the processing object included in the first processing unit (GOS) of the processing object. That is, the three-dimensional data encoding device 100 does not refer to the second processing unit (SPC) included in the first processing unit (GOS) different from the first processing unit (GOS) of the processing object.

[0227] Furthermore, when the first processing unit (GOS) of the processing object is an open GOS, the second processing unit (SPC) of the processing object contained in the first processing unit (GOS) of the processing object is encoded with reference to other second processing units (SPCs) contained in the first processing unit (GOS) of the processing object, or a second processing unit (SPC) contained in a first processing unit (GOS) different from the first processing unit (GOS) of the processing object.

[0228] Furthermore, the three-dimensional data encoding device 100 selects one as the type of the second processing unit (SPC) of the processing object from among the first type (I-SPC) that does not refer to other second processing units (SPC), the second type (P-SPC) that refers to one other second processing unit (SPC), and the third type that refers to two other second processing units (SPC), and encodes the second processing unit (SPC) of the processing object according to the selected type.

[0229] Next, the configuration and operation flow of the three-dimensional data decoding device according to the present embodiment will be described. Figure 8 It is a block diagram of the three-dimensional data decoding device 200 according to this embodiment. Fig. 9 : is a flowchart showing an example of the operation of the three-dimensional data decoding device 200 .

[0230] Figure 8 The three-dimensional data decoding device 200 shown generates decoded three-dimensional data 212 by decoding the encoded three-dimensional data 211. Here, the encoded three-dimensional data 211 is, for example, the encoded three-dimensional data 112 generated by the three-dimensional data encoding device 100. The three-dimensional data decoding device 200 includes an acquisition unit 201, a decoding start GOS determination unit 202, a decoding SPC determination unit 203, and a decoding unit 204.

[0231] First, the acquisition unit 201 acquires the encoded three-dimensional data 211 (S201). Next, the decoding start GOS determination unit 202 determines the GOS to be decoded (S202). Specifically, the decoding start GOS determination unit 202 refers to the meta information stored in the encoded three-dimensional data 211 or separately from the encoded three-dimensional data, and determines the GOS including the spatial position, object, or SPC corresponding to the time at which decoding is started as the GOS to be decoded.

[0232] Next, the decoded SPC determination unit 203 determines the type (I, P, B) of the SPC to be decoded in the GOS (S203). For example, the decoded SPC determination unit 203 determines (1) whether to decode only the I-SPC, (2) whether to decode the I-SPC and the P-SPC, or (3) whether to decode all types. In addition, when the type of the SPC to be decoded is predetermined, such as when all SPCs are decoded, this step may not be performed.

[0233] Next, the decoding unit 204 obtains the SPC that is the first in the decoding order (the same as the encoding order) in the GOS, and the address position that starts in the encoded three-dimensional data 211, obtains the encoded data of the first SPC from the address position, and decodes each SPC in sequence from the first SPC (S204). And, the above-mentioned address position is stored in meta information, etc.

[0234] In this way, the three-dimensional data decoding device 200 decodes the decoded three-dimensional data 212. Specifically, the three-dimensional data decoding device 200 generates the decoded three-dimensional data 212 of the first processing unit (GOS) as a random access unit by decoding each of the encoded three-dimensional data 211 of the first processing unit (GOS) corresponding to the three-dimensional coordinates. More specifically, the three-dimensional data decoding device 200 decodes each of the plurality of second processing units (SPC) in each first processing unit (GOS). And, the three-dimensional data decoding device 200 decodes each of the plurality of third processing units (VLM) in each second processing unit (SPC).

[0235] The meta information for random access is described below. This meta information is generated by the three-dimensional data encoding device 100 and included in the encoded three-dimensional data 112 (211).

[0236] In conventional random access of two-dimensional moving images, decoding is started from the head frame of a random access unit near a designated time. However, in the world space, random access to (coordinates or objects, etc.) is also envisioned in addition to time.

[0237] Therefore, in order to realize random access to at least three elements, namely coordinates, objects, and time, a table is prepared in which each element is associated with the index number of the GOS. Furthermore, the index number of the GOS is associated with the address of the I-SPC at the beginning of the GOS. Fig.10 An example of a table included in the meta information is shown. Fig.10 Of all the tables shown, at least one table may be used.

[0238] As an example, random access starting from a coordinate is described below. When accessing the coordinates (x2, y2, z2), the coordinate-GOS table is first referenced, and it can be known that the location with the coordinates (x2, y2, z2) is included in the second GOS. Next, the GOS address table is referenced, and since it can be known that the address of the first I-SPC in the second GOS is addr(2), the decoding unit 204 obtains data from this address and starts decoding.

[0239] In addition, the address may be an address in a logical format or a physical address of a HDD or a memory. Furthermore, information for identifying a file segment may be used instead of an address. For example, a file segment is a unit after segmenting one or more GOSs.

[0240] Furthermore, when the object spans across multiple GOSs, the GOSs to which the multiple objects belong may also be shown in the object GOS table. If the multiple GOSs are closed GOSs, the encoding device and the decoding device can perform encoding or decoding in parallel. In addition, if the multiple GOSs are open GOSs, the compression efficiency can be further improved by having the multiple GOSs refer to each other.

[0241] Examples of objects include people, animals, cars, bicycles, traffic lights, or buildings that serve as landmarks on land, etc. For example, when encoding in world space, the three-dimensional data encoding device 100 extracts feature points unique to the object from a three-dimensional point cloud, etc., detects the object based on the feature points, and can set the detected object as a random access point.

[0242] In this way, the three-dimensional data encoding device 100 generates the first information, which indicates the plurality of first processing units (GOS) and the three-dimensional coordinates corresponding to each of the plurality of first processing units (GOS). And the encoded three-dimensional data 112 (211) includes the first information. And the first information further indicates at least one of the object, time, and data storage destination corresponding to each of the plurality of first processing units (GOS).

[0243] The three-dimensional data decoding device 200 obtains the first information from the encoded three-dimensional data 211 , uses the first information to determine the first processing unit of the encoded three-dimensional data 211 corresponding to the specified three-dimensional coordinates, object or time, and decodes the encoded three-dimensional data 211 .

[0244] Other examples of meta information are described below. In addition to the meta information for random access, the three-dimensional data encoding device 100 can also generate and store the following meta information. Furthermore, the three-dimensional data decoding device 200 can also use this meta information during decoding.

[0245] In the case where three-dimensional data is used as map information, a profile is specified according to the purpose, and information indicating the profile may be included in the meta-information. For example, a profile for urban areas or suburbs is specified, or a profile for flying objects is specified, and the maximum or minimum size of the world space, SPC or VLM is defined respectively. For example, in the profile for urban areas, more detailed information is required than in the suburbs, so the minimum size of VLM is set smaller.

[0246] The meta information may also include a tag value indicating the type of the object. The tag value corresponds to the VLM, SPC, or GOS that constitutes the object. The tag value may be set according to the type of the object, for example, a tag value of "0" indicates a "person", a tag value of "1" indicates a "car", and a tag value of "2" indicates a "traffic light". Alternatively, when the type of the object is difficult to determine or does not need to be determined, a tag value indicating the size, or whether it is a dynamic object or a static object, etc. may be used.

[0247] Furthermore, the meta-information may include information indicating the range of the spatial region occupied by the world space.

[0248] Furthermore, the meta-information may be used as header information common to the entire stream of coded data or to a plurality of SPCs such as an SPC in a GOS, and may store the size of the SPC or VXL.

[0249] Furthermore, the meta-information may include identification information of a distance sensor or a camera used in generating the point cloud, or information indicating the positional accuracy of a point group in the point cloud.

[0250] Also, the meta information may include information showing whether the world space is composed of only static objects or contains dynamic objects.

[0251] Modifications of this embodiment will be described below.

[0252] The encoding device or decoding device may encode or decode two or more different SPCs or GOSs in parallel. The GOSs to be encoded or decoded in parallel can be determined based on meta information indicating the spatial position of the GOSs.

[0253] In the case where three-dimensional data is used as a spatial map for the movement of a vehicle or a flying object, or such a spatial map is generated, the encoding device or decoding device can encode or decode the GOS or SPC contained in the space determined based on GPS, path information, or a zoom factor, etc.

[0254] Furthermore, the decoding device may also start decoding from the space close to its own position or walking path. The encoding device or decoding device may also make the priority of the space far from its own position or walking path lower than the priority of the space close to it, and perform encoding or decoding. Here, lowering the priority means lowering the processing order, lowering the resolution (screening post-processing), or lowering the image quality (improving the coding efficiency. For example, increasing the quantization step size), etc.

[0255] Furthermore, when decoding the coded data which is hierarchically coded in space, the decoding device may decode only the lower layer.

[0256] Furthermore, the decoding device may start decoding from a lower layer according to the zoom ratio or purpose of the map.

[0257] Furthermore, in applications such as estimating the self-position of a car or robot during automatic driving or identifying an object, the encoding device or decoding device may also reduce the resolution of an area outside the area within a specified height from the road surface (the area to be identified) for encoding or decoding.

[0258] Furthermore, the encoding device may also encode the point clouds representing the indoor and outdoor spatial shapes separately, for example, by separating the GOS representing the indoor space (indoor GOS) from the GOS representing the outdoor space (outdoor GOS), so that when the decoding device uses the encoded data, it can select the GOS to be decoded according to the viewpoint position.

[0259] Furthermore, the encoding device may encode the indoor GOS and the outdoor GOS whose coordinates are close to each other by placing them adjacent to each other in the coded stream. For example, the encoding device may correspond the identifiers of the two and store information indicating that the corresponding identifiers are established in the coded stream or in meta-information stored separately. Accordingly, the decoding device can refer to the information in the meta-information to identify the indoor GOS and the outdoor GOS whose coordinates are close to each other.

[0260] Furthermore, the encoding device may switch the size of the GOS or SPC between the indoor GOS and the outdoor GOS. For example, the encoding device sets the size of the GOS to be smaller indoors than outdoors. Furthermore, the encoding device may change the accuracy of extracting feature points from the point cloud or the accuracy of object detection between the indoor GOS and the outdoor GOS.

[0261] Furthermore, the encoding device may add information for the decoding device to distinguish and display dynamic objects from static objects to the encoded data. Accordingly, the decoding device can represent the dynamic objects by combining them with red frames or explanatory texts. In addition, the decoding device may replace the dynamic objects and represent them with only red frames or explanatory texts. Furthermore, the decoding device may represent more detailed object categories. For example, a car may be represented by a red frame, and a person may be represented by a yellow frame.

[0262] Furthermore, the encoding device or decoding device may determine whether to perform encoding or decoding by treating the dynamic object and the static object as different SPCs or GOSs according to the frequency of occurrence of the dynamic object or the ratio of the static object to the dynamic object. For example, when the frequency of occurrence or the ratio of the dynamic object exceeds a threshold, the SPC or GOS in which the dynamic object and the static object are mixed is allowed, and when the frequency of occurrence or the ratio of the dynamic object does not exceed the threshold, the SPC or GOS in which the dynamic object and the static object are mixed is not allowed.

[0263] When the dynamic object is detected from the two-dimensional image information of the camera instead of the point cloud, the encoding device can obtain the information (frame or text, etc.) for identifying the detection result and the object position separately, and encode the information as part of the three-dimensional encoded data. In this case, the decoding device overlaps the auxiliary information (frame or text) representing the dynamic object with the decoding result of the static object.

[0264] Furthermore, the coding device may change the density of VXL or VLM according to the complexity of the shape of the static object, etc. For example, the coding device sets VXL or VLM to be denser when the shape of the static object is more complex. Furthermore, the coding device may determine the quantization step size when quantizing the spatial position or color information according to the density of VXL or VLM. For example, the coding device sets the quantization step size to be smaller when the VXL or VLM is denser.

[0265] As described above, the encoding device or decoding device according to the present embodiment performs spatial encoding or decoding in a spatial unit having coordinate information.

[0266] Furthermore, the encoding device and the decoding device perform encoding or decoding in volume units in space. The volume includes voxels which are the minimum units corresponding to the position information.

[0267] Furthermore, the encoding device and the decoding device establish a table of correspondence between each element of spatial information including coordinates, objects, and time and GOP, or a table of correspondence between each element, so as to establish correspondence between arbitrary elements to perform encoding or decoding. Furthermore, the decoding device determines the coordinates using the value of the selected element, determines the volume, voxel, or space based on the coordinates, and decodes the space including the volume or voxel, or the determined space.

[0268] Then, the encoding device determines a volume, voxel, or space that can be selected by an element by extracting feature points or recognizing an object, and encodes it as a volume, voxel, or space that can be randomly accessed.

[0269] Spaces are divided into three types: I-SPC which can be encoded or decoded as a single space, P-SPC which can be encoded or decoded with reference to any one of the processed spaces, and B-SPC which can be encoded or decoded with reference to any two of the processed spaces.

[0270] One or more volumes correspond to static objects or dynamic objects. Spaces containing static objects and spaces containing dynamic objects are encoded or decoded as different GOSs. That is, SPCs containing static objects and SPCs containing dynamic objects are assigned to different GOSs.

[0271] Dynamic objects are encoded or decoded for each object and correspond to one or more spaces containing only static objects. That is, multiple dynamic objects are encoded separately, and the encoded data of multiple dynamic objects obtained correspond to SPCs containing only static objects.

[0272] The encoding device and the decoding device increase the priority of the I-SPC in the GOS to perform encoding or decoding. For example, the encoding device performs encoding in a manner that reduces the degradation of the I-SPC (after decoding, the original three-dimensional data can be reproduced more faithfully). And, the decoding device, for example, only decodes the I-SPC.

[0273] The encoding device may change the frequency of using the I-SPC according to the density or value (number) of objects in the world space to perform encoding. That is, the encoding device changes the frequency of selecting the I-SPC according to the number or density of objects included in the three-dimensional data. For example, the encoding device increases the frequency of using the I space when the density of objects in the world space is greater.

[0274] Furthermore, the encoding device sets the random access point in units of GOS, and stores information indicating the spatial area corresponding to the GOS in the header information.

[0275] The coding device, for example, adopts a default value as the spatial size of the GOS. In addition, the coding device may also change the size of the GOS according to the value (number) or density of the objects or dynamic objects. For example, the coding device sets the spatial size of the GOS to be smaller when the objects or dynamic objects are denser or more in number.

[0276] Furthermore, the space or volume includes a group of feature points derived from information obtained by sensors such as a depth sensor, a gyroscope, or a camera. The coordinates of the feature points are set to the center position of the voxel. Furthermore, by subdividing the voxels, the position information can be highly accurate.

[0277] The feature point group is derived using a plurality of pictures. The plurality of pictures have at least the following two types of time information: actual time information and the same time information in the plurality of pictures corresponding to the space (for example, encoding time for rate control, etc.).

[0278] Furthermore, encoding or decoding is performed in units of GOS including one or more spaces.

[0279] The encoding device and the decoding device refer to the space in the processed GOS to predict the P space or the B space in the GOS to be processed.

[0280] Alternatively, the encoding device and the decoding device predict the P space or B space in the GOS to be processed by using the processed space in the GOS to be processed without referring to different GOSs.

[0281] Furthermore, the encoding device and the decoding device transmit or receive the encoded stream in units of a world space including one or more GOSs.

[0282] Furthermore, the GOS has a layer structure in at least one direction in the world space, and the encoding device and the decoding device perform encoding or decoding from the lower layer. For example, the GOS that can be randomly accessed belongs to the lowest layer. The GOS belonging to the upper layer only refers to the GOS belonging to the layer below the same layer. That is, the GOS is spatially divided in a predetermined direction, including a plurality of layers each having one or more SPCs. The encoding device and the decoding device perform encoding or decoding for each SPC by referring to the SPC included in the layer that is the same layer as the SPC or the layer below the SPC.

[0283] Furthermore, the encoding device and the decoding device continuously encode or decode the GOS in the world space unit including the plurality of GOS. The encoding device and the decoding device write or read information indicating the order (direction) of encoding or decoding as metadata. That is, the encoded data includes information indicating the encoding order of the plurality of GOS.

[0284] Furthermore, the encoding device and the decoding device encode or decode two or more different spaces or GOS in parallel.

[0285] Furthermore, the encoding device and the decoding device encode or decode the space information (coordinates, size, etc.) of the space or the GOS.

[0286] Furthermore, the encoding device and the decoding device encode or decode the space or GOS included in the specific space determined based on external information related to the own position and / or area size, such as GPS, path information, or magnification.

[0287] The encoding device or decoding device performs encoding or decoding by giving a lower priority to a space far from the own position than to a space close to the own position.

[0288] The encoding device sets a direction in the world space according to the magnification or purpose, and encodes the GOS having a layer structure in the direction. And the decoding device preferentially decodes the GOS having a layer structure in the direction of the world space set according to the magnification or purpose from the lower layer.

[0289] The encoding device changes the feature point extraction, object recognition accuracy, or space area size contained in the indoor and outdoor spaces. However, the encoding device and the decoding device encode or decode the indoor GOS and the outdoor GOS with close coordinates adjacent to each other in the world space, and also encode or decode these identifiers in correspondence.

[0290] (Implementation Method 2)

[0291] When point cloud coded data is used in actual devices or services, it is desirable to transmit and receive required information according to the application in order to reduce network bandwidth. However, such a function does not exist in the existing 3D data coding structure, and therefore there is no corresponding coding method.

[0292] What will be described in this embodiment is a three-dimensional data encoding method and a three-dimensional data encoding device for providing the function of sending and receiving required information according to the purpose in the encoded data of a three-dimensional point cloud, as well as a three-dimensional data decoding method and a three-dimensional data decoding device for decoding the encoded data.

[0293] A voxel (VXL) having a certain feature value or more is defined as a feature voxel (FVXL), and a world space (WLD) formed by the FVXL is defined as a sparse world space (SWLD). Fig.11The sparse world space and the configuration example of the world space are shown. SWLD includes: FGOS, which is a GOS composed of FVXL; FSPC, which is an SPC composed of FVXL; and FVLM, which is a VLM composed of FVXL. The data structure and prediction structure of FGOS, FSPC and FVLM can be the same as those of GOS, SPC and VLM.

[0294] The feature quantity refers to a feature quantity that expresses the three-dimensional position information of VXL or the visible light information of the position of VXL, and in particular, a feature quantity that can detect more corners and edges of a three-dimensional object. Specifically, the feature quantity is a three-dimensional feature quantity or a visible light feature quantity described below, but other than this, any feature quantity can be used as long as it expresses the position, brightness, or color information of VXL.

[0295] As the three-dimensional feature, a SHOT feature (Signature of Histograms of Orientations), a PFH feature (Point Feature Histograms), or a PPF feature (Point Pair Feature) is used.

[0296] The SHOT feature is obtained by dividing the area around the VXL, calculating the inner product between the reference point and the normal vector of the divided area, and converting it into a histogram. The SHOT feature has a high dimension and high feature expression.

[0297] The PFH feature is obtained by selecting a plurality of two-point groups near VXL, calculating the normal vector etc. from the two points, and histogramming them. Since the PFH feature is a histogram feature, it is robust to small amounts of interference and has a high feature expression.

[0298] The PPF feature is a feature calculated according to the VXL of two points using a normal vector, etc. Since all VXLs are used in this PPF feature, it is robust to occlusion.

[0299] Furthermore, as feature quantities of visible light, SIFT (Scale-Invariant Feature Transform), SURF (Speeded Up Robust Features), or HOG (Histogram of Oriented Gradients) that utilize information such as brightness gradient information of an image can be used.

[0300] SWLD is generated by calculating the above-mentioned feature amount from each VXL of WLD and extracting FVXL. Here, SWLD may be updated every time WLD is updated, or may be updated regularly after a certain period of time has passed regardless of the update timing of WLD.

[0301] SWLD can be generated for each feature quantity. For example, as shown in SWLD1 based on SHOT feature quantity and SWLD2 based on SIFT feature quantity, SWLD can be generated for each feature quantity and used according to the purpose. In addition, the feature quantity of each FVXL calculated can be stored as feature quantity information in each FVXL.

[0302] Next, a method of using the sparse world space (SWLD) will be described. Since the SWLD only includes feature voxels (FVXL), the data size is generally smaller than that of the WLD including all VXLs.

[0303] In applications that use feature quantities to achieve a certain purpose, by using SWLD information instead of WLD, it is possible to reduce the time required to read from the hard disk, and to reduce the bandwidth and transmission time during network transmission. For example, as map information, WLD and SWLD are stored in the server in advance, and the map information to be sent is switched to WLD or SWLD according to the needs of the client, thereby reducing the network bandwidth and transmission time. A specific example is shown below.

[0304] Fig.12 as well as Fig.13 The following shows the use cases of SWLD and WLD. Fig.12 As shown, when the client 1 as a vehicle-mounted device needs map information for the purpose of determining its own position, the client 1 sends a request for obtaining map data for estimating its own position to the server (S301). The server sends SWLD to the client 1 according to the acquisition request (S302). The client 1 uses the received SWLD to determine its own position (S303). At this time, the client 1 obtains the VXL information of the surrounding area of ​​the client 1 through various methods such as a distance sensor such as a rangefinder, a stereo camera, or a combination of multiple monocular cameras, and estimates its own position information based on the obtained VXL information and SWLD. Here, the own position information includes the three-dimensional position information and orientation of the client 1.

[0305] like Fig.13As shown, when the client 2 as a vehicle-mounted device needs map information for the purpose of drawing a three-dimensional map or the like, the client 2 sends a request for obtaining map data for drawing a map to the server (S311). The server sends a WLD to the client 2 in accordance with the acquisition request (S312). The client 2 uses the received WLD to draw a map (S313). At this time, the client 2 uses, for example, an image taken by itself with a visible light camera and the WLD obtained from the server to create a conceptual image, and draws the created image on a screen such as a car navigation system.

[0306] As described above, the server sends SWLD to the client when the feature amount of each VXL is mainly needed, such as estimating the own position, and sends WLD to the client when detailed VXL information is needed, such as drawing a map. This makes it possible to efficiently send and receive map data.

[0307] In addition, the client can determine which one of SWLD and WLD it needs and request the server to send SWLD or WLD. And the server can determine which one of SWLD and WLD should be sent according to the status of the client or the network.

[0308] Next, a method of switching the transmission and reception between the sparse world space (SWLD) and the world space (WLD) will be described.

[0309] The reception of WLD or SWLD can be switched according to the network bandwidth. Fig.14 An example of operation in this case is shown. For example, when a low-speed network that can use the network bandwidth in an LTE (Long Term Evolution) environment is used, when the client accesses the server via the low-speed network (S321), the SWLD as map information is obtained from the server (S322). In addition, when a high-speed network with sufficient network bandwidth is used in a WiFi environment, the client accesses the server via the high-speed network (S323) and obtains the WLD from the server (S324). Accordingly, the client can obtain appropriate map information according to the network bandwidth of the client.

[0310] Specifically, the client receives SWLD via LTE outdoors, and acquires WLD via WiFi when entering a facility or the like indoors. This allows the client to acquire more detailed map information indoors.

[0311] In this way, the client can request WLD or SWLD from the server according to the frequency band of the network it uses. Alternatively, the client can send information showing the frequency band of the network it uses to the server, and the server can send appropriate data (WLD or SWLD) to the client according to the information. Alternatively, the server can determine the network bandwidth of the client and send appropriate data (WLD or SWLD) to the client.

[0312] Furthermore, the reception of WLD or SWLD can be switched according to the moving speed. Fig.15 An example of operation in this case is shown. For example, when the client moves at high speed (S331), the client receives SWLD from the server (S332). In addition, when the client moves at low speed (S333), the client receives WLD from the server (S334). Accordingly, the client can suppress the network bandwidth and obtain map information according to the speed. Specifically, when the client is driving on a highway, by receiving SWLD with a small amount of data, the map information can be updated at a roughly appropriate speed. In addition, when the client is driving on a general road, by receiving WLD, more detailed map information can be obtained.

[0313] In this way, the client can request WLD or SWLD from the server according to its own moving speed. Alternatively, the client can send information showing its own moving speed to the server, and the server can send appropriate data (WLD or SWLD) to the client according to the information. Alternatively, the server can determine the moving speed of the client and send appropriate data (WLD or SWLD) to the client.

[0314] Furthermore, the client may first obtain the SWLD from the server, and then obtain the WLD of the important area. For example, when the client obtains map data, it first obtains the general map information with the SWLD, selects the area where the features such as buildings, signs, or people appear more, and then obtains the WLD of the selected area. In this way, the client can suppress the amount of data received from the server and obtain the detailed information of the required area.

[0315] Furthermore, the server may create SWLDs for each object based on the WLD, and the client may receive them separately according to the purpose. In this way, the network bandwidth can be suppressed. For example, the server pre-identifies a person or a car from the WLD, and creates a SWLD for the person and a SWLD for the car. When the client wants to obtain information about the people around it, it receives the SWLD for the person, and when it wants to obtain information about the car, it receives the SWLD for the car. Furthermore, the type of such SWLD can be distinguished based on information (flag or type, etc.) attached to the header, etc.

[0316] Next, the configuration and operation flow of the three-dimensional data encoding device (for example, a server) according to the present embodiment will be described. Fig.16 It is a block diagram of the three-dimensional data encoding device 400 involved in this embodiment. Fig.17 This is a flowchart of the three-dimensional data encoding process performed by the three-dimensional data encoding device 400.

[0317] Fig.16 The three-dimensional data encoding device 400 shown in the figure generates encoded three-dimensional data 413 and 414 as encoded streams by encoding input three-dimensional data 411. Here, the encoded three-dimensional data 413 is encoded three-dimensional data corresponding to WLD, and the encoded three-dimensional data 414 is encoded three-dimensional data corresponding to SWLD. The three-dimensional data encoding device 400 includes: an acquisition unit 401, a coding area determination unit 402, a SWLD extraction unit 403, a WLD encoding unit 404, and a SWLD encoding unit 405.

[0318] like Fig.17 As shown, first, the acquisition unit 401 acquires input three-dimensional data 411 which is point group data in a three-dimensional space (S401).

[0319] Next, the encoding region determination unit 402 determines the encoding target spatial region based on the spatial region where the point cloud data exists ( S402 ).

[0320] Next, the SWLD extraction unit 403 defines the spatial region of the encoding object as WLD, and calculates the feature value from each VXL included in the WLD. In addition, the SWLD extraction unit 403 extracts the VXL whose feature value is greater than a predetermined threshold value, defines the extracted VXL as FVXL, and adds the FVXL to the SWLD to generate the extracted three-dimensional data 412 (S403). That is, the extracted three-dimensional data 412 whose feature value is greater than the threshold value is extracted from the input three-dimensional data 411.

[0321] Next, the WLD encoding unit 404 generates encoded three-dimensional data 413 corresponding to the WLD by encoding the input three-dimensional data 411 corresponding to the WLD (S404). At this time, the WLD encoding unit 404 adds information for distinguishing that the encoded three-dimensional data 413 is a stream containing the WLD to the header of the encoded three-dimensional data 413.

[0322] Then, the SWLD encoding unit 405 generates encoded three-dimensional data 414 corresponding to the SWLD by encoding the extracted three-dimensional data 412 corresponding to the SWLD (S405). At this time, the SWLD encoding unit 405 adds information for distinguishing that the encoded three-dimensional data 414 is a stream containing the SWLD to the header of the encoded three-dimensional data 414.

[0323] Furthermore, the processing order of the process of generating the encoded three-dimensional data 413 and the process of generating the encoded three-dimensional data 414 may be reversed from the above. Furthermore, part or all of the above-mentioned processes may be executed in parallel.

[0324] Information assigned to the header of the encoded three-dimensional data 413 and 414 is defined as a parameter such as "world_type". When world_type = 0, it means that the stream contains WLD, and when world_type = 1, it means that the stream contains SWLD. In the case of defining more other categories, the assigned value can be increased, such as world_type = 2. In addition, a specific flag can be included in one of the encoded three-dimensional data 413 and 414. For example, the encoded three-dimensional data 414 can be assigned a flag indicating that the stream contains SWLD. In this case, the decoding device can determine whether it is a stream containing WLD or a stream containing SWLD based on the presence or absence of the flag.

[0325] Furthermore, the encoding method used by the WLD encoding unit 404 when encoding the WLD may be different from the encoding method used by the SWLD encoding unit 405 when encoding the SWLD.

[0326] For example, since data is decimated in SWLD, the correlation with surrounding data may be lower than that in WLD. Therefore, in the encoding method for SWLD, inter-frame prediction is prioritized over intra-frame prediction and inter-frame prediction in comparison with the encoding method for WLD.

[0327] Furthermore, the encoding method used for SWLD and the encoding method used for WLD may have different methods of expressing the three-dimensional position. For example, the three-dimensional position of FVXL may be expressed by three-dimensional coordinates in FWLD, and the three-dimensional position may be expressed by an octree described later in WLD, or vice versa.

[0328] Furthermore, the SWLD coding unit 405 performs coding in such a manner that the data size of the coded three-dimensional data 414 of SWLD is smaller than the data size of the coded three-dimensional data 413 of WLD. For example, as described above, the correlation between data in SWLD may be reduced compared to that in WLD. Accordingly, the coding efficiency is reduced, and the data size of the coded three-dimensional data 414 may be larger than the data size of the coded three-dimensional data 413 of WLD. Therefore, when the data size of the obtained coded three-dimensional data 414 is larger than the data size of the coded three-dimensional data 413 of WLD, the SWLD coding unit 405 performs re-coding to generate the coded three-dimensional data 414 with a reduced data size again.

[0329] For example, the SWLD extraction unit 403 generates the extracted three-dimensional data 412 again with the number of extracted feature points reduced, and the SWLD encoding unit 405 encodes the extracted three-dimensional data 412. Alternatively, the degree of quantization in the SWLD encoding unit 405 may be made coarser. For example, in the octree structure described later, the degree of quantization may be made coarser by rounding the data of the lowest layer.

[0330] Furthermore, when the data size of the SWLD encoded three-dimensional data 414 cannot be made smaller than the data size of the WLD encoded three-dimensional data 413, the SWLD encoding unit 405 may not generate the SWLD encoded three-dimensional data 414. Alternatively, the WLD encoded three-dimensional data 413 may be copied to the SWLD encoded three-dimensional data 414. That is, the WLD encoded three-dimensional data 413 may be used directly as the SWLD encoded three-dimensional data 414.

[0331] Next, the configuration and operation flow of the three-dimensional data decoding device (eg, client) according to the present embodiment will be described. Fig.18 It is a block diagram of a three-dimensional data decoding device 500 according to this embodiment. Fig.19 3D data decoding processing performed by the 3D data decoding apparatus 500 is shown in FIG.

[0332] Fig.18 The three-dimensional data decoding device 500 shown generates decoded three-dimensional data 512 or 513 by decoding the encoded three-dimensional data 511. Here, the encoded three-dimensional data 511 is, for example, the encoded three-dimensional data 413 or 414 generated by the three-dimensional data encoding device 400.

[0333] The three-dimensional data decoding device 500 includes an acquisition unit 501 , a header analysis unit 502 , a WLD decoding unit 503 , and a SWLD decoding unit 504 .

[0334] like Fig.19 As shown, first, the acquisition unit 501 acquires the encoded three-dimensional data 511 (S501). Then, the header analysis unit 502 analyzes the header of the encoded three-dimensional data 511 to determine whether the encoded three-dimensional data 511 is a stream containing WLD or a stream containing SWLD (S502). For example, the determination is made by referring to the above-mentioned world_type parameter.

[0335] When the coded three-dimensional data 511 is a stream including WLD ("Yes" in S503), the WLD decoding unit 503 decodes the coded three-dimensional data 511 to generate decoded three-dimensional data 512 of WLD (S504). In addition, when the coded three-dimensional data 511 is a stream including SWLD ("No" in S503), the SWLD decoding unit 504 decodes the coded three-dimensional data 511 to generate decoded three-dimensional data 513 of SWLD (S505).

[0336] Also, similar to the encoding device, the decoding method used by the WLD decoding unit 503 when decoding the WLD may be different from the decoding method used by the SWLD decoding unit 504 when decoding the SWLD. For example, in the decoding method for SWLD, priority may be given to the inter-frame prediction of intra-frame prediction and inter-frame prediction over the decoding method for WLD.

[0337] Furthermore, the decoding method for SWLD and the decoding method for WLD may have different representation methods for the three-dimensional position. For example, the three-dimensional position of FVXL may be represented by three-dimensional coordinates in SWLD and by an octree described later in WLD, or vice versa.

[0338] Next, octree expression as a method of expressing three-dimensional positions will be described. VXL data included in three-dimensional data is converted into an octree structure and encoded. Fig. 20 An example of a VXL of a WLD is shown. Fig.21 Shows Fig. 20 The octree structure of WLD is shown in Figure 1. Fig. 20 In the example shown, there are three VXLs 1 to 3 that are VXLs (hereinafter, effective VXLs) that include point groups. Fig.21 As shown, the octree structure consists of nodes and leaf nodes. Each node has a maximum of 8 nodes or leaf nodes. Each leaf node has VXL information. Here, Fig.21 Among the leaf nodes shown, leaf nodes 1, 2, and 3 represent Fig. 20 VXL1, VXL2, VXL3 shown.

[0339] Specifically, each node and leaf node corresponds to a three-dimensional position. Fig. 20 The block corresponding to node 1 is divided into 8 blocks, and among the 8 blocks, the block including the valid VXL is set as a node, and the other blocks are set as leaf nodes. The block corresponding to the node is further divided into 8 nodes or leaf nodes, and this process is repeated the same number of times as the number of levels in the tree structure. And all the blocks at the bottom are set as leaf nodes.

[0340] and, Fig. 22 Shows from Fig. 20 An example of SWLD generated by WLD is shown. Fig. 20 The feature extraction results of VXL1 and VXL2 shown are determined to be FVXL1 and FVXL2 and are included in SWLD. In addition, VXL3 is not determined to be FVXL and is therefore not included in SWLD. Fig.23 Shows Fig. 22 The octree structure of SWLD is shown in Figure 1. Fig.23 In the octree structure shown, Fig.21 The leaf node 3 shown, which is equivalent to VXL3, is deleted. Fig.21 The node 3 shown has no valid VXL and is changed to a leaf node. In this way, generally speaking, the number of leaf nodes of SWLD is smaller than that of WLD, and the encoded three-dimensional data of SWLD is also smaller than that of WLD.

[0341] Modifications of this embodiment will be described below.

[0342] For example, when a client such as a vehicle-mounted device estimates its own position, it receives SWLD from a server, uses SWLD to estimate its own position, and performs obstacle detection. It then uses various methods such as distance sensors such as rangefinders, stereo cameras, or a combination of multiple monocular cameras to perform obstacle detection based on the three-dimensional information of the surrounding area obtained by itself.

[0343] In general, it is difficult to include VXL data of a flat area in SWLD. For this reason, the server maintains a subsampled world space (SubWLD) that is a downsampled version of WLD for detecting stationary obstacles, and can send SWLD and SubWLD to the client. This can suppress network bandwidth while enabling the client to estimate its own position and detect obstacles.

[0344] Furthermore, when the client quickly depicts three-dimensional map data, it is convenient if the map information is in a grid structure. Therefore, the server can generate a grid based on the WLD and store it in advance as a grid world space (MWLD). For example, when the client needs to perform a rough three-dimensional depiction, the MWLD is received, and when a detailed three-dimensional depiction is required, the WLD is received. In this way, the network bandwidth can be suppressed.

[0345] Furthermore, although the server sets the VXL with a feature value above the threshold value as FVXL from each VXL, FVXL can also be calculated by different methods. For example, if the server determines that VXL, VLM, SPC, or GOS constituting a signal or intersection is required for self-position estimation, driving assistance, or automatic driving, it can be included in SWLD as FVXL, FVLM, FSPC, FGOS. Furthermore, the above judgment can be performed manually. In addition, FVXL obtained by the above method can be added to FVXL set based on the feature value. That is, the SWLD extraction unit 403 can further extract data corresponding to an object with predetermined attributes from the input three-dimensional data 411 as extracted three-dimensional data 412.

[0346] Furthermore, labels different from feature quantities may be assigned to situations that require these purposes. The server may separately maintain FVXL required for self-position estimation such as signals or intersections, driving assistance, or autonomous driving as an upper layer of SWLD (e.g., lane world space).

[0347] Furthermore, the server may also attach attributes to the VXL in the WLD in random access units or specified units. Attributes include, for example, information indicating whether the location is required or not required for estimating the location, or information indicating whether traffic information such as signals or intersections is important. Furthermore, attributes may also include the correspondence between lane information (GDF: Geographic Data Files, etc.) and features (intersections or roads, etc.).

[0348] Furthermore, as a method of updating WLD or SWLD, the following method can be adopted.

[0349] Update information showing changes in people, construction, or street trees (trajectory orientation) is uploaded to the server as a point group or metadata. The server updates the WLD based on the upload, and then updates the SWLD using the updated WLD.

[0350] Furthermore, when the client detects a mismatch between the three-dimensional information generated by itself and the three-dimensional information received from the server when estimating its own position, the client can send the three-dimensional information generated by itself to the server together with the update notification. In this case, the server updates the SWLD using the WLD. If the SWLD is not updated, the server determines that the WLD itself is old.

[0351] Furthermore, although information for distinguishing between WLD and SWLD is added as the header information of the coded stream, when there are multiple world spaces such as a grid world space or a lane world space, information for distinguishing them may be added to the header information. Furthermore, when there are multiple SWLDs with different feature quantities, information for distinguishing them may also be added to the header information.

[0352] Furthermore, although SWLD is composed of FVXL, it may also include VXL that is not determined to be FVXL. For example, SWLD may include adjacent VXL used when calculating the feature quantity of FVXL. Accordingly, even if each FVXL of SWLD does not have feature quantity information attached, the client can calculate the feature quantity of FVXL when receiving SWLD. In addition, at this time, SWLD may include information for distinguishing whether each VXL is FVXL or VXL.

[0353] As described above, the three-dimensional data encoding device 400 extracts extracted three-dimensional data 412 (second three-dimensional data) whose feature value is above a threshold value from the input three-dimensional data 411 (first three-dimensional data), and generates encoded three-dimensional data 414 (first encoded three-dimensional data) by encoding the extracted three-dimensional data 412.

[0354] According to this, the three-dimensional data encoding device 400 generates the encoded three-dimensional data 414 obtained by encoding the data whose feature quantity is greater than the threshold value. In this way, the amount of data can be reduced compared to the case where the input three-dimensional data 411 is directly encoded. Therefore, the three-dimensional data encoding device 400 can reduce the amount of data during transmission.

[0355] Furthermore, the three-dimensional data encoding device 400 further encodes the input three-dimensional data 411 to generate encoded three-dimensional data 413 (second encoded three-dimensional data).

[0356] According to this, the three-dimensional data encoding device 400 can selectively transmit the encoded three-dimensional data 413 and the encoded three-dimensional data 414 according to the purpose of use, for example.

[0357] Furthermore, the extracted three-dimensional data 412 is encoded by a first encoding method, and the input three-dimensional data 411 is encoded by a second encoding method different from the first encoding method.

[0358] Accordingly, the three-dimensional data encoding device 400 can adopt appropriate encoding methods for the input three-dimensional data 411 and the extracted three-dimensional data 412 respectively.

[0359] Furthermore, in the first encoding method, compared with the second encoding method, priority is given to the inter-frame prediction between the intra-frame prediction and the inter-frame prediction.

[0360] According to this, the three-dimensional data encoding device 400 can increase the priority of inter-frame prediction for the extracted three-dimensional data 412 where the correlation between adjacent data is likely to be low.

[0361] Furthermore, the first encoding method and the second encoding method have different methods of expressing the three-dimensional position. For example, in the second encoding method, the three-dimensional position is expressed by an octree, while in the first encoding method, the three-dimensional position is expressed by three-dimensional coordinates.

[0362] According to this, the three-dimensional data encoding device 400 can adopt a more appropriate three-dimensional position expression method for three-dimensional data having different data numbers (number of VXLs or FVXLs).

[0363] Furthermore, at least one of the coded three-dimensional data 413 and 414 includes an identifier indicating whether the coded three-dimensional data is coded three-dimensional data obtained by encoding the input three-dimensional data 411 or coded three-dimensional data obtained by encoding a part of the input three-dimensional data 411. That is, the identifier indicates whether the coded three-dimensional data is the coded three-dimensional data 413 of the WLD or the coded three-dimensional data 414 of the SWLD.

[0364] Based on this, the decoding device can easily determine whether the acquired encoded three-dimensional data is the encoded three-dimensional data 413 or the encoded three-dimensional data 414.

[0365] Furthermore, the three-dimensional data encoding device 400 encodes the extracted three-dimensional data 412 in such a manner that the data amount of the encoded three-dimensional data 414 is smaller than the data amount of the encoded three-dimensional data 413 .

[0366] According to this, the three-dimensional data encoding device 400 can make the data amount of the encoded three-dimensional data 414 smaller than the data amount of the encoded three-dimensional data 413 .

[0367] Furthermore, the three-dimensional data encoding device 400 further extracts data corresponding to an object having a predetermined attribute from the input three-dimensional data 411 as extracted three-dimensional data 412. For example, the object having a predetermined attribute is an object required for self-position estimation, driving assistance, or automatic driving, such as a signal or an intersection.

[0368] Accordingly, the three-dimensional data encoding device 400 can generate the encoded three-dimensional data 414 including the data required by the decoding device.

[0369] Furthermore, the three-dimensional data encoding device 400 (server) further transmits one of the encoded three-dimensional data 413 and 414 to the client according to the status of the client.

[0370] Accordingly, the three-dimensional data encoding device 400 can send appropriate data according to the status of the client.

[0371] Furthermore, the status of the client includes the communication status of the client (eg, network bandwidth) or the moving speed of the client.

[0372] Furthermore, the three-dimensional data encoding device 400 further transmits one of the encoded three-dimensional data 413 and 414 to the client according to the request of the client.

[0373] Thereby, the three-dimensional data encoding device 400 can send appropriate data according to the request of the client.

[0374] Furthermore, the three-dimensional data decoding device 500 according to the present embodiment decodes the encoded three-dimensional data 413 or 414 generated by the three-dimensional data encoding device 400 .

[0375] That is, the three-dimensional data decoding device 500 decodes the encoded three-dimensional data 414 obtained by encoding the extracted three-dimensional data 412 whose feature amount extracted from the input three-dimensional data 411 is greater than the threshold value by the first decoding method. Furthermore, the three-dimensional data decoding device 500 decodes the encoded three-dimensional data 413 obtained by encoding the input three-dimensional data 411 by using a second decoding method different from the first decoding method.

[0376] According to this, the three-dimensional data decoding device 500 can selectively receive the encoded three-dimensional data 414 and the encoded three-dimensional data 413 obtained by encoding data having a feature value greater than a threshold value, for example, according to the purpose of use. According to this, the three-dimensional data decoding device 500 can reduce the amount of data during transmission. Furthermore, the three-dimensional data decoding device 500 can adopt appropriate decoding methods for the input three-dimensional data 411 and the extracted three-dimensional data 412, respectively.

[0377] Furthermore, in the first decoding method, compared with the second decoding method, priority is given to the inter prediction between the intra prediction and the inter prediction.

[0378] According to this, the three-dimensional data decoding apparatus 500 can increase the priority of inter-frame prediction for extracting three-dimensional data in which the correlation between adjacent data is likely to be low.

[0379] Furthermore, the first decoding method and the second decoding method use different methods for expressing the three-dimensional position. For example, the second decoding method expresses the three-dimensional position by an octree, while the first decoding method expresses the three-dimensional position by three-dimensional coordinates.

[0380] According to this, the three-dimensional data decoding apparatus 500 can adopt a more appropriate three-dimensional position expression method for three-dimensional data having different data numbers (number of VXLs or FVXLs).

[0381] Furthermore, at least one of the encoded three-dimensional data 413 and 414 includes an identifier indicating whether the encoded three-dimensional data is encoded three-dimensional data obtained by encoding the input three-dimensional data 411 or encoded three-dimensional data obtained by encoding a part of the input three-dimensional data 411. The three-dimensional data decoding device 500 identifies the encoded three-dimensional data 413 and 414 by referring to the identifier.

[0382] Based on this, the three-dimensional data decoding device 500 can easily determine whether the obtained encoded three-dimensional data is the encoded three-dimensional data 413 or the encoded three-dimensional data 414 .

[0383] Furthermore, the 3D data decoding device 500 notifies the server of the state of the client (the 3D data decoding device 500). The 3D data decoding device 500 receives one of the encoded 3D data 413 and 414 transmitted from the server according to the state of the client.

[0384] Accordingly, the three-dimensional data decoding apparatus 500 can receive appropriate data according to the status of the client.

[0385] Furthermore, the status of the client includes the communication status of the client (eg, network bandwidth) or the moving speed of the client.

[0386] Furthermore, the three-dimensional data decoding apparatus 500 further requests one of the encoded three-dimensional data 413 and 414 from the server, and receives one of the encoded three-dimensional data 413 and 414 transmitted from the server in accordance with the request.

[0387] Thereby, the three-dimensional data decoding device 500 can receive appropriate data according to the application.

[0388] (Implementation method 3)

[0389] In this embodiment, a method for transmitting and receiving three-dimensional data between vehicles will be described. For example, the three-dimensional data is transmitted and received between the own vehicle and surrounding vehicles.

[0390] Fig.24 This is a block diagram of a three-dimensional data production device 620 according to this embodiment. The three-dimensional data production device 620 is included in the vehicle, for example, and produces denser third three-dimensional data 636 by synthesizing the received second three-dimensional data 635 with the first three-dimensional data 632 produced by the three-dimensional data production device 620.

[0391] The three-dimensional data creation device 620 includes a three-dimensional data creation unit 621 , a request range determination unit 622 , a search unit 623 , a reception unit 624 , a decoding unit 625 , and a synthesis unit 626 .

[0392] First, the three-dimensional data creation unit 621 uses sensor information 631 detected by a sensor of the own vehicle to create first three-dimensional data 632. Next, the request range determination unit 622 determines a request range, which is a three-dimensional space range where the created first three-dimensional data 632 does not have enough data.

[0393] Next, the search unit 623 searches for surrounding vehicles that have three-dimensional data of the requested range, and sends request range information 633 showing the requested range to the surrounding vehicles determined by the search. Next, the receiving unit 624 receives the encoded three-dimensional data 634 (S624) as the encoded stream of the requested range from the surrounding vehicles. In addition, the search unit 623 can indiscriminately issue requests to all vehicles existing in the determined range, and receive the encoded three-dimensional data 634 from the responding party. In addition, the search unit 623 is not limited to vehicles, and can also issue requests to objects such as traffic lights or signs, and receive the encoded three-dimensional data 634 from the object.

[0394] Next, the decoder 625 decodes the received encoded three-dimensional data 634 to obtain second three-dimensional data 635. Next, the synthesizer 626 synthesizes the first three-dimensional data 632 and the second three-dimensional data 635 to create denser third three-dimensional data 636.

[0395] Next, the configuration and operation of the three-dimensional data transmitting device 640 according to this embodiment will be described. Fig.25 It is a block diagram of the three-dimensional data transmitting device 640.

[0396] The three-dimensional data sending device 640 is included in the above-mentioned surrounding vehicles, for example, and processes the fifth three-dimensional data 652 produced by the surrounding vehicles into the sixth three-dimensional data 654 requested by the own vehicle, and generates encoded three-dimensional data 634 by encoding the sixth three-dimensional data 654, and sends the encoded three-dimensional data 634 to the own vehicle.

[0397] The three-dimensional data transmitting device 640 includes a three-dimensional data creating unit 641 , a receiving unit 642 , an extracting unit 643 , a coding unit 644 , and a transmitting unit 645 .

[0398] First, the three-dimensional data creation unit 641 creates fifth three-dimensional data 652 using sensor information 651 detected by sensors provided in surrounding vehicles. Next, the receiving unit 642 receives the requested range information 633 transmitted from the own vehicle.

[0399] Next, the extraction unit 643 extracts the three-dimensional data of the requested range indicated by the requested range information 633 from the fifth three-dimensional data 652, and processes the fifth three-dimensional data 652 into the sixth three-dimensional data 654. Next, the encoding unit 644 encodes the sixth three-dimensional data 654, thereby generating the encoded three-dimensional data 634 as an encoded stream. Then, the transmission unit 645 transmits the encoded three-dimensional data 634 to the own vehicle.

[0400] In addition, although an example is described here in which the own vehicle has the three-dimensional data creation device 620 and the surrounding vehicles have the three-dimensional data transmission device 640, each vehicle may also have the functions of the three-dimensional data creation device 620 and the three-dimensional data transmission device 640.

[0401] (Implementation 4)

[0402] In this embodiment, the operation related to abnormal situations in the estimation of the own position based on the three-dimensional map will be described.

[0403] The use of autonomous movement of moving objects such as automatic driving of automobiles, robots, or flying objects such as drones will expand in the future. An example of a method for achieving such autonomous movement is a method in which a moving object estimates its position in a three-dimensional map (self-position estimation) and drives according to the map.

[0404] The self-position estimation is achieved by matching the three-dimensional map with the three-dimensional information around the own vehicle obtained by sensors such as the rangefinder (LIDAR, etc.) or stereo camera mounted on the own vehicle (hereinafter referred to as the own vehicle detection three-dimensional data), and estimating the own vehicle position within the three-dimensional map.

[0405] As shown in the HD map proposed by HERE, a 3D map is not only a 3D point cloud, but also includes 2D map data such as road and intersection shape information, and information such as congestion and accidents that change in real time. The 3D map is composed of multiple layers such as 3D data, 2D data, and metadata that changes in real time, and the device can obtain only the required data or refer to the required data.

[0406] The point cloud data may be the above-mentioned SWLD, or may include point group data that is not a feature point. Furthermore, the transmission and reception of the point cloud data is basically performed in one or more random access units.

[0407] As a matching method for the three-dimensional map and the three-dimensional data detected by the own vehicle, the following method can be adopted. For example, the device compares the shapes of the point groups in each other's point clouds and determines the parts with high similarity between the feature points as the same position. In addition, when the three-dimensional map is composed of SWLD, the device compares and matches the feature points constituting the SWLD with the three-dimensional feature points extracted from the three-dimensional data detected by the own vehicle.

[0408] Here, in order to estimate the own position with high accuracy, the following (A) and (B) need to be met: (A) a three-dimensional map and three-dimensional data for detecting the own vehicle are available, and (B) their accuracy meets a predetermined benchmark. However, in the following abnormal situation, (A) or (B) cannot be met.

[0409] (1) A three-dimensional map cannot be obtained through the communication path.

[0410] (2) There is no three-dimensional map or the obtained three-dimensional map is damaged.

[0411] (3) The sensor of the own vehicle fails, or due to bad weather, the accuracy of the three-dimensional data generated by the own vehicle is insufficient.

[0412] The following describes the operation for dealing with these abnormal situations. Although the operation is described below using a vehicle as an example, the following method can also be applied to all moving objects that move autonomously, such as robots and drones.

[0413] The following describes the configuration and operation of the three-dimensional information processing device according to the present embodiment for detecting abnormalities in three-dimensional data corresponding to a three-dimensional map or a vehicle. Fig.26 It is a block diagram showing a configuration example of a three-dimensional information processing device 700 according to this embodiment.

[0414] The three-dimensional information processing device 700 is mounted on a mobile object such as a motor vehicle. Fig.26 As shown, the three-dimensional information processing device 700 includes a three-dimensional map acquisition unit 701 , a vehicle detection data acquisition unit 702 , an abnormality determination unit 703 , a countermeasure operation determination unit 704 , and an operation control unit 705 .

[0415] In addition, the three-dimensional information processing device 700 may also include a camera for obtaining a two-dimensional image, or may include a sensor for one-dimensional data using ultrasonic waves or lasers, which is used to detect structural objects or moving objects around the vehicle, which is not shown in the figure. In addition, the three-dimensional information processing device 700 may also include a communication unit (not shown) for obtaining a three-dimensional map through a mobile communication network such as 4G or 5G, or communication between vehicles, or communication between roads and vehicles.

[0416] The three-dimensional map acquisition unit 701 acquires a three-dimensional map 711 near the driving route. For example, the three-dimensional map acquisition unit 701 acquires the three-dimensional map 711 through a mobile communication network, communication between vehicles, or communication between roads and vehicles.

[0417] Next, the own vehicle detection data acquisition unit 702 acquires the own vehicle detection three-dimensional data 712 based on the sensor information. For example, the own vehicle detection data acquisition unit 702 generates the own vehicle detection three-dimensional data 712 based on the sensor information acquired by the sensor included in the own vehicle.

[0418] Next, the abnormality determination unit 703 detects an abnormality by performing a predetermined check on at least one of the obtained three-dimensional map 711 and the vehicle detection three-dimensional data 712. That is, the abnormality determination unit 703 determines whether at least one of the obtained three-dimensional map 711 and the vehicle detection three-dimensional data 712 is abnormal.

[0419] When an abnormal situation is detected, the countermeasure action determination unit 704 determines a countermeasure action for the abnormal situation. Next, the action control unit 705 controls the operation of each processing unit required for executing the countermeasure action, such as the three-dimensional map acquisition unit 701.

[0420] In addition, when no abnormality is detected, the three-dimensional information processing device 700 ends the processing.

[0421] The three-dimensional information processing device 700 estimates the own position of the vehicle having the three-dimensional information processing device 700 using the three-dimensional map 711 and the own vehicle detection three-dimensional data 712. Then, the three-dimensional information processing device 700 uses the result of the own position estimation to make the vehicle automatically drive.

[0422] Accordingly, the three-dimensional information processing device 700 obtains map data (three-dimensional map 711) including the first three-dimensional position information via the channel. For example, the first three-dimensional position information is encoded in units of partial spaces having three-dimensional coordinate information, and the first three-dimensional position information includes a plurality of random access units, each of which is a collection of more than one partial space and can be independently decoded. For example, the first three-dimensional position information is data (SWLD) encoded with a feature point whose three-dimensional feature quantity is above a predetermined threshold.

[0423] Furthermore, the three-dimensional information processing device 700 generates the second three-dimensional position information (the vehicle detection three-dimensional data 712) based on the information detected by the sensor. Next, the three-dimensional information processing device 700 performs an abnormality determination process on the first three-dimensional position information or the second three-dimensional position information to determine whether the first three-dimensional position information or the second three-dimensional position information is abnormal.

[0424] When determining that the first three-dimensional position information or the second three-dimensional position information is abnormal, the three-dimensional information processing device 700 determines a countermeasure for the abnormality. Next, the three-dimensional information processing device 700 executes control required for executing the countermeasure.

[0425] According to this, the three-dimensional information processing device 700 can detect abnormality in the first three-dimensional position information or the second three-dimensional position information and perform a corresponding operation.

[0426] (Implementation method 5)

[0427] In this embodiment, a method of transmitting three-dimensional data to a rear vehicle and the like will be described.

[0428] Fig. 27 1 is a block diagram showing a configuration example of a three-dimensional data production device 810 according to the present embodiment. The three-dimensional data production device 810 is mounted on a vehicle, for example. The three-dimensional data production device 810 transmits and receives three-dimensional data with external traffic cloud monitoring, a vehicle in front, or a vehicle behind, and produces and accumulates the three-dimensional data.

[0429] The three-dimensional data production device 810 includes: a data receiving unit 811, a communication unit 812, a receiving control unit 813, a format conversion unit 814, multiple sensors 815, a three-dimensional data production unit 816, a three-dimensional data synthesis unit 817, a three-dimensional data storage unit 818, a communication unit 819, a sending control unit 820, a format conversion unit 821, and a data sending unit 822.

[0430] The data receiving unit 811 receives three-dimensional data 831 from traffic cloud monitoring or the vehicle ahead. The three-dimensional data 831 includes, for example, a point cloud including information on areas that cannot be detected by the vehicle's sensor 815, visible light images, depth information, sensor position information, or speed information.

[0431] The communication unit 812 communicates with the traffic cloud monitoring or the vehicle ahead, and sends a data transmission request or the like to the traffic cloud monitoring or the vehicle ahead.

[0432] The reception control unit 813 exchanges information such as a corresponding format with the communication partner via the communication unit 812, and establishes communication with the communication partner.

[0433] The format conversion unit 814 generates three-dimensional data 832 by performing format conversion or the like on the three-dimensional data 831 received by the data reception unit 811. Furthermore, when the three-dimensional data 831 is compressed or encoded, the format conversion unit 814 performs decompression or decoding processing.

[0434] The plurality of sensors 815 are a group of sensors such as LiDAR, a visible light camera, or an infrared camera that obtains information outside the vehicle and generates sensor information 833. For example, when the sensor 815 is a laser sensor such as LiDAR, the sensor information 833 is three-dimensional data such as a point cloud (point group data). In addition, the number of sensors 815 may not be multiple.

[0435] The three-dimensional data generating unit 816 generates three-dimensional data 834 based on the sensor information 833. The three-dimensional data 834 includes, for example, point cloud, visible light image, depth information, sensor position information, or speed information.

[0436] The three-dimensional data synthesis unit 817 synthesizes the three-dimensional data 832 produced by traffic cloud monitoring or the front vehicle, etc., into the three-dimensional data 834 produced based on the sensor information 833 of the own vehicle, thereby constructing three-dimensional data 835 that also includes the space in front of the front vehicle that cannot be detected by the sensor 815 of the own vehicle.

[0437] The three-dimensional data accumulation unit 818 accumulates the generated three-dimensional data 835 and the like.

[0438] The communication unit 819 communicates with the traffic cloud monitoring or the vehicle behind, and sends a data transmission request and the like to the traffic cloud monitoring or the vehicle behind.

[0439] The transmission control unit 820 exchanges information such as the corresponding format with the communication partner via the communication unit 819 to establish communication with the communication partner. In addition, the transmission control unit 820 determines the transmission area of ​​the space of the three-dimensional data to be transmitted based on the three-dimensional data construction information of the three-dimensional data 832 generated by the three-dimensional data synthesis unit 817 and the data transmission request from the communication partner.

[0440] Specifically, the transmission control unit 820 determines the transmission area including the space in front of the vehicle that cannot be detected by the sensor of the rear vehicle according to the data transmission request from the traffic cloud monitoring or the rear vehicle. In addition, the transmission control unit 820 determines the transmission area by judging whether the space that can be transmitted or the transmitted space has been updated based on the three-dimensional data construction information. For example, the transmission control unit 820 determines the area that is both specified by the data transmission request and the area where the corresponding three-dimensional data 835 exists as the transmission area. In addition, the transmission control unit 820 notifies the format corresponding to the communication partner and the transmission area to the format conversion unit 821.

[0441] The format conversion unit 821 generates three-dimensional data 837 by converting the three-dimensional data 836 of the transmission area in the three-dimensional data 835 accumulated in the three-dimensional data accumulation unit 818 into a format corresponding to the receiving side. In addition, the format conversion unit 821 may compress or encode the three-dimensional data 837 to reduce the data amount.

[0442] The data transmission unit 822 transmits the three-dimensional data 837 to the traffic cloud monitoring or the rear vehicle. The three-dimensional data 837 includes, for example, a point cloud in front of the vehicle including information on the area that becomes a blind spot for the rear vehicle, a visible light image, depth information, or sensor position information.

[0443] In addition, although the example in which the format conversion units 814 and 821 perform format conversion is described here, format conversion may not be performed.

[0444] With this configuration, the three-dimensional data production device 810 obtains three-dimensional data 831 of an area that cannot be detected by the sensor 815 of the own vehicle from the outside, and generates three-dimensional data 835 by synthesizing the three-dimensional data 831 with three-dimensional data 834 based on sensor information 833 detected by the sensor 815 of the own vehicle. In this way, the three-dimensional data production device 810 can generate three-dimensional data of a range that cannot be detected by the sensor 815 of the own vehicle.

[0445] In addition, the three-dimensional data production device 810 can send three-dimensional data of the space in front of its own vehicle that cannot be detected by the sensors of the rear vehicle to the traffic cloud monitoring or the rear vehicle, etc. according to the data sending request from the traffic cloud monitoring or the rear vehicle.

[0446] (Implementation 6)

[0447] In the fifth embodiment, a client device such as a vehicle sends three-dimensional data to another vehicle or a server such as a traffic cloud monitoring device. In this embodiment, the client device sends sensor information obtained by the sensor to the server or other client devices.

[0448] First, the configuration of a system according to this embodiment will be described. Fig.28 1 shows the structure of the system for transmitting and receiving three-dimensional maps and sensor information according to the present embodiment. The system includes a server 901 and client devices 902A and 902B. In addition, when the client devices 902A and 902B are not specifically distinguished, they are also referred to as client devices 902.

[0449] The client device 902 is, for example, an in-vehicle device mounted on a mobile body such as a vehicle. The server 901 is, for example, a traffic cloud monitoring system, and can communicate with a plurality of client devices 902 .

[0450] The server 901 transmits the three-dimensional map composed of the point cloud to the client device 902. In addition, the composition of the three-dimensional map is not limited to the point cloud, and can also be expressed by other three-dimensional data such as a mesh structure.

[0451] The client device 902 sends the sensor information obtained by the client device 902 to the server 901. The sensor information includes, for example, at least one of LiDAR information, visible light images, infrared images, depth images, sensor position information, and speed information.

[0452] The data sent and received between the server 901 and the client device 902 may be compressed when it is desired to reduce the data, and may not be compressed when it is desired to maintain the accuracy of the data. When compressing the data, a three-dimensional compression method based on an octree may be used in the point cloud, for example. In addition, a two-dimensional image compression method may be used in the visible light image, the infrared image, and the depth image. The two-dimensional image compression method is, for example, MPEG-4AVC or HEVC standardized by MPEG.

[0453] Furthermore, the server 901 transmits the three-dimensional map managed by the server 901 to the client device 902 in accordance with the three-dimensional map transmission request from the client device 902. In addition, the server 901 may transmit the three-dimensional map without waiting for the three-dimensional map transmission request from the client device 902. For example, the server 901 may broadcast the three-dimensional map to one or more client devices 902 in a predetermined space. Furthermore, the server 901 may transmit the three-dimensional map suitable for the location of the client device 902 at regular intervals to the client device 902 that has received a transmission request once. Furthermore, the server 901 may transmit the three-dimensional map to the client device 902 every time the three-dimensional map managed by the server 901 is updated.

[0454] The client device 902 sends a three-dimensional map transmission request to the server 901. For example, when the client device 902 wants to estimate its own position while driving, the client device 902 sends a three-dimensional map transmission request to the server 901.

[0455] In addition, in the following cases, the client device 902 may also issue a three-dimensional map transmission request to the server 901. In the case where the three-dimensional map held by the client device 902 is relatively old, the client device 902 may also issue a three-dimensional map transmission request to the server 901. For example, in the case where the client device 902 obtains the three-dimensional map and a certain period of time has passed, the client device 902 may also issue a three-dimensional map transmission request to the server 901.

[0456] Alternatively, the client device 902 may send a three-dimensional map transmission request to the server 901 before a certain time when the client device 902 is about to leave the space shown on the three-dimensional map held by the client device 902. For example, the client device 902 may send a three-dimensional map transmission request to the server 901 when the client device 902 is within a predetermined distance from the boundary of the space shown on the three-dimensional map held by the client device 902. Furthermore, when the moving path and moving speed of the client device 902 are known, the time when the client device 902 leaves the space shown on the three-dimensional map held by the client device 902 may be predicted based on the known moving path and moving speed.

[0457] When the error between the three-dimensional data generated by the client device 902 based on the sensor information and the position of the three-dimensional map is greater than a certain range, the client device 902 may send a request to the server 901 to send the three-dimensional map.

[0458] The client device 902 transmits the sensor information to the server 901 in accordance with the sensor information transmission request transmitted from the server 901. In addition, the client device 902 may transmit the sensor information to the server 901 without waiting for the sensor information transmission request from the server 901. For example, when the client device 902 has received a sensor information transmission request from the server 901 once, the client device 902 may periodically transmit the sensor information to the server 901 within a certain period. Furthermore, when the error between the three-dimensional data generated by the client device 902 based on the sensor information and the position of the three-dimensional map obtained from the server 901 is greater than a certain range, the client device 902 may determine that there is a possibility that the three-dimensional map around the client device 902 has changed, and transmit the determination result together with the sensor information to the server 901.

[0459] The server 901 issues a request to send sensor information to the client device 902. For example, the server 901 receives location information of the client device 902 such as GPS from the client device 902. When the server 901 determines that the client device 902 is close to a space with little information in the three-dimensional map managed by the server 901 based on the location information of the client device 902, the server 901 issues a request to send sensor information to the client device 902 in order to regenerate the three-dimensional map. In addition, the server 901 may issue a request to send sensor information when it is desired to update the three-dimensional map, when it is desired to check the road conditions such as when snow is accumulated or when a disaster occurs, when it is desired to check the congestion conditions or the accident conditions.

[0460] Furthermore, the client device 902 may set the data volume of the sensor information to be sent to the server 901 according to the communication state or the frequency band when receiving the sensor information sending request received from the server 901. Setting the data volume of the sensor information to be sent to the server 901 means, for example, increasing or decreasing the data itself or selecting an appropriate compression method.

[0461] Fig.29 is a block diagram showing an example of the configuration of the client device 902. The client device 902 receives a three-dimensional map composed of a point cloud or the like from the server 901, and estimates the position of the client device 902 itself based on the three-dimensional data produced based on the sensor information of the client device 902. The client device 902 then transmits the obtained sensor information to the server 901.

[0462] The client device 902 includes: a data receiving unit 1011, a communication unit 1012, a receiving control unit 1013, a format conversion unit 1014, multiple sensors 1015, a three-dimensional data production unit 1016, a three-dimensional image processing unit 1017, a three-dimensional data storage unit 1018, a format conversion unit 1019, a communication unit 1020, a sending control unit 1021, and a data sending unit 1022.

[0463] The data receiving unit 1011 receives a three-dimensional map 1031 from the server 901. The three-dimensional map 1031 is data including a point cloud such as WLD or SWLD. The three-dimensional map 1031 may include either compressed data or uncompressed data.

[0464] The communication unit 1012 communicates with the server 901 and transmits a data transmission request (for example, a transmission request of a three-dimensional map) and the like to the server 901 .

[0465] The reception control unit 1013 exchanges information such as the corresponding format with the communication partner via the communication unit 1012, and establishes communication with the communication partner.

[0466] The format conversion unit 1014 generates a three-dimensional map 1032 by performing format conversion on the three-dimensional map 1031 received by the data receiving unit 1011. Furthermore, the format conversion unit 1014 performs decompression or decoding processing when the three-dimensional map 1031 is compressed or encoded. In addition, the format conversion unit 1014 does not perform decompression or decoding processing when the three-dimensional map 1031 is non-compressed data.

[0467] The plurality of sensors 1015 are a group of sensors mounted on the client device 902 such as LiDAR, a visible light camera, an infrared camera, or a depth sensor for obtaining information outside the vehicle, and generate sensor information 1033. For example, when the sensor 1015 is a laser sensor such as LiDAR, the sensor information 1033 is three-dimensional data such as a point cloud (point group data). In addition, the number of sensors 1015 may not be multiple.

[0468] The three-dimensional data generating unit 1016 generates three-dimensional data 1034 of the surroundings of the own vehicle based on the sensor information 1033. For example, the three-dimensional data generating unit 1016 generates point cloud data having color information of the surroundings of the own vehicle using information obtained by LiDAR and visible light images obtained by a visible light camera.

[0469] The three-dimensional image processing unit 1017 uses the received three-dimensional map 1032 such as the point cloud and the three-dimensional data 1034 of the surroundings of the own vehicle generated based on the sensor information 1033 to perform the own vehicle's own position estimation processing, etc. Alternatively, the three-dimensional image processing unit 1017 may synthesize the three-dimensional map 1032 and the three-dimensional data 1034 to produce three-dimensional data 1035 of the surroundings of the own vehicle, and use the produced three-dimensional data 1035 to perform the own position estimation processing.

[0470] The three-dimensional data storage unit 1018 stores the three-dimensional map 1032 , the three-dimensional data 1034 , the three-dimensional data 1035 , and the like.

[0471] The format conversion unit 1019 generates sensor information 1037 by converting the sensor information 1033 into a format corresponding to the receiving side. In addition, the format conversion unit 1019 can reduce the amount of data by compressing or encoding the sensor information 1037. In addition, when format conversion is not required, the format conversion unit 1019 can omit the processing. In addition, the format conversion unit 1019 can control the amount of data to be transmitted according to the designation of the transmission range.

[0472] The communication unit 1020 communicates with the server 901 , and receives a data transmission request (a sensor information transmission request) and the like from the server 901 .

[0473] The transmission control unit 1021 exchanges information such as a corresponding format with the communication partner via the communication unit 1020, thereby establishing communication.

[0474] The data transmission unit 1022 transmits the sensor information 1037 to the server 901. The sensor information 1037 includes information obtained by multiple sensors 1015, such as information obtained by LiDAR, a brightness image (visible light image) obtained by a visible light camera, an infrared image obtained by an infrared camera, a depth image obtained by a depth sensor, sensor position information, and speed information.

[0475] Next, the configuration of the server 901 will be described. Fig.30 1 is a block diagram showing an example of the configuration of the server 901. The server 901 receives sensor information sent from the client device 902, and creates three-dimensional data based on the received sensor information. The server 901 updates the three-dimensional map managed by the server 901 using the created three-dimensional data. Furthermore, the server 901 sends the updated three-dimensional map to the client device 902 in accordance with the three-dimensional map sending request from the client device 902.

[0476] The server 901 includes: a data receiving unit 1111, a communication unit 1112, a receiving control unit 1113, a format conversion unit 1114, a three-dimensional data production unit 1116, a three-dimensional data synthesis unit 1117, a three-dimensional data storage unit 1118, a format conversion unit 1119, a communication unit 1120, a sending control unit 1121, and a data sending unit 1122.

[0477] The data receiving unit 1111 receives sensor information 1037 from the client device 902. The sensor information 1037 includes, for example, information obtained by LiDAR, a brightness image (visible light image) obtained by a visible light camera, an infrared image obtained by an infrared camera, a depth image obtained by a depth sensor, sensor position information, and speed information.

[0478] The communication unit 1112 communicates with the client device 902 and transmits a data transmission request (for example, a sensor information transmission request) and the like to the client device 902 .

[0479] The reception control unit 1113 exchanges information such as the corresponding format with the communication partner via the communication unit 1112, thereby establishing communication.

[0480] When the received sensor information 1037 is compressed or encoded, the format conversion unit 1114 performs decompression or decoding processing to generate the sensor information 1132. When the sensor information 1037 is uncompressed data, the format conversion unit 1114 does not perform decompression or decoding processing.

[0481] The three-dimensional data generating unit 1116 generates three-dimensional data 1134 of the surroundings of the client device 902 based on the sensor information 1132. For example, the three-dimensional data generating unit 1116 generates point cloud data having color information of the surroundings of the client device 902 using information obtained by LiDAR and visible light images obtained by a visible light camera.

[0482] The three-dimensional data synthesis unit 1117 synthesizes the three-dimensional data 1134 generated based on the sensor information 1132 with the three-dimensional map 1135 managed by the server 901 , thereby updating the three-dimensional map 1135 .

[0483] The three-dimensional data storage unit 1118 stores the three-dimensional map 1135 and the like.

[0484] The format conversion unit 1119 generates the three-dimensional map 1031 by converting the three-dimensional map 1135 into a format corresponding to the receiving side. In addition, the format conversion unit 1119 can also reduce the amount of data by compressing or encoding the three-dimensional map 1135. In addition, when format conversion is not required, the format conversion unit 1119 can also omit the processing. In addition, the format conversion unit 1119 can control the amount of data sent according to the designation of the sending range.

[0485] The communication unit 1120 communicates with the client device 902 , and receives a data transmission request (a transmission request of a three-dimensional map) and the like from the client device 902 .

[0486] The transmission control unit 1121 exchanges information such as a corresponding format with the communication partner via the communication unit 1120 to establish communication.

[0487] The data transmission unit 1122 transmits the three-dimensional map 1031 to the client device 902. The three-dimensional map 1031 is data including a point cloud such as WLD or SWLD. The three-dimensional map 1031 may include either compressed data or uncompressed data.

[0488] Next, the workflow of the client device 902 will be described. Fig.31 1 is a flowchart showing the operation performed by the client device 902 when obtaining a three-dimensional map.

[0489] First, the client device 902 requests the server 901 to send a three-dimensional map (point cloud, etc.) (S1001). At this time, the client device 902 also sends the location information of the client device 902 obtained by GPS, etc., and accordingly, the server 901 can be requested to send a three-dimensional map related to the location information.

[0490] Next, the client device 902 receives the three-dimensional map from the server 901 (S1002). If the received three-dimensional map is compressed data, the client device 902 decodes the received three-dimensional map to generate a non-compressed three-dimensional map (S1003).

[0491] Next, the client device 902 creates three-dimensional data 1034 of the surroundings of the client device 902 based on the sensor information 1033 obtained from the plurality of sensors 1015 (S1004). Next, the client device 902 estimates its own position using the three-dimensional map 1032 received from the server 901 and the three-dimensional data 1034 created based on the sensor information 1033 (S1005).

[0492] Fig.32 101 is a flowchart showing the operation of the client device 902 when transmitting sensor information. First, the client device 902 receives a request to transmit sensor information from the server 901 (S1011). The client device 902 that has received the transmission request transmits the sensor information 1037 to the server 901 (S1012). In addition, when the sensor information 1033 includes a plurality of information obtained by a plurality of sensors 1015, the client device 902 compresses each information in a compression method suitable for each information, thereby generating the sensor information 1037.

[0493] Next, the workflow of the server 901 is described. Fig.33 1 is a flowchart showing the operation of the server 901 when acquiring sensor information. First, the server 901 requests the client device 902 to send sensor information (S1021). Next, the server 901 receives the sensor information 1037 sent from the client device 902 in accordance with the request (S1022). Next, the server 901 uses the received sensor information 1037 to create three-dimensional data 1134 (S1023). Next, the server 901 reflects the created three-dimensional data 1134 on the three-dimensional map 1135 (S1024).

[0494] Fig.341031 is a flowchart showing the work of the server 901 when sending a three-dimensional map. First, the server 901 receives a request to send a three-dimensional map from the client device 902 (S1031). The server 901 that has received the request to send a three-dimensional map sends the three-dimensional map 1031 to the client device 902 (S1032). At this time, the server 901 can extract a three-dimensional map in the vicinity corresponding to the location information of the client device 902, and send the extracted three-dimensional map. Furthermore, the server 901 can compress the three-dimensional map composed of the point cloud, for example, using an octree compression method, and send the compressed three-dimensional map.

[0495] Hereinafter, modified examples of the present embodiment will be described.

[0496] The server 901 uses the sensor information 1037 received from the client device 902 to create three-dimensional data 1134 near the location of the client device 902. Next, the server 901 matches the created three-dimensional data 1134 with a three-dimensional map 1135 of the same area managed by the server 901, and calculates the difference between the three-dimensional data 1134 and the three-dimensional map 1135. When the difference is greater than a predetermined threshold, the server 901 determines that some abnormality has occurred around the client device 902. For example, when the ground subsidence occurs due to a natural disaster such as an earthquake, it can be considered that there will be a large difference between the three-dimensional map 1135 managed by the server 901 and the three-dimensional data 1134 created based on the sensor information 1037.

[0497] The sensor information 1037 may also include at least one of the type of sensor, the performance of the sensor, and the model of the sensor. In addition, a category ID corresponding to the performance of the sensor may be added to the sensor information 1037. For example, when the sensor information 1037 is information obtained by LiDAR, an identifier may be assigned in consideration of the performance of the sensor, for example, category 1 may be assigned to a sensor that can obtain information with an accuracy of several mm, category 2 may be assigned to a sensor that can obtain information with an accuracy of several cm, and category 3 may be assigned to a sensor that can obtain information with an accuracy of several m. In addition, the server 901 may estimate the performance information of the sensor from the model of the client device 902. For example, when the client device 902 is mounted on a vehicle, the server 901 may determine the specification information of the sensor according to the model of the vehicle. In this case, the server 901 may obtain the information of the model of the vehicle in advance, or include the information in the sensor information. In addition, the server 901 may switch the degree of correction of the three-dimensional data 1134 produced using the sensor information 1037 using the obtained sensor information 1037. For example, when the sensor performance is high accuracy (category 1), the server 901 does not perform correction on the three-dimensional data 1134. When the sensor performance is low accuracy (category 3), the server 901 applies correction suitable for the accuracy of the sensor to the three-dimensional data 1134. For example, the server 901 increases the degree (intensity) of correction as the accuracy of the sensor is lower.

[0498] The server 901 may also simultaneously send a request to send sensor information to multiple client devices 902 existing in a certain space. When the server 901 receives multiple sensor information from multiple client devices 902, it is not necessary to use all the sensor information to create the three-dimensional data 1134. For example, the server 901 may select the sensor information to be used according to the performance of the sensor. For example, when the server 901 updates the three-dimensional map 1135, it may select high-precision sensor information (category 1) from the received multiple sensor information and use the selected sensor information to create the three-dimensional data 1134.

[0499] The server 901 is not limited to servers such as traffic cloud monitoring, but can also be other client devices (car-mounted). Fig.35 The system configuration in this case is shown.

[0500] For example, the client device 902C sends a request to send sensor information to the client device 902A that is nearby, and obtains the sensor information from the client device 902A. Then, the client device 902C uses the obtained sensor information of the client device 902A to create three-dimensional data, and updates the three-dimensional map of the client device 902C. In this way, the client device 902C can utilize the performance of the client device 902C to generate a three-dimensional map of the space that can be obtained from the client device 902A. For example, when the performance of the client device 902C is high, this situation can be considered to occur.

[0501] In this case, the client device 902A that has provided the sensor information is given the right to obtain the high-precision three-dimensional map generated by the client device 902C. The client device 902A receives the high-precision three-dimensional map from the client device 902C in accordance with the right.

[0502] Furthermore, the client device 902C may send a request to send sensor information to multiple client devices 902 (client device 902A and client device 902B) in the vicinity. If the sensor of the client device 902A or the client device 902B is high-performance, the client device 902C can use the sensor information obtained by the high-performance sensor to create three-dimensional data.

[0503] Fig.36 1 is a block diagram showing the functional configuration of the server 901 and the client device 902. The server 901 includes, for example, a 3D map compression / decoding processing unit 1201 that compresses and decodes a 3D map and a sensor information compression / decoding processing unit 1202 that compresses and decodes sensor information.

[0504] The client device 902 includes: a three-dimensional map decoding processing unit 1211, and a sensor information compression processing unit 1212. The three-dimensional map decoding processing unit 1211 receives the coded data of the compressed three-dimensional map, decodes the coded data and obtains the three-dimensional map. The sensor information compression processing unit 1212 does not compress the three-dimensional data produced by the obtained sensor information, but compresses the sensor information itself, and sends the compressed coded data of the sensor information to the server 901. According to this structure, the client device 902 can keep the processing unit (device or LSI) for decoding the three-dimensional map (point cloud, etc.) inside, without having to keep the processing unit for compressing the three-dimensional data of the three-dimensional map (point cloud, etc.) inside. In this way, the cost and power consumption of the client device 902 can be suppressed.

[0505] As described above, the client device 902 involved in this embodiment is mounted on the mobile body, and generates three-dimensional data 1034 of the surroundings of the mobile body based on the sensor information 1033 indicating the surrounding conditions of the mobile body obtained by the sensor 1015 mounted on the mobile body. The client device 902 estimates the own position of the mobile body using the generated three-dimensional data 1034. The client device 902 transmits the obtained sensor information 1033 to the server 901 or another mobile body 902.

[0506] Based on this, the client device 902 transmits the sensor information 1033 to the server 901 or the like. In this way, the amount of data to be transmitted may be reduced compared to the case of transmitting three-dimensional data. Furthermore, since it is not necessary to perform processing such as compression or encoding of three-dimensional data on the client device 902, the amount of processing on the client device 902 can be reduced. Therefore, the client device 902 can reduce the amount of data to be transmitted or simplify the configuration of the device.

[0507] Furthermore, the client device 902 further sends a request to send a three-dimensional map to the server 901, and receives the three-dimensional map 1031 from the server 901. The client device 902 estimates its own position using the three-dimensional data 1034 and the three-dimensional map 1032.

[0508] Furthermore, the sensor information 1033 includes at least one of information obtained by a laser sensor, a brightness image (visible light image), an infrared image, a depth image, position information of the sensor, and speed information of the sensor.

[0509] Also, the sensor information 1033 includes information showing the performance of the sensor.

[0510] Furthermore, the client device 902 encodes or compresses the sensor information 1033, and transmits the encoded or compressed sensor information 1037 to the server 901 or another mobile body 902. Thus, the client device 902 can reduce the amount of data to be transmitted.

[0511] For example, the client device 902 includes a processor and a memory, and the processor performs the above-mentioned processing using the memory.

[0512] Furthermore, the server 901 according to the present embodiment can communicate with the client device 902 mounted on the mobile body, and receive sensor information 1037 indicating the surrounding conditions of the mobile body obtained by the sensor 1015 mounted on the mobile body from the client device 902. The server 901 generates three-dimensional data 1134 of the surroundings of the mobile body based on the received sensor information 1037.

[0513] Accordingly, the server 901 uses the sensor information 1037 sent from the client device 902 to create the three-dimensional data 1134. In this way, compared with the case where the client device 902 sends the three-dimensional data, it is possible to reduce the amount of data to be sent. In addition, since it is not necessary to perform processing such as compression or encoding of the three-dimensional data on the client device 902, the processing amount of the client device 902 can be reduced. In this way, the server 901 can reduce the amount of data to be transmitted or simplify the configuration of the device.

[0514] Furthermore, the server 901 further transmits a request for transmitting sensor information to the client device 902 .

[0515] Furthermore, the server 901 further updates the three-dimensional map 1135 using the created three-dimensional data 1134 , and transmits the three-dimensional map 1135 to the client device 902 in response to a transmission request for the three-dimensional map 1135 from the client device 902 .

[0516] Furthermore, the sensor information 1037 includes at least one of information obtained by a laser sensor, a brightness image (visible light image), an infrared image, a depth image, position information of the sensor, and speed information of the sensor.

[0517] Also, the sensor information 1037 includes information showing the performance of the sensor.

[0518] Furthermore, the server 901 further calibrates the three-dimensional data according to the performance of the sensor. Accordingly, the three-dimensional data production method can improve the quality of the three-dimensional data.

[0519] Furthermore, when receiving sensor information, the server 901 receives a plurality of sensor information 1037 from a plurality of client devices 902, and selects the sensor information 1037 used in the production of the three-dimensional data 1134 based on a plurality of information indicating the performance of the sensors included in the plurality of sensor information 1037. In this way, the server 901 can improve the quality of the three-dimensional data 1134.

[0520] Furthermore, the server 901 decodes or decompresses the received sensor information 1137, and creates three-dimensional data 1134 based on the decoded or decompressed sensor information 1132. In this way, the server 901 can reduce the amount of data to be transmitted.

[0521] For example, the server 901 includes a processor and a memory, and the processor performs the above-mentioned processing using the memory.

[0522] (Implementation 7)

[0523] In this embodiment, a method for encoding and a method for decoding three-dimensional data using an inter-frame prediction process will be described.

[0524] Fig.37 3D data encoding device 1300 according to the present embodiment is a block diagram. The 3D data encoding device 1300 generates a coded bit stream (hereinafter also simply referred to as a bit stream) as a coded signal by encoding 3D data. Fig.37 As shown, the three-dimensional data encoding device 1300 includes: a division unit 1301, a subtraction unit 1302, a transformation unit 1303, a quantization unit 1304, an inverse quantization unit 1305, an inverse transformation unit 1306, an addition unit 1307, a reference volume memory 1308, an intra-frame prediction unit 1309, a reference space memory 1310, an inter-frame prediction unit 1311, a prediction control unit 1312, and an entropy coding unit 1313.

[0525] The division unit 1301 divides each space (SPC) included in the three-dimensional data into a plurality of volumes (VLM) as coding units. Furthermore, the division unit 1301 performs octree representation (Octree) on the voxels in each volume. In addition, the division unit 1301 may make the space and the volume the same size and perform octree representation on the space. Furthermore, the division unit 1301 may also attach information required for octree representation (depth information, etc.) to the header of the bitstream, etc.

[0526] The subtraction unit 1302 calculates the difference between the volume (coding target volume) output from the division unit 1301 and the prediction volume generated by intra prediction or inter prediction described later, and outputs the calculated difference as a prediction residual to the transformation unit 1303 . Fig.38 An example of calculating the prediction residual is shown. The bit strings of the encoding target volume and the prediction volume shown here are, for example, position information indicating the positions of three-dimensional points (for example, point clouds) included in the volume.

[0527] The following describes the octree representation and the scanning order of voxels. The volume is transformed into an octree structure (octreeization) and then encoded. The octree structure consists of nodes and leaf nodes. Each node has 8 nodes or leaf nodes, and each leaf node has voxel (VXL) information. Fig.39 An example of the configuration of a volume including a plurality of voxels is shown. Fig.40 Shows the Fig.39 The volume shown is transformed into an example of an octree structure. Here, Fig.40 The leaf nodes 1, 2, and 3 in the leaf nodes shown represent Fig.39 The voxels VXL1, VXL2, and VXL3 shown express the VXL (hereinafter referred to as effective VXL) including the point group.

[0528] The octree is represented by a binary sequence of 0 and 1. For example, when a node or valid VXL is set to a value of 1 and the others are set to a value of 0, each node and leaf node is assigned Fig.40 The binary sequence shown in FIG. Then, the binary sequence is scanned in a width-first or depth-first scanning order. For example, in the case of a width-first scan, Fig.41 The binary sequence shown in A. When the depth-first scan is performed, we get Fig.41 The binary sequence shown in B. The binary sequence obtained by this scanning is encoded by entropy coding, so that the amount of information is reduced.

[0529] Next, the depth information in the octree representation is explained. The depth in the octree representation is used to control the granularity of the point cloud information contained in the volume. If the depth is set to a large level, the point cloud information can be reproduced at a finer level, but the amount of data used to represent the nodes and leaf nodes will increase. On the contrary, if the depth is set to a small level, although the amount of data can be reduced, multiple point cloud information in different positions and colors will be regarded as the same position and the same color, so the original information of the point cloud information will be lost.

[0530] For example, Fig.42 Shows the Fig.40 The octree with depth = 2 is shown as an example of expressing an octree with depth = 1. Fig.42 The octree shown is Fig.40 The octree shown has a small amount of data. Fig.42 The octree shown is similar to Fig.42 Compared with the octree shown, the number of bits after binary serialization is less. Fig.40 The leaf nodes 1 and 2 shown in the figure become Fig.41 The leaf node 1 shown is shown. That is, Fig.40 The leaf node 1 and the leaf node 2 shown are information of different locations.

[0531] Fig.43 Shown with Fig.42 The volume corresponding to the octree shown. Fig.39 The VXL1 and VXL2 shown are Fig.43 In this case, the three-dimensional data encoding device 1300 corresponds to VXL12 shown in FIG. Fig.39 The color information of VXL1 and VXL2 shown in the figure generates Fig.43For example, the three-dimensional data encoding device 1300 calculates the color information of VXL1 and VXL2 as the color information of VXL12 using the average value, median value, or weighted average value. In this way, the three-dimensional data encoding device 1300 can control the reduction of the data amount by changing the depth of the octree.

[0532] The three-dimensional data encoding device 1300 may also use any one of the world space units, space units, and volume units to set the depth information of the octree. In addition, at this time, the three-dimensional data encoding device 1300 may also attach the depth information to the header information of the world space, the header information of the space, or the header information of the volume. In addition, the same value may be used as the depth information in all world spaces, spaces, and volumes at different times. In this case, the three-dimensional data encoding device 1300 may also attach the depth information to the header information that manages the world space at all times.

[0533] When the voxel contains color information, the transformation unit 1303 applies a frequency transformation such as an orthogonal transformation to the prediction residual of the color information of the voxels in the volume. For example, the transformation unit 1303 scans the prediction residual in a certain scanning order to produce a one-dimensional arrangement. Thereafter, the transformation unit 1303 transforms the one-dimensional arrangement into the frequency domain by applying a one-dimensional orthogonal transformation to the produced one-dimensional arrangement. Accordingly, when the value of the prediction residual in the volume is close, the value of the frequency component of the low frequency band becomes larger, and the value of the frequency component of the high frequency band becomes smaller. Therefore, the quantization unit 1304 can more effectively reduce the amount of coding.

[0534] Furthermore, the transformation unit 1303 may use orthogonal transformation of two or more dimensions instead of one-dimensional orthogonal transformation. For example, the transformation unit 1303 maps the prediction residual into a two-dimensional arrangement in a certain scanning order, and applies a two-dimensional orthogonal transformation to the obtained two-dimensional arrangement. Furthermore, the transformation unit 1303 may select the orthogonal transformation method to be used from a plurality of orthogonal transformation methods. In this case, the three-dimensional data encoding device 1300 attaches information indicating which orthogonal transformation method is used to the bit stream. Furthermore, the transformation unit 1303 may select the orthogonal transformation method to be used from a plurality of orthogonal transformation methods with different dimensions. In this case, the three-dimensional data encoding device 1300 attaches information indicating which dimensional orthogonal transformation method is used to the bit stream.

[0535] For example, the transformation unit 1303 matches the scanning order of the prediction residual with the scanning order (width-first or depth-first, etc.) in the octree within the volume. Accordingly, since it is not necessary to attach information showing the scanning order of the prediction residual to the bitstream, the additional overhead can be reduced. In addition, the transformation unit 1303 may also apply a scanning order different from the scanning order of the octree. In this case, the three-dimensional data encoding device 1300 attaches information showing the scanning order of the prediction residual to the bitstream. Accordingly, the three-dimensional data encoding device 1300 can efficiently encode the prediction residual. In addition, the three-dimensional data encoding device 1300 may attach information (flag, etc.) indicating whether the scanning order of the octree is applicable to the bitstream, and when the scanning order of the octree is not applicable, the information showing the scanning order of the prediction residual is attached to the bitstream.

[0536] The conversion unit 1303 may convert not only the prediction residual of the color information but also other attribute information of the voxel. For example, the conversion unit 1303 may convert and encode information such as reflectivity obtained when a point cloud is obtained by LiDAR or the like.

[0537] The transform unit 1303 may skip the process when the space does not have attribute information such as color information. Furthermore, the three-dimensional data encoding device 1300 may add information (flag) indicating whether to skip the process of the transform unit 1303 to the bit stream.

[0538] The quantization unit 1304 quantizes the frequency components of the prediction residual generated by the transformation unit 1303 using the quantization control parameters to generate quantization coefficients. The amount of information is reduced accordingly. The generated quantization coefficients are output to the entropy coding unit 1313. The quantization unit 1304 can control the quantization control parameters according to world space units, space units, or volume units. At this time, the three-dimensional data encoding device 1300 attaches the quantization control parameters to the respective header information, etc. In addition, the quantization unit 1304 can also change the weight according to the frequency components of each prediction residual to perform quantization control. For example, the quantization unit 1304 can perform detailed quantization on the low-frequency components and coarse quantization on the high-frequency components. In this case, the three-dimensional data encoding device 1300 can attach parameters representing the weights of each frequency component to the header.

[0539] The quantization unit 1304 may skip the process when the space does not have attribute information such as color information. In addition, the three-dimensional data encoding device 1300 may add information (flag) indicating whether the process of the quantization unit 1304 is skipped to the bit stream.

[0540] The inverse quantization unit 1305 inversely quantizes the quantization coefficients generated by the quantization unit 1304 using the quantization control parameters, thereby generating inverse quantization coefficients of the prediction residual, and outputs the generated inverse quantization coefficients to the inverse transformation unit 1306 .

[0541] The inverse transform unit 1306 generates an inverse transform applied prediction residual by applying inverse transform to the inverse quantized coefficients generated by the inverse quantization unit 1305. Since the inverse transform applied prediction residual is a prediction residual generated after quantization, it may not be completely consistent with the prediction residual output by the transform unit 1303.

[0542] The adding unit 1307 adds the prediction residual after inverse transformation generated by the inverse transform unit 1306 and the prediction volume generated by intra-frame prediction or inter-frame prediction described later, which is used in generating the prediction residual before quantization, to generate a reconstructed volume. The reconstructed volume is stored in the reference volume memory 1308 or the reference space memory 1310.

[0543] The intra prediction unit 1309 generates a predicted volume of the encoding target volume using the attribute information of the adjacent volume stored in the reference volume memory 1308. The attribute information includes the color information or reflectivity of the voxel. The intra prediction unit 1309 generates a predicted value of the color information or reflectivity of the encoding target volume.

[0544] Fig.44 1309 is a diagram for explaining the operation of the intra prediction unit 1309. For example, Fig.44 As shown, the intra-frame prediction unit 1309 generates a predicted volume of the encoding target volume (volume idx=3) based on the adjacent volume (volume idx=0). Here, volume idx is identifier information added to the volume in the space, and different values ​​are assigned to each volume. The order of assigning volume idx can be the same as the encoding order or different from the encoding order. For example, as Fig.44 The intra prediction unit 1309 uses the average value of the color information of the voxels contained in the volume idx=0 which is the adjacent volume as the predicted value of the color information of the encoding object volume shown. In this case, the prediction residual is generated by subtracting the predicted value of the color information from the color information of each voxel contained in the encoding object volume. The processing after the transformation unit 1303 is performed on the prediction residual. And, in this case, the three-dimensional data encoding device 1300 adds the adjacent volume information and the prediction mode information to the bit stream. Here, the adjacent volume information is information showing the adjacent volume used in the prediction, for example, the volume idx showing the adjacent volume used in the prediction. And the prediction mode information shows the mode used in the generation of the prediction volume. The mode refers to, for example, an average value mode for generating a prediction value based on the average value of the voxels in the adjacent volume, or an intermediate value mode for generating a prediction value based on the intermediate value of the voxels in the adjacent volume.

[0545] The intra-frame prediction unit 1309 may also generate a prediction volume based on a plurality of adjacent volumes. Fig.44 In the configuration shown, the intra prediction unit 1309 generates prediction volume 0 based on volume idx=0, and generates prediction volume 1 based on volume idx=1. Then, the intra prediction unit 1309 generates the final prediction volume by averaging the prediction volume 0 and the prediction volume 1. In this case, the three-dimensional data encoding device 1300 may also attach multiple volume idxs of the multiple volumes used in generating the prediction volume to the bitstream.

[0546] Fig.45 The inter-frame prediction process involved in this embodiment is shown in the mode. The inter-frame prediction unit 1311 uses the coded space at a different time T_LX to encode (inter-frame prediction) for the space (SPC) at a certain time T_Cur. In this case, the inter-frame prediction unit 1311 applies rotation and translation processing to the coded space at different time T_LX to perform encoding processing.

[0547] Furthermore, the three-dimensional data encoding device 1300 adds RT information related to the rotation and translation processing of the space at the different time T_LX to the bitstream. The different time T_LX is, for example, the time T_L0 before the certain time T_Cur. At this time, the three-dimensional data encoding device 1300 may also add RT information RT_L0 related to the rotation and translation processing of the space at the time T_L0 to the bitstream.

[0548] Alternatively, the different time T_LX is, for example, time T_L1 after the certain time T_Cur. In this case, the three-dimensional data encoding apparatus 1300 may add RT information RT_L1 about the spatial rotation and translation processing applied at time T_L1 to the bit stream.

[0549] Alternatively, the inter prediction unit 1311 performs encoding by referring to spaces at different times T_L0 and T_L1 (bi-prediction). In this case, the three-dimensional data encoding device 1300 may add both RT information RT_L0 and RT_L1 about rotation and translation applied to the space to the bitstream.

[0550] In addition, although T_L0 is set as the time before T_Cur and T_L1 is set as the time after T_Cur, it is not limited to this. For example, T_L0 and T_L1 can both be the time before T_Cur. Or, T_L0 and T_L1 can both be the time after T_Cur.

[0551] Furthermore, the three-dimensional data encoding device 1300 may add RT information related to the rotation and translation applied to each space to the bitstream when encoding with reference to multiple spaces at different times. For example, the three-dimensional data encoding device 1300 manages the multiple encoded spaces referred to by two reference lists (L0 list and L1 list). When the first reference space in the L0 list is set to L0R0, the second reference space in the L0 list is set to L0R1, the first reference space in the L1 list is set to L1R0, and the second reference space in the L1 list is set to L1R1, the three-dimensional data encoding device 1300 adds RT information RT_L0R0 of L0R0, RT information RT_L0R1 of L0R1, RT information RT_L1R0 of L1R0, and RT information RT_L1R1 of L1R1 to the bitstream. For example, the three-dimensional data encoding device 1300 adds these RT information to the header of the bitstream, etc.

[0552] Furthermore, the three-dimensional data encoding device 1300 may determine whether rotation and translation are applicable to each reference space when encoding with reference to a plurality of reference spaces at different times. At this time, the three-dimensional data encoding device 1300 may attach information indicating whether rotation and translation are applicable to each reference space (RT applicable flag, etc.) to the header information of the bitstream, etc. For example, the three-dimensional data encoding device 1300 calculates RT information and ICP error value using the ICP (Interactive Closest Point) algorithm according to the encoding object space and each reference space to be referred to. When the ICP error value is below a predetermined certain value, the three-dimensional data encoding device 1300 determines that rotation and translation are not required and sets the RT applicable flag to OFF (invalid). In addition, when the ICP error value is larger than the above-mentioned certain value, the three-dimensional data encoding device 1300 sets the RT applicable flag to ON (valid) and attaches the RT information to the bitstream.

[0553] Fig.46 An example of a syntax in which RT information and an RT applicable flag are attached to a header is shown. In addition, the number of bits allocated to each syntax can be determined based on the range that the syntax can take. For example, when the number of reference spaces included in the reference list L0 is 8, 3 bits can be allocated in MaxRefSpc_l0. The number of allocated bits can be changed according to the values ​​that each syntax can take, or the number of allocated bits can be fixed without being affected by the possible values. In the case of fixing the number of allocated bits, the three-dimensional data encoding device 1300 can attach the fixed number of bits to other header information.

[0554] Here, Fig.46MaxRefSpc_10 shown shows the number of reference spaces included in reference list L0. RT_flag_10[i] is the RT application flag of reference space i in reference list L0. When RT_flag_10[i] is 1, rotation and translation are applied to reference space i. When RT_flag_10[i] is 0, rotation and translation are not applied to reference space i.

[0555] R_l0[i] and T_l0[i] are RT information of reference space i in reference list L0. R_l0[i] is rotation information of reference space i in reference list L0. The rotation information indicates the content of the applied rotation processing, such as a rotation matrix or quaternion. T_l0[i] is translation information of reference space i in reference list L0. The translation information indicates the content of the applied translation processing, such as a translation vector.

[0556] MaxRefSpc_l1 indicates the number of reference spaces included in the reference list L1. RT_flag_l1[i] is the RT application flag for the reference space i in the reference list L1. When RT_flag_l1[i] is 1, rotation and translation are applied to the reference space i. When RT_flag_l1[i] is 0, rotation and translation are not applied to the reference space i.

[0557] R_l1[i] and T_l1[i] are RT information of reference space i in reference list L1. R_l1[i] is rotation information of reference space i in reference list L1. The rotation information indicates the content of the applied rotation processing, such as a rotation matrix or quaternion. T_l1[i] is translation information of reference space i in reference list L1. The translation information indicates the content of the applied translation processing, such as a translation vector.

[0558] The inter-frame prediction unit 1311 generates a predicted volume of the encoding target volume using the information of the encoded reference space stored in the reference space memory 1310. As described above, before generating the predicted volume of the encoding target volume, the inter-frame prediction unit 1311 uses the ICP (Interactive Closest Point) algorithm to obtain RT information in the encoding target space and the reference space in order to make the positional relationship between the encoding target space and the reference space as a whole close. Then, the inter-frame prediction unit 1311 uses the obtained RT information to apply rotation and translation processing to the reference space to obtain the reference space B. Thereafter, the inter-frame prediction unit 1311 uses the information in the reference space B to generate a predicted volume of the encoding target volume in the encoding target space. Here, the three-dimensional data encoding device 1300 attaches the RT information used to obtain the reference space B to the header information of the encoding target space, etc.

[0559] In this way, the inter-frame prediction unit 1311 applies rotation and translation processing to the reference space, so as to make the overall positional relationship between the encoding object space and the reference space close, and then uses the information of the reference space to generate a prediction volume, thereby improving the accuracy of the prediction volume. In addition, since the prediction residual can be suppressed, the amount of encoding can be reduced. In addition, although an example of using the encoding object space and the reference space to perform ICP is shown here, it is not limited to this. For example, in order to reduce the amount of processing, the inter-frame prediction unit 1311 can also use at least one of the encoding object space from which the number of voxels or point clouds is extracted, and the reference space from which the number of voxels or point clouds is extracted to perform ICP, thereby obtaining RT information.

[0560] Furthermore, when the ICP error value obtained from the ICP result is smaller than a predetermined first threshold value, that is, when the positional relationship between the encoding object space and the reference space is close, the inter-frame prediction unit 1311 may determine that rotation and translation processing are not required, and does not perform rotation and translation. In this case, the three-dimensional data encoding device 1300 may not add RT information to the bitstream, thereby suppressing additional overhead.

[0561] Furthermore, when the ICP error value is greater than a predetermined second threshold, the inter-frame prediction unit 1311 determines that the shape change in space is large, and intra-frame prediction can be applied to all volumes of the encoding object space. Hereinafter, the space to which intra-frame prediction is applied is referred to as intra-frame space. Furthermore, the second threshold is a value greater than the above-mentioned first threshold. Furthermore, it is not limited to ICP, and any method can be applied as long as the method of obtaining RT information from two voxel sets or two point cloud sets.

[0562] Furthermore, when the three-dimensional data contains attribute information such as shape or color, the inter-frame prediction unit 1311 searches, for example, a volume in the reference space that is closest to the shape or color attribute information of the encoding target volume as a prediction volume of the encoding target volume in the encoding target space. Furthermore, the reference space is, for example, a reference space that has been subjected to the above-mentioned rotation and translation processing. The inter-frame prediction unit 1311 generates a prediction volume based on the volume (reference volume) obtained by the search. Fig.47 is a diagram for explaining the generation of the prediction volume. Fig.47When the encoding target volume (volume idx=0) shown in the figure is encoded by using inter-frame prediction, the reference volumes in the reference space are scanned in sequence while searching for the volume with the smallest prediction residual, that is, the difference between the encoding target volume and the reference volume. The inter-frame prediction unit 1311 selects the volume with the smallest prediction residual as the prediction volume. The prediction residual between the encoding target volume and the prediction volume is encoded by the processing after the transformation unit 1303. Here, the prediction residual refers to the difference between the attribute information of the encoding target volume and the attribute information of the prediction volume. In addition, the three-dimensional data encoding device 1300 adds the volume idx of the reference volume in the reference space referred to as the prediction volume to the header of the bit stream, etc.

[0563] exist Fig.47 In the example shown, the reference volume idx=4 of the reference space L0R0 is selected as the prediction volume of the encoding target volume. Then, the prediction residual between the encoding target volume and the reference volume and the reference volume idx=4 are encoded and added to the bit stream.

[0564] In addition, although the description here is made by taking the generation of the predicted volume of the attribute information as an example, the same processing can be performed on the predicted volume of the position information.

[0565] The prediction control unit 1312 controls whether intra-frame prediction or inter-frame prediction is used to encode the encoding target volume. Here, a mode including intra-frame prediction and inter-frame prediction is referred to as a prediction mode. For example, the prediction control unit 1312 calculates the prediction residual when the encoding target volume is predicted by intra-frame prediction and the prediction residual when it is predicted by inter-frame prediction as an evaluation value, and selects the prediction mode with the smaller evaluation value. Alternatively, the prediction control unit 1312 may apply orthogonal transformation, quantization, and entropy coding to the prediction residual of intra-frame prediction and the prediction residual of inter-frame prediction, respectively, to calculate the actual coding amount, and select the prediction mode using the calculated coding amount as the evaluation value. In addition, overhead information (reference volume idx information, etc.) other than the prediction residual may be added to the evaluation value. In addition, the prediction control unit 1312 may also generally select intra-frame prediction when the encoding target space is predetermined to be encoded in the intra-frame space.

[0566] The entropy coding unit 1313 generates a coded signal (coded bit stream) by performing variable length coding on the quantized coefficients input from the quantization unit 1304. Specifically, the entropy coding unit 1313 binarizes the quantized coefficients and performs arithmetic coding on the obtained binary signal, for example.

[0567] Next, a three-dimensional data decoding device that decodes the encoded signal generated by the three-dimensional data encoding device 1300 will be described. Fig.4814 is a block diagram of a three-dimensional data decoding device 1400 according to the present embodiment. The three-dimensional data decoding device 1400 includes an entropy decoding unit 1401, an inverse quantization unit 1402, an inverse transformation unit 1403, an addition unit 1404, a reference volume memory 1405, an intra-frame prediction unit 1406, a reference space memory 1407, an inter-frame prediction unit 1408, and a prediction control unit 1409.

[0568] The entropy decoding unit 1401 performs variable length decoding on the coded signal (coded bit stream). For example, the entropy decoding unit 1401 performs arithmetic decoding on the coded signal to generate a binary signal, and generates a quantized coefficient based on the generated binary signal.

[0569] The inverse quantization unit 1402 performs inverse quantization on the quantized coefficients input from the entropy decoding unit 1401 using a quantization parameter added to a bit stream or the like, thereby generating inverse quantized coefficients.

[0570] The inverse transform unit 1403 generates a prediction residual by performing an inverse transform on the inverse quantized coefficient input from the inverse quantization unit 1402. For example, the inverse transform unit 1403 generates a prediction residual by performing an inverse orthogonal transform on the inverse quantized coefficient based on information added to the bit stream.

[0571] The adder 1404 adds the prediction residual generated by the inverse transform unit 1403 and the prediction volume generated by intra prediction or inter prediction to generate a reconstructed volume. The reconstructed volume is output as decoded three-dimensional data and stored in the reference volume memory 1405 or the reference space memory 1407.

[0572] The intra prediction unit 1406 generates a prediction volume by intra prediction using the reference volume in the reference volume memory 1405 and the information added to the bitstream. Specifically, the intra prediction unit 1406 obtains the prediction mode information and the adjacent volume information (e.g., volume idx) added to the bitstream, and generates a prediction volume using the adjacent volume indicated by the adjacent volume information in the mode indicated by the prediction mode information. In addition, the details of these processes are the same as the processes of the intra prediction unit 1309 described above, except that the information added to the bitstream is used.

[0573] The inter-frame prediction unit 1408 generates a prediction volume through inter-frame prediction using the reference space in the reference space memory 1407 and the information attached to the bitstream. Specifically, the inter-frame prediction unit 1408 uses the RT information of each reference space attached to the bitstream, applies rotation and translation processing to the reference space, and generates a prediction volume using the applied reference space. In addition, when the RT applicable flag of each reference space exists in the bitstream, the inter-frame prediction unit 1408 applies rotation and translation processing to the reference space according to the RT applicable flag. In addition, the details of the above-mentioned processing are the same as the processing of the above-mentioned inter-frame prediction unit 1311, except for using the information attached to the bitstream.

[0574] Whether to decode the decoding target volume by intra prediction or inter prediction is controlled by the prediction control unit 1409. For example, the prediction control unit 1409 selects intra prediction or inter prediction according to information indicating the prediction mode to be used, which is added to the bit stream. In addition, when it is predetermined that the decoding target space is decoded as the intra space, the prediction control unit 1409 may normally select the intra prediction.

[0575] The following is a description of a variation of the present embodiment. In the present embodiment, although the application of rotation and translation in spatial units is described as an example, rotation and translation in finer units may also be applied. For example, the three-dimensional data encoding device 1300 may divide the space into subspaces, and apply rotation and translation in subspace units. In this case, the three-dimensional data encoding device 1300 generates RT information according to each subspace, and attaches the generated RT information to the header of the bitstream, etc. In addition, the three-dimensional data encoding device 1300 may use volume units as encoding units to apply rotation and translation. In this case, the three-dimensional data encoding device 1300 generates RT information in encoding volume units, and attaches the generated RT information to the header of the bitstream, etc. In addition, the above may be combined. That is, the three-dimensional data encoding device 1300 may apply rotation and translation in finer units after applying rotation and translation in large units. For example, the three-dimensional data encoding device 1300 may apply rotation and translation in spatial units, and apply different rotations and translations to each of the multiple volumes contained in the obtained space.

[0576] Furthermore, although the present embodiment is described by taking the application of rotation and translation to the reference space as an example, it is not limited thereto. For example, the three-dimensional data encoding device 1300 may apply scaling processing to change the size of the three-dimensional data. Furthermore, the three-dimensional data encoding device 1300 may also apply any one or two of rotation, translation, and scaling. Furthermore, as described above, when the processing is applied in different units in multiple stages, the type of processing applied in each unit may be different. For example, rotation and translation may be applied in the spatial unit, and translation may be applied in the volume unit.

[0577] Note that these modifications are also applicable to the three-dimensional data decoding device 1400 .

[0578] As described above, the three-dimensional data encoding device 1300 according to this embodiment performs the following processing. Fig.48 3D data encoding apparatus 1300 performs an inter-frame prediction process.

[0579] First, the three-dimensional data encoding device 1300 generates predicted position information (e.g., predicted volume) using the position information of three-dimensional points included in the target three-dimensional data (e.g., encoding target space) and the reference three-dimensional data (e.g., reference space) at different times (S1301). Specifically, the three-dimensional data encoding device 1300 generates the predicted position information by applying rotation and translation processing to the position information of the three-dimensional points included in the reference three-dimensional data.

[0580] In addition, the three-dimensional data encoding device 1300 performs rotation and translation processing with a first unit (e.g., space), and generates predicted position information with a second unit (e.g., volume) that is smaller than the first unit. For example, the three-dimensional data encoding device 1300 searches for a volume in which the difference between the encoding object volume and the position information contained in the encoding object space is the smallest from among a plurality of volumes contained in the reference space after the rotation and translation processing, and uses the obtained volume as the predicted volume. In addition, the three-dimensional data encoding device 1300 may perform the rotation and translation processing and the generation of the predicted position information with the same unit.

[0581] Furthermore, the three-dimensional data encoding device 1300 may apply a first rotation and translation process to the position information of the three-dimensional points contained in the reference three-dimensional data using a first unit (e.g., space), and may apply a second rotation and translation process to the position information of the three-dimensional points obtained by the first rotation and translation process using a second unit (e.g., volume) that is smaller than the first unit, thereby generating predicted position information.

[0582] Here, the position information of the three-dimensional point and the predicted position information are as follows: Fig.41As shown, the octree structure is used for representation. For example, the position information and predicted position information of the three-dimensional point are represented in a scanning order that prioritizes the width among the depth and width in the octree structure. Alternatively, the position information and predicted position information of the three-dimensional point are represented in a scanning order that prioritizes the depth and width among the depth and width in the octree structure.

[0583] And, if Fig.46 As shown, the three-dimensional data encoding device 1300 encodes the RT applicable flag indicating whether the rotation and translation processing is applied to the position information of the three-dimensional point included in the reference three-dimensional data. That is, the three-dimensional data encoding device 1300 generates a coded signal (coded bit stream) including the RT applicable flag. In addition, the three-dimensional data encoding device 1300 encodes the RT information indicating the content of the rotation and translation processing. That is, the three-dimensional data encoding device 1300 generates a coded signal (coded bit stream) including the RT information. Alternatively, the three-dimensional data encoding device 1300 may encode the RT information when the RT applicable flag indicates that the rotation and translation processing is applicable, and may not encode the RT information when the RT applicable flag indicates that the rotation and translation processing is not applicable.

[0584] The three-dimensional data includes, for example, position information of three-dimensional points and attribute information (color information, etc.) of each three-dimensional point. The three-dimensional data encoding device 1300 generates predicted attribute information using the attribute information of the three-dimensional points included in the reference three-dimensional data (S1302).

[0585] Next, the three-dimensional data encoding device 1300 uses the predicted position information to encode the position information of the three-dimensional points included in the object three-dimensional data. Fig.38 As shown, differential position information which is a difference between the position information of the three-dimensional point included in the target three-dimensional data and the predicted position information is calculated (S1303).

[0586] Furthermore, the three-dimensional data encoding device 1300 uses the predicted attribute information to encode the attribute information of the three-dimensional points included in the target three-dimensional data. For example, the three-dimensional data encoding device 1300 calculates the difference between the attribute information of the three-dimensional points included in the target three-dimensional data and the predicted attribute information, that is, differential attribute information (S1304). Next, the three-dimensional data encoding device 1300 transforms and quantizes the calculated differential attribute information (S1305).

[0587] Finally, the 3D data encoding device 1300 encodes (for example, entropy encoding) the differential position information and the quantized differential attribute information (S1306). That is, the 3D data encoding device 1300 generates an encoded signal (encoded bit stream) including the differential position information and the differential attribute information.

[0588] In addition, when the three-dimensional data does not include attribute information, the three-dimensional data encoding device 1300 may not perform steps S1302, S1304, and S1305. In addition, the three-dimensional data encoding device 1300 may only encode the position information of the three-dimensional point or encode the attribute information of the three-dimensional point.

[0589] and, Fig.49 The processing sequence shown is only an example and is not limited thereto. For example, since the processing of location information (S1301, S1303) and the processing of attribute information (S1302, S1304, S1305) are independent of each other, they can be executed in any order, or some of them can be processed in parallel.

[0590] As described above, in this embodiment, the three-dimensional data encoding device 1300 generates predicted position information using the position information of the three-dimensional points included in the target three-dimensional data and the reference three-dimensional data at different times, and encodes the difference between the position information of the three-dimensional points included in the target three-dimensional data and the predicted position information, that is, the differential position information. Accordingly, since the amount of data of the encoded signal can be reduced, the encoding efficiency can be improved.

[0591] Furthermore, in this embodiment, the three-dimensional data encoding device 1300 generates predicted attribute information using attribute information of three-dimensional points included in the reference three-dimensional data, and encodes the difference between the attribute information of three-dimensional points included in the target three-dimensional data and the predicted attribute information, that is, the differential attribute information. Accordingly, since the amount of data of the encoded signal can be reduced, the encoding efficiency can be improved.

[0592] For example, the three-dimensional data encoding device 1300 includes a processor and a memory, and the processor performs the above-mentioned processing using the memory.

[0593] Fig.48 3D data decoding apparatus 1400 performs an inter-frame prediction process.

[0594] First, the three-dimensional data decoding apparatus 1400 decodes (eg, performs entropy decoding) the difference position information and the difference attribute information according to the coded signal (coded bit stream) ( S1401 ).

[0595] Furthermore, the three-dimensional data decoding device 1400 decodes the RT applicable flag indicating whether the rotation and translation processing is applied to the position information of the three-dimensional point included in the reference three-dimensional data according to the coded signal. Furthermore, the three-dimensional data decoding device 1400 decodes the RT information indicating the content of the rotation and translation processing. In addition, the three-dimensional data decoding device 1400 decodes the RT information when the RT applicable flag indicates that the rotation and translation processing is applied, and does not decode the RT information when the RT applicable flag indicates that the rotation and translation processing is not applied.

[0596] Next, the three-dimensional data decoding apparatus 1400 performs inverse quantization and inverse transformation on the decoded differential attribute information ( S1402 ).

[0597] Next, the three-dimensional data decoding device 1400 generates predicted position information (e.g., predicted volume) using the position information of the three-dimensional points included in the target three-dimensional data (e.g., decoding target space) and the reference three-dimensional data (e.g., reference space) at different times (S1403). Specifically, the three-dimensional data decoding device 1400 generates the predicted position information by applying rotation and translation processing to the position information of the three-dimensional points included in the reference three-dimensional data.

[0598] More specifically, when the RT application flag indicates that the rotation and translation processing is applicable, the three-dimensional data decoding device 1400 applies the rotation and translation processing to the position information of the three-dimensional point included in the reference three-dimensional data indicated by the RT information. And when the RT application flag indicates that the rotation and translation processing is not applicable, the three-dimensional data decoding device 1400 does not apply the rotation and translation processing to the position information of the three-dimensional point included in the reference three-dimensional data.

[0599] In addition, the three-dimensional data decoding device 1400 may perform the rotation and translation processing in a first unit (e.g., space), and may generate the predicted position information in a second unit (e.g., volume) that is smaller than the first unit. In addition, the three-dimensional data decoding device 1400 may also perform the rotation and translation processing, and the generation of the predicted position information in the same unit.

[0600] Furthermore, the three-dimensional data decoding device 1400 may apply a first rotation and translation process with a first unit (e.g., space) to the position information of the three-dimensional points included in the reference three-dimensional data, and may apply a second rotation and translation process with a second unit (e.g., volume) that is smaller than the first unit to the position information of the three-dimensional points obtained by the first rotation and translation process, thereby generating predicted position information.

[0601] Here, the position information of the three-dimensional point and the predicted position information are as follows: Fig.41As shown, the octree structure is used for representation. For example, the position information and predicted position information of the three-dimensional point are represented in a scanning order that prioritizes the width among the depth and width in the octree structure. Alternatively, the position information and predicted position information of the three-dimensional point are represented in a scanning order that prioritizes the depth among the depth and width in the octree structure.

[0602] The three-dimensional data decoding apparatus 1400 generates predicted attribute information using attribute information of three-dimensional points included in the reference three-dimensional data ( S1404 ).

[0603] Next, the three-dimensional data decoding device 1400 decodes the coded position information contained in the coded signal by using the predicted position information, thereby restoring the position information of the three-dimensional point contained in the target three-dimensional data. Here, the coded position information is, for example, differential position information, and the three-dimensional data decoding device 1400 restores the position information of the three-dimensional point contained in the target three-dimensional data by adding the differential position information to the predicted position information (S1405).

[0604] Furthermore, the three-dimensional data decoding device 1400 decodes the coded attribute information included in the coded signal by using the predicted attribute information, thereby restoring the attribute information of the three-dimensional point included in the target three-dimensional data. Here, the coded attribute information is, for example, differential attribute information, and the three-dimensional data decoding device 1400 restores the attribute information of the three-dimensional point included in the target three-dimensional data by adding the differential attribute information to the predicted attribute information (S1406).

[0605] Alternatively, if the three-dimensional data does not include attribute information, the three-dimensional data decoding device 1400 may not perform steps S1402, S1404, and S1406. Furthermore, the three-dimensional data decoding device 1400 may only decode the position information of the three-dimensional point or decode the attribute information of the three-dimensional point.

[0606] and, Fig.50 The order of processing shown is an example and is not limited to this. For example, since the processing of location information (S1403, S1405) and the processing of attribute information (S1402, S1404, S1406) are independent of each other, they can be performed in any order, and some of them can be processed in parallel.

[0607] (Implementation 8)

[0608] In this embodiment, a method of expressing three-dimensional points (point cloud) in encoding of three-dimensional data will be described.

[0609] Fig.51 This is a block diagram showing the configuration of a three-dimensional data distribution system according to the present embodiment. Fig.51 The distribution system shown includes a server 1501 and multiple clients 1502 .

[0610] The server 1501 includes a storage unit 1511 and a control unit 1512. The storage unit 1511 stores a coded three-dimensional map 1513 which is coded three-dimensional data.

[0611] Fig.52 An example of the composition of a bit stream for encoding a three-dimensional map 1513 is shown. The three-dimensional map is divided into a plurality of sub-maps, and each sub-map is encoded. A random access header (RA) including sub-coordinate information is added to each sub-map. The sub-coordinate information is used to improve the coding efficiency of the sub-map. The sub-coordinate information shows the sub-coordinate of the sub-map. The sub-coordinate is the coordinate of the sub-map based on the reference coordinate. In addition, a three-dimensional map including a plurality of sub-maps is referred to as an entire map. And, in the entire map, the coordinates (e.g., the origin) that will serve as a reference are referred to as reference coordinates. That is, the sub-coordinate is the coordinate of the sub-map in the coordinate system of the entire map. In other words, the sub-coordinate shows the deviation between the coordinate system of the entire map and the coordinate system of the sub-map. And, the coordinate in the coordinate system of the entire map based on the reference coordinate is referred to as an entire coordinate. The coordinate in the coordinate system of the sub-map based on the sub-coordinate is referred to as a differential coordinate.

[0612] The client 1502 sends a message to the server 1501. The message includes the location information of the client 1502. The control unit 1512 included in the server 1501 obtains the bit stream of the submap at the location closest to the location of the client 1502 based on the location information included in the received message. The bit stream of the submap includes sub-coordinate information and is sent to the client 1502. The decoder 1521 included in the client 1502 uses the sub-coordinate information to obtain the overall coordinates of the submap based on the reference coordinates. The application 1522 included in the client 1502 uses the obtained overall coordinates of the submap to execute the application related to its own location.

[0613] Furthermore, the submap shows a part of the entire map. The sub-coordinates are the coordinates of the position of the submap in the reference coordinate space of the entire map. For example, it is considered that there is a submap A of AA and a submap B of AB in the entire map of A. When the vehicle wants to refer to the map of AA, it starts decoding from submap A, and when it wants to refer to the map of AB, it starts decoding from submap B. Here, the submap is a random access point. Specifically, A is Osaka Prefecture, AA is Osaka City, AB is Takatsuki City, etc.

[0614] Each sub-map is sent to the client together with the sub-coordinate information. The sub-coordinate information is included in the header information of each sub-map or in the transmission data packet.

[0615] The reference coordinates that serve as the reference coordinates for the sub-coordinate information of each sub-map may be added to header information of a space higher than the sub-map, such as header information of the entire map.

[0616] A submap can be composed of one space (SPC). Also, a submap can be composed of multiple SPCs.

[0617] Furthermore, a submap may also include a GOS (Group of Space). Furthermore, a submap may also be composed of a world space. For example, in the case where there are multiple objects in a submap, if the multiple objects are assigned to different SPCs, the submap is composed of multiple SPCs. Furthermore, if the multiple objects are assigned to one SPC, the submap is composed of one SPC.

[0618] Next, the effect of improving the coding efficiency when the sub-coordinate information is used will be described. Fig.53 is a diagram used to illustrate this effect. Fig.53 To encode the three-dimensional point A at a position far from the reference coordinates shown in the figure, a larger number of bits are required. Here, the distance between the sub-coordinates and the three-dimensional point A is shorter than the distance between the reference coordinates and the three-dimensional point A. Therefore, compared with the case of encoding the coordinates of the three-dimensional point A based on the reference coordinates, the encoding efficiency can be improved when encoding the coordinates of the three-dimensional point A based on the sub-coordinates. In addition, the bit stream of the submap includes the sub-coordinate information. By sending the bit stream of the submap and the reference coordinates to the decoding side (client), the overall coordinates of the submap can be restored on the decoding side.

[0619] Fig.54 This is a flowchart of the processing performed by the server 1501 which is the transmission side of the submap.

[0620] First, the server 1501 receives a message including the location information of the client 1502 from the client 1502 (S1501). The control unit 1512 obtains the coded bit stream of the submap based on the location information of the client from the storage unit 1511 (S1502). Then, the server 1501 sends the coded bit stream of the submap and the reference coordinates to the client 1502 (S1503).

[0621] Fig.55 This is a flowchart of the processing performed by the client 1502 which is the receiving side of the submap.

[0622] First, the client 1502 receives the coded bit stream and reference coordinates of the submap sent from the server 1501 (S1511). Next, the client 1502 decodes the coded bit stream to obtain the submap and sub-coordinate information (S1512). Next, the client 1502 uses the reference coordinates and sub-coordinates to restore the differential coordinates in the submap to the overall coordinates (S1513).

[0623] Next, a syntax example of information related to the submap is described. In the encoding of the submap, the three-dimensional data encoding device calculates the differential coordinates by subtracting the sub-coordinates from the coordinates of each point cloud (three-dimensional point). Then, the three-dimensional data encoding device encodes the differential coordinates into a bit stream as the value of each point cloud. And, the encoding device encodes the sub-coordinate information showing the sub-coordinates as the header information of the bit stream. Based on this, the three-dimensional data decoding device can obtain the overall coordinates of each point cloud. For example, the three-dimensional data encoding device is included in the server 1501, and the three-dimensional data decoding device is included in the client 1502.

[0624] Fig.56 An example of the syntax of a submap is shown. Fig.56 The NumOfPoint shown indicates the number of point clouds included in the submap. sub_coordinate_x, sub_coordinate_y, and sub_coordinate_z are sub-coordinate information. sub_coordinate_x indicates the x-coordinate of the sub-coordinate. sub_coordinate_y indicates the y-coordinate of the sub-coordinate. sub_coordinate_z indicates the z-coordinate of the sub-coordinate.

[0625] Furthermore, diff_x[i], diff_y[i], and diff_z[i] are the differential coordinates of the i-th point cloud in the submap. diff_x[i] represents the difference between the x-coordinate of the i-th point cloud in the submap and the x-coordinate of the submap. diff_y[i] represents the difference between the y-coordinate of the i-th point cloud in the submap and the y-coordinate of the submap. diff_z[i] represents the difference between the z-coordinate of the i-th point cloud in the submap and the z-coordinate of the submap.

[0626] The three-dimensional data decoding device decodes point_cloud[i]_x, point_cloud[i]_y, and point_cloud[i]_z, which are the overall coordinates of the i-th point cloud, using the following formula. point_cloud[i]_x is the x coordinate of the overall coordinates of the i-th point cloud. point_cloud[i]_y is the y coordinate of the overall coordinates of the i-th point cloud. point_cloud[i]_z is the z coordinate of the overall coordinates of the i-th point cloud.

[0627] point_cloud[i]_x=sub_coordinate_x+diff_x[i]

[0628] point_cloud[i]_y=sub_coordinate_y+diff_y[i]

[0629] point_cloud[i]_z=sub_coordinate_z+diff_z[i]

[0630] Next, the applicable switching process of octree coding is described. When performing sub-map coding, the three-dimensional data coding device either selects octree representation to encode each point cloud (hereinafter referred to as octree coding (octree coding)), or selects to encode the difference value with the sub-coordinate (hereinafter referred to as non-octree coding (non-octree coding)). Fig.57 This operation is shown in a pattern. For example, when the number of point clouds in a submap is greater than a predetermined threshold, the three-dimensional data encoding device applies octree encoding to the submap. When the number of point clouds in a submap is less than the above threshold, the three-dimensional data encoding device applies non-octree encoding to the submap. Accordingly, the three-dimensional data encoding device appropriately selects whether to use octree encoding or non-octree encoding according to the shape and density of the objects contained in the submap, thereby improving the encoding efficiency.

[0631] Furthermore, the three-dimensional data encoding device adds information indicating whether octree encoding or non-octree encoding is applicable to the submap (hereinafter referred to as octree encoding application information) to the header of the submap, etc. Based on this, the three-dimensional data decoding device can determine whether the bit stream is a bit stream obtained by octree encoding the submap or a bit stream obtained by non-octree encoding the submap.

[0632] Furthermore, the three-dimensional data encoding device can calculate the encoding efficiency when octree encoding and non-octree encoding are applied to the same point cloud respectively, and apply the encoding method with higher encoding efficiency to the sub-map.

[0633] Fig.58 An example of the syntax of a submap in the case of performing such switching is shown. Fig.58 The coding_type shown is information indicating the coding type, and is the above-mentioned octree coding application information. coding_type=00 indicates that octree coding is applied. coding_type=01 indicates that non-octree coding is applied. coding_type=10 or 11 indicates that other coding methods other than the above are applied.

[0634] In the case where the encoding type is non-octree encoding (non_octree), the submap includes NumOfPoint, and sub-coordinate information (sub_coordinate_x, sub_coordinate_y, and sub_coordinate_z).

[0635] When the encoding type is octree encoding (octree), the submap includes octree_info. Octree_info is information required in octree encoding, such as depth information.

[0636] In the case where the encoding type is non-octree encoding (non_octree), the submap includes differential coordinates (diff_x[i], diff_y[i], and diff_z[i]).

[0637] When the encoding type is octree encoding (octree), the submap includes encoding data related to the octree encoding, namely octree_data.

[0638] In addition, although an example in which the xyz coordinate system is used as the coordinate system of the point cloud is shown here, a polar coordinate system may also be used.

[0639] Fig.59 3D data encoding process performed by the 3D data encoding device. First, the 3D data encoding device calculates the number of point clouds in the submap of the processing object, that is, the object submap (S1521). Then, the 3D data encoding device determines whether the calculated number of point clouds is above a predetermined threshold (S1522).

[0640] When the number of point clouds is greater than the threshold value (Yes in S1522), the three-dimensional data encoding device applies octree encoding to the object submap (S1523). Furthermore, the three-dimensional point data encoding device adds octree encoding application information indicating that octree encoding is applied to the object submap to the header of the bitstream (S1525).

[0641] In addition, when the number of point clouds is lower than the threshold value (No in S1522), the three-dimensional data encoding device applies non-octree encoding to the object submap (S1524). Furthermore, the three-dimensional point data encoding device adds octree encoding application information indicating that non-octree encoding is applied to the object submap to the header of the bitstream (S1525).

[0642] Fig.603D data decoding processing performed by a 3D data decoding device. First, the 3D data decoding device decodes octree coding application information from the header of the bit stream (S1531). Next, the 3D data decoding device determines whether the coding type applied to the object submap is octree coding based on the decoded octree coding application information (S1532).

[0643] When the encoding type indicated by the octree encoding application information is octree encoding ("Yes" in S1532), the three-dimensional data decoding device uses octree decoding to decode the object submap (S1533). In addition, when the encoding type indicated by the octree encoding application information is non-octree encoding ("No" in S1532), the three-dimensional data decoding device uses non-octree decoding to decode the object submap (S1534).

[0644] Modifications of this embodiment will be described below. Figure 61 to Figure 63 The operation of a modified example of the switching process of the encoding type is shown in a schematic manner.

[0645] like Fig.61 As shown, the three-dimensional data encoding device can select whether to apply octree encoding or non-octree encoding according to each space. In this case, the three-dimensional data encoding device adds octree encoding application information to the header of the space. Based on this, the three-dimensional data decoding device can determine whether to apply octree encoding according to each space. And, in this case, the three-dimensional data encoding device sets sub-coordinates according to each space, and encodes the difference value after subtracting the value of the sub-coordinate from the coordinates of each point cloud in the space.

[0646] According to this, since the three-dimensional data encoding device can appropriately switch whether to apply octree encoding according to the shape of the object in the space or the number of point clouds, the encoding efficiency can be improved.

[0647] And, if Fig.62 As shown, the three-dimensional data encoding device can select whether to apply octree encoding or non-octree encoding for each volume. In this case, the three-dimensional data encoding device adds octree encoding application information to the head of the volume. Based on this, the three-dimensional data decoding device can determine whether to apply octree encoding for each volume. And, in this case, the three-dimensional data encoding device sets sub-coordinates for each volume, and encodes the difference value after subtracting the value of the sub-coordinate from the coordinates of each point cloud in the volume.

[0648] Accordingly, since the three-dimensional data encoding device can appropriately switch whether to apply octree encoding according to the shape of the object or the number of point clouds within the volume, the encoding efficiency can be improved.

[0649] Furthermore, in the above description, an example of encoding the difference after subtracting the sub-coordinates from the coordinates of each point cloud is shown as non-octree encoding, but the present invention is not limited thereto, and any encoding method other than octree encoding may be used for encoding. Fig.63 As shown, the three-dimensional data encoding device may not use the difference with the sub-coordinates, but may use the method of encoding the value of the point cloud within the sub-map, space, or volume itself (hereinafter referred to as original coordinate encoding) as non-octree encoding.

[0650] In this case, the 3D data encoding device stores information indicating that original coordinate encoding is applied to the object space (submap, space, or volume) in the header, thereby enabling the 3D data decoding device to determine whether original coordinate encoding is applied to the object space.

[0651] Furthermore, when the original coordinate encoding is applied, the three-dimensional data encoding device may encode the original coordinates without applying quantization and arithmetic encoding. Furthermore, the three-dimensional data encoding device may encode the original coordinates with a predetermined fixed bit length. Accordingly, the three-dimensional data encoding device can generate a stream with a certain bit length at a certain timing.

[0652] Furthermore, in the above description, although an example of encoding the difference obtained by subtracting the sub-coordinates from the coordinates of each point cloud is shown as non-octree encoding, the present invention is not limited to this.

[0653] For example, the three-dimensional data encoding device may encode the difference values ​​between the coordinates of each point cloud in sequence. Fig.64 is a diagram used to illustrate the operation of this situation. Fig.64 In the example shown, when the three-dimensional data encoding device encodes the point cloud PA, the sub-coordinates are used as predicted coordinates, and the difference between the coordinates of the point cloud PA and the predicted coordinates is encoded. Furthermore, when the three-dimensional data encoding device encodes the point cloud PB, the coordinates of the point cloud PA are used as predicted coordinates, and the difference between the point cloud PB and the predicted coordinates is encoded. Furthermore, when the three-dimensional data encoding device encodes the point cloud PC, the point cloud PB is used as predicted coordinates, and the difference between the point cloud PB and the predicted coordinates is encoded. In this way, the three-dimensional data encoding device can set a scanning order for multiple point clouds, and encode the difference between the coordinates of the object point cloud of the processing object and the coordinates of the previous point cloud in the scanning order relative to the object point cloud.

[0654] Furthermore, in the above description, the sub-coordinates are the coordinates of the lower left front corner of the sub-map, but the position of the sub-coordinates is not limited thereto. Figure 65 to Figure 671 shows another example of the position of the sub-coordinate. The sub-coordinate setting position can be set to any coordinate in the object space (sub-map, space, or volume). That is, as described above, the sub-coordinate can be the coordinate of the lower left front corner. Fig.65 As shown, the sub-coordinates can also be the coordinates of the center of the object space. Fig.66 As shown, the sub-coordinates may also be the coordinates of the upper right rear corner of the object space. Furthermore, the sub-coordinates are not limited to the coordinates of the lower left front or upper right rear corner of the object space, and may be the coordinates of any corner in the object space.

[0655] Furthermore, the sub-coordinates can also be set to the same position as the coordinates of a point cloud in the object space (sub-map, space, or volume). Fig.67 In the example shown, the coordinates of the sub-coordinates coincide with the coordinates of the point cloud PD.

[0656] Furthermore, although the present embodiment shows an example of switching between octree coding and non-octree coding, the present invention is not limited thereto. For example, the three-dimensional data encoding device may also switch between using a tree structure other than an octree or a non-tree structure other than the tree structure. For example, other tree structures refer to a kd tree that is divided by a plane perpendicular to one of the coordinate axes. In addition, any other tree structure may be used.

[0657] Furthermore, although the example of encoding the coordinate information of the point cloud is shown in this embodiment, it is not limited to this. The three-dimensional data encoding device may also encode color information, three-dimensional feature quantities, or feature quantities of visible light, etc. in the same way as the coordinate information. For example, the three-dimensional data encoding device may also set the average value of the color information of each point cloud in the sub-map as sub-color information, and encode the difference between the color information of each point cloud and the sub-color information.

[0658] Furthermore, although the present embodiment shows an example of selecting a coding method (octree coding or non-octree coding) with high coding efficiency according to the number of point clouds, etc., the present invention is not limited thereto. For example, the three-dimensional data coding device on the server side may retain in advance the bit stream of the point cloud encoded by the octree coding, the bit stream of the point cloud encoded by the non-octree coding, and the bit stream of the point cloud encoded by both methods, and switch the bit stream sent to the three-dimensional data decoding device according to the communication environment or the processing capacity of the three-dimensional data decoding device.

[0659] Fig.68 An example of the syntax of volume when switching the application of octree encoding is shown. Fig.68 The syntax shown is similar to Fig.58 The syntax shown is basically the same, and the difference is that each piece of information is information in volume units. Specifically, NumOfPoint shows the number of point clouds included in the volume. sub_coordinate_x, sub_coordinate_y, and sub_coordinate_z are sub-coordinate information of the volume.

[0660] Furthermore, diff_x[i], diff_y[i], and diff_z[i] are the differential coordinates of the i-th point cloud in the volume. diff_x[i] represents the differential value between the x-coordinate of the i-th point cloud in the volume and the x-coordinate of the sub-coordinate. diff_y[i] represents the differential value between the y-coordinate of the i-th point cloud in the volume and the y-coordinate of the sub-coordinate. diff_z[i] represents the differential value between the z-coordinate of the i-th point cloud in the volume and the z-coordinate of the sub-coordinate.

[0661] In addition, when the relative position of the volume in the space can be calculated, the three-dimensional data encoding device may not include the sub-coordinate information in the head of the volume. That is, the three-dimensional data encoding device may calculate the relative position of the volume in the space without including the sub-coordinate information in the head, and use the calculated position as the sub-coordinate of each volume.

[0662] As described above, the three-dimensional data encoding device involved in this embodiment determines whether to encode the object space unit among multiple space units (such as submaps, spaces or volumes) included in the three-dimensional data using an octree structure (for example, Fig.59 For example, when the number of three-dimensional points included in the object space unit is greater than a predetermined threshold, the three-dimensional data encoding device determines to encode the object space unit with an octree structure. Furthermore, when the number of three-dimensional points included in the object space unit is less than the threshold, the three-dimensional data encoding device determines not to encode the object space unit with an octree structure.

[0663] When it is determined that the object space unit is encoded with an octree structure ("Yes" in S1522), the three-dimensional data encoding device encodes the object space unit with an octree structure (S1523). And, when it is determined that the object space unit is not encoded with an octree structure ("No" in S1522), the three-dimensional data encoding device encodes the object space unit in a manner different from the octree structure (S1524). For example, as a different manner, the three-dimensional data encoding device encodes the coordinates of the three-dimensional points included in the object space unit. Specifically, as a different manner, the three-dimensional data encoding device encodes the difference between the reference coordinates of the object space unit and the coordinates of the three-dimensional points included in the object space unit.

[0664] Next, the three-dimensional data encoding apparatus adds information indicating whether the target space unit is encoded in the octree structure to the bit stream (S1525).

[0665] Accordingly, the three-dimensional data encoding device can reduce the data amount of the encoded signal, thereby improving the encoding efficiency.

[0666] For example, the three-dimensional data encoding device includes a processor and a memory, and the processor performs the above-mentioned processing using the memory.

[0667] Furthermore, the three-dimensional data decoding device according to the present embodiment decodes information indicating whether to decode an object space unit among a plurality of object space units (for example, a submap, a space, or a volume) included in the three-dimensional data in an octree structure from a bit stream (for example, Fig.60 When the above information indicates that the target space unit is to be decoded in an octree structure (Yes in S1532), the three-dimensional data decoding apparatus decodes the target space unit in an octree structure (S1533).

[0668] When the above information indicates that the object space unit is not to be decoded in the octree structure (No in S1532), the three-dimensional data decoding device decodes the object space unit in a manner different from the octree structure (S1534). For example, the three-dimensional data decoding device decodes the coordinates of the three-dimensional point included in the object space unit in a different manner. Specifically, the three-dimensional data decoding device decodes the difference between the reference coordinates of the object space unit and the coordinates of the three-dimensional point included in the object space unit in a different manner.

[0669] Therefore, the three-dimensional data decoding device can reduce the data amount of the encoded signal, thereby improving the encoding efficiency.

[0670] For example, the three-dimensional data decoding device includes a processor and a memory, and the processor performs the above-mentioned processing using the memory.

[0671] (Implementation method 9)

[0672] In this embodiment, a coding method of a tree structure such as an octree structure is described.

[0673] By identifying important areas and decoding the three-dimensional data of the important areas preferentially, efficiency can be improved.

[0674] Fig.69This is a diagram showing an example of an important area in a three-dimensional map. An important area is, for example, an area containing a certain number of three-dimensional points with large feature values ​​among the three-dimensional points in the three-dimensional map. Alternatively, an important area may be, for example, an area containing a certain number of three-dimensional points required for a vehicle-mounted client to estimate its own position. Alternatively, an important area may be an area of ​​a face in a three-dimensional model of a person. In this way, important areas can be defined for each application, and important areas can be switched according to the application.

[0675] In this embodiment, occupancy coding and location coding are used as a method of expressing an octree structure, etc. In addition, the bit string obtained by occupancy coding is called an occupancy code. The bit string obtained by location coding is called a location code.

[0676] Fig.70 is a diagram showing an example of an occupancy code. Fig.70 Example of occupancy code representing a quadtree structure. Fig.70 In the example, an occupancy code is assigned to each node. Each occupancy code indicates whether a three-dimensional point is contained in the child nodes or leaf nodes of each node. For example, in the case of a quadtree, the information indicating whether the four child nodes or leaf nodes of each node respectively contain three-dimensional points is represented by a 4-bit occupancy code. In addition, in the case of an octree, the information indicating whether the eight child nodes or leaf nodes of each node respectively contain three-dimensional points is represented by an 8-bit occupancy code. In addition, here, for the sake of simplicity, the quadtree structure is used as an example, but the same is also applicable to the octree structure. For example, Fig.70 As shown, the occupancy code is based on Fig.40 In the occupancy code, since the information of multiple three-dimensional points is decoded in a fixed order, it is not possible to preferentially decode the information of any three-dimensional point. In addition, the occupancy code can also be based on Fig.40 A bit string that performs a depth-first scan of nodes and leaf nodes as described in, etc.

[0677] The following describes position coding. By using position codes, important parts of the octree structure can be directly decoded. In addition, important three-dimensional points in deep layers can be efficiently encoded.

[0678] Fig.71 is a diagram for explaining position coding, and is a diagram showing an example of a quadtree structure. Fig.71In the example shown, three-dimensional points A to I are represented by a quadtree structure. In addition, three-dimensional points A and C are important three-dimensional points included in the important area.

[0679] Fig.72 Yes means Fig.71 A diagram showing occupancy codes and position codes of important three-dimensional points A and C in a quadtree structure is shown.

[0680] In position coding, in a tree structure, the indexes of nodes existing in the path to the leaf node to which the object 3D point to be coded belongs and the indexes of the leaf nodes are coded. Here, the index is a numerical value assigned to each node and leaf node. In other words, the index refers to an identifier for identifying multiple child nodes of the object node. Fig.71 As shown, in the case of a quadtree, the index represents any one of 0 to 3.

[0681] For example, in Fig.71 In the quadtree structure shown, when leaf node A is the object three-dimensional point, leaf node A is expressed as 0→2→1→0→1→2→1. Here, in the case of the right figure, the maximum value of each index is 4 (which can be expressed in 2 bits), so the number of bits required for the position code of leaf node A is 7×2bit=14bit. When leaf node C is the encoding object, the number of bits required is also 14 bits. In addition, in the case of an octree, since the maximum value of each index is 8 (which can be expressed in 3 bits), the required number of bits can be calculated as 3bit×the depth of the leaf node. In addition, the three-dimensional data encoding device can also perform entropy encoding after binarizing each index to reduce the amount of data.

[0682] In addition, if Fig.72 As shown in FIG. 1 , in the occupancy code, in order to decode the leaf nodes A and C, it is necessary to decode all the nodes in the upper layer. On the other hand, in the position code, only the data of the leaf nodes A and C can be decoded. Fig.72 As shown, by using the position code, the number of bits can be reduced compared to the occupancy code.

[0683] In addition, if Fig.72 As shown, the amount of code can be further reduced by performing dictionary compression such as LZ77 on part or all of the position code.

[0684] Next, an example of applying position encoding to three-dimensional points (point cloud) obtained by LiDAR will be described. Fig.73This is a diagram showing an example of a three-dimensional point obtained by LiDAR. The three-dimensional points obtained by LiDAR are sparse. That is, when the three-dimensional point is represented by an occupancy code, the number of zero values ​​increases. In addition, a high three-dimensional accuracy is requested for the three-dimensional point. That is, the hierarchy of the octree structure becomes deeper.

[0685] Fig.74 is a diagram showing an example of such a sparse deep octree structure. Fig.74 The occupancy code of the octree structure shown is 136 bits (=8 bits × 17 nodes). In addition, since the depth is 6, there are 6 three-dimensional points, so the position code is 3 bits × 6 × 6 = 108 bits. That is, the position code can reduce the amount of code by 20% relative to the occupancy code. In this way, by applying position coding to the sparse deep octree structure, the amount of code can be reduced.

[0686] The following describes the code amounts of the occupancy code and the position code. When the depth of the octree structure is 10, the maximum number of three-dimensional points is 8. 10 = 1073741824. In addition, the number of bits of the occupancy code of the octree structure is L o It is represented by the following.

[0687] L o =8+8 2 +…+8 10 =127133512 bits

[0688] Therefore, the number of bits per three-dimensional point is 1.143 bits. In addition, in the occupancy code, the number of bits does not change even if the number of three-dimensional points included in the octree structure changes.

[0689] On the other hand, in the position code, the number of bits of each three-dimensional point directly affects the depth of the octree structure. Specifically, the number of bits of the position code of each three-dimensional point is 3 bits×depth 10=30 bits.

[0690] Therefore, the number of bits of the position code of the octree structure is L l It is represented by the following.

[0691] L l =30×N

[0692] Here, N is the number of three-dimensional points included in the octree structure.

[0693] Therefore, when N<L o When / 30=40904450.4, that is, when the number of three-dimensional points is less than 40904450, the code amount of the position code becomes smaller than the code amount of the occupancy code (L l <L o ).

[0694] In this way, when there are few three-dimensional points, the amount of code of the position code is smaller than the amount of code of the occupancy code, and when there are many three-dimensional points, the amount of code of the position code is larger than the amount of code of the occupancy code.

[0695] Therefore, the 3D data encoding device may also switch between position coding and occupancy coding according to the number of input 3D points. In this case, the 3D data encoding device may also attach information indicating which of the position coding and occupancy coding is used to the header information of the bit stream, etc.

[0696] Hybrid coding combining position coding and occupancy coding is described below. When coding a dense important area, hybrid coding combining position coding and occupancy coding is effective. Fig.75 is a diagram showing this example. Fig.75 In the example shown, important three-dimensional points are densely arranged. In this case, the three-dimensional data encoding device performs position encoding on the shallow upper layer and uses occupancy encoding on the lower layer. Specifically, position encoding is used until the deepest common node, and occupancy encoding is used at a position deeper than the deepest common node. Here, the deepest common node refers to the deepest node among the nodes that are the common ancestors of multiple important three-dimensional points.

[0697] Next, hybrid coding that prioritizes compression efficiency will be described. The three-dimensional data coding device may switch between position coding and occupancy coding according to a rule predetermined in octree coding.

[0698] Fig.76 is a diagram showing an example of the rule. First, the three-dimensional data encoding device confirms the proportion of nodes containing three-dimensional points in each level (depth). When the proportion is higher than a predetermined threshold, the three-dimensional data encoding device performs occupancy encoding on several nodes in the upper layer of the object level. For example, the three-dimensional data encoding device applies occupancy encoding to the levels from the object level to the deepest common node.

[0699] For example, in Fig.76 In the example shown, the ratio of nodes including three-dimensional points in level 3 is higher than the threshold value. Therefore, the three-dimensional data encoding device applies occupancy encoding to the second and third levels from the third level to the deepest common node, and applies position encoding to the other first and fourth levels.

[0700] The calculation method of the above threshold value is explained. In one layer of the octree structure, there is one root node and eight child nodes. Therefore, in occupancy coding, 8 bits are required to encode one layer of the octree structure. On the other hand, in position coding, 3 bits are required for each child node containing a three-dimensional point. Therefore, in the case where the number of nodes containing a three-dimensional point is greater than 2, occupancy coding is more effective than position coding. That is, in this case, the threshold value is 2.

[0701] Hereinafter, a configuration example of a bit stream generated by the above-mentioned position coding, occupancy coding, or hybrid coding will be described.

[0702] Fig.77 is a diagram showing an example of a bit stream generated by position coding. Fig.77 As shown in FIG. 1 , the bit stream generated by position coding includes a header and a plurality of position codes, each of which is for a three-dimensional point.

[0703] With this configuration, the three-dimensional data decoding device can decode a plurality of three-dimensional points with high accuracy. Fig.77 This shows an example of a bitstream in the case of a quadtree structure. In the case of an octree structure, each index can take a value from 0 to 7.

[0704] In addition, the three-dimensional data encoding device may also perform entropy encoding after binarizing the column (string) of the index representing a three-dimensional point. For example, when the column of the index is 0121, the three-dimensional data encoding device may binarize 0121 into 00011001 and perform arithmetic encoding on the bit string.

[0705] Fig.78 is a diagram showing an example of a bit stream generated by hybrid coding in the case of including important three-dimensional points. Fig.78 As shown, the upper layer position code, the lower layer occupancy code of the important 3D points, and the lower layer occupancy code of the non-important 3D points other than the important 3D points are sequentially configured. Fig.78 The position code length shown indicates the code size of the position code that follows. In addition, the occupancy code size indicates the code size of the occupancy code that follows.

[0706] With this configuration, the three-dimensional data decoding device can select different decoding plans according to the application program.

[0707] In addition, the encoded data of the important three-dimensional points is stored near the beginning of the bit stream, and the encoded data of the unimportant three-dimensional points not included in the important area is stored after the encoded data of the important three-dimensional points.

[0708] Fig.79 It means by Fig.78A diagram showing a tree structure of occupancy codes for important three-dimensional points is shown. Fig.80 It means by Fig.78 The tree structure of the occupancy code representation of the non-significant 3D points is shown in FIG. Fig.79 As shown, the information related to non-important 3D points is excluded from the occupancy code of the important 3D points. Specifically, since nodes 0 and 3 at depth 5 do not contain important 3D points, nodes 0 and 3 are assigned a value of 0 indicating that they do not contain 3D points.

[0709] On the other hand, Fig.80 As shown, in the occupancy code of the non-important 3D points, information related to the important 3D points is excluded. Specifically, since the node 1 at the depth 5 does not contain the non-important 3D points, the value 0 indicating that the node 1 does not contain the 3D points is assigned to the node 1.

[0710] In this way, the 3D data encoding device divides the original tree structure into a first tree structure including important 3D points and a second tree structure including unimportant 3D points, and independently encodes the occupancy of the first tree structure and the second tree structure. Thus, the 3D data decoding device can preferentially decode important 3D points.

[0711] Next, a configuration example of a bit stream generated by hybrid coding with emphasis on efficiency will be described. Fig.81 FIG. 1 is a diagram showing an example of the structure of a bit stream generated by hybrid coding with an emphasis on efficiency. Fig.81 As shown, for each subtree, the subtree root node position, occupancy code amount and occupancy code are configured in sequence. Fig.81 The subtree position shown is the position code of the root node of the subtree.

[0712] In the above-mentioned configuration, when only one of position coding and occupancy coding is applied to the octree structure, the following holds.

[0713] When the length of the position code of the root node of a subtree is equal to the depth of the octree structure, the subtree has no child nodes. That is, the position code is applied to the entire tree structure.

[0714] When the root node of the subtree is equal to the root node of the octree structure, occupancy encoding is applied to the entire tree structure.

[0715] For example, based on the above-mentioned rule, the three-dimensional data decoding device can determine whether the position code or the occupancy code is included in the bit stream.

[0716] Furthermore, the bitstream may include coding mode information indicating which of position coding, occupancy coding, and hybrid coding is used. Fig.82 is a diagram showing an example of a bit stream in this case. Fig.82 As shown, 2-bit coding mode information indicating the coding mode is appended to the bitstream.

[0717] In addition, (1) the "number of 3D points" in the position coding indicates the number of subsequent 3D points. Also, (2) the "amount of occupancy code" in the occupancy coding indicates the amount of the subsequent occupancy code. Also, (3) the "number of important subtrees" in the hybrid coding (important 3D points) indicates the number of subtrees containing important 3D points. Furthermore, (4) the "number of occupancy subtrees" in the hybrid coding (efficiency emphasis) indicates the number of subtrees after occupancy coding.

[0718] Next, a syntax example used for switching the application of occupancy coding and position coding will be described. Fig.83 It is a figure showing this syntax example.

[0719] Fig.83 The shown isleaf is a flag indicating whether the object node is a leaf node. isleaf = 1 indicates that the object node is a leaf node, and isleaf = 0 indicates that the object node is not a leaf node but a node.

[0720] When the object node is a leaf node, point_flag is appended to the bitstream. point_flag is a flag indicating whether the object node (leaf node) contains 3D points. point_flag = 1 indicates that the object node contains 3D points, and point_flag = 0 indicates that the object node does not contain 3D points.

[0721] When the object node is not a leaf node, coding_type is appended to the bitstream. coding_type is coding type information indicating the applicable coding type. coding_type = 00 indicates that position coding is applied, coding_type = 01 indicates that occupancy coding is applied, and coding_type = 10 or 11 indicates that other coding methods are applied, etc.

[0722] When the coding type is position coding, numPoint, num_idx[i], and idx[i][j] are appended to the bitstream.

[0723] numPoint indicates the number of 3D points for which position coding is performed. num_idx[i] indicates the number (depth) of the index from the object node to 3D point i. When all the 3D points for which position coding is performed are at the same depth, num_idx[i] are all the same value. Therefore, it can also be that, before the Fig.83 shown for statement (for(i = 0; i < numPoint; i++) {}), num_idx is defined as a common value.

[0724] Idx[i][j] represents the value of the j-th index among the indices from the object node to the three-dimensional point i. In the case of an octree, the number of bits of idx[i][j] is 3 bits.

[0725] In addition, as described above, an index refers to an identifier for identifying multiple child nodes of an object node. In the case of an octree, idx[i][j] represents any one of 0 to 7. In the case of an octree, there are 8 child nodes, each of which corresponds to each of the 8 sub-blocks obtained by spatially dividing the object block corresponding to the object node into 8. Therefore, idx[i][j] may also be information indicating the three-dimensional position of the sub-block corresponding to the child node. For example, idx[i][j] may also be a total of 3 bits of information including 1 bit of information indicating the position of each of the x, y, and z of the sub-block.

[0726] When the encoding type is occupancy encoding, occupancy_code is added to the bit stream. Occupancy_code is the occupancy code of the target node. In the case of an octree, occupancy_code is an 8-bit bit string such as a bit string "00101000".

[0727] When the value of the (i+1)th bit of occupancy_code is 1, the process moves to the child node. That is, the child node is set as the next target node, and a bit string is recursively generated.

[0728] In this embodiment, an example of indicating the end of an octree by attaching leaf node information (isleaf, point_flag) to a bitstream is shown, but it is not necessarily limited to this. For example, the three-dimensional data encoding device may attach the maximum depth (depth) from the start node (root node) of the occupancy code to the end (leaf node) where a three-dimensional point exists to the head of the start node. Then, the three-dimensional data encoding device may also recursively stringify the information of the child nodes while increasing the depth from the start node, and judge that the leaf node has been reached at the time point when the depth becomes the maximum depth. In addition, the three-dimensional data encoding device may attach the information indicating the maximum depth to the initial node where coding_type becomes the occupancy code, or to the start node (root node) of the octree.

[0729] As described above, the three-dimensional data encoding device may also add information for switching between occupancy coding and position coding in the bit stream as header information of each node.

[0730] In addition, the three-dimensional data encoding device may also perform entropy encoding on the coding_type, numPoint, num_idx, idx, and occupancy_code of each node generated by the above method. For example, the three-dimensional data encoding device performs arithmetic encoding after binarizing each value.

[0731] In addition, in the above syntax, the case of using a depth-first bit string of an octree structure as an occupancy code is exemplified, but it is not necessarily limited to this. The three-dimensional data encoding device may also use a width-first bit string of an octree structure as an occupancy code. When the three-dimensional data encoding device uses a width-first bit string, it may also add information for switching occupancy coding and position coding in the bit stream as header information of each node.

[0732] In this embodiment, an octree structure is used as an example, but it is not necessarily limited to this. The above method can also be applied to N-ary trees (N is an integer greater than 2) such as quadtrees and hexadecimal trees or other tree structures.

[0733] An example of a flow of a coding process for switching between application of occupancy coding and position coding will be described below. Fig.84 This is a flowchart of the encoding process according to this embodiment.

[0734] First, the three-dimensional data encoding device uses an octree structure to represent a plurality of three-dimensional points contained in the three-dimensional data (S1601). Next, the three-dimensional data encoding device sets the root node in the octree structure as the object node (S1602). Next, the three-dimensional data encoding device generates a bit string of the octree structure by performing node encoding processing on the object node (S1603). Next, the three-dimensional data encoding device generates a bit stream by performing entropy encoding on the generated bit string (S1604).

[0735] Fig.85 This is a flowchart of the node encoding process (S1603). First, the three-dimensional data encoding device determines whether the object node is a leaf node (S1611). If the object node is not a leaf node ("No" in S1611), the three-dimensional data encoding device sets the leaf node flag (isleaf) to 0 and adds the leaf node flag to the bit string (S1612).

[0736] Next, the three-dimensional data encoding device determines whether the number of child nodes including the three-dimensional point is greater than a predetermined threshold value (S1613). In addition, the three-dimensional data encoding device may also add the threshold value to the bit string.

[0737] When the number of child nodes including a three-dimensional point is greater than a predetermined threshold (Yes in S1613), the three-dimensional data encoding device sets the encoding type (coding_type) to occupancy coding and adds the encoding type to the bit string (S1614).

[0738] Next, the three-dimensional data encoding device sets the occupancy coding information and adds the occupancy coding information to the bit string. Specifically, the three-dimensional data encoding device generates an occupancy code of the target node and adds the occupancy code to the bit string (S1615).

[0739] Next, the three-dimensional data encoding device sets the next target node according to the occupancy code (S1616). Specifically, the three-dimensional data encoding device sets the unprocessed child node with the occupancy code "1" as the next target node.

[0740] Next, the three-dimensional data encoding device performs node encoding processing on the newly set object node (S1617). Fig.85 Processing shown.

[0741] If all child nodes have not been processed (No in S1618), the process after step S1616 is performed again. On the other hand, if all child nodes have been processed (Yes in S1618), the three-dimensional data encoding device ends the node encoding process.

[0742] In addition, in step S1613, when the number of child nodes including three-dimensional points is below a predetermined threshold ("No" in S1613), the three-dimensional data encoding device sets the encoding type to position encoding and appends the encoding type to the bit string (S1619).

[0743] Next, the three-dimensional data encoding device sets position coding information and adds the position coding information to the bit string. Specifically, the three-dimensional data encoding device generates a position code and adds the position code to the bit string (S1620). The position code includes numPoint, num_idx, and idx.

[0744] In addition, in step S1611, when the object node is a leaf node ("Yes" in S1611), the three-dimensional data encoding device sets the leaf node flag to 1, and adds the leaf node flag to the bit string (S1621). In addition, the three-dimensional data encoding device sets a point flag (point_flag) indicating whether the leaf node includes a three-dimensional point, and adds the point flag to the bit string (S1622).

[0745] Next, an example of a decoding process flow for switching between application of occupancy coding and position coding will be described. Fig.85 This is a flowchart of the decoding process according to this embodiment.

[0746] The three-dimensional data decoding device generates a bit string by entropy decoding the bit stream (S1631). Then, the three-dimensional data decoding device restores the octree structure by performing node decoding processing on the obtained bit string (S1632). Then, the three-dimensional data decoding device generates a three-dimensional point according to the restored octree structure (S1633).

[0747] Fig.87 1632 is a flowchart of the node decoding process. First, the three-dimensional data decoding device obtains (decodes) a leaf node flag (isleaf) from the bit string (S1641). Next, the three-dimensional data decoding device determines whether the target node is a leaf node based on the leaf node flag (S1642).

[0748] If the target node is not a leaf node (No in S1642), the three-dimensional data decoding device obtains the coding type (coding_type) from the bit string (S1643). The three-dimensional data decoding device determines whether the coding type is occupancy coding (S1644).

[0749] When the encoding type is occupancy encoding (Yes in S1644), the three-dimensional data decoding device obtains occupancy encoding information from the bit string. Specifically, the three-dimensional data decoding device obtains an occupancy code from the bit string (S1645).

[0750] Next, the three-dimensional data decoding apparatus sets the next target node according to the occupancy code (S1646). Specifically, the three-dimensional data decoding apparatus sets the unprocessed child node with the occupancy code "1" as the next target node.

[0751] Next, the three-dimensional data decoding device performs node decoding processing on the newly set object node (S1647). Fig.87 Processing shown.

[0752] If the processing of all child nodes is not completed (No in S1648), the processing after step S1646 is performed again. On the other hand, if the processing of all child nodes is completed (Yes in S1648), the three-dimensional data decoding device ends the node decoding processing.

[0753] In addition, when the encoding type is position encoding in step S1644 (No in S1644), the three-dimensional data decoding device obtains position encoding information from the bit string. Specifically, the three-dimensional data decoding device obtains a position code from the bit string (S1649). The position code includes numPoint, num_idx, and idx.

[0754] Furthermore, when the target node is a leaf node in step S1642 (Yes in S1642), the three-dimensional data decoding apparatus obtains a point flag (point_flag) which is information indicating whether the leaf node includes a three-dimensional point from the bit string (S1650).

[0755] In addition, in this embodiment, an example of switching the encoding type for each node is shown, but it is not necessarily limited to this. The encoding type can also be fixed in units of volume, space, or world space. In this case, the three-dimensional data encoding device can also attach the encoding type information to the header information of the volume, space, or world space.

[0756] As described above, the three-dimensional data encoding device of the present embodiment generates first information of an N-ary tree structure (N is an integer greater than or equal to 2) representing a plurality of three-dimensional points included in three-dimensional data in a first manner (position coding), and generates a bit stream including the first information. The first information includes three-dimensional point information (position code) corresponding to each of the plurality of three-dimensional points. Each three-dimensional point information includes an index (idx) corresponding to each of the plurality of layers in the N-ary tree structure. Each index indicates a sub-block to which the corresponding three-dimensional point belongs among the N sub-blocks belonging to the corresponding layer.

[0757] In other words, each piece of three-dimensional point information indicates a path to the corresponding three-dimensional point in the N-ary tree structure, and each index indicates a child node included in the path among the N child nodes belonging to the corresponding layer (node).

[0758] Thus, the three-dimensional data encoding method can generate a bit stream that can selectively decode three-dimensional points.

[0759] For example, the three-dimensional point information (position code) includes information (num_idx) indicating the number of indexes included in the three-dimensional point information. In other words, the information indicates the depth (number of layers) to the corresponding three-dimensional point in the N-ary tree structure.

[0760] For example, the first information includes information (numPoint) indicating the number of three-dimensional point information included in the first information. In other words, the information indicates the number of three-dimensional points included in the N-ary tree structure.

[0761] For example, N is 8 and the index is 3 bits.

[0762] For example, a three-dimensional data encoding device has: a first encoding mode for generating first information; and a second encoding mode for generating second information (occupancy code) representing an N-ary tree structure in a second manner (occupancy coding), and generating a bit stream including the second information. The second information includes a plurality of 1-bit information corresponding to each of a plurality of sub-blocks belonging to a plurality of layers in the N-ary tree structure and indicating whether a three-dimensional point exists in the corresponding sub-block.

[0763] For example, the 3D data encoding device uses the first encoding mode when the number of the 3D points is less than a predetermined threshold, and uses the second encoding mode when the number of the 3D points is greater than the threshold. Thus, the 3D data encoding device can reduce the amount of code in the bitstream.

[0764] For example, the first information and the second information include information (coding mode information) indicating whether the information represents the N-ary tree structure in the first manner or the N-ary tree structure in the second manner.

[0765] For example, Fig.75 As shown in FIG. 1 , the three-dimensional data encoding device uses the first encoding mode in a part of the N-ary tree structure and uses the second encoding mode in another part of the N-ary tree structure.

[0766] For example, the three-dimensional data encoding device includes a processor and a memory, and the processor performs the above-mentioned processing using the memory.

[0767] In addition, the three-dimensional data decoding device of the present embodiment obtains first information (position code) of an N-ary tree structure (N is an integer greater than or equal to 2) that represents a plurality of three-dimensional points included in the three-dimensional data in a first manner (position coding) from a bit stream. The first information includes three-dimensional point information (position code) corresponding to each of the plurality of three-dimensional points. Each three-dimensional point information includes an index (idx) corresponding to each of the plurality of layers in the N-ary tree structure. Each index indicates a sub-block to which the corresponding three-dimensional point belongs among the N sub-blocks belonging to the corresponding layer.

[0768] In other words, each piece of three-dimensional point information indicates a path to the corresponding three-dimensional point in the N-ary tree structure, and each index indicates a child node included in the path among the N child nodes belonging to the corresponding layer (node).

[0769] The three-dimensional data decoding device further uses the three-dimensional point information to restore the three-dimensional point corresponding to the three-dimensional point information.

[0770] Thus, the three-dimensional data decoding device can selectively decode three-dimensional points from a bit stream.

[0771] For example, the three-dimensional point information (position code) includes information (num_idx) indicating the number of indexes included in the three-dimensional point information. In other words, the information indicates the depth (number of layers) to the corresponding three-dimensional point in the N-ary tree structure.

[0772] For example, the first information includes information (numPoint) indicating the number of three-dimensional point information included in the first information. In other words, the information indicates the number of three-dimensional points included in the N-ary tree structure.

[0773] For example, N is 8 and the index is 3 bits.

[0774] For example, the three-dimensional data decoding device further obtains second information (occupancy code) representing the N-ary tree structure in a second manner (occupancy coding) from the bit stream. The three-dimensional data decoding device uses the second information to restore the plurality of three-dimensional points. The second information includes a plurality of 1-bit information corresponding to each of the plurality of sub-blocks belonging to the plurality of layers in the N-ary tree structure and indicating whether the three-dimensional point exists in the corresponding sub-block.

[0775] For example, the first information and the second information include information (coding mode information) indicating whether the information represents the N-ary tree structure in the first manner or the N-ary tree structure in the second manner.

[0776] For example, Fig.75 As shown in the above, a part of the N-ary tree structure is represented by the first method, and another part of the N-ary tree structure is represented by the second method.

[0777] For example, the three-dimensional data decoding device includes a processor and a memory, and the processor performs the above-mentioned processing using the memory.

[0778] (Implementation 10)

[0779] In this embodiment, another example of a method for encoding a tree structure such as an octree structure will be described. Fig.88 is a diagram showing an example of a tree structure related to this embodiment. Fig.88 An example showing a quadtree structure.

[0780] The leaf nodes that contain three-dimensional points are called valid leaf nodes, and the leaf nodes that do not contain three-dimensional points are called invalid leaf nodes. The branches whose number of valid leaf nodes is above the threshold are called dense branches. The branches whose number of valid leaf nodes is less than the threshold are called sparse branches.

[0781] The three-dimensional data encoding device calculates the number of three-dimensional points included in each branch (ie, the number of valid leaf nodes) in a certain layer of the tree structure. Fig.88This shows an example where the threshold is 5. In this example, there are two branches in layer 1. Since the left branch contains 7 3D points, the left branch is determined to be a dense branch. Since the right branch contains 2 3D points, the right branch is determined to be a sparse branch.

[0782] Fig.89 For example, this is a diagram showing an example of the number of valid leaf nodes (3D points) that each branch of layer 5 has. Fig.89 The horizontal axis of represents the identification number, i.e., index, of the branch of layer 5. Fig.89 As shown in Fig. 3, a certain branch contains significantly more 3D points than other branches. In such a dense branch, occupancy encoding is more effective than in a sparse branch.

[0783] The following describes how to apply occupancy coding and position coding. Fig.90 is a diagram showing the relationship between the number of three-dimensional points (the number of valid leaf nodes) included in each branch of layer 5 and the applied encoding method. Fig.90 As shown, the three-dimensional data encoding device applies occupancy coding to dense branches and applies position coding to sparse branches, thereby improving the encoding efficiency.

[0784] Fig.91 is a diagram showing an example of a dense branching region in LiDAR data. Fig.91 As shown, the density of three-dimensional points calculated based on the number of three-dimensional points included in each branch is different depending on the area.

[0785] In addition, by separating dense 3D points (branches) from sparse 3D points (branches), there are the following advantages. The closer to the LiDAR sensor, the higher the density of the 3D points. Therefore, by separating the branches according to the sparseness, it is possible to perform zoning in the distance direction. Such zoning is effective in specific applications. In addition, for sparse branches, it is effective to use methods other than occupancy coding.

[0786] In this embodiment, the three-dimensional data encoding device separates an input three-dimensional point group into two or more sub-three-dimensional point groups, and applies a different encoding method to each sub-three-dimensional point group.

[0787] For example, the three-dimensional data encoding device separates the input three-dimensional point group into a sub-three-dimensional point group A including dense branches (dense three-dimensional point group: dense cloud) and a sub-three-dimensional point group B including sparse branches (sparse three-dimensional point group: sparse cloud). Fig.92 It means from Fig.88 FIG. 1 is a diagram showing an example of a sub-three-dimensional point group A (a dense three-dimensional point group) including dense branches separated by a tree structure. Fig.93 It means from Fig.88 FIG. 1 is a diagram showing an example of a sub-three-dimensional point group B (sparse three-dimensional point group) including sparse branches separated by the tree structure shown.

[0788] Next, the three-dimensional data encoding device encodes the sub-three-dimensional point group A through occupancy encoding, and encodes the sub-three-dimensional point group B through position encoding.

[0789] In addition, an example of applying different encoding methods (occupancy coding and position coding) as different encoding methods is shown here, but for example, the three-dimensional data encoding device can also use the same encoding method for sub-three-dimensional point group A and sub-three-dimensional point group B, and make the parameters used in the encoding different between sub-three-dimensional point group A and sub-three-dimensional point group B.

[0790] The following describes the flow of three-dimensional data encoding processing performed by the three-dimensional data encoding device. Fig.94 This is a flowchart of a three-dimensional data encoding process performed by the three-dimensional data encoding device according to the present embodiment.

[0791] First, the three-dimensional data encoding device separates the input three-dimensional point group into sub-three-dimensional point groups (S1701). The three-dimensional data encoding device can perform the separation automatically or based on information input by the user. For example, the range of the sub-three-dimensional point group can also be specified by the user. In addition, as an example of automatic operation, for example, when the input data is LiDAR data, the three-dimensional data encoding device uses the distance information to each point group for separation. Specifically, the three-dimensional data encoding device separates the point group within a certain range from the measurement location from the point group outside the range. In addition, the three-dimensional data encoding device can also use the information of important areas and unimportant areas for separation.

[0792] Next, the three-dimensional data encoding device encodes the sub-three-dimensional point group A by method A, thereby generating encoded data (encoded bit stream) (S1702). In addition, the three-dimensional data encoding device encodes the sub-three-dimensional point group B by method B, thereby generating encoded data (S1703). In addition, the three-dimensional data encoding device may also encode the sub-three-dimensional point group B by method A. In this case, the three-dimensional data encoding device encodes the sub-three-dimensional point group B using a parameter different from the encoding parameter used in the encoding of the sub-three-dimensional point group A. For example, the parameter may also be a quantization parameter. For example, the three-dimensional data encoding device encodes the sub-three-dimensional point group B using a quantization parameter larger than the quantization parameter used in the encoding of the sub-three-dimensional point group A. In this case, the three-dimensional data encoding device may also add information indicating the quantization parameter used in the encoding of the sub-three-dimensional point group to the header of the encoded data of each sub-three-dimensional point group.

[0793] Next, the three-dimensional data encoding device generates a bit stream by combining the encoded data obtained in step S1702 and the encoded data obtained in step S1703 ( S1704 ).

[0794] Furthermore, the three-dimensional data encoding device may encode information for decoding each sub-three-dimensional point group as header information of the bit stream. For example, the three-dimensional data encoding device may encode the following information.

[0795] The header information may also include information indicating the number of sub-3D points to be encoded. In this example, the information indicates 2.

[0796] The header information may also include information indicating the number of 3D points included in each sub-3D point group and the encoding method. In this example, the information indicates the number of 3D points included in sub-3D point group A, the encoding method (method A) applied to sub-3D point group A, the number of 3D points included in sub-3D point group B, and the encoding method (method B) applied to sub-3D point group B.

[0797] The header information may include information for identifying the start position or the end position of the encoded data of each sub-three-dimensional point group.

[0798] Furthermore, the three-dimensional data encoding device may encode the sub-three-dimensional point group A and the sub-three-dimensional point group B in parallel. Alternatively, the three-dimensional data encoding device may encode the sub-three-dimensional point group A and the sub-three-dimensional point group B sequentially.

[0799] In addition, the method of separation into sub-three-dimensional point groups is not limited to the above. For example, the three-dimensional data encoding device changes the separation method, uses each of a plurality of separation methods for encoding, and calculates the encoding efficiency of the encoded data obtained using each separation method. And, the three-dimensional data encoding device selects the separation method with the highest encoding efficiency. For example, the three-dimensional data encoding device may also separate the three-dimensional point group in each of a plurality of layers, calculate the encoding efficiency in each case, select the separation method with the highest encoding efficiency (i.e., the layer to be separated), and use the selected separation method to generate a sub-three-dimensional point group for encoding.

[0800] Furthermore, the 3D data encoding device may also arrange the encoding information of the more important sub-3D point groups closer to the beginning of the bit stream when combining the encoded data. Thus, the 3D data decoding device can obtain the important information only by decoding the beginning bit stream, so the important information can be obtained earlier.

[0801] Next, the flow of the three-dimensional data decoding process performed by the three-dimensional data decoding device will be described. Fig.95 This is a flowchart of a three-dimensional data decoding process performed by the three-dimensional data decoding device according to the present embodiment.

[0802] First, the 3D data decoding device obtains, for example, a bit stream generated by the 3D data encoding device. Next, the 3D data decoding device separates the encoded data of the sub-3D point group A from the encoded data of the sub-3D point group B from the obtained bit stream (S1711). Specifically, the 3D data decoding device decodes information for decoding each sub-3D point group from header information of the bit stream, and uses the information to separate the encoded data of each sub-3D point group.

[0803] Next, the three-dimensional data decoding device obtains a sub-three-dimensional point group A by decoding the coded data of the sub-three-dimensional point group A using technique A (S1712). In addition, the three-dimensional data decoding device obtains a sub-three-dimensional point group B by decoding the coded data of the sub-three-dimensional point group B using technique B (S1713). Next, the three-dimensional data decoding device combines the sub-three-dimensional point group A with the sub-three-dimensional point group B (S1714).

[0804] In addition, the three-dimensional data decoding apparatus may decode the three-dimensional point group sub-group A and the three-dimensional point group sub-group B in parallel. Alternatively, the three-dimensional data decoding apparatus may decode the three-dimensional point group sub-group A and the three-dimensional point group sub-group B sequentially.

[0805] In addition, the 3D data decoding device may also decode the required sub-3D point group. For example, the 3D data decoding device may also decode the sub-3D point group A, but not decode the sub-3D point group B. For example, when the sub-3D point group A is a 3D point group included in an important area of ​​the LiDAR data, the 3D data decoding device decodes the 3D point group of the important area. The 3D point group of the important area is used to estimate the self-position of the vehicle, etc.

[0806] Next, a specific example of the encoding process according to this embodiment will be described. Fig.96 This is a flowchart of a three-dimensional data encoding process performed by the three-dimensional data encoding device according to the present embodiment.

[0807] First, the three-dimensional data encoding device separates the input three-dimensional points into a sparse three-dimensional point group and a dense three-dimensional point group (S1721). Specifically, the three-dimensional data encoding device counts the number of valid leaf nodes of the branches of a certain layer of the octree structure. The three-dimensional data encoding device sets each branch as a dense branch or a sparse branch according to the number of valid leaf nodes of each branch. In addition, the three-dimensional data encoding device generates a sub-three-dimensional point group (dense three-dimensional point group) that collects the dense branches and a sub-three-dimensional point group (sparse three-dimensional point group) that collects the sparse branches.

[0808] Next, the three-dimensional data encoding device generates encoded data by encoding the sparse three-dimensional point group (S1722). For example, the three-dimensional data encoding device encodes the sparse three-dimensional point group using position encoding.

[0809] In addition, the three-dimensional data encoding device generates encoded data by encoding the dense three-dimensional point group (S1723). For example, the three-dimensional data encoding device encodes the dense three-dimensional point group using occupancy coding.

[0810] Next, the three-dimensional data encoding device generates a bit stream by combining the encoded data of the sparse three-dimensional point group obtained in step S1722 and the encoded data of the dense three-dimensional point group obtained in step S1723 ( S1724 ).

[0811] Furthermore, the three-dimensional data encoding device may encode information for decoding a sparse three-dimensional point group and a dense three-dimensional point group as header information of a bit stream. For example, the three-dimensional data encoding device may encode the following information.

[0812] The header information may also include information indicating the number of sub-three-dimensional point groups to be encoded. In this example, the information indicates 2.

[0813] The header information may also include information indicating the number of 3D points included in each sub-3D point group and the encoding method. In this example, the information indicates the number of 3D points included in the sparse 3D point group, the encoding method (position encoding) applied to the sparse 3D point group, the number of 3D points included in the dense 3D point group, and the encoding method (occupancy encoding) applied to the dense 3D point group.

[0814] The header information may also include information for identifying the start position or end position of the encoded data of each sub-three-dimensional point group. In this example, the information indicates at least one of the start position and end position of the encoded data of the sparse three-dimensional point group and the start position and end position of the encoded data of the dense three-dimensional point group.

[0815] In addition, the three-dimensional data encoding device may encode the sparse three-dimensional point group and the dense three-dimensional point group in parallel. Alternatively, the three-dimensional data encoding device may encode the sparse three-dimensional point group and the dense three-dimensional point group sequentially.

[0816] Next, a specific example of three-dimensional data decoding processing will be described. Fig.97 This is a flowchart of a three-dimensional data decoding process performed by the three-dimensional data decoding device according to the present embodiment.

[0817] First, the three-dimensional data decoding device obtains, for example, a bit stream generated by the above-mentioned three-dimensional data encoding device. Next, the three-dimensional data decoding device separates the obtained bit stream into sparse three-dimensional point group encoding data and dense three-dimensional point group encoding data (S1731). Specifically, the three-dimensional data decoding device decodes information used to decode each sub-three-dimensional point group from the header information of the bit stream, and uses the information to separate the encoding data of each sub-three-dimensional point group. In this example, the three-dimensional data decoding device uses the header information to separate the sparse three-dimensional point group and the dense three-dimensional point group encoding data from the bit stream.

[0818] Next, the three-dimensional data decoding device obtains a sparse three-dimensional point group by decoding the coded data of the sparse three-dimensional point group (S1732). For example, the three-dimensional data decoding device decodes the sparse three-dimensional point group using position decoding for decoding the position-coded coded data.

[0819] In addition, the 3D data decoding device obtains a dense 3D point group by decoding the coded data of the dense 3D point group (S1733). For example, the 3D data decoding device decodes the dense 3D point group using occupancy decoding for decoding the coded data coded by occupancy.

[0820] Next, the three-dimensional data decoding apparatus combines the sparse three-dimensional point group obtained in step S1732 with the dense three-dimensional point group obtained in step S1733 ( S1734 ).

[0821] In addition, the three-dimensional data decoding device may decode the sparse three-dimensional point group and the dense three-dimensional point group in parallel, or the three-dimensional data decoding device may decode the sparse three-dimensional point group and the dense three-dimensional point group sequentially.

[0822] In addition, the 3D data decoding device may also decode a part of the required sub-3D point group. For example, the 3D data decoding device may also decode a dense 3D point group without decoding sparse 3D data. For example, in the case where the dense 3D point group is a 3D point group included in an important area of ​​LiDAR data, the 3D data decoding device decodes the 3D point group of the important area. The 3D point group of the important area is used to estimate the self-position of the vehicle, etc.

[0823] Fig.98 First, the three-dimensional data encoding device generates a sparse three-dimensional point group and a dense three-dimensional point group by separating an input three-dimensional point group into a sparse three-dimensional point group and a dense three-dimensional point group (S1741).

[0824] Next, the three-dimensional data encoding device generates encoded data by encoding the dense three-dimensional point group (S1742). In addition, the three-dimensional data encoding device generates encoded data by encoding the sparse three-dimensional point group (S1743). Finally, the three-dimensional data encoding device generates a bit stream by combining the encoded data of the sparse three-dimensional point group obtained in step S1742 with the encoded data of the dense three-dimensional point group obtained in step S1743 (S1744).

[0825] Fig.99 This is a flowchart of the decoding process related to the present embodiment. First, the three-dimensional data decoding device extracts the encoded data of the dense three-dimensional point group and the encoded data of the sparse three-dimensional point group from the bit stream (S1751). Next, the three-dimensional data decoding device obtains the decoded data of the dense three-dimensional point group by decoding the encoded data of the dense three-dimensional point group (S1752). In addition, the three-dimensional data decoding device obtains the decoded data of the sparse three-dimensional point group by decoding the encoded data of the sparse three-dimensional point group (S1753). Next, the three-dimensional data decoding device generates a three-dimensional point group by combining the decoded data of the dense three-dimensional point group obtained in step S1752 with the decoded data of the sparse three-dimensional point group obtained in step S1753 (S1754).

[0826] In addition, the three-dimensional data encoding device and the three-dimensional data decoding device may encode or decode first the dense three-dimensional point group and the sparse three-dimensional point group. In addition, the encoding process or the decoding process may be performed in parallel by a plurality of processors.

[0827] In addition, the three-dimensional data encoding device may also encode one of a dense three-dimensional point group and a sparse three-dimensional point group. For example, when the dense three-dimensional point group contains important information, the three-dimensional data encoding device extracts the dense three-dimensional point group and the sparse three-dimensional point group from the input three-dimensional point group, encodes the dense three-dimensional point group, and does not encode the sparse three-dimensional point group. Thus, the three-dimensional data encoding device can attach important information to the stream while suppressing the bit amount. For example, between a server and a client, when there is a request from the client to send the three-dimensional point group information around the client to the server, the server encodes the important information around the client as a dense three-dimensional point group and sends it to the client. Thus, the server can send the information requested by the client while suppressing the network bandwidth.

[0828] In addition, the three-dimensional data decoding device may decode one of a dense three-dimensional point group and a sparse three-dimensional point group. For example, when the dense three-dimensional point group contains important information, the three-dimensional data decoding device decodes the dense three-dimensional point group and does not decode the sparse three-dimensional point group. Thus, the three-dimensional data decoding device can obtain the required information while suppressing the processing load of the decoding process.

[0829] Fig.100 yes Fig.98 Flowchart of separation processing (S1741) of three-dimensional points shown in FIG. First, the three-dimensional data encoding device sets the layer L and the threshold TH (S1761). In addition, the three-dimensional data encoding device may also add information indicating the set layer L and threshold TH to the bit stream. That is, the three-dimensional data encoding device may also generate a bit stream including information indicating the set layer L and threshold TH.

[0830] Next, the three-dimensional data encoding device moves the position of the processing target from the root node of the octree to the first branch of the layer L. That is, the three-dimensional data encoding device selects the first branch of the layer L as the processing target branch (S1762).

[0831] Next, the three-dimensional data encoding device counts the number of valid leaf nodes of the branch of the processing object of layer L (S1763). When the number of valid leaf nodes of the branch of the processing object is greater than the threshold value TH ("Yes" of S1764), the three-dimensional data encoding device registers the branch of the processing object as a dense branch to the dense three-dimensional point group (S1765). On the other hand, when the number of valid leaf nodes of the branch of the processing object is less than the threshold value TH ("No" of S1764), the three-dimensional data encoding device registers the branch of the processing object as a sparse branch to the sparse three-dimensional point group (S1766).

[0832] When the processing of all branches of layer L is not completed (No in S1767), the three-dimensional data encoding device moves the position of the processing object to the next branch of layer L. That is, the three-dimensional data encoding device selects the next branch of layer L as the branch of the processing object (S1768). And the three-dimensional data encoding device performs the processing after step S1763 on the selected next branch of the processing object.

[0833] The above-mentioned processing is repeated until the processing of all branches of layer L is completed ("Yes" in S1767).

[0834] In addition, in the above description, the layer L and the threshold TH are pre-set, but they are not necessarily limited to this. For example, the three-dimensional data encoding device sets a plurality of styles of groups of layer L and threshold TH, uses each group to generate a dense three-dimensional point group and a sparse three-dimensional point group, and encodes them separately. The three-dimensional data encoding device uses the group of layer L and threshold TH with the highest encoding efficiency of the generated encoded data among the plurality of groups, and finally encodes the dense three-dimensional point group and the sparse three-dimensional point group. This can improve the encoding efficiency. In addition, the three-dimensional data encoding device can also calculate the layer L and the threshold TH, for example. For example, the three-dimensional data encoding device can also set the value of half the maximum value of the layer contained in the tree structure for the layer L. In addition, the three-dimensional data encoding device can also set the value of half the total number of the plurality of three-dimensional points contained in the tree structure as the threshold TH.

[0835] In addition, in the above description, an example of classifying the input three-dimensional point group into two types, a dense three-dimensional point group and a sparse three-dimensional point group, is described, but the three-dimensional data encoding device may also classify the input three-dimensional point group into three or more three-dimensional point groups. For example, when the number of valid leaf nodes of the branch of the processing object is greater than the threshold value TH1, the three-dimensional data encoding device classifies the branch of the processing object into the first dense three-dimensional point group, and when the number of valid leaf nodes of the branch of the processing object is less than the first threshold value TH1 and greater than the second threshold value TH2, the branch of the processing object is classified as the second dense three-dimensional point group. When the number of valid leaf nodes of the branch of the processing object is less than the second threshold value TH2 and greater than the third threshold value TH3, the three-dimensional data encoding device classifies the branch of the processing object into the first sparse three-dimensional point group, and when the number of valid leaf nodes of the branch of the processing object is less than the threshold value TH3, the branch of the processing object is classified as the second sparse three-dimensional point group.

[0836] Hereinafter, a syntax example of the encoded data of the three-dimensional point group according to the present embodiment will be described. Fig.101 pc_header() is, for example, header information of a plurality of input three-dimensional points.

[0837] Fig.101 The num_sub_pc shown indicates the number of sub-3D point groups. numPoint[i] indicates the number of 3D points contained in the i-th sub-3D point group. coding_type[i] is the coding type information indicating the coding type (coding method) applied to the i-th sub-3D point group. For example, coding_type=...

Claims

1. A three-dimensional data encoding method, wherein: entropy encoding a bit string of an N-ary tree structure representing a plurality of three-dimensional points included in the three-dimensional data using a context set selected from a plurality of context sets, where N is an integer greater than or equal to 2; The bit string contains N bits of information for each node in the N-ary tree structure; The N bits of information include N 1-bit information indicating whether a three-dimensional point exists in each of the N child nodes of the corresponding node; In each of the plurality of context sets, a context is provided for each bit of the N bits of information; In the entropy coding, each bit of the N-bit information is entropy coded using the context set for the bit in the selected context set.

2. The three-dimensional data encoding method according to claim 1, wherein: In the entropy encoding, the context set to be used is selected from the plurality of context sets based on whether a three-dimensional point exists in each of a plurality of neighboring nodes adjacent to a target node.

3. The three-dimensional data encoding method according to claim 2, wherein: In the entropy coding, Selecting the context set based on a configuration pattern representing the configuration positions of the adjacent nodes where the three-dimensional points exist among the plurality of adjacent nodes; For configuration patterns among the configuration patterns that become the same configuration pattern through rotation, the same context set is selected.

4. The three-dimensional data encoding method according to claim 1, wherein: In the entropy encoding, the context set to be used is selected from the plurality of context sets based on the layer to which the object node belongs.

5. The three-dimensional data encoding method according to claim 1, wherein: In the entropy encoding, the context set to be used is selected from the plurality of context sets based on a normal vector of an object node.

6. A three-dimensional data decoding method, wherein: entropy decoding a bit string of an N-ary tree structure representing a plurality of three-dimensional points included in the three-dimensional data using a context set selected from a plurality of context sets, where N is an integer greater than or equal to 2; The bit string contains N bits of information for each node in the N-ary tree structure; The N bits of information include N 1-bit information indicating whether a three-dimensional point exists in each of the N child nodes of the corresponding node; In each of the plurality of context sets, a context is provided for each bit of the N bits of information; In the entropy decoding, each bit of the N-bit information is entropy decoded using the context set for the bit in the selected context set.

7. The three-dimensional data decoding method according to claim 6, wherein: In the entropy decoding, the context set to be used is selected from the plurality of context sets based on whether or not a three-dimensional point exists in each of a plurality of neighboring nodes adjacent to the object node.

8. The three-dimensional data decoding method according to claim 7, wherein: In the entropy decoding, Selecting the context set based on a configuration pattern representing the configuration positions of the adjacent nodes where the three-dimensional points exist among the plurality of adjacent nodes; For configuration patterns among the configuration patterns that become the same configuration pattern through rotation, the same context set is selected.

9. The three-dimensional data decoding method according to claim 6, wherein: In the entropy decoding, the context set to be used is selected from the plurality of context sets based on the layer to which the object node belongs.

10. The three-dimensional data decoding method according to claim 6, wherein: In the entropy decoding, the context set to be used is selected from the plurality of context sets based on a normal vector of the object node.

11. A three-dimensional data encoding device, wherein: have: Processor; and Memory; The processor uses the memory to entropy encode a bit string of an N-ary tree structure representing a plurality of three-dimensional points included in the three-dimensional data using a context set selected from a plurality of context sets, where N is an integer greater than or equal to 2; The bit string contains N bits of information for each node in the N-ary tree structure; The N bits of information include N 1-bit information indicating whether a three-dimensional point exists in each of the N child nodes of the corresponding node; In each of the plurality of context sets, a context is provided for each bit of the N bits of information; In the entropy coding, each bit of the N-bit information is entropy coded using the context set for the bit in the selected context set.

12. A three-dimensional data decoding device, wherein: have: Processor; and Memory; The processor uses the memory to entropy decode a bit string of an N-ary tree structure representing a plurality of three-dimensional points included in the three-dimensional data using a context set selected from a plurality of context sets, where N is an integer greater than or equal to 2; The bit string contains N bits of information for each node in the N-ary tree structure; The N bits of information include N 1-bit information indicating whether a three-dimensional point exists in each of the N child nodes of the corresponding node; In each of the plurality of context sets, a context is provided for each bit of the N bits of information; In the entropy decoding, each bit of the N-bit information is entropy decoded using the context set for the bit in the selected context set.

Citation Information

Patent Citations

  • Map display device

    WO2014020663A1