Three-dimensional data encoding method, three-dimensional data decoding method, three-dimensional data encoding device, and three-dimensional data decoding device
By using the bit string of N forktree structure in three-dimensional data encoding for entropy encoding and switching the encoding table according to the configuration of different nodes and normal vectors, the problem of low 3D data encoding efficiency in the prior art is solved, and more efficient data compression and coding performance is achieved.
Patent Information
- Application Number
- CN202510489731.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2018-01-26
- Filing Date
- 2019-01-24
- Publication Date
- 2025-06-13
AI Technical Summary
The prior art is relatively low in encoding three-dimensional data, and it is difficult to effectively compress large-scale three-dimensional point cloud data.
The bit string with an N forktree structure is used for entropy encoding, and the encoding efficiency is improved by switching the context and encoding table. The specific method includes using multiple encoding tables to select the most suitable encoding table for encoding according to the configuration of different nodes and normal vectors.
It significantly improves the efficiency of three-dimensional data encoding, reduces the compression of data volume, and improves the encoding and decoding performance.
Smart Images

Figure CN120147553A_ABST
Abstract
Description
[0001] This application is a divisional of a patent application for an invention titled "3D Data Encoding Method, 3D Data Decoding Method, 3D Data Encoding Apparatus, and 3D Data Decoding Apparatus" with an application date of January 24, 2019, an application number of 201980013461.3. Technical Field
[0002] The present disclosure relates to a 3D data encoding method, a 3D data decoding method, a 3D data encoding apparatus, and a 3D data decoding apparatus. Background Art
[0003] In large fields such as computer vision, map information, monitoring, infrastructure inspection, or video distribution for autonomous operation of automobiles or robots, devices or services that make flexible use of 3D data will be popularized in the future. 3D data is obtained by various methods such as distance sensors like rangefinders, stereo cameras, or combinations of multiple monocular cameras.
[0004] As a representation method of 3D data, there is a representation method called point cloud, which represents the shape of a 3D structure by a point group in a 3D space (for example, refer to Non-Patent Document 1). In the point cloud, the positions and colors of the point group are stored. Although it is expected that the point cloud will become the mainstream as a representation method of 3D data, the data volume of the point group is very large. Therefore, in the storage or transmission of 3D data, like in the case of 2D moving images (as an example, MPEG-4 AVC or HEVC standardized by MPEG), data volume compression is required through encoding.
[0005] In addition, for the compression of point clouds, there is support from some publicly available libraries (such as the PointCloud Library) that perform point cloud correlation processing.
[0006] In addition, there is a well-known technology that uses 3D map data to retrieve facilities around a vehicle and display them (for example, refer to Patent Document 1).
[0007] Prior Art Documents
[0008] Patent Documents
[0009] Patent Document 1: International Publication No. 2014 / 020663 Summary of the Invention
[0010] Problems to be Solved by the Invention
[0011] It is desired to improve the encoding efficiency in the encoding of 3D data.
[0012] An object of the present disclosure is to provide a three-dimensional data encoding method, a three-dimensional data decoding method, a three-dimensional data encoding apparatus, or a three-dimensional data decoding apparatus capable of improving the encoding efficiency.
[0013] Means for solving the problem
[0014] Regarding a three-dimensional data encoding method according to an aspect of the present disclosure, entropy encoding is performed on a bit string of an N-ary tree structure representing a plurality of three-dimensional points included in three-dimensional data using first to Nth contexts, where N is an integer of 2 or more. The bit string includes information of first to Nth bits for each node in the N-ary tree structure, and the information of the first to Nth bits includes N pieces of 1-bit information indicating whether three-dimensional points exist in respective N child nodes corresponding to the node. In the entropy encoding, the information of the first to Nth bits is entropy encoded using the first to Nth contexts respectively corresponding to the first to Nth bits.
[0015] Regarding a three-dimensional data decoding method according to an aspect of the present disclosure, entropy decoding is performed on a bit string of an N-ary tree structure representing a plurality of three-dimensional points included in three-dimensional data using first to Nth contexts, where N is an integer of 2 or more. The bit string includes information of first to Nth bits for each node in the N-ary tree structure, and the information of the first to Nth bits includes N pieces of 1-bit information indicating whether three-dimensional points exist in respective N child nodes corresponding to the node. In the entropy decoding, the information of the first to Nth bits is entropy decoded using the first to Nth contexts respectively corresponding to the first to Nth bits.
[0016] Effect of the invention
[0017] The present disclosure can provide a three-dimensional data encoding method, a three-dimensional data decoding method, a three-dimensional data encoding apparatus, or a three-dimensional data decoding apparatus capable of improving the encoding efficiency. Description of the drawings
[0018] Figure 1 Shows the configuration of encoded three-dimensional data according to Embodiment 1.
[0019] Figure 2 Shows an example of a prediction structure between SPCs belonging to the lowest layer of GOS according to Embodiment 1.
[0020] Figure 3 Shows an example of an inter-layer prediction structure according to Embodiment 1.
[0021] Figure 4 Shows an example of the encoding order of GOS according to Embodiment 1.
[0022] Figure 5Shows an example of the encoding order of GOS related to Embodiment 1.
[0023] Figure 6 Is a block diagram of the three-dimensional data encoding device related to Embodiment 1.
[0024] Figure 7 Is a flowchart of the encoding process related to Embodiment 1.
[0025] Figure 8 Is a block diagram of the three-dimensional data decoding device related to Embodiment 1.
[0026] Figure 9 Is a flowchart of the decoding process related to Embodiment 1.
[0027] Figure 10 Shows an example of the meta information related to Embodiment 1.
[0028] Figure 11 Shows a configuration example of SWLD related to Embodiment 2.
[0029] Figure 12 Shows an operation example of the server and the client related to Embodiment 2.
[0030] Figure 13 Shows an operation example of the server and the client related to Embodiment 2.
[0031] Figure 14 Shows an operation example of the server and the client related to Embodiment 2.
[0032] Figure 15 Shows an operation example of the server and the client related to Embodiment 2.
[0033] Figure 16 Is a block diagram of the three-dimensional data encoding device related to Embodiment 2.
[0034] Figure 17 Is a flowchart of the encoding process related to Embodiment 2.
[0035] Figure 18 Is a block diagram of the three-dimensional data decoding device related to Embodiment 2.
[0036] Figure 19 Is a flowchart of the decoding process related to Embodiment 2.
[0037] Figure 20 Shows a configuration example of WLD related to Embodiment 2.
[0038] Figure 21Shows an example of the octree structure of the WLD related to Embodiment 2.
[0039] Figure 22 Shows a configuration example of the SWLD related to Embodiment 2.
[0040] Figure 23 Shows an example of the octree structure of the SWLD related to Embodiment 2.
[0041] Figure 24 Is a block diagram of the three-dimensional data production device related to Embodiment 3.
[0042] Figure 25 Is a block diagram of the three-dimensional data transmission device related to Embodiment 3.
[0043] Figure 26 Is a block diagram of the three-dimensional information processing device related to Embodiment 4.
[0044] Figure 27 Is a block diagram of the three-dimensional data production device related to Embodiment 5.
[0045] Figure 28 Shows the configuration of the system related to Embodiment 6.
[0046] Figure 29 Is a block diagram of the client device related to Embodiment 6.
[0047] Figure 30 Is a block diagram of the server related to Embodiment 6.
[0048] Figure 31 Is a flowchart of the three-dimensional data production process performed by the client device related to Embodiment 6.
[0049] Figure 32 Is a flowchart of the sensor information transmission process performed by the client device related to Embodiment 6.
[0050] Figure 33 Is a flowchart of the three-dimensional data production process performed by the server related to Embodiment 6.
[0051] Figure 34 Is a flowchart of the three-dimensional map transmission process performed by the server related to Embodiment 6.
[0052] Figure 35 Shows the configuration of a modified example of the system related to Embodiment 6.
[0053] Figure 36 Shows the configuration of the server and client device related to Embodiment 6.
[0054] Figure 37 It is a block diagram of the three-dimensional data encoding device according to Embodiment 7.
[0055] Figure 38 An example of the prediction residual according to Embodiment 7 is shown.
[0056] Figure 39 An example of the volume according to Embodiment 7 is shown.
[0057] Figure 40 An example of the octree representation of the volume according to Embodiment 7 is shown.
[0058] Figure 41 An example of the bit string of the volume according to Embodiment 7 is shown.
[0059] Figure 42 An example of the octree representation of the volume according to Embodiment 7 is shown.
[0060] Figure 43 An example of the volume according to Embodiment 7 is shown.
[0061] Figure 44 It is a diagram for explaining the intra prediction process according to Embodiment 7.
[0062] Figure 45 It is a diagram for explaining the rotation and translation processes according to Embodiment 7.
[0063] Figure 46 An example of the syntax of the RT application flag and RT information according to Embodiment 7 is shown.
[0064] Figure 47 It is a diagram for explaining the inter prediction process according to Embodiment 7.
[0065] Figure 48 It is a block diagram of the three-dimensional data decoding device according to Embodiment 7.
[0066] Figure 49 It is a flowchart of the three-dimensional data encoding process performed by the three-dimensional data encoding device according to Embodiment 7.
[0067] Figure 50 It is a flowchart of the three-dimensional data decoding process performed by the three-dimensional data decoding device according to Embodiment 7.
[0068] Figure 51 The configuration of the distribution system according to Embodiment 8 is shown.
[0069] Figure 52 An example of the configuration of the bitstream of the encoded three-dimensional map according to Embodiment 8 is shown.
[0070] Figure 53 It is a diagram for explaining the improvement effect of the coding efficiency related to Embodiment 8.
[0071] Figure 54 It is a flowchart of the processing performed by the server related to Embodiment 8.
[0072] Figure 55 It is a flowchart of the processing performed by the client related to Embodiment 8.
[0073] Figure 56 It shows a syntactic example of the sub - map related to Embodiment 8.
[0074] Figure 57 It shows the switching process of the coding type related to Embodiment 8 in a schematic manner.
[0075] Figure 58 It shows a syntactic example of the sub - map related to Embodiment 8.
[0076] Figure 59 It is a flowchart of the three - dimensional data encoding process related to Embodiment 8.
[0077] Figure 60 It is a flowchart of the three - dimensional data decoding process related to Embodiment 8.
[0078] Figure 61 It shows the operation of a modified example of the switching process of the coding type related to Embodiment 8 in a schematic manner.
[0079] Figure 62 It shows the operation of a modified example of the switching process of the coding type related to Embodiment 8 in a schematic manner.
[0080] Figure 63 It shows the operation of a modified example of the switching process of the coding type related to Embodiment 8 in a schematic manner.
[0081] Figure 64 It shows the operation of a modified example of the calculation process of the difference value related to Embodiment 8 in a schematic manner.
[0082] Figure 65 It shows the operation of a modified example of the calculation process of the difference value related to Embodiment 8 in a schematic manner.
[0083] Figure 66 It shows the operation of a modified example of the calculation process of the difference value related to Embodiment 8 in a schematic manner.
[0084] Figure 67 It shows the operation of a modified example of the calculation process of the difference value related to Embodiment 8 in a schematic manner.
[0085] Figure 68 Shows a syntactic example of the volume related to Embodiment 8.
[0086] Figure 69 Is a diagram showing an example of an important area related to Embodiment 9.
[0087] Figure 70 Is a diagram showing an example of an occupancy code related to Embodiment 9.
[0088] Figure 71 Is a diagram showing an example of a quadtree structure related to Embodiment 9.
[0089] Figure 72 Is a diagram showing an example of an occupancy code and a position code related to Embodiment 9.
[0090] Figure 73 Is a diagram showing an example of three-dimensional points obtained in the LiDAR related to Embodiment 9.
[0091] Figure 74 Is a diagram showing an example of an octree structure related to Embodiment 9.
[0092] Figure 75 Is a diagram showing an example of hybrid coding related to Embodiment 9.
[0093] Figure 76 Is a diagram for explaining a method of switching between position coding and occupancy coding related to Embodiment 9.
[0094] Figure 77 Is a diagram showing an example of a bitstream of position coding related to Embodiment 9.
[0095] Figure 78 Is a diagram showing an example of a bitstream of hybrid coding related to Embodiment 9.
[0096] Figure 79 Is a diagram showing a tree structure of the occupancy code of important three-dimensional points related to Embodiment 9.
[0097] Figure 80 Is a diagram showing a tree structure of the occupancy code of non-important three-dimensional points related to Embodiment 9.
[0098] Figure 81 Is a diagram showing an example of a bitstream of hybrid coding related to Embodiment 9.
[0099] Figure 82 Is a diagram showing an example of a bitstream including coding mode information related to Embodiment 9.
[0100] Figure 83 It is a diagram showing a syntactic example related to Embodiment 9.
[0101] Figure 84 It is a flowchart of the encoding process related to Embodiment 9.
[0102] Figure 85 It is a flowchart of the node encoding process related to Embodiment 9.
[0103] Figure 86 It is a flowchart of the decoding process related to Embodiment 9.
[0104] Figure 87 It is a flowchart of the node decoding process related to Embodiment 9.
[0105] Figure 88 It is a diagram showing an example of the tree structure related to Embodiment 10.
[0106] Figure 89 It is a diagram showing an example of the number of valid leaf nodes in each branch related to Embodiment 10.
[0107] Figure 90 It is a diagram showing an application example of the encoding method related to Embodiment 10.
[0108] Figure 91 It is a diagram showing an example of a dense branch area related to Embodiment 10.
[0109] Figure 92 It is a diagram showing an example of a dense three-dimensional point group related to Embodiment 10.
[0110] Figure 93 It is a diagram showing an example of a sparse three-dimensional point group related to Embodiment 10.
[0111] Figure 94 It is a flowchart of the encoding process related to Embodiment 10.
[0112] Figure 95 It is a flowchart of the decoding process related to Embodiment 10.
[0113] Figure 96 It is a flowchart of the encoding process related to Embodiment 10.
[0114] Figure 97 It is a flowchart of the decoding process related to Embodiment 10.
[0115] Figure 98 It is a flowchart of the encoding process related to Embodiment 10.
[0116] Figure 99It is a flowchart of the decoding process according to Embodiment 10.
[0117] Figure 100 It is a flowchart showing the separation process of three-dimensional points according to Embodiment 10.
[0118] Figure 101 It is a diagram showing a syntactic example according to Embodiment 10.
[0119] Figure 102 It is a diagram showing an example of a dense branch according to Embodiment 10.
[0120] Figure 103 It is a diagram showing an example of a sparse branch according to Embodiment 10.
[0121] Figure 104 It is a flowchart of the encoding process of a modification example according to Embodiment 10.
[0122] Figure 105 It is a flowchart of the decoding process of a modification example according to Embodiment 10.
[0123] Figure 106 It is a flowchart of the separation process of three-dimensional points of a modification example according to Embodiment 10.
[0124] Figure 107 It is a diagram showing a syntactic example of a modification example according to Embodiment 10.
[0125] Figure 108 It is a flowchart of the encoding process according to Embodiment 10.
[0126] Figure 109 It is a flowchart of the decoding process according to Embodiment 10.
[0127] Figure 110 It is a diagram showing an example of a tree structure according to Embodiment 11.
[0128] Figure 111 It is a diagram showing an example of an occupancy code according to Embodiment 11.
[0129] Figure 112 It is a diagram schematically showing the operation of a three-dimensional data encoding device according to Embodiment 11.
[0130] Figure 113 It is a diagram showing an example of geometric information according to Embodiment 11.
[0131] Figure 114 It is a diagram showing a selection example of an encoding table using geometric information according to Embodiment 11.
[0132] Figure 115It is a diagram showing a selection example of a coding table indicating usage structure information regarding Embodiment 11.
[0133] Figure 116 It is a diagram showing a selection example of a coding table indicating usage attribute information regarding Embodiment 11.
[0134] Figure 117 It is a diagram showing a selection example of a coding table indicating usage attribute information regarding Embodiment 11.
[0135] Figure 118 It is a diagram showing a structural example of a bitstream regarding Embodiment 11.
[0136] Figure 119 It is a diagram showing an example of a coding table regarding Embodiment 11.
[0137] Figure 120 It is a diagram showing an example of a coding table regarding Embodiment 11.
[0138] Figure 121 It is a diagram showing a structural example of a bitstream regarding Embodiment 11.
[0139] Figure 122 It is a diagram showing an example of a coding table regarding Embodiment 11.
[0140] Figure 123 It is a diagram showing an example of a coding table regarding Embodiment 11.
[0141] Figure 124 It is a diagram showing an example of the bit numbers of occupancy rate codes regarding Embodiment 11.
[0142] Figure 125 It is a flowchart of the coding process for using geometric information regarding Embodiment 11.
[0143] Figure 126 It is a flowchart of the decoding process for using geometric information regarding Embodiment 11.
[0144] Figure 127 It is a flowchart of the coding process for using structure information regarding Embodiment 11.
[0145] Figure 128 It is a flowchart of the decoding process for using structure information regarding Embodiment 11.
[0146] Figure 129 It is a flowchart of the coding process for using attribute information regarding Embodiment 11.
[0147] Figure 130 It is a flowchart of the decoding process for using attribute information regarding Embodiment 11.
[0148] Figure 131 This is a flowchart of the encoding table selection process using geometric information according to Embodiment 11.
[0149] Figure 132 This is a flowchart of the encoding table selection process using structural information according to Embodiment 11.
[0150] Figure 133 This is a flowchart of the encoding table selection process using attribute information according to Embodiment 11.
[0151] Figure 134 This is a block diagram of the three-dimensional data encoding device according to Embodiment 11.
[0152] Figure 135 This is a block diagram of the three-dimensional data decoding device according to Embodiment 11. Detailed implementation manners
[0153] Regarding a three-dimensional data encoding method according to one aspect of the present disclosure, an entropy encoding is performed on a bit string representing an N-ary tree structure (where N is an integer of 2 or more) of a plurality of three-dimensional points included in three-dimensional data using an encoding table selected from a plurality of encoding tables; the bit string includes N bits of information for each node in the N-ary tree structure; the N bits of information include N 1-bit pieces of information indicating whether there is a three-dimensional point in each of the N child nodes corresponding to the node; in each of the plurality of encoding tables, contexts are provided for each bit of the N bits of information; in the entropy encoding, each bit of the N bits of information is entropy encoded using the context set for the bit in the selected encoding table.
[0154] Thus, this three-dimensional data encoding method can improve the encoding efficiency by switching the context for each bit.
[0155] For example, in the entropy encoding, the encoding table to be used may also be selected from the plurality of encoding tables based on whether there is a three-dimensional point in each of a plurality of adjacent nodes adjacent to the target node.
[0156] Thus, this three-dimensional data encoding method can improve the encoding efficiency by switching the encoding table based on whether there is a three-dimensional point in the adjacent nodes.
[0157] For example, in the entropy encoding, the encoding table may also be selected based on a configuration pattern indicating the configuration positions of the adjacent nodes having three-dimensional points among the plurality of adjacent nodes; for configuration patterns that become the same configuration pattern by rotation in the configuration patterns, the same encoding table is selected.
[0158] Thus, this three-dimensional data encoding method can suppress an increase in the encoding table.
[0159] For example, in the entropy encoding, the encoding table to be used may also be selected from the multiple encoding tables based on the layer to which the object node belongs.
[0160] Thus, by switching the encoding table based on the layer to which the object node belongs, this three-dimensional data encoding method can improve the encoding efficiency.
[0161] For example, in the entropy encoding, the encoding table to be used may also be selected from the multiple encoding tables based on the normal vector of the object node.
[0162] Thus, by switching the encoding table based on the normal vector, this three-dimensional data encoding method can improve the encoding efficiency.
[0163] In addition, regarding a three-dimensional data decoding method according to an aspect of the present disclosure, an entropy decoding is performed on a bit string of an N-ary tree structure (where N is an integer of 2 or more) representing a plurality of three-dimensional points included in three-dimensional data using an encoding table selected from multiple encoding tables; each node in the N-ary tree structure of the bit string includes N bits of information; the N bits of information include N one-bit pieces of information indicating whether there is a three-dimensional point in each of the N child nodes corresponding to the node; in each of the multiple encoding tables, contexts are set for each bit of the N bits of information; in the entropy decoding, each bit of the N bits of information is entropy decoded using the context set for the bit in the selected encoding table.
[0164] Thus, this three-dimensional data decoding method can improve the encoding efficiency by switching the context for each bit.
[0165] For example, in the entropy decoding, the encoding table to be used may also be selected from the multiple encoding tables based on whether there is a three-dimensional point in each of a plurality of adjacent nodes adjacent to the object node.
[0166] Thus, by switching the encoding table based on whether there is a three-dimensional point in the adjacent nodes, this three-dimensional data decoding method can improve the encoding efficiency.
[0167] For example, in the entropy decoding, the encoding table may also be selected based on a configuration pattern indicating the configuration positions of the adjacent nodes having three-dimensional points among the plurality of adjacent nodes; for configuration patterns that become the same configuration pattern by rotation in the configuration pattern, the same encoding table is selected.
[0168] Thus, this three-dimensional data decoding method can suppress an increase in the encoding table.
[0169] For example, in the entropy decoding, the encoding table to be used may also be selected from the multiple encoding tables based on the layer to which the object node belongs.
[0170] Thus, the 3D data decoding method can improve the encoding efficiency by switching the encoding table based on the layer to which the object node belongs.
[0171] For example, in the entropy decoding, the encoding table to be used can also be selected from the multiple encoding tables based on the normal vector of the object's node.
[0172] Thus, the 3D data decoding method can improve the encoding efficiency by switching the encoding table based on the normal vector.
[0173] In addition, a 3D data encoding apparatus according to an aspect of the present disclosure includes a processor and a memory; the processor uses the memory to perform the following processing: entropy-encoding a bit string representing an N-ary tree structure (where N is an integer of 2 or more) of a plurality of 3D points included in 3D data using an encoding table selected from a plurality of encoding tables; the bit string includes N bits of information for each node in the N-ary tree structure; the N bits of information include N 1-bit pieces of information indicating whether there are 3D points in each of the N child nodes corresponding to the node; in each of the plurality of encoding tables, contexts are set for the respective bits of the N bits of information; in the entropy encoding, each bit of the N bits of information is entropy-encoded using the context set for the bit in the selected encoding table.
[0174] Thus, the 3D data encoding apparatus can improve the encoding efficiency by switching the context for each bit.
[0175] In addition, a 3D data decoding apparatus according to an aspect of the present disclosure includes a processor and a memory; the processor uses the memory to perform the following processing: entropy-decoding a bit string representing an N-ary tree structure (where N is an integer of 2 or more) of a plurality of 3D points included in 3D data using an encoding table selected from a plurality of encoding tables; the bit string includes N bits of information for each node in the N-ary tree structure; the N bits of information include N 1-bit pieces of information indicating whether there are 3D points in each of the N child nodes corresponding to the node; in each of the plurality of encoding tables, contexts are set for the respective bits of the N bits of information; in the entropy decoding, each bit of the N bits of information is entropy-decoded using the context set for the bit in the selected encoding table.
[0176] Thus, the 3D data decoding apparatus can improve the encoding efficiency by switching the context for each bit.
[0177] In addition, these general or specific aspects can be implemented by a system, a method, an integrated circuit, a computer program, or a recording medium such as a computer-readable CD-ROM, and can be implemented by any combination of a system, a method, an integrated circuit, a computer program, and a recording medium.
[0178] The following describes the embodiments with reference to the accompanying drawings. In addition, all the embodiments to be described below are specific examples showing the present disclosure. The numerical values, shapes, materials, constituent elements, arrangement positions and connection forms of the constituent elements, steps, the order of steps, etc. shown in the following embodiments are all examples, and the gist thereof is not to limit the present disclosure. And, among the constituent elements of the following embodiments, the constituent elements not described in the technical solution showing the uppermost concept are described as optional constituent elements.
[0179] (Embodiment 1)
[0180] First, the data structure of the encoded three-dimensional data (hereinafter also referred to as encoded data) related to this embodiment will be described. Figure 1 The configuration of the encoded three-dimensional data related to this embodiment is shown.
[0181] In this embodiment, the three-dimensional space is divided into spaces (SPC) corresponding to pictures in the encoding of moving images, and the three-dimensional data is encoded in units of space. The space is further divided into volumes (VLM) corresponding to macroblocks and the like in moving image encoding, and prediction and transformation are performed in units of VLM. A volume includes a plurality of voxels (VXL) which are the smallest units corresponding to position coordinates. In addition, prediction means, similar to the prediction performed in a two-dimensional image, referring to other processing units, generating predicted three-dimensional data similar to the processing unit to be processed, and encoding the difference between the predicted three-dimensional data and the processing unit to be processed. And this prediction includes not only spatial prediction referring to other prediction units at the same time, but also temporal prediction referring to prediction units at different times.
[0182] For example, when a three-dimensional data encoding device (hereinafter also referred to as an encoding device) encodes a three-dimensional space represented by point cloud data such as a point cloud, it encodes each point of the point cloud or a plurality of points included in a voxel together according to the size of the voxel. If the voxel is subdivided, the three-dimensional shape of the point cloud can be represented with high precision, and if the size of the voxel is increased, the three-dimensional shape of the point cloud can be represented roughly.
[0183] In addition, although the case where the three-dimensional data is a point cloud is described as an example below, the three-dimensional data is not limited to the point cloud and can be any form of three-dimensional data.
[0184] Furthermore, voxels with a hierarchical structure can be utilized. In this case, in the n-th hierarchy, it is possible to sequentially show whether there are sampling points in the hierarchies below the (n - 1)-th hierarchy (the lower layer of the n-th hierarchy). For example, when decoding only the n-th hierarchy, if there are sampling points in the hierarchies below the (n - 1)-th hierarchy, it is possible to perform decoding by considering that there are sampling points at the center of the voxels in the n-th hierarchy.
[0185] Furthermore, the encoding device obtains point cloud data through a distance sensor, a stereo camera, a monocular camera, a gyroscope, or an inertial sensor, etc.
[0186] Regarding space, similar to the encoding of moving images, it is at least classified into any one of the following three prediction structures: an intra-frame space (I-SPC) that can be decoded independently, a predictive space (P-SPC) that can only be referenced unidirectionally, and a bidirectional space (B-SPC) that can be referenced bidirectionally. And space has two types of time information: a decoding time and a display time.
[0187] Furthermore, as Figure 1 shown, as a processing unit including multiple spaces, there is a GOS (Group Of Space) as a random access unit. Moreover, as a processing unit including multiple GOSs, there is a world space (WLD).
[0188] The spatial region occupied by the world space is associated with the absolute position on the earth through GPS or latitude and longitude information, etc. This position information is stored as meta-information. In addition, the meta-information can be included in the encoded data or transmitted separately from the encoded data.
[0189] Furthermore, within a GOS, all SPCs can be three-dimensionally adjacent, or there can be SPCs that are not three-dimensionally adjacent to other SPCs.
[0190] In addition, hereinafter, the processes such as encoding, decoding, or referencing corresponding to the three-dimensional data included in processing units such as GOS, SPC, or VLM will also be simply referred to as encoding, decoding, or referencing the processing unit. And the three-dimensional data included in the processing unit includes at least one group of spatial positions such as three-dimensional coordinates and characteristic values such as color information.
[0191] Next, the prediction structure of SPCs in a GOS will be described. Multiple SPCs within the same GOS, or multiple VLMs within the same SPC, although they occupy different spaces from each other, hold the same time information (decoding time and display time).
[0192] Also, within a GOS, the SPC that is the first in decoding order is an I-SPC. And there are two types of GOSs in the GOS, namely a closed GOS and an open GOS. A closed GOS is a GOS in which all SPCs within the GOS can be decoded when starting decoding from the initial I-SPC. In an open GOS, within the GOS, some SPCs that are earlier in display time than the initial I-SPC refer to different GOSs and can only be decoded in those GOSs.
[0193] In addition, in encoded data such as map information, there are cases where the WLD is decoded in the direction opposite to the encoding order. If there are dependencies between GOSs, it is difficult to perform reverse reproduction. Therefore, in such cases, a closed GOS is basically adopted.
[0194] Also, the GOS has a layer structure in the height direction, and encoding or decoding is performed sequentially starting from the SPCs in the bottom layer.
[0195] Figure 2 An example of the prediction structure between SPCs belonging to the bottommost layer of the GOS is shown. Figure 3 An example of the inter-layer prediction structure is shown.
[0196] There is more than one I-SPC within the GOS. Although there are objects such as people, animals, cars, bicycles, traffic lights, or buildings that serve as land marks in a three-dimensional space, it is particularly effective when encoding small-sized objects as I-SPCs. For example, when a three-dimensional data decoding device (hereinafter also referred to as a decoding device) decodes a GOS with a low processing amount or at high speed, it only decodes the I-SPCs within the GOS.
[0197] Also, the encoding device can switch the encoding interval or occurrence frequency of the I-SPCs according to the density of the objects within the WLD.
[0198] And, in Figure 3 In the configuration shown, the encoding device or the decoding device encodes or decodes multiple layers sequentially starting from the lower layer (layer 1). Accordingly, for example, for an automatically moving vehicle or the like, the priority of data near the ground with a large amount of information can be increased.
[0199] In addition, in encoded data used in a drone or the like, within the GOS, encoding or decoding can be performed sequentially starting from the SPCs in the upper layer in the height direction.
[0200] Also, the encoding device or the decoding device can encode or decode multiple layers in such a way that the decoding device can generally grasp the GOS and gradually increase the resolution. For example, the encoding device or the decoding device can perform encoding or decoding in the order of layer 3, 8, 1, 9...
[0201] Next, the corresponding methods for static objects and dynamic objects will be described.
[0202] In a three-dimensional space, there are static objects or scenes such as buildings or roads (collectively referred to as static objects hereafter), and dynamic objects such as vehicles or people (referred to as dynamic objects hereafter). Detection of objects can be performed separately by extracting feature points from data of point clouds, or images captured by a stereo camera or the like. Here, an example of an encoding method for dynamic objects will be described.
[0203] The first method is a method of encoding without distinguishing between static objects and dynamic objects. The second method is a method of distinguishing between static objects and dynamic objects by identification information.
[0204] For example, GOS is used as an identification unit. In this case, the GOS including the SPC constituting the static object and the GOS including the SPC constituting the dynamic object are distinguished within the encoded data or by identification information stored separately from the encoded data.
[0205] Alternatively, SPC is used as an identification unit. In this case, only the SPC including the VLM constituting the static object and the SPC including the VLM constituting the dynamic object are distinguished by the above-mentioned identification information.
[0206] Alternatively, VLM or VXL can be used as an identification unit. In this case, the VLM or VXL including the static object and the VLM or VXL including the dynamic object are distinguished by the above-mentioned identification information.
[0207] Moreover, the encoding device can encode the dynamic object as one or more VLMs or SPCs, and encode the VLM or SPC including the static object and the SPC including the dynamic object as different GOSs. And when the size of the GOS becomes variable according to the size of the dynamic object, the size of the GOS is stored separately as meta information.
[0208] And, the encoding device encodes the static object and the dynamic object independently of each other, and for the world space composed of the static object, the dynamic object can be overlapped. At this time, the dynamic object is composed of one or more SPCs, and each SPC corresponds to one or more SPCs of the static object overlapping the SPC. In addition, the dynamic object may not be represented by SPC, and may be represented by one or more VLMs or VXLs.
[0209] And, the encoding device can encode the static object and the dynamic object as different streams.
[0210] Also, the encoding device can also generate a GOS including one or more SPCs that constitute a dynamic object. Moreover, the encoding device can set the GOS (GOS_M) including the dynamic object and the GOS of the static object corresponding to the spatial region of GOS_M to be the same size (occupy the same spatial region). In this way, overlapping processing can be performed in units of GOS.
[0211] The P-SPC or B-SPC constituting the dynamic object can also refer to the SPCs included in different encoded GOSs. The position of the dynamic object changes over time. In the case where the same dynamic object is encoded as GOSs at different times, cross-GOS reference is effective from the viewpoint of compression ratio.
[0212] Also, the above first method and second method can be switched according to the use of the encoded data. For example, when the encoded three-dimensional data is applied as a map, since it is desired to separate from the dynamic object, the encoding device adopts the second method. In addition, when the encoding device encodes the three-dimensional data of an event such as a concert or sports, if there is no need to separate the dynamic object, the first method is adopted.
[0213] Also, the decoding time and display time of the GOS or SPC can be stored in the encoded data or as meta-information. And the time information of the static object can all be the same. At this time, the actual decoding time and display time can be determined by the decoding device. Or, as the decoding time, different values are assigned to each GOS or SPC, and as the display time, the same value can be assigned to all. Moreover, as shown in the decoder mode in dynamic image coding such as HRD (Hypothetical Reference Decoder) of HEVC, the decoder has a buffer of a specified size, and as long as the bitstream is read at a specified bit rate according to the decoding time, a model that will not be damaged and can be guaranteed to be decoded can be imported.
[0214] Next, the configuration of GOSs in the world space will be described. The coordinates of the three-dimensional space in the world space are represented by three mutually orthogonal coordinate axes (x-axis, y-axis, z-axis). By setting a specified rule in the encoding order of GOSs, GOSs that are spatially adjacent can be encoded continuously in the encoded data. For example, in Figure 4 the example shown, the GOSs in the xz plane are encoded continuously. After the encoding of all the GOSs in one xz plane is completed, the value of the y-axis is updated. That is, as the encoding progresses, the world space extends in the y-axis direction. And the index number of the GOS is set as the encoding order.
[0215] Here, the three-dimensional space of the world space corresponds one-to-one with the absolute coordinates such as GPS, or latitude and longitude. Alternatively, the three-dimensional space can be represented by the relative position with respect to a preset reference position. The directions of the x-axis, y-axis, and z-axis of the three-dimensional space are represented as direction vectors determined based on latitude and longitude, etc., and this direction vector is stored together with the encoded data as meta information.
[0216] Moreover, the size of the GOS is set to be fixed, and the encoding device stores this size as meta information. Moreover, the size of the GOS can be switched according to, for example, whether it is in the city, indoors, outdoors, etc. That is, the size of the GOS can be switched according to the quantity or nature of the object having the value as information. Alternatively, the encoding device can appropriately switch the size of the GOS or the interval of the I-SPC within the GOS according to the density of the object, etc. within the same world space. For example, when the density of the object is higher, the encoding device sets the size of the GOS to be smaller and the interval of the I-SPC within the GOS to be shorter.
[0217] In Figure 5 the example, in the area of the 3rd to 10th GOS, since the density of the object is high, in order to achieve fine-grained random access, the GOS is subdivided. Moreover, the 7th to 10th GOSs are respectively located behind the 3rd to 6th GOSs.
[0218] Next, the configuration and the working process of the three-dimensional data encoding device according to the present embodiment will be described. Figure 6 is a block diagram of the three-dimensional data encoding device 100 according to the present embodiment. Figure 7 is a flowchart showing an example of the operation of the three-dimensional data encoding device 100.
[0219] Figure 6 The three-dimensional data encoding device 100 shown generates encoded three-dimensional data 112 by encoding three-dimensional data 111. This three-dimensional data encoding device 100 includes: an acquisition unit 101, an encoding area determination unit 102, a division unit 103, and an encoding unit 104.
[0220] As Figure 7 shown, first, the acquisition unit 101 acquires three-dimensional data 111 as point cloud data (S101).
[0221] Next, the encoding area determination unit 102 determines the area of the encoding object from the space area corresponding to the acquired point cloud data (S102). For example, the encoding area determination unit 102 determines the space area around the position as the area of the encoding object according to the position of the user or the vehicle.
[0222] Next, the division unit 103 divides the point cloud data included in the region of the object to be encoded into respective processing units. Here, the processing units are the above-mentioned GOS, SPC, etc. And the region of the object to be encoded corresponds to the above-mentioned world space, for example. Specifically, the division unit 103 divides the point cloud data into processing units according to the size of the GOS set in advance, the presence or size of dynamic objects (S103). And the division unit 103 determines the start position of the SPC that becomes the beginning in the encoding order in each GOS.
[0223] Next, the encoding unit 104 generates the encoded three-dimensional data 112 by sequentially encoding a plurality of SPCs within each GOS (S104).
[0224] In addition, here, after dividing the region of the object to be encoded into GOS and SPC, an example of encoding each GOS is shown, but the order of processing is not limited to the above. For example, after determining the composition of one GOS, the GOS can be encoded, and then the order of determining the composition of the GOS, etc. can be carried out.
[0225] In this way, the three-dimensional data encoding device 100 generates the encoded three-dimensional data 112 by encoding the three-dimensional data 111. Specifically, the three-dimensional data encoding device 100 divides the three-dimensional data into random access units, that is, divides it into first processing units (GOS) corresponding to three-dimensional coordinates respectively, divides the first processing units (GOS) into a plurality of second processing units (SPC), and divides the second processing units (SPC) into a plurality of third processing units (VLM). And the third processing unit (VLM) includes one or more voxels (VXL), and the voxel (VXL) is the smallest unit corresponding to the position information.
[0226] Next, the three-dimensional data encoding device 100 generates the encoded three-dimensional data 112 by encoding each of the plurality of first processing units (GOS). Specifically, the three-dimensional data encoding device 100 encodes each of the plurality of second processing units (SPC) in each first processing unit (GOS). And the three-dimensional data encoding device 100 encodes each of the plurality of third processing units (VLM) in each second processing unit (SPC).
[0227] For example, when the first processing unit (GOS) of the processing object is a closed GOS, for the second processing unit (SPC) of the processing object included in the first processing unit (GOS) of the processing object, encoding is performed with reference to other second processing units (SPCs) included in the first processing unit (GOS) of the processing object. That is, the three-dimensional data encoding device 100 does not refer to the second processing units (SPCs) included in the first processing unit (GOS) different from the first processing unit (GOS) of the processing object.
[0228] Further, when the first processing unit (GOS) of the processing object is an open GOS, for the second processing unit (SPC) of the processing object included in the first processing unit (GOS) of the processing object, encoding is performed with reference to other second processing units (SPCs) included in the first processing unit (GOS) of the processing object or the second processing units (SPCs) included in the first processing unit (GOS) different from the first processing unit (GOS) of the processing object.
[0229] Further, the three-dimensional data encoding device 100 selects one of the first type (I-SPC) of other second processing units (SPCs), the second type (P-SPC) of one other second processing unit (SPC), and the third type of two other second processing units (SPCs) as the type of the second processing unit (SPC) of the processing object, and encodes the second processing unit (SPC) of the processing object according to the selected type.
[0230] Next, the configuration and operation process of the three-dimensional data decoding device according to the present embodiment will be described. Figure 8 It is a block diagram of the three-dimensional data decoding device 200 according to the present embodiment. Figure 9 It is a flowchart showing an operation example of the three-dimensional data decoding device 200.
[0231] Figure 8 The three-dimensional data decoding device 200 shown generates decoded three-dimensional data 212 by decoding the encoded three-dimensional data 211. Here, the encoded three-dimensional data 211 is, for example, the encoded three-dimensional data 112 generated by the three-dimensional data encoding device 100. The three-dimensional data decoding device 200 includes: an acquisition unit 201, a decoding start GOS determination unit 202, a decoding SPC determination unit 203, and a decoding unit 204.
[0232] First, acquisition unit 201 acquires encoded three-dimensional data 211 (S201). Next, decoding start GOS determination unit 202 determines the GOS to be decoded (S202). Specifically, decoding start GOS determination unit 202 refers to the meta information within the encoded three-dimensional data 211 or stored separately from the encoded three-dimensional data, and determines the GOS including the spatial position, object, or SPC corresponding to the time at which decoding starts as the GOS to be decoded.
[0233] Next, decoding SPC determination unit 203 determines the type (I, P, B) of the SPC to be decoded within the GOS (S203). For example, decoding SPC determination unit 203 determines (1) whether to decode only I-SPC, (2) whether to decode I-SPC and P-SPC, (3) whether to decode all types. Additionally, in cases where the type of SPC to be decoded, such as decoding all SPCs, is predefined, this step may not be performed.
[0234] Next, decoding unit 204 acquires the SPC that is the first in the decoding order (same as the encoding order) within the GOS, the address position where it starts within the encoded three-dimensional data 211, acquires the encoded data of the first SPC from this address position, and decodes each SPC in sequence starting from the first SPC (S204). And the above address position is stored in meta information or the like.
[0235] In this way, three-dimensional data decoding device 200 decodes decoded three-dimensional data 212. Specifically, three-dimensional data decoding device 200 generates decoded three-dimensional data 212 of the first processing unit (GOS) as a random access unit by decoding each of the encoded three-dimensional data 211 of the first processing unit (GOS) corresponding to the three-dimensional coordinates respectively. More specifically, three-dimensional data decoding device 200 decodes each of the multiple second processing units (SPCs) within each first processing unit (GOS). And three-dimensional data decoding device 200 decodes each of the multiple third processing units (VLM) within each second processing unit (SPC).
[0236] The meta information for random access is described below. This meta information is generated by three-dimensional data encoding device 100 and is included in the encoded three-dimensional data 112 (211).
[0237] In the random access of conventional two-dimensional moving images, decoding starts from the first frame of the random access unit near the specified time. However, in the world space, random access is envisioned not only for time but also for (coordinates or objects, etc.).
[0238] Therefore, in order to achieve at least random access to the three elements of coordinates, objects, and time, a table is prepared in which the index numbers of each element are corresponded to the GOS. Moreover, the index number of the GOS is corresponded to the address of the I-SPC that is the start of the GOS. Figure 10 An example of the table included in the meta information is shown. In addition, it is not necessary to use Figure 10 all the tables shown. At least one table can be used.
[0239] Hereinafter, as an example, random access starting from coordinates will be described. When accessing the coordinates (x2, y2, z2), first referring to the coordinate-GOS table, it can be known that the location with the coordinates (x2, y2, z2) is included in the second GOS. Then, referring to the GOS address table, since it can be known that the address of the I-SPC at the start of the second GOS is addr(2), the decoding unit 204 obtains data from this address and starts decoding.
[0240] In addition, the address can be an address in the logical format or a physical address of an HDD or a memory. Also, information for determining a file segment can be used instead of the address. For example, a file segment is a unit obtained by segmenting one or more GOSs, etc.
[0241] And, in the case where an object spans multiple GOSs, the GOSs to which the multiple objects belong can also be shown in the object GOS table. If the multiple GOSs are closed GOSs, the encoding device and the decoding device can perform encoding or decoding in parallel. In addition, if the multiple GOSs are open GOSs, by referring to each other among the multiple GOSs, the compression efficiency can be further improved.
[0242] Examples of objects include people, animals, cars, bicycles, traffic lights, or buildings that are land marks. For example, when the three-dimensional data encoding device 100 encodes in the world space, characteristic points unique to the object are extracted from a three-dimensional point cloud or the like, the object is detected based on the characteristic points, and the detected object can be set as a random access point.
[0243] In this way, the three-dimensional data encoding device 100 generates the first information, which shows multiple first processing units (GOSs) and the three-dimensional coordinates corresponding to each of the multiple first processing units (GOSs). And, encoding the three-dimensional data 112(211) includes this first information. And, the first information further shows at least one of the object, time, and data storage destination corresponding to each of the multiple first processing units (GOSs).
[0244] The three-dimensional data decoding device 200 obtains the first information from the encoded three-dimensional data 211, uses the first information to determine the encoded three-dimensional data 211 of the first processing unit corresponding to the specified three-dimensional coordinates, object, or time, and decodes the encoded three-dimensional data 211.
[0245] Examples of other meta-information will be described below. In addition to the meta-information for random access, the three-dimensional data encoding device 100 can also generate and store the following meta-information. Moreover, the three-dimensional data decoding device 200 can also use this meta-information during decoding.
[0246] In the case of using three-dimensional data as map information, etc., a profile is defined according to the usage, and the information indicating the profile can be included in the meta-information. For example, a profile for urban areas or suburbs is defined, or a profile for flying objects is defined, and the maximum or minimum size of the world space, SPC, or VLM is defined respectively. For example, in the profile for urban areas, more detailed information is required than in the suburbs, so the minimum size of the VLM is set smaller.
[0247] The meta-information can also include a tag value indicating the type of object. This tag value corresponds to the VLM, SPC, or GOS that constitutes the object. The tag value can be set according to the type of object. For example, the tag value "0" represents "person", the tag value "1" represents "car", and the tag value "2" represents "traffic light". Or, in the case where the type of object is difficult to judge or does not need to be judged, a tag value indicating properties such as size, or whether it is a dynamic object or a static object can also be used.
[0248] Moreover, the meta-information can also include information indicating the range of the spatial region occupied by the world space.
[0249] Moreover, the meta-information can also store the size of the SPC or VXL as the header information shared by the entire stream of encoded data or multiple SPCs such as SPCs within the GOS.
[0250] Moreover, the meta-information can also include identification information such as a distance sensor or camera used in the generation of the point cloud, or information indicating the position accuracy of the point group within the point cloud.
[0251] Moreover, the meta-information can include information indicating whether the world space consists only of static objects or contains dynamic objects.
[0252] A modification example of this embodiment will be described below.
[0253] An encoding device or a decoding device can encode or decode two or more SPCs or GOSs that are different from each other in parallel. The GOSs encoded or decoded in parallel can be determined based on meta-information indicating the spatial positions of the GOSs, etc.
[0254] In cases where three-dimensional data is used as a spatial map when a vehicle or a flying object moves, or when such a spatial map is generated, etc., the encoding device or the decoding device can encode or decode GOSs or SPCs included in a space determined based on GPS, path information, or zoom ratio, etc.
[0255] Moreover, the decoding device can also start decoding sequentially from the space close to its own position or walking path. The encoding device or the decoding device can also perform encoding or decoding by making the priority of the space far from its own position or walking path lower than that of the space close to it. Here, reducing the priority means reducing the processing order, reducing the resolution (post-screening processing), or reducing the image quality (increasing the encoding efficiency. For example, increasing the quantization step size), etc.
[0256] Moreover, when the decoding device decodes the encoded data hierarchically encoded in a space, it can also decode only the lower hierarchy.
[0257] Moreover, the decoding device can also start decoding from the lower hierarchy first according to the zoom ratio or use of the map.
[0258] Moreover, in applications such as self-position estimation or object recognition performed during the automatic driving of an automobile or a robot, the encoding device or the decoding device can also reduce the resolution of the area outside the area within a specified height from the road surface (the area to be recognized) to perform encoding or decoding.
[0259] Moreover, the encoding device can also encode the point clouds representing the spatial shapes of the indoor and outdoor spaces independently. For example, by separating the GOS representing the indoor (indoor GOS) from the GOS representing the outdoor (outdoor GOS), the decoding device can select the GOS to be decoded according to the viewpoint position when using the encoded data.
[0260] Moreover, the encoding device can make the indoor GOS and the outdoor GOS with close coordinates adjacent in the encoding stream to perform encoding. For example, the encoding device corresponds their identifiers and stores the information indicating the identifiers that are corresponding in the encoding stream or in the meta-information stored separately. Accordingly, the decoding device can refer to the information in the meta-information to identify the indoor GOS and the outdoor GOS with close coordinates.
[0261] Furthermore, the encoding device can also switch the size of GOS or SPC between indoor GOS and outdoor GOS. For example, the encoding device sets the size of GOS to be smaller indoors than outdoors. Also, the encoding device can change the accuracy when extracting feature points from the point cloud or the accuracy of object detection, etc., between indoor GOS and outdoor GOS.
[0262] Moreover, the encoding device can attach information for the decoding device to distinguish and display dynamic objects from static objects to the encoded data. Accordingly, the decoding device can combine and represent dynamic objects with a red frame or explanatory text, etc. In addition, the decoding device can also represent only with a red frame or explanatory text instead of the dynamic object. And the decoding device can represent more detailed object categories. For example, a car can be represented by a red frame, and a person can be represented by a yellow frame.
[0263] Furthermore, the encoding device or the decoding device can determine whether to encode or decode dynamic objects and static objects as different SPCs or GOSs according to the appearance frequency of dynamic objects, or the ratio of static objects to dynamic objects, etc. For example, when the appearance frequency or ratio of dynamic objects exceeds a threshold, an SPC or GOS in which dynamic objects and static objects are mixed is allowed, and when the appearance frequency or ratio of dynamic objects does not exceed the threshold, an SPC or GOS in which dynamic objects and static objects are mixed is not allowed.
[0264] When a dynamic object is detected not from the point cloud but from the two-dimensional image information of a camera, the encoding device can separately obtain information (such as a frame or text) for identifying the detection result and the object position, and encode these as part of the three-dimensional encoded data. In this case, for the decoding result of static objects, the decoding device overlays and displays auxiliary information (frame or text) representing the dynamic object.
[0265] Also, the encoding device can change the density of VXL or VLM according to the complexity of the shape of static objects, etc. For example, the more complex the shape of the static object, the denser the encoding device sets VXL or VLM. Moreover, the encoding device can determine the quantization step, etc., when quantizing spatial position or color information according to the density of VXL or VLM. For example, the denser VXL or VLM is, the smaller the encoding device sets the quantization step.
[0266] As described above, the encoding device or the decoding device according to this embodiment performs spatial encoding or decoding in a spatial unit with coordinate information.
[0267] Moreover, the encoding device and the decoding device perform encoding or decoding in units of volume within the space. The volume includes voxels, which are the smallest units corresponding to position information.
[0268] Further, the encoding device and the decoding device perform encoding or decoding by establishing a correspondence table between each element of spatial information including coordinates, objects, time, etc. and GOPs, or a correspondence table between elements, to establish a correspondence between any elements. Further, the decoding device determines coordinates using the value of the selected element, and determines a volume, voxel, or space based on the coordinates, and decodes the space including the volume or voxel, or the determined space.
[0269] Further, the encoding device determines a volume, voxel, or space that can be selected by an element through feature point extraction or object recognition, and encodes it as a volume, voxel, or space that can be randomly accessed.
[0270] The space is divided into three types, namely: I-SPC that can be encoded or decoded by the space alone, P-SPC that encodes or decodes with reference to any one processed space, and B-SPC that encodes or decodes with reference to any two processed spaces.
[0271] One or more volumes correspond to static objects or dynamic objects. The space containing static objects and the space containing dynamic objects are encoded or decoded as different GOSs respectively. That is, the SPC containing static objects and the SPC containing dynamic objects are assigned to different GOSs.
[0272] Dynamic objects are encoded or decoded for each object and correspond to one or more spaces containing only static objects. That is, multiple dynamic objects are encoded separately, and the encoded data of the multiple dynamic objects obtained corresponds to the SPC containing only static objects.
[0273] The encoding device and the decoding device perform encoding or decoding by increasing the priority of I-SPC within GOS. For example, the encoding device performs encoding in a manner that reduces the degradation of I-SPC (after decoding, the original three-dimensional data can be reproduced more faithfully). Further, the decoding device decodes only I-SPC, for example.
[0274] The encoding device can change the frequency of using I-SPC according to the density or value (quantity) of objects in the world space to perform encoding. That is, the encoding device changes the frequency of selecting I-SPC according to the quantity or density of objects included in the three-dimensional data. For example, the encoding device increases the usage frequency of the I space as the density of objects in the world space increases.
[0275] Further, the encoding device sets random access points in units of GOS, and stores information indicating the spatial region corresponding to GOS in the header information.
[0276] The encoding device, for example, uses a default value as the spatial size of the GOS. Additionally, the encoding device can also change the size of the GOS according to the value (quantity) or density of the object or dynamic object. For example, when the object or dynamic object is denser or the quantity is larger, the encoding device sets the spatial size of the GOS to be smaller.
[0277] Moreover, the space or volume includes a feature point group derived using information obtained from sensors such as depth sensors, gyroscopes, or cameras. The coordinates of the feature points are set as the center positions of the voxels. And through the subdivision of the voxels, high-precision position information can be achieved.
[0278] The feature point group is derived using multiple pictures. The multiple pictures have at least the following two types of time information, namely: actual time information, and the same time information in multiple pictures corresponding to the space (for example, the encoding time for rate control, etc.).
[0279] And encoding or decoding is performed in units of GOS including one or more spaces.
[0280] The encoding device and the decoding device predict the P space or B space in the GOS of the processing object by referring to the spaces in the processed GOS.
[0281] Alternatively, the encoding device and the decoding device do not refer to different GOSs, and use the processed spaces in the GOS of the processing object to predict the P space or B space in the GOS of the processing object.
[0282] And the encoding device and the decoding device send or receive the encoded stream in units of a world space including one or more GOSs.
[0283] And the GOS has a layer structure in at least one direction in the world space. The encoding device and the decoding device perform encoding or decoding starting from the lower layer. For example, the GOS that can be randomly accessed belongs to the lowest layer. The GOS belonging to the upper layer only refers to the GOSs belonging to the layers below the same layer. That is, the GOS is spatially divided in a predefined direction and includes multiple layers each having one or more SPCs. The encoding device and the decoding device perform encoding or decoding for each SPC by referring to the SPCs included in the same layer as or lower than the SPC.
[0284] And the encoding device and the decoding device continuously perform encoding or decoding on the GOSs within the unit of the world space including multiple GOSs. The encoding device and the decoding device write or read the information indicating the encoding or decoding order (direction) as metadata. That is, the encoded data includes the information indicating the encoding order of multiple GOSs.
[0285] Further, an encoding device and a decoding device encode or decode two or more different spaces or GOSs in parallel with each other.
[0286] Further, the encoding device and the decoding device encode or decode spatial information (coordinates, size, etc.) of a space or GOS.
[0287] Further, the encoding device and the decoding device encode or decode a space or GOS included in a specific space determined based on external information such as GPS, path information, or magnification related to its own position or / and area size.
[0288] The encoding device or the decoding device encodes or decodes by making the priority of a space far from its own position lower than that of a space close to its own position.
[0289] The encoding device sets a direction in the world space according to magnification or use, and encodes a GOS having a layer structure in that direction. And the decoding device preferentially decodes from the lower layer for a GOS having a layer structure in a direction in the world space set according to magnification or use.
[0290] The encoding device changes the feature point extraction, object recognition accuracy, or spatial area size, etc. included in the indoor and outdoor spaces. However, the encoding device and the decoding device encode or decode by making the indoor GOS and the outdoor GOS with close coordinates adjacent in the world space, and also encode or decode by corresponding these identifiers.
[0291] (Embodiment 2)
[0292] When using the encoded data of the point cloud for an actual device or service, in order to suppress network bandwidth, it is desired to transmit and receive the required information according to use. However, such a function does not exist in the encoding structure of three-dimensional data so far, and thus there is no encoding method corresponding thereto.
[0293] What will be described in this embodiment is a three-dimensional data encoding method and a three-dimensional data encoding device for providing a function of transmitting and receiving required information according to use in the encoded data of three-dimensional point cloud, and also a three-dimensional data decoding method and a three-dimensional data decoding device for decoding the encoded data.
[0294] A voxel (VXL) having a feature amount equal to or more than a certain value is defined as a feature voxel (FVXL), and a world space (WLD) composed of FVXLs is defined as a sparse world space (SWLD). Figure 11Shows a sparse world space and a composition example of the world space. In SWLD, it includes: FGOS, which is a GOS composed of FVXL; FSPC, which is an SPC composed of FVXL; and FVLM, which is a VLM composed of FVXL. The data structures and prediction structures of FGOS, FSPC, and FVLM can be the same as those of GOS, SPC, and VLM.
[0295] The feature quantity refers to a feature quantity that represents the three-dimensional position information of VXL or the visible light information of the VXL position. In particular, more feature quantities can be detected at the corners and edges of three-dimensional objects. Specifically, although the feature quantity is the three-dimensional feature quantity or the visible light feature quantity described below, as long as it is a feature quantity that represents the position, brightness, or color information of VXL, it can be any feature quantity.
[0296] As the three-dimensional feature quantity, the SHOT feature quantity (Signature of Histograms of OrienTations), the PFH feature quantity (Point Feature Histograms), or the PPF feature quantity (Point Pair Feature) is adopted.
[0297] The SHOT feature quantity is obtained by dividing the periphery of VXL and calculating the inner product of the normal vector of the reference point and the divided area, and then performing histogramming. The SHOT feature quantity has the characteristics of high dimensionality and high feature expressiveness.
[0298] The PFH feature quantity is obtained by selecting multiple two-point groups near VXL, calculating the normal vector, etc. based on these two points, and then performing histogramming. Since the PFH feature quantity is a histogram feature, it is robust to a small amount of interference and has the characteristic of high feature expressiveness.
[0299] The PPF feature quantity is a feature quantity calculated using the normal vector, etc. according to two VXLs. In this PPF feature quantity, since all VXLs are used, it is robust to occlusion.
[0300] Moreover, as the feature quantity of visible light, SIFT (Scale-Invariant Feature Transform), SURF (Speeded Up Robust Features), or HOG (Histogram of Oriented Gradients), etc. that adopt information such as the luminance gradient information of the image can be used.
[0301] The SWLD is generated by calculating the above-described feature amounts from each VXL of the WLD and extracting the FVXL. Here, the SWLD can be updated each time the WLD is updated, or it can be updated periodically after a certain period of time regardless of the update timing of the WLD.
[0302] The SWLD can be generated for each feature amount. For example, as shown by SWLD1 based on the SHOT feature amount and SWLD2 based on the SIFT feature amount, the SWLD can be generated separately for each feature amount, and the SWLD can be distinguished and used according to the purpose. Also, the feature amounts of the calculated FVXLs can be held as feature amount information in the respective FVXLs.
[0303] Next, the method of using the sparse world space (SWLD) will be described. Since the SWLD only contains the feature voxels (FVXL), the data size is generally smaller than that of the WLD including all the VXLs.
[0304] In an application that uses feature amounts to achieve a certain purpose, by using the information of the SWLD instead of the WLD, it is possible to suppress the read time from the hard disk, and it is also possible to suppress the bandwidth and transmission time during network transmission. For example, as map information, the WLD and the SWLD are held in advance in the server, and by switching the transmitted map information to the WLD or the SWLD according to the demand from the client, it is possible to suppress the network bandwidth and the transmission time. Specific examples are shown below.
[0305] Figure 12 And Figure 13 Examples of using the SWLD and the WLD are shown. As Figure 12 shown, when the client 1 as an in-vehicle device needs map information for its own position determination, the client 1 sends a request for obtaining map data for its own position estimation to the server (S301). The server sends the SWLD to the client 1 according to the acquisition request (S302). The client 1 uses the received SWLD to determine its own position (S303). At this time, the client 1 obtains the VXL information around the client 1 by various methods such as a distance sensor such as a rangefinder, a stereo camera, or a combination of multiple monocular cameras, and estimates its own position information based on the obtained VXL information and the SWLD. Here, the own position information includes the three-dimensional position information of the client 1 and the orientation, etc.
[0306] As Figure 13As shown in the figure, when the client 2, which is a vehicle-mounted device, needs map information for map drawing purposes such as a 3D map, the client 2 sends a request for obtaining map data for map drawing to the server (S311). The server sends the WLD to the client 2 according to this acquisition request (S312). The client 2 uses the received WLD for map drawing (S313). At this time, the client 2, for example, uses an image captured by its own visible light camera or the like, and the WLD obtained from the server to create a concept image, and depicts the created image on a screen such as a car navigation system.
[0307] As described above, the server sends the SWLD to the client in applications mainly requiring the feature quantities of each VXL for self-position estimation, and sends the WLD to the client in cases where detailed VXL information is required, such as map drawing. Accordingly, map data can be efficiently transmitted and received.
[0308] In addition, the client can determine which of the SWLD and the WLD it needs and request the server to send the SWLD or the WLD. And the server can judge which of the SWLD or the WLD should be sent according to the status of the client or the network.
[0309] Next, a method for switching the transmission and reception of the sparse world space (SWLD) and the world space (WLD) will be described.
[0310] The reception of the WLD or the SWLD can be switched according to the network bandwidth. Figure 14 A working example in this case is shown. For example, when a low-speed network with a network bandwidth such as in an LTE (Long Term Evolution) environment can be used, when the client accesses the server via the low-speed network (S321), it obtains the SWLD as map information from the server (S322). In addition, when a high-speed network with a surplus network bandwidth such as in a WiFi environment is used, the client accesses the server via the high-speed network (S323) and obtains the WLD from the server (S324). Accordingly, the client can obtain appropriate map information according to the network bandwidth of the client.
[0311] Specifically, the client receives the SWLD via LTE outdoors, and when entering indoors such as a facility, it obtains the WLD via WiFi. Accordingly, the client can obtain more detailed map information indoors.
[0312] In this way, the client can request the WLD or SWLD from the server according to the frequency band of the network it uses. Alternatively, the client can send the information indicating the frequency band of the network it uses to the server, and the server sends the appropriate data (WLD or SWLD) to the client according to this information. Alternatively, the server can determine the network bandwidth of the client and send the appropriate data (WLD or SWLD) to the client.
[0313] Moreover, the reception of the WLD or SWLD can be switched according to the moving speed. Figure 15 A working example in this case is shown. For example, when the client is moving at high speed (S331), the client receives the SWLD from the server (S332). In addition, when the client is moving at low speed (S333), the client receives the WLD from the server (S334). Accordingly, the client can both suppress the network bandwidth and obtain map information according to the speed. Specifically, when the client is driving on a highway, by receiving the SWLD with less data volume, the map information can be updated at an appropriate speed approximately. In addition, when the client is driving on an ordinary road, by receiving the WLD, more detailed map information can be obtained.
[0314] In this way, the client can request the WLD or SWLD from the server according to its own moving speed. Alternatively, the client can send the information indicating its own moving speed to the server, and the server sends the appropriate data (WLD or SWLD) to the client according to this information. Alternatively, the server can determine the moving speed of the client and send the appropriate data (WLD or SWLD) to the client.
[0315] Also, it can be that the client first obtains the SWLD from the server and then obtains the WLD of the important areas therein. For example, when the client obtains map data, it first obtains the general map information in the form of SWLD, screens out the areas where features such as buildings, signs, or people appear more frequently, and then obtains the WLD of the screened areas. Accordingly, the client can both suppress the amount of received data from the server and obtain the detailed information of the required areas.
[0316] Also, it can be that the server respectively creates the SWLD for each object according to the WLD, and the client receives them respectively according to the usage. Accordingly, the network bandwidth can be suppressed. For example, the server pre-identifies people or vehicles from the WLD and creates the SWLD of people and the SWLD of vehicles. When the client wants to obtain information about the people around, it receives the SWLD of people, and when it wants to obtain information about vehicles, it receives the SWLD of vehicles. And the types of such SWLD can be distinguished according to the information (flags or types, etc.) attached to the head, etc.
[0317] Next, the configuration and operation process of the three-dimensional data encoding device (such as a server) according to this embodiment will be described. Figure 16 is a block diagram of the three-dimensional data encoding device 400 according to this embodiment. Figure 17 is a flowchart of the three-dimensional data encoding process performed by the three-dimensional data encoding device 400.
[0318] Figure 16 The three-dimensional data encoding device 400 shown generates encoded three-dimensional data 413 and 414 as an encoded stream by encoding the input three-dimensional data 411. Here, the encoded three-dimensional data 413 is the encoded three-dimensional data corresponding to the WLD, and the encoded three-dimensional data 414 is the encoded three-dimensional data corresponding to the SWLD. The three-dimensional data encoding device 400 includes: an acquisition unit 401, an encoding region determination unit 402, an SWLD extraction unit 403, a WLD encoding unit 404, and an SWLD encoding unit 405.
[0319] As Figure 17 shown, first, the acquisition unit 401 acquires the input three-dimensional data 411 (S401) as point cloud data in a three-dimensional space.
[0320] Next, the encoding region determination unit 402 determines the spatial region to be encoded according to the spatial region where the point cloud data exists (S402).
[0321] Next, the SWLD extraction unit 403 defines the spatial region to be encoded as the WLD, calculates the feature amount according to each VXL included in the WLD. And the SWLD extraction unit 403 extracts the VXL whose feature amount is above a preset threshold, defines the extracted VXL as the FVXL, and generates the extracted three-dimensional data 412 by adding the FVXL to the SWLD (S403). That is, the extracted three-dimensional data 412 whose feature amount is above the threshold is extracted from the input three-dimensional data 411.
[0322] Next, the WLD encoding unit 404 encodes the input three-dimensional data 411 corresponding to the WLD to generate the encoded three-dimensional data 413 corresponding to the WLD (S404). At this time, the WLD encoding unit 404 attaches information for distinguishing that the encoded three-dimensional data 413 is a stream containing the WLD to the head of the encoded three-dimensional data 413.
[0323] And, the SWLD encoding unit 405 encodes the extracted three-dimensional data 412 corresponding to the SWLD to generate the encoded three-dimensional data 414 corresponding to the SWLD (S405). At this time, the SWLD encoding unit 405 attaches information for distinguishing that the encoded three-dimensional data 414 is a stream containing the SWLD to the head of the encoded three-dimensional data 414.
[0324] Moreover, the processing order of the process of generating the encoded three-dimensional data 413 and the process of generating the encoded three-dimensional data 414 may be opposite to the above. Also, a part or all of the above processes may be executed in parallel.
[0325] Information assigned to the headers of the encoded three-dimensional data 413 and 414 is defined as a parameter such as "world_type", for example. When world_type = 0, it indicates that the stream contains WLD, and when world_type = 1, it indicates that the stream contains SWLD. When defining more other categories, the assigned value can be increased like world_type = 2. Also, a specific flag may be included in one of the encoded three-dimensional data 413 and 414. For example, the encoded three-dimensional data 414 may be assigned a flag indicating that the stream contains SWLD. In this case, the decoding device can determine whether the stream contains WLD or SWLD based on the presence or absence of the flag.
[0326] Moreover, the encoding method used by the WLD encoding unit 404 when encoding WLD may be different from the encoding method used by the SWLD encoding unit 405 when encoding SWLD.
[0327] For example, since SWLD data is decimated, the correlation with surrounding data may be lower than that of WLD. Therefore, in the encoding method for SWLD, inter-frame prediction among intra-frame prediction and inter-frame prediction is prioritized compared to the encoding method for WLD.
[0328] Also, it may be that the representation method of the three-dimensional position is different between the encoding method for SWLD and the encoding method for WLD. For example, it may be that the three-dimensional position of FVXL is represented by three-dimensional coordinates in FWLD, and the three-dimensional position is represented by an octree described later in WLD, and vice versa.
[0329] Moreover, the SWLD encoding unit 405 encodes in such a way that the data size of the encoded three-dimensional data 414 of SWLD is smaller than the data size of the encoded three-dimensional data 413 of WLD. As described above, for example, the correlation between data may be reduced in SWLD compared to WLD. Accordingly, the encoding efficiency is reduced, and the data size of the encoded three-dimensional data 414 may be larger than the data size of the encoded three-dimensional data 413 of WLD. Therefore, when the data size of the obtained encoded three-dimensional data 414 is larger than the data size of the encoded three-dimensional data 413 of WLD, the SWLD encoding unit 405 performs re-encoding to regenerate the encoded three-dimensional data 414 with a reduced data size.
[0330] For example, the SWLD extraction unit 403 regenerates the extracted 3D data 412 with a reduced number of extracted feature points, and the SWLD encoding unit 405 encodes the extracted 3D data 412. Alternatively, the quantization level in the SWLD encoding unit 405 can be made coarser. For example, in the octree structure described later, by rounding the data in the bottom layer, the quantization level can be made coarser.
[0331] Moreover, when the SWLD encoding unit 405 cannot make the data size of the encoded 3D data 414 of SWLD smaller than the data size of the encoded 3D data 413 of WLD, the SWLD encoding unit 405 may not generate the encoded 3D data 414 of SWLD. Alternatively, the encoded 3D data 413 of WLD can be copied to the encoded 3D data 414 of SWLD. That is, the encoded 3D data 413 of WLD can be directly used as the encoded 3D data 414 of SWLD.
[0332] Next, the configuration and operation process of the 3D data decoding device (e.g., client) according to the present embodiment will be described. Figure 18 FIG. is a block diagram of the 3D data decoding device 500 according to the present embodiment. Figure 19 FIG. is a flowchart of the 3D data decoding process performed by the 3D data decoding device 500.
[0333] Figure 18 The shown 3D data decoding device 500 decodes the encoded 3D data 511 to generate the decoded 3D data 512 or 513. Here, the encoded 3D data 511 is, for example, the encoded 3D data 413 or 414 generated by the 3D data encoding device 400.
[0334] The 3D data decoding device 500 includes: an acquisition unit 501, a header analysis unit 502, a WLD decoding unit 503, and an SWLD decoding unit 504.
[0335] As shown in Figure 19 First, the acquisition unit 501 acquires the encoded 3D data 511 (S501). Next, the header analysis unit 502 analyzes the header of the encoded 3D data 511 to determine whether the encoded 3D data 511 includes a WLD stream or an SWLD stream (S502). For example, the determination is made with reference to the above-mentioned world_type parameter.
[0336] In the case where the encoded three-dimensional data 511 is a stream including WLD (Yes in S503), the WLD decoding unit 503 decodes the encoded three-dimensional data 511 to generate decoded three-dimensional data 512 of WLD (S504). On the other hand, in the case where the encoded three-dimensional data 511 is a stream including SWLD (No in S503), the SWLD decoding unit 504 decodes the encoded three-dimensional data 511 to generate decoded three-dimensional data 513 of SWLD (S505).
[0337] Moreover, similar to the encoding device, the decoding method used by the WLD decoding unit 503 when decoding WLD and the decoding method used by the SWLD decoding unit 504 when decoding SWLD may be different. For example, in the decoding method for SWLD, inter-frame prediction in intra-frame prediction and inter-frame prediction may be prioritized compared to the decoding method for WLD.
[0338] Furthermore, in the decoding method for SWLD and the decoding method for WLD, the representation methods of three-dimensional positions may be different. For example, in SWLD, the three-dimensional position of FVXL may be represented by three-dimensional coordinates, and in WLD, the three-dimensional position may be represented by an octree described later, and vice versa.
[0339] Next, the octree representation as a method for representing three-dimensional positions will be described. The VXL data included in the three-dimensional data is converted into an octree structure and then encoded. Figure 20 An example of VXL of WLD is shown. Figure 21 Shows Figure 20 The octree structure of the WLD shown. In Figure 20 the example shown, there are three VXLs 1 to 3 as VXLs (hereinafter, valid VXLs) including point groups. As Figure 21 shown, the octree structure is composed of nodes and leaf nodes. Each node has a maximum of 8 nodes or leaf nodes. Each leaf node has VXL information. Here, Figure 21 among the leaf nodes shown, leaf nodes 1, 2, and 3 represent Figure 20 the VXL1, VXL2, and VXL3 shown, respectively.
[0340] Specifically, each node and leaf node correspond to a three-dimensional position. Node 1 corresponds to Figure 20 all the blocks shown. The block corresponding to Node 1 is divided into 8 blocks. Among the 8 blocks, the blocks including valid VXLs are set as nodes, and the other blocks are set as leaf nodes. The block corresponding to the node is further divided into 8 nodes or leaf nodes, and the number of times this process is repeated is the same as the number of levels in the tree structure. And all the blocks in the bottom layer are set as leaf nodes.
[0341] Moreover,Figure 22 shows an example of the SWLD generated from the Figure 20 shown WLD. Figure 20 The results of the feature quantity extraction of the VXL1 and VXL2 shown are judged as FVXL1 and FVXL2 and added to the SWLD. In addition, since VXL3 is not judged as FVXL, it is not included in the SWLD. Figure 23 shows Figure 22 the octree structure of the SWLD shown. In Figure 23 the octree structure shown, Figure 21 the leaf node 3 corresponding to VXL3 shown is deleted. Accordingly, Figure 21 the node 3 shown has no valid VXL and is changed to a leaf node. In this way, generally speaking, the number of leaf nodes of the SWLD is smaller than that of the WLD, and the encoded three-dimensional data of the SWLD is also smaller than that of the WLD.
[0342] The following describes a modification example of the present embodiment.
[0343] For example, it may also be the case where, when a client such as an in-vehicle device estimates its own position, it receives the SWLD from the server, uses the SWLD to estimate its own position, and performs obstacle detection. In this case, various methods such as a distance sensor such as a rangefinder, a stereo camera, or a combination of multiple monocular cameras are used to perform obstacle detection based on the three-dimensional information of the surrounding area obtained by itself.
[0344] And, generally speaking, it is difficult to include VXL data of a flat area in the SWLD. For this reason, the server maintains a downsampled world space (SubWLD) obtained by downsampling the WLD for the detection of stationary obstacles, and can send the SWLD and the SubWLD to the client. Accordingly, both the network bandwidth can be suppressed and the client side can perform its own position estimation and obstacle detection.
[0345] And, when the client quickly depicts three-dimensional map data, it may be convenient if the map information is in a grid structure. Then, the server can generate a grid based on the WLD and pre-maintain it as a grid world space (MWLD). For example, when the client needs to perform rough three-dimensional rendering, it receives the MWLD, and when it needs to perform detailed three-dimensional rendering, it receives the WLD. Accordingly, the network bandwidth can be suppressed.
[0346] Also, although the server sets the VXLs with feature amounts above the threshold as FVXLs from each VXL, the FVXLs can also be calculated by different methods. For example, if the server determines that VXLs, VLMs, SPCs, or GOSs that make up signals or intersections are required for its own position estimation, driving assistance, or autonomous driving, etc., they can be included in the SWLD as FVXLs, FVLMs, FSPCs, or FGOSs. Also, the above determination can be made manually. In addition, the FVXLs obtained by the above method can be added to the FVXLs, etc. set based on the feature amount. That is, the SWLD extraction unit 403 can further extract, from the input three-dimensional data 411, the data corresponding to the object with a predetermined attribute as the extracted three-dimensional data 412.
[0347] Also, different labels can be assigned to the situations required for these uses, different from the feature amounts. The server can separately hold the FVXLs required for its own position estimation, driving assistance, or autonomous driving, such as signals or intersections, as the upper layer of the SWLD (for example, the lane world space).
[0348] Also, the server can attach attributes to the VXLs in the WLD according to a random access unit or a specified unit. The attributes include, for example, information indicating whether it is required or not required for its own position estimation, or information indicating whether it is important as traffic information such as a signal or an intersection. Also, the attributes can include the correspondence relationship with the Feature (intersection or road, etc.) in the lane information (GDF: Geographic Data Files, etc.).
[0349] Also, as a method for updating the WLD or SWLD, the following method can be adopted.
[0350] Update information showing changes in people, construction, or street trees (facing the trajectory), etc. is loaded into the server as a point cloud or metadata. The server updates the WLD based on this loading, and then uses the updated WLD to update the SWLD.
[0351] Also, when the client detects a mismatch between the three-dimensional information generated by itself during its own position estimation and the three-dimensional information received from the server, the three-dimensional information generated by itself can be sent to the server together with an update notification. In this case, the server uses the WLD to update the SWLD. If the SWLD is not updated, the server determines that the WLD itself is old.
[0352] Also, as the header information of the encoded stream, although information for distinguishing between WLD and SWLD is attached, for example, in a case where there are multiple world spaces such as a grid world space or a lane world space, information for distinguishing them can be attached to the header information. Also, in a case where there are multiple SWLDs with different feature amounts, information for distinguishing them separately can also be attached to the header information.
[0353] Also, although SWLD is composed of FVXL, it can also include VXL that is not determined to be FVXL. For example, SWLD can include adjacent VXL used when calculating the feature amount of FVXL. Accordingly, even when no feature amount information is attached to each FVXL of SWLD, the client can calculate the feature amount of FVXL when receiving SWLD. Also, at this time, SWLD can include information for distinguishing whether each VXL is FVXL or VXL.
[0354] As described above, the three-dimensional data encoding device 400 extracts the extracted three-dimensional data 412 (second three-dimensional data) whose feature amount is equal to or greater than the threshold from the input three-dimensional data 411 (first three-dimensional data), and generates the encoded three-dimensional data 414 (first encoded three-dimensional data) by encoding the extracted three-dimensional data 412.
[0355] Accordingly, the three-dimensional data encoding device 400 generates the encoded three-dimensional data 414 obtained by encoding the data whose feature amount is equal to or greater than the threshold. In this way, compared with the case of directly encoding the input three-dimensional data 411, the data amount can be reduced. Therefore, the three-dimensional data encoding device 400 can reduce the data amount during transmission.
[0356] Also, the three-dimensional data encoding device 400 further generates the encoded three-dimensional data 413 (second encoded three-dimensional data) by encoding the input three-dimensional data 411.
[0357] Accordingly, the three-dimensional data encoding device 400 can selectively transmit the encoded three-dimensional data 413 and the encoded three-dimensional data 414, for example, according to the usage purpose.
[0358] Also, the extracted three-dimensional data 412 is encoded by the first encoding method, and the input three-dimensional data 411 is encoded by the second encoding method different from the first encoding method.
[0359] Accordingly, the three-dimensional data encoding device 400 can adopt appropriate encoding methods for the input three-dimensional data 411 and the extracted three-dimensional data 412 respectively.
[0360] Also, in the first encoding method, among intra prediction and inter prediction, inter prediction is prioritized compared with the second encoding method.
[0361] Accordingly, the three-dimensional data encoding device 400 can increase the priority of inter-frame prediction for the extracted three-dimensional data 412 where the correlation between adjacent data is likely to be low.
[0362] Moreover, in the first encoding method and the second encoding method, the representation methods of three-dimensional positions are different. For example, in the second encoding method, the three-dimensional position is represented by an octree, and in the first encoding method, the three-dimensional position is represented by three-dimensional coordinates.
[0363] Accordingly, the three-dimensional data encoding device 400 can adopt a more appropriate representation method of three-dimensional positions for three-dimensional data with different numbers of data (the number of VXL or FVXL).
[0364] Moreover, in at least one of the encoded three-dimensional data 413 and 414, there is an identifier indicating whether the encoded three-dimensional data is the encoded three-dimensional data obtained by encoding the input three-dimensional data 411 or the encoded three-dimensional data obtained by encoding a part of the input three-dimensional data 411. That is, this identifier indicates whether the encoded three-dimensional data is the encoded three-dimensional data 413 of WLD or the encoded three-dimensional data 414 of SWLD.
[0365] Accordingly, the decoding device can easily determine whether the acquired encoded three-dimensional data is the encoded three-dimensional data 413 or the encoded three-dimensional data 414.
[0366] Moreover, the three-dimensional data encoding device 400 encodes the extracted three-dimensional data 412 in such a way that the data amount of the encoded three-dimensional data 414 is less than the data amount of the encoded three-dimensional data 413.
[0367] Accordingly, the three-dimensional data encoding device 400 can make the data amount of the encoded three-dimensional data 414 less than the data amount of the encoded three-dimensional data 413.
[0368] Moreover, the three-dimensional data encoding device 400 further extracts, as the extracted three-dimensional data 412, the data corresponding to an object having a predetermined attribute from the input three-dimensional data 411. For example, an object having a predetermined attribute refers to an object required in self-position estimation, driving assistance, or autonomous driving, etc., such as a signal or an intersection.
[0369] Accordingly, the three-dimensional data encoding device 400 can generate the encoded three-dimensional data 414 including the data required by the decoding device.
[0370] Moreover, the three-dimensional data encoding device 400 (server) further sends one of the encoded three-dimensional data 413 and 414 to the client according to the state of the client.
[0371] Accordingly, the three-dimensional data encoding device 400 can send appropriate data according to the state of the client.
[0372] Moreover, the state of the client includes the communication status of the client (e.g., network bandwidth) or the moving speed of the client.
[0373] Furthermore, the three-dimensional data encoding device 400 further sends one of the encoded three-dimensional data 413 and 414 to the client according to the request of the client.
[0374] Accordingly, the three-dimensional data encoding device 400 can send appropriate data according to the request of the client.
[0375] In addition, the three-dimensional data decoding device 500 according to the present embodiment decodes the encoded three-dimensional data 413 or 414 generated by the above three-dimensional data encoding device 400.
[0376] That is, the three-dimensional data decoding device 500 decodes the encoded three-dimensional data 414 obtained by encoding the extracted three-dimensional data 412 whose feature amount extracted from the input three-dimensional data 411 is above the threshold value by the first decoding method. And the three-dimensional data decoding device 500 decodes the encoded three-dimensional data 413 obtained by encoding the input three-dimensional data 411 by using a second decoding method different from the first decoding method.
[0377] Accordingly, the three-dimensional data decoding device 500 can selectively receive, for example, according to the usage purpose, etc., the encoded three-dimensional data 414 and the encoded three-dimensional data 413 obtained by encoding the data whose feature amount is above the threshold value. Accordingly, the three-dimensional data decoding device 500 can reduce the amount of data during transmission. Moreover, the three-dimensional data decoding device 500 can adopt appropriate decoding methods for the input three-dimensional data 411 and the extracted three-dimensional data 412 respectively.
[0378] In addition, in the first decoding method, among intra prediction and inter prediction, inter prediction is prioritized compared to the second decoding method.
[0379] Accordingly, the three-dimensional data decoding device 500 can increase the priority of inter prediction for the extracted three-dimensional data whose correlation between adjacent data is likely to become low.
[0380] Moreover, in the first decoding method and the second decoding method, the representation methods of three-dimensional positions are different. For example, in the second decoding method, the three-dimensional position is represented by an octree, and in the first decoding method, the three-dimensional position is represented by three-dimensional coordinates.
[0381] Accordingly, the three-dimensional data decoding device 500 can adopt a more appropriate representation method of three-dimensional positions for three-dimensional data with different numbers of data (the number of VXL or FVXL).
[0382] Further, at least one of the encoded three-dimensional data 413 and 414 includes an identifier indicating whether the encoded three-dimensional data is obtained by encoding the input three-dimensional data 411 or by encoding a part of the input three-dimensional data 411. The three-dimensional data decoding device 500 identifies the encoded three-dimensional data 413 and 414 with reference to the identifier.
[0383] Accordingly, the three-dimensional data decoding device 500 can easily determine whether the obtained encoded three-dimensional data is the encoded three-dimensional data 413 or the encoded three-dimensional data 414.
[0384] Further, the three-dimensional data decoding device 500 further notifies the server of the state of the client (the three-dimensional data decoding device 500). The three-dimensional data decoding device 500 receives one of the encoded three-dimensional data 413 and 414 sent from the server according to the state of the client.
[0385] Accordingly, the three-dimensional data decoding device 500 can receive appropriate data according to the state of the client.
[0386] The state of the client includes the communication status of the client (e.g., network bandwidth) or the moving speed of the client.
[0387] Further, the three-dimensional data decoding device 500 further requests one of the encoded three-dimensional data 413 and 414 from the server and receives one of the encoded three-dimensional data 413 and 414 sent from the server according to the request.
[0388] Accordingly, the three-dimensional data decoding device 500 can receive appropriate data corresponding to the use.
[0389] (Embodiment 3)
[0390] In this embodiment, a method for transmitting and receiving three-dimensional data between vehicles is described. For example, three-dimensional data is transmitted and received between the host vehicle and surrounding vehicles.
[0391] Figure 24 It is a block diagram of a three-dimensional data production device 620 according to this embodiment. The three-dimensional data production device 620 is included in the host vehicle, for example, and produces denser third three-dimensional data 636 by synthesizing the received second three-dimensional data 635 and the first three-dimensional data 632 produced by the three-dimensional data production device 620.
[0392] The three-dimensional data production device 620 includes: a three-dimensional data production unit 621, a request range determination unit 622, a search unit 623, a reception unit 624, a decoding unit 625, and a synthesis unit 626.
[0393] First, the three-dimensional data production unit 621 produces the first three-dimensional data 632 by using the sensor information 631 detected by the sensors equipped on its own vehicle. Next, the request range determination unit 622 determines the request range, which refers to the three-dimensional space range where the data in the produced first three-dimensional data 632 is insufficient.
[0394] Next, the search unit 623 searches for surrounding vehicles that hold the three-dimensional data within the request range, and sends the request range information 633 indicating the request range to the surrounding vehicles determined through the search. Next, the receiving unit 624 receives the encoded three-dimensional data 634 (S624), which is the encoded stream of the request range, from the surrounding vehicles. In addition, the search unit 623 can send requests to all vehicles existing within the determined range without discrimination, and receive the encoded three-dimensional data 634 from the responding counterparts. Also, the search unit 623 is not limited to vehicles, and can also send requests to objects such as traffic lights or signs, and receive the encoded three-dimensional data 634 from such objects.
[0395] Next, the received encoded three-dimensional data 634 is decoded by the decoding unit 625 to obtain the second three-dimensional data 635. Next, the first three-dimensional data 632 and the second three-dimensional data 635 are synthesized by the synthesis unit 626 to produce a denser third three-dimensional data 636.
[0396] Next, the configuration and operation of the three-dimensional data transmission device 640 according to the present embodiment will be described. Figure 25 It is a block diagram of the three-dimensional data transmission device 640.
[0397] The three-dimensional data transmission device 640 is included, for example, in the above-mentioned surrounding vehicles, processes the fifth three-dimensional data 652 produced by the surrounding vehicles into the sixth three-dimensional data 654 requested by its own vehicle, generates the encoded three-dimensional data 634 by encoding the sixth three-dimensional data 654, and sends the encoded three-dimensional data 634 to its own vehicle.
[0398] The three-dimensional data transmission device 640 includes: a three-dimensional data production unit 641, a receiving unit 642, an extraction unit 643, an encoding unit 644, and a sending unit 645.
[0399] First, the three-dimensional data production unit 641 produces the fifth three-dimensional data 652 by using the sensor information 651 detected by the sensors equipped on the surrounding vehicles. Next, the receiving unit 642 receives the request range information 633 sent from its own vehicle.
[0400] Next, the extraction unit 643 extracts the three-dimensional data of the requested range represented by the request range information 633 from the fifth three-dimensional data 652, and processes the fifth three-dimensional data 652 into the sixth three-dimensional data 654. Next, the encoding unit 644 encodes the sixth three-dimensional data 654 to generate the encoded three-dimensional data 634 as an encoded stream. Then, the transmission unit 645 transmits the encoded three-dimensional data 634 to its own vehicle.
[0401] In addition, here, although an example in which the own vehicle is equipped with the three-dimensional data production device 620 and the surrounding vehicles are equipped with the three-dimensional data transmission device 640 has been described, each vehicle may also have the functions of the three-dimensional data production device 620 and the three-dimensional data transmission device 640.
[0402] (Embodiment 4)
[0403] In this embodiment, the operation related to abnormal conditions in the own position estimation based on the three-dimensional map will be described.
[0404] The use of autonomous movement of moving bodies such as the autonomous driving of motor vehicles, robots, or flying objects such as drones will expand in the future. As an example of a method for realizing such autonomous movement, there is a method in which the moving body estimates its own position in the three-dimensional map (own position estimation) and travels according to the map.
[0405] The own position estimation is achieved by matching the three-dimensional map with the three-dimensional information around the own vehicle (hereinafter referred to as the own vehicle detection three-dimensional data) obtained by sensors such as a distance measuring instrument (LIDAR, etc.) or a stereo camera mounted on the own vehicle, and estimating the position of the own vehicle in the three-dimensional map.
[0406] As shown in the HD map proposed by HERE Corporation, the three-dimensional map can include not only three-dimensional point clouds, but also two-dimensional map data such as the shapes of roads and intersections, or information that changes in real time such as congestion and accidents. The three-dimensional map is composed of multiple layers such as three-dimensional data, two-dimensional data, and metadata that changes in real time. The device can obtain only the required data or can also refer to the required data.
[0407] The data of the point cloud can be the above-mentioned SWLD or can also include point group data that is not feature points. And the transmission and reception of the data of the point cloud are basically performed in one or more random access units.
[0408] As a method for matching a three-dimensional map with three-dimensional data detected by the vehicle itself, the following method can be adopted. For example, the device compares the shapes of point groups in the point clouds of each other, and determines the part with a high similarity between feature points as the same position. Moreover, when the three-dimensional map is composed of SWLD, the device compares the feature points constituting the SWLD with the three-dimensional feature points extracted from the three-dimensional data detected by the vehicle itself and performs matching.
[0409] Here, in order to estimate the vehicle's own position with high accuracy, the following (A) and (B) need to be satisfied. (A) A three-dimensional map and three-dimensional data detected by the vehicle itself can already be obtained. (B) Their accuracies satisfy a predetermined standard. However, in the following abnormal situations, (A) or (B) cannot be satisfied.
[0410] (1) The three-dimensional map cannot be obtained through the communication path.
[0411] (2) There is no three-dimensional map, or the obtained three-dimensional map is damaged.
[0412] (3) The sensors of the vehicle itself malfunction, or due to bad weather, the generation accuracy of the three-dimensional data detected by the vehicle itself is insufficient.
[0413] The operations for coping with these abnormal situations will be described below. Although the operations are described below taking a vehicle as an example, the following methods can also be applied to all moving objects such as robots or drones that perform autonomous movement.
[0414] What will be described below is the configuration and operations of the three-dimensional information processing device according to the present embodiment for coping with abnormal situations in the three-dimensional map or the three-dimensional data detected by the vehicle itself. Figure 26 It is a block diagram showing a configuration example of the three-dimensional information processing device 700 according to the present embodiment.
[0415] The three-dimensional information processing device 700 is mounted on a moving object such as a motor vehicle, for example. As Figure 26 shown, the three-dimensional information processing device 700 includes: a three-dimensional map acquisition unit 701, a vehicle's own detection data acquisition unit 702, an abnormal situation determination unit 703, a coping operation determination unit 704, and an operation control unit 705.
[0416] In addition, the three-dimensional information processing device 700 may include a camera that acquires a two-dimensional image, or may include a two-dimensional or one-dimensional sensor (not shown) such as a sensor that uses ultrasonic waves or lasers to detect a structural object or a moving object around the vehicle itself. Moreover, the three-dimensional information processing device 700 may include a communication unit (not shown) that is used to acquire a three-dimensional map through a mobile communication network such as 4G or 5G, or vehicle-to-vehicle communication, or road-to-vehicle communication.
[0417] The three-dimensional map acquisition unit 701 acquires a three-dimensional map 711 near the driving route. For example, the three-dimensional map acquisition unit 701 acquires the three-dimensional map 711 through a mobile communication network, vehicle-to-vehicle communication, or road-to-vehicle communication.
[0418] Next, the own vehicle detection data acquisition unit 702 acquires own vehicle detection three-dimensional data 712 based on sensor information. For example, the own vehicle detection data acquisition unit 702 generates own vehicle detection three-dimensional data 712 based on the sensor information obtained by the sensors equipped on the own vehicle.
[0419] Next, the abnormal situation determination unit 703 detects an abnormal situation by performing a pre-determined check on at least one of the acquired three-dimensional map 711 and the own vehicle detection three-dimensional data 712. That is, the abnormal situation determination unit 703 determines whether at least one of the acquired three-dimensional map 711 and the own vehicle detection three-dimensional data 712 is abnormal.
[0420] When an abnormal situation is detected, the response work determination unit 704 determines the response work for the abnormal situation. Next, the work control unit 705 controls the work of each processing unit required in the implementation of the response work, such as the three-dimensional map acquisition unit 701.
[0421] In addition, when no abnormal situation is detected, the three-dimensional information processing device 700 ends the processing.
[0422] Furthermore, the three-dimensional information processing device 700 estimates the own position of the vehicle equipped with the three-dimensional information processing device 700 by using the three-dimensional map 711 and the own vehicle detection three-dimensional data 712. Next, the three-dimensional information processing device 700 uses the result of the own position estimation to make the vehicle perform autonomous driving.
[0423] Accordingly, the three-dimensional information processing device 700 acquires map data (three-dimensional map 711) including first three-dimensional position information via a channel. For example, the first three-dimensional position information is encoded in units of partial spaces having three-dimensional coordinate information, the first three-dimensional position information includes a plurality of random access units, each of the plurality of random access units is an aggregate of one or more partial spaces, and can be independently decoded. For example, the first three-dimensional position information is data (SWLD) in which feature points where three-dimensional feature amounts exceed a specified threshold are encoded.
[0424] Further, the three-dimensional information processing device 700 generates second three-dimensional position information (own vehicle detection three-dimensional data 712) based on the information detected by the sensor. Next, the three-dimensional information processing device 700 determines whether the first three-dimensional position information or the second three-dimensional position information is abnormal by performing an abnormality determination process on the first three-dimensional position information or the second three-dimensional position information.
[0425] When the three-dimensional information processing device 700 determines that the first three-dimensional position information or the second three-dimensional position information is abnormal, it determines a response operation for the abnormality. Next, the three-dimensional information processing device 700 performs control required for the implementation of the response operation.
[0426] Accordingly, the three-dimensional information processing device 700 can detect an abnormality in the first three-dimensional position information or the second three-dimensional position information and can perform a response operation.
[0427] (Embodiment 5)
[0428] In the present embodiment, a method for transmitting three-dimensional data to a following vehicle and the like will be described.
[0429] Figure 27 FIG. is a block diagram showing a configuration example of a three-dimensional data production device 810 according to the present embodiment. The three-dimensional data production device 810 is mounted on a vehicle, for example. The three-dimensional data production device 810 transmits and receives three-dimensional data to and from external traffic cloud monitoring, a preceding vehicle, or a following vehicle, and at the same time produces and stores the three-dimensional data.
[0430] The three-dimensional data production device 810 includes: a data reception unit 811, a communication unit 812, a reception control unit 813, a format conversion unit 814, a plurality of sensors 815, a three-dimensional data production unit 816, a three-dimensional data synthesis unit 817, a three-dimensional data storage unit 818, a communication unit 819, a transmission control unit 820, a format conversion unit 821, and a data transmission unit 822.
[0431] The data reception unit 811 receives three-dimensional data 831 from traffic cloud monitoring or a preceding vehicle. The three-dimensional data 831 includes, for example, point clouds, visible light images, depth information, sensor position information, or speed information that contains information on areas that cannot be detected by the sensors 815 of the own vehicle.
[0432] The communication unit 812 communicates with traffic cloud monitoring or a preceding vehicle and sends a data transmission request or the like to traffic cloud monitoring or a preceding vehicle.
[0433] The reception control unit 813 exchanges information such as corresponding formats with the communication partner via the communication unit 812 to establish communication with the communication partner.
[0434] The format conversion unit 814 generates the three-dimensional data 832 by performing format conversion or the like on the three-dimensional data 831 received by the data reception unit 811. Further, when the three-dimensional data 831 is compressed or encoded, the format conversion unit 814 performs decompression or decoding processing.
[0435] The plurality of sensors 815 are a group of sensors such as LiDAR, visible light cameras, or infrared cameras that obtain information about the exterior of the vehicle, and generate sensor information 833. For example, when the sensor 815 is a laser sensor such as LiDAR, the sensor information 833 is three-dimensional data such as point clouds (point group data). Additionally, the sensor 815 may not be plural.
[0436] The three-dimensional data creation unit 816 generates three-dimensional data 834 based on the sensor information 833. The three-dimensional data 834 includes, for example, information such as point clouds, visible light images, depth information, sensor position information, or speed information.
[0437] The three-dimensional data synthesis unit 817 synthesizes the three-dimensional data 832 created by traffic cloud monitoring or a preceding vehicle or the like into the three-dimensional data 834 created based on the sensor information 833 of its own vehicle, thereby enabling the construction of three-dimensional data 835 that also includes the space in front of the preceding vehicle that cannot be detected by the sensors 815 of its own vehicle.
[0438] The three-dimensional data storage unit 818 stores the generated three-dimensional data 835 and the like.
[0439] The communication unit 819 communicates with traffic cloud monitoring or a following vehicle, and sends a data transmission request or the like to traffic cloud monitoring or a following vehicle.
[0440] The transmission control unit 820 exchanges information such as the corresponding format with the communication partner via the communication unit 819 to establish communication with the communication partner. Further, the transmission control unit 820 determines the transmission area of the space of the three-dimensional data to be transmitted based on the three-dimensional data construction information of the three-dimensional data 832 generated by the three-dimensional data synthesis unit 817 and the data transmission request from the communication partner.
[0441] Specifically, the transmission control unit 820 determines the transmission area that includes the space in front of its own vehicle that cannot be detected by the sensors of the following vehicle in accordance with the data transmission request from traffic cloud monitoring or a following vehicle. Further, the transmission control unit 820 determines the transmission area by judging, based on the three-dimensional data construction information, whether there is an update to the space that can be transmitted or the space that has already been transmitted. For example, the transmission control unit 820 determines as the transmission area the area that is both specified by the data transmission request and where the corresponding three-dimensional data 835 exists. Further, the transmission control unit 820 notifies the format conversion unit 821 of the format corresponding to the communication partner and the transmission area.
[0442] The format conversion unit 821 generates the three-dimensional data 837 by converting the three-dimensional data 836 of the transmission area in the three-dimensional data 835 stored in the three-dimensional data storage unit 818 into a format corresponding to the receiving side. Additionally, the format conversion unit 821 can also compress or encode the three-dimensional data 837 to reduce the data volume.
[0443] The data transmission unit 822 transmits the three-dimensional data 837 to the traffic cloud monitoring or the following vehicle. This three-dimensional data 837 includes, for example, the point cloud in front of the own vehicle containing information on the area that is a blind spot for the following vehicle, visible light images, depth information, or sensor position information, etc.
[0444] In addition, although the format conversion units 814 and 821 are taken as examples for format conversion and the like here, format conversion may not be performed.
[0445] With this configuration, the three-dimensional data production device 810 obtains the three-dimensional data 831 of the area that cannot be detected by the sensor 815 of the own vehicle from the outside, and generates the three-dimensional data 835 by synthesizing the three-dimensional data 831 and the three-dimensional data 834 based on the sensor information 833 detected by the sensor 815 of the own vehicle. Accordingly, the three-dimensional data production device 810 can generate the three-dimensional data of the range that cannot be detected by the sensor 815 of the own vehicle.
[0446] Moreover, the three-dimensional data production device 810 can send the three-dimensional data of the space in front of the own vehicle that cannot be detected by the sensors of the following vehicle to the traffic cloud monitoring or the following vehicle, etc., in accordance with the data transmission request from the traffic cloud monitoring or the following vehicle.
[0447] (Embodiment 6)
[0448] In the example to be described in Embodiment 5, the client device such as a vehicle sends the three-dimensional data to other vehicles or servers such as traffic cloud monitoring. In this embodiment, the client device sends the sensor information obtained by the sensor to the server or other client devices.
[0449] First, the configuration of the system according to this embodiment will be described. Figure 28 The configuration of the three-dimensional map and the transceiver system of the sensor information according to this embodiment is shown. This system includes a server 901, and client devices 902A and 902B. Additionally, when not making a special distinction between the client devices 902A and 902B, they are also denoted as the client device 902.
[0450] The client device 902 is, for example, an in-vehicle device mounted on a moving body such as a vehicle. The server 901 is, for example, a traffic cloud monitor or the like and can communicate with multiple client devices 902.
[0451] The server 901 sends a three-dimensional map composed of point clouds to the client device 902. Additionally, the composition of the three-dimensional map is not limited to point clouds and can also be represented by other three-dimensional data such as a mesh structure.
[0452] The client device 902 sends the sensor information obtained by the client device 902 to the server 901. The sensor information includes, for example, at least one of the information obtained by LiDAR, visible light images, infrared images, depth images, sensor position information, and speed information.
[0453] Regarding the data transmitted and received between the server 901 and the client device 902, it can be compressed when wanting to reduce the data, and can be not compressed when wanting to maintain the accuracy of the data. When compressing the data, for example, a three-dimensional compression method based on an octree can be adopted in the point cloud. And, a two-dimensional image compression method can be adopted in visible light images, infrared images, and depth images. The two-dimensional image compression method is, for example, MPEG-4 AVC or HEVC standardized by MPEG.
[0454] And, the server 901 sends the three-dimensional map managed by the server 901 to the client device 902 in accordance with the sending request of the three-dimensional map from the client device 902. Additionally, the server 901 can also send the three-dimensional map without waiting for the sending request of the three-dimensional map from the client device 902. For example, the server 901 can also broadcast the three-dimensional map to one or more client devices 902 in a pre-specified space. And, the server 901 can also send a three-dimensional map suitable for the position of the client device 902 to the client device 902 that has received a sending request once at regular intervals. And, the server 901 can also send the three-dimensional map to the client device 902 whenever the three-dimensional map managed by the server 901 is updated.
[0455] The client device 902 sends a sending request for the three-dimensional map to the server 901. For example, when the client device 902 wants to estimate its own position while driving, the client device 902 sends the sending request for the three-dimensional map to the server 901.
[0456] In addition, the client device 902 may also send a request to the server 901 to send a 3D map under the following circumstances. When the 3D map held by the client device 902 is relatively old, the client device 902 may also send a request to the server 901 to send a 3D map. For example, when a certain period of time has passed since the client device 902 obtained the 3D map, the client device 902 may also send a request to the server 901 to send a 3D map.
[0457] It may also be that, a certain moment before the client device 902 is about to leave the space shown in the 3D map held by the client device 902, the client device 902 sends a request to the server 901 to send a 3D map. For example, it may also be that when the client device 902 is within a pre-specified distance from the boundary of the space shown in the 3D map held by the client device 902, the client device 902 sends a request to the server 901 to send a 3D map. Moreover, when the movement path and movement speed of the client device 902 are known, the moment when the client device 902 leaves the space shown in the 3D map held by the client device 902 can be predicted based on the known movement path and movement speed.
[0458] When the error in the position comparison between the 3D data generated by the client device 902 based on sensor information and the 3D map is above a certain range, the client device 902 may send a request to the server 901 to send a 3D map.
[0459] The client device 902 sends the sensor information to the server 901 in accordance with the request from the server 901 to send the sensor information. In addition, the client device 902 may also send the sensor information to the server 901 without waiting for the request from the server 901 to send the sensor information. For example, when the client device 902 has received a request from the server 901 to send the sensor information once, it may regularly send the sensor information to the server 901 within a certain period. It may also be that when the error in the position comparison between the 3D data generated by the client device 902 based on sensor information and the 3D map obtained from the server 901 is above a certain range, the client device 902 determines that there is a possibility that the 3D map around the client device 902 has changed, and sends this judgment result together with the sensor information to the server 901.
[0460] Server 901 sends a request to client device 902 to send sensor information. For example, server 901 receives location information of client device 902 such as GPS from client device 902. When server 901 determines, based on the location information of client device 902, that client device 902 is approaching a space with less information in the three-dimensional map managed by server 901, in order to regenerate the three-dimensional map, server 901 sends a request to client device 902 to send sensor information. Also, server 901 may send a request to send sensor information when it wants to update the three-dimensional map, when it wants to confirm road conditions such as during snow accumulation or disasters, or when it wants to confirm traffic jams or accident situations.
[0461] Also, client device 902 may set the amount of sensor information to be sent to server 901 according to the communication state or frequency band at the time of receiving the request to send sensor information from server 901. Setting the amount of sensor information to be sent to server 901, for example, means increasing or decreasing the data itself, or selecting an appropriate compression method.
[0462] Figure 29 It is a block diagram showing a configuration example of client device 902. Client device 902 receives a three-dimensional map composed of point clouds, etc. from server 901, and estimates its own position based on three-dimensional data created from the sensor information of client device 902. Then, client device 902 sends the acquired sensor information to server 901.
[0463] Client device 902 includes: data receiving unit 1011, communication unit 1012, reception control unit 1013, format conversion unit 1014, multiple sensors 1015, three-dimensional data creation unit 1016, three-dimensional image processing unit 1017, three-dimensional data storage unit 1018, format conversion unit 1019, communication unit 1020, transmission control unit 1021, and data transmission unit 1022.
[0464] Data receiving unit 1011 receives three-dimensional map 1031 from server 901. Three-dimensional map 1031 is data including point clouds such as WLD or SWLD. Three-dimensional map 1031 may include either compressed data or uncompressed data.
[0465] Communication unit 1012 communicates with server 901 and sends a data transmission request (for example, a request to send a three-dimensional map) to server 901.
[0466] Reception control unit 1013 exchanges information such as corresponding formats with the communication partner via communication unit 1012 to establish communication with the communication partner.
[0467] The format conversion unit 1014 generates a 3D map 1032 by performing format conversion and the like on the 3D map 1031 received by the data reception unit 1011. Moreover, when the 3D map 1031 is compressed or encoded, the format conversion unit 1014 performs decompression or decoding processing. In addition, when the 3D map 1031 is uncompressed data, the format conversion unit 1014 does not perform decompression or decoding processing.
[0468] The plurality of sensors 1015 are a group of sensors mounted on the client device 902 such as a LiDAR, a visible light camera, an infrared camera, or a depth sensor, which are used to obtain information about the outside of the vehicle, and generate sensor information 1033. For example, when the sensor 1015 is a laser sensor such as a LiDAR, the sensor information 1033 is 3D data such as point cloud (point group data). In addition, the number of sensors 1015 may not be plural.
[0469] The 3D data production unit 1016 produces 3D data 1034 around its own vehicle according to the sensor information 1033. For example, the 3D data production unit 1016 uses the information obtained by the LiDAR and the visible light image obtained by the visible light camera to produce point cloud data with color information around its own vehicle.
[0470] The 3D image processing unit 1017 performs its own position estimation processing and the like of its own vehicle by using the received 3D map 1032 such as point cloud and the 3D data 1034 around its own vehicle generated according to the sensor information 1033. In addition, it is also possible that the 3D image processing unit 1017 synthesizes the 3D map 1032 and the 3D data 1034 to produce 3D data 1035 around its own vehicle, and uses the produced 3D data 1035 to perform its own position estimation processing.
[0471] The 3D data storage unit 1018 stores the 3D map 1032, the 3D data 1034, the 3D data 1035, and the like.
[0472] The format conversion unit 1019 generates sensor information 1037 by converting the sensor information 1033 into a format corresponding to the receiving side. In addition, the format conversion unit 1019 may reduce the data amount by compressing or encoding the sensor information 1037. Moreover, when format conversion is not required, the format conversion unit 1019 may omit the processing. And the format conversion unit 1019 can control the data amount of the data transmitted according to the specified transmission range.
[0473] The communication unit 1020 communicates with the server 901 and receives a data transmission request (a transmission request for sensor information) and the like from the server 901.
[0474] The transmission control unit 1021 exchanges information such as corresponding formats with the communication partner via the communication unit 1020 to establish communication.
[0475] The data transmission unit 1022 transmits the sensor information 1037 to the server 901. The sensor information 1037 includes, for example, information obtained by LiDAR, luminance images (visible light images) obtained by visible light cameras, infrared images obtained by infrared cameras, depth images obtained by depth sensors, sensor position information, and speed information, etc., which are obtained by a plurality of sensors 1015.
[0476] Next, the configuration of the server 901 will be described. Figure 30 It is a block diagram showing a configuration example of the server 901. The server 901 receives the sensor information sent from the client device 902, and creates three-dimensional data based on the received sensor information. The server 901 updates the three-dimensional map managed by the server 901 using the created three-dimensional data. And the server 901 sends the updated three-dimensional map to the client device 902 according to the transmission request of the three-dimensional map from the client device 902.
[0477] The server 901 includes: a data reception unit 1111, a communication unit 1112, a reception control unit 1113, a format conversion unit 1114, a three-dimensional data creation unit 1116, a three-dimensional data synthesis unit 1117, a three-dimensional data storage unit 1118, a format conversion unit 1119, a communication unit 1120, a transmission control unit 1121, and a data transmission unit 1122.
[0478] The data reception unit 1111 receives the sensor information 1037 from the client device 902. The sensor information 1037 includes, for example, information obtained by LiDAR, luminance images (visible light images) obtained by visible light cameras, infrared images obtained by infrared cameras, depth images obtained by depth sensors, sensor position information, and speed information, etc.
[0479] The communication unit 1112 communicates with the client device 902 and sends a data transmission request (for example, a transmission request for sensor information) to the client device 902.
[0480] The reception control unit 1113 exchanges information such as corresponding formats with the communication partner via the communication unit 1112 to establish communication.
[0481] When the received sensor information 1037 is compressed or encoded, the format conversion unit 1114 generates the sensor information 1132 by performing decompression or decoding processing. Additionally, when the sensor information 1037 is uncompressed data, the format conversion unit 1114 does not perform decompression or decoding processing.
[0482] Based on the sensor information 1132, the three-dimensional data creation unit 1116 creates three-dimensional data 1134 of the surroundings of the client device 902. For example, the three-dimensional data creation unit 1116 uses the information obtained by LiDAR and the visible light image obtained by the visible light camera to create point cloud data with color information of the surroundings of the client device 902.
[0483] The three-dimensional data synthesis unit 1117 synthesizes the three-dimensional data 1134 created based on the sensor information 1132 with the three-dimensional map 1135 managed by the server 901, thereby updating the three-dimensional map 1135.
[0484] The three-dimensional data storage unit 1118 stores the three-dimensional map 1135 and the like.
[0485] The format conversion unit 1119 generates the three-dimensional map 1031 by converting the three-dimensional map 1135 into a format corresponding to the receiving side. Additionally, the format conversion unit 1119 can also reduce the data volume by compressing or encoding the three-dimensional map 1135. And when format conversion is not required, the format conversion unit 1119 can also omit the processing. And the format conversion unit 1119 can control the data volume to be sent according to the specified sending range.
[0486] The communication unit 1120 communicates with the client device 902 and receives a data transmission request (a transmission request for the three-dimensional map) and the like from the client device 902.
[0487] The transmission control unit 1121 exchanges information such as the corresponding format with the communication partner via the communication unit 1120 to establish communication.
[0488] The data transmission unit 1122 sends the three-dimensional map 1031 to the client device 902. The three-dimensional map 1031 is data including point clouds such as WLD or SWLD. Either compressed data or uncompressed data may be included in the three-dimensional map 1031.
[0489] Next, the operation flow of the client device 902 will be described. Figure 31 It is a flowchart showing the operations when the client device 902 obtains the three-dimensional map.
[0490] First, the client device 902 requests the server 901 to send a three-dimensional map (such as a point cloud) (S1001). At this time, the client device 902 also sends the position information of the client device 902 obtained through GPS or the like. Accordingly, the client device 902 can request the server 901 to send a three-dimensional map related to the position information.
[0491] Next, the client device 902 receives the three-dimensional map from the server 901 (S1002). If the received three-dimensional map is compressed data, the client device 902 decodes the received three-dimensional map to generate an uncompressed three-dimensional map (S1003).
[0492] Next, the client device 902 creates three-dimensional data 1034 of the periphery of the client device 902 based on the sensor information 1033 obtained from the plurality of sensors 1015 (S1004). Next, the client device 902 estimates the own position of the client device 902 using the three-dimensional map 1032 received from the server 901 and the three-dimensional data 1034 created based on the sensor information 1033 (S1005).
[0493] Figure 32 It is a flowchart showing the operation when the client device 902 sends sensor information. First, the client device 902 receives a request to send sensor information from the server 901 (S1011). The client device 902 that has received the send request sends the sensor information 1037 to the server 901 (S1012). In addition, when the sensor information 1033 includes a plurality of information obtained through the plurality of sensors 1015, the client device 902 compresses each information in a compression method suitable for each information, thereby generating the sensor information 1037.
[0494] Next, the operation process of the server 901 will be described. Figure 33 It is a flowchart showing the operation when the server 901 obtains sensor information. First, the server 901 requests the client device 902 to send sensor information (S1021). Next, the server 901 receives the sensor information 1037 sent from the client device 902 in accordance with the request (S1022). Next, the server 901 creates three-dimensional data 1134 using the received sensor information 1037 (S1023). Next, the server 901 reflects the created three-dimensional data 1134 in the three-dimensional map 1135 (S1024).
[0495] Figure 34It is a flowchart showing the operations when the server 901 sends a 3D map. First, the server 901 receives a 3D map sending request from the client device 902 (S1031). The server 901 that has received the 3D map sending request sends the 3D map 1031 to the client device 902 (S1032). At this time, the server 901 can extract the 3D map in its vicinity corresponding to the location information of the client device 902 and send the extracted 3D map. And it can be that the server 901 compresses the 3D map composed of point clouds, for example, using a compression method such as an octree, and sends the compressed 3D map.
[0496] Hereinafter, a modified example of the present embodiment will be described.
[0497] The server 901 uses the sensor information 1037 received from the client device 902 to create 3D data 1134 near the location of the client device 902. Next, the server 901 matches the created 3D data 1134 with the 3D map 1135 of the same area managed by the server 901 and calculates the difference between the 3D data 1134 and the 3D map 1135. When the difference is equal to or greater than a predetermined threshold, the server 901 determines that some abnormality has occurred around the client device 902. For example, when the ground surface sinks due to a natural disaster such as an earthquake, a large difference may be considered to occur between the 3D map 1135 managed by the server 901 and the 3D data 1134 created based on the sensor information 1037.
[0498] The sensor information 1037 may also include at least one of the type of the sensor, the performance of the sensor, and the model of the sensor. Also, it may be that a category ID corresponding to the performance of the sensor is attached to the sensor information 1037. For example, when the sensor information 1037 is information obtained by LiDAR, it is possible to consider assigning identifiers according to the performance of the sensor. For example, category 1 is assigned to a sensor capable of obtaining information with an accuracy of several millimeters, category 2 is assigned to a sensor capable of obtaining information with an accuracy of several centimeters, and category 3 is assigned to a sensor capable of obtaining information with an accuracy of several meters. Also, the server 901 may estimate the performance information of the sensor from the model of the client device 902. For example, when the client device 902 is mounted on a vehicle, the server 901 can determine the specification information of the sensor based on the model of the vehicle. In this case, the server 901 may obtain the information of the vehicle model in advance or include this information in the sensor information. Also, it may be that the server 901 uses the obtained sensor information 1037 to switch the degree of correction for the three-dimensional data 1134 created using the sensor information 1037. For example, when the sensor performance is high accuracy (category 1), the server 901 does not perform correction on the three-dimensional data 1134. When the sensor performance is low accuracy (category 3), the server 901 applies correction suitable for the accuracy of the sensor to the three-dimensional data 1134. For example, the server 901 increases the degree (intensity) of correction as the accuracy of the sensor decreases.
[0499] The server 901 may also send a request to send sensor information to multiple client devices 902 existing in a certain space simultaneously. When the server 901 receives multiple sensor information from the multiple client devices 902, it is not necessary to use all the sensor information for the creation of the three-dimensional data 1134. For example, the sensor information to be used can be selected according to the performance of the sensor. For example, when updating the three-dimensional map 1135, the server 901 can select high-accuracy sensor information (category 1) from the received multiple sensor information and use the selected sensor information to create the three-dimensional data 1134.
[0500] The server 901 is not limited to servers such as traffic cloud monitoring, and may also be other client devices (in-vehicle). Figure 35 The system configuration in this case is shown.
[0501] For example, the client device 902C sends a request to the nearby client device 902A to send sensor information, and obtains the sensor information from the client device 902A. Then, the client device 902C uses the obtained sensor information of the client device 902A to create three-dimensional data and update the three-dimensional map of the client device 902C. In this way, the client device 902C can utilize the performance of the client device 902C to generate a three-dimensional map of the space that can be obtained from the client device 902A. For example, this situation can be considered when the performance of the client device 902C is high.
[0502] Moreover, in this case, the client device 902A that provided the sensor information is given the right to obtain the highly accurate three-dimensional map generated by the client device 902C. The client device 902A receives the highly accurate three-dimensional map from the client device 902C according to this right.
[0503] Alternatively, the client device 902C may send a request to send sensor information to multiple nearby client devices 902 (client device 902A and client device 902B). When the sensor of the client device 902A or the client device 902B has high performance, the client device 902C can use the sensor information obtained through this high-performance sensor to create three-dimensional data.
[0504] Figure 36 It is a block diagram showing the functional configuration of the server 901 and the client device 902. The server 901 includes, for example: a three-dimensional map compression / decoding processing unit 1201 that compresses and decodes a three-dimensional map, and a sensor information compression / decoding processing unit 1202 that compresses and decodes sensor information.
[0505] The client device 902 includes: a three-dimensional map decoding processing unit 1211 and a sensor information compression processing unit 1212. The three-dimensional map decoding processing unit 1211 receives the encoded data of the compressed three-dimensional map, decodes the encoded data, and obtains the three-dimensional map. The sensor information compression processing unit 1212 does not compress the three-dimensional data created from the obtained sensor information, but compresses the sensor information itself and sends the encoded data of the compressed sensor information to the server 901. According to this configuration, the client device 902 can keep the processing unit (device or LSI) for decoding the three-dimensional map (point cloud, etc.) inside, without having to keep the processing unit for compressing the three-dimensional data of the three-dimensional map (point cloud, etc.) inside. In this way, the cost and power consumption of the client device 902 can be suppressed.
[0506] As described above, the client device 902 according to the present embodiment is mounted on a moving body, and creates three-dimensional data 1034 of the surroundings of the moving body based on sensor information 1033 obtained by a sensor 1015 mounted on the moving body and showing the surrounding conditions of the moving body. The client device 902 estimates its own position of the moving body by using the created three-dimensional data 1034. The client device 902 sends the obtained sensor information 1033 to the server 901 or other moving bodies 902.
[0507] Accordingly, the client device 902 sends the sensor information 1033 to the server 901 or the like. In this way, there is a possibility that the amount of data to be transmitted can be reduced compared with the case of sending three-dimensional data. And since it is not necessary to perform processing such as compression or encoding of three-dimensional data on the client device 902, the amount of processing of the client device 902 can be reduced. Therefore, the client device 902 can achieve a reduction in the amount of transmitted data or a simplification of the device configuration.
[0508] In addition, the client device 902 further sends a transmission request for a three-dimensional map to the server 901, and receives a three-dimensional map 1031 from the server 901. The client device 902 estimates its own position by using the three-dimensional data 1034 and the three-dimensional map 1032 in the estimation of its own position.
[0509] Moreover, the sensor information 1033 includes at least one of information obtained by a laser sensor, a luminance image (visible light image), an infrared image, a depth image, position information of the sensor, and speed information of the sensor.
[0510] And the sensor information 1033 includes information showing the performance of the sensor.
[0511] In addition, the client device 902 encodes or compresses the sensor information 1033, and sends the encoded or compressed sensor information 1037 to the server 901 or other moving bodies 902 in the transmission of the sensor information. Accordingly, the client device 902 can reduce the amount of data transmitted.
[0512] For example, the client device 902 includes a processor and a memory, and the processor performs the above processing by using the memory.
[0513] And the server 901 according to the present embodiment can communicate with the client device 902 mounted on a moving body, and receives sensor information 1037 obtained by a sensor 1015 mounted on the moving body and showing the surrounding conditions of the moving body. The server 901 creates three-dimensional data 1134 of the surroundings of the moving body based on the received sensor information 1037.
[0514] Accordingly, the server 901 uses the sensor information 1037 sent from the client device 902 to create three-dimensional data 1134. In this way, compared with the case where the client device 902 sends three-dimensional data, there is a possibility of reducing the amount of data to be sent. Also, since it is not necessary to perform processes such as compression or encoding of three-dimensional data on the client device 902, the processing amount of the client device 902 can be reduced. In this way, the server 901 can achieve a reduction in the amount of transmitted data or a simplification of the device configuration.
[0515] Furthermore, the server 901 further sends a transmission request for the sensor information to the client device 902.
[0516] Furthermore, the server 901 further uses the created three-dimensional data 1134 to update the three-dimensional map 1135, and sends the three-dimensional map 1135 to the client device 902 in accordance with a transmission request for the three-dimensional map 1135 from the client device 902.
[0517] Moreover, the sensor information 1037 includes at least one of information obtained by a laser sensor, a luminance image (visible light image), an infrared image, a depth image, the position information of the sensor, and the speed information of the sensor.
[0518] Moreover, the sensor information 1037 includes information indicating the performance of the sensor.
[0519] Furthermore, the server 901 corrects the three-dimensional data according to the performance of the sensor. Accordingly, this method for creating three-dimensional data can improve the quality of the three-dimensional data.
[0520] Moreover, in receiving the sensor information, the server 901 receives a plurality of sensor information 1037 from a plurality of client devices 902, and selects the sensor information 1037 to be used in creating the three-dimensional data 1134 according to a plurality of information indicating the performance of the sensors included in the plurality of sensor information 1037. Accordingly, the server 901 can improve the quality of the three-dimensional data 1134.
[0521] Moreover, the server 901 decodes or decompresses the received sensor information 1037, and creates three-dimensional data 1134 according to the decoded or decompressed sensor information 1132. Accordingly, the server 901 can reduce the amount of transmitted data.
[0522] For example, the server 901 includes a processor and a memory, and the processor uses the memory to perform the above processing.
[0523] (Embodiment 7)
[0524] In this embodiment, a method for encoding and decoding three-dimensional data using inter-frame prediction processing will be described.
[0525] Figure 37 FIG. is a block diagram of a three-dimensional data encoding apparatus 1300 according to this embodiment. The three-dimensional data encoding apparatus 1300 generates an encoded bitstream (hereinafter also simply referred to as a bitstream) as an encoded signal by encoding three-dimensional data. As Figure 37 shown, the three-dimensional data encoding apparatus 1300 includes: a partitioning unit 1301, a subtraction unit 1302, a transformation unit 1303, a quantization unit 1304, an inverse quantization unit 1305, an inverse transformation unit 1306, an addition unit 1307, a reference volume memory 1308, an intra-frame prediction unit 1309, a reference space memory 1310, an inter-frame prediction unit 1311, a prediction control unit 1312, and an entropy encoding unit 1313.
[0526] The partitioning unit 1301 divides each space (SPC) included in the three-dimensional data into a plurality of volumes (VLM) as encoding units. Further, the partitioning unit 1301 performs octree representation (octree conversion) on the voxels within each volume. Additionally, the partitioning unit 1301 may make the space and the volume the same size and perform octree representation on the space. Moreover, the partitioning unit 1301 may attach information (such as depth information) required for octree conversion to the head of the bitstream or the like.
[0527] The subtraction unit 1302 calculates the difference between the volume (encoding target volume) output from the partitioning unit 1301 and the predicted volume generated by intra-frame prediction or inter-frame prediction described later, and outputs the calculated difference as a prediction residual to the transformation unit 1303. Figure 38 FIG. shows an example of calculating the prediction residual. Additionally, the bit strings of the encoding target volume and the predicted volume shown here are, for example, position information indicating the positions of three-dimensional points (such as point clouds) included in the volume.
[0528] Hereinafter, octree representation and the voxel scanning order will be described. After the volume is transformed into an octree structure (octree conversion), it is encoded. The octree structure is composed of nodes and leaf nodes. Each node has eight nodes or leaf nodes, and each leaf node has voxel (VXL) information. Figure 39 FIG. shows a configuration example of a volume including a plurality of voxels. Figure 40 FIG. shows Figure 39 an example of transforming the volume shown in FIG. into an octree structure. Here, Figure 40 among the leaf nodes shown in FIG., leaf nodes 1, 2, and 3 respectively represent Figure 39 the voxels VXL1, VXL2, and VXL3 shown in FIG., and represent VXLs (hereinafter referred to as valid VXLs) including point groups.
[0529] An octree is represented by a binary sequence of 0s and 1s, for example. For example, when a node or a valid VXL is set to the value 1 and the rest are set to the value 0, the binary sequence shown is assigned to each node and leaf node. Then, the binary sequence is scanned in a breadth-first or depth-first scan order. For example, when scanned in breadth-first order, the binary sequence shown in A is obtained. When scanned in depth-first order, the binary sequence shown in B is obtained. The binary sequence obtained by this scan is encoded by entropy coding, thereby reducing the amount of information. Figure 40 Next, the depth information in the octree representation will be described. The depth in the octree representation is used for controlling up to which granularity the point cloud information contained in the volume is maintained. If the depth is set large, the point cloud information can be reproduced at a finer level, but the amount of data for representing the nodes and leaf nodes will increase. On the contrary, if the depth is set small, although the amount of data can be reduced, point cloud information at multiple different positions and with different colors will be regarded as being at the same position and having the same color, so the information originally possessed by the point cloud information will be lost. Figure 41 For example, Figure 41 shows an example of representing the octree with a depth of 2 shown in
[0530] as an octree with a depth of 1.
[0531] The octree shown in Figure 42 has less data volume than the octree shown in Figure 40 . That is, Figure 42 the octree shown in Figure 40 has fewer bits after binary serialization than the octree shown in Figure 42 . Here, Figure 42 the leaf node 1 and leaf node 2 shown in Figure 40 become represented by the leaf node 1 shown in Figure 41 . That is, the information that the leaf node 1 and leaf node 2 shown in Figure 40 are at different positions is lost.
[0532] Figure 43 shows the volume corresponding to the octree shown in Figure 42 . Figure 39 The VXL1 and VXL2 shown in Figure 43 correspond to the VXL12 shown in Figure 39 . In this case, the three-dimensional data encoding device 1300 generates Figure 43The color information of VXL12 shown. For example, the three-dimensional data encoding device 1300 calculates the average value, median value, weighted average value, etc. of the color information of VXL1 and VXL2 as the color information of VXL12. In this way, the three-dimensional data encoding device 1300 can control the reduction of the data amount by changing the depth of the octree.
[0533] The three-dimensional data encoding device 1300 can also set the depth information of the octree using any one of the world space unit, space unit, and volume unit. And at this time, the three-dimensional data encoding device 1300 can also attach the depth information to the header information of the world space, the header information of the space, or the header information of the volume. Also, the same value can be used as the depth information in all world spaces, spaces, and volumes at different times. In this case, the three-dimensional data encoding device 1300 can also attach the depth information to the header information for managing the world space for all times.
[0534] In the case where the voxel contains color information, the transformation unit 1303 applies a frequency transformation such as an orthogonal transformation to the prediction residual of the color information of the voxel in the volume. For example, the transformation unit 1303 scans the prediction residual in a certain scan order to create a one-dimensional arrangement. After that, the transformation unit 1303 transforms the created one-dimensional arrangement into the frequency domain by applying a one-dimensional orthogonal transformation. Accordingly, when the values of the prediction residuals in the volume are close, the values of the frequency components in the low-frequency band become larger, and the values of the frequency components in the high-frequency band become smaller. Therefore, the quantization unit 1304 can more effectively reduce the encoding amount.
[0535] Also, the transformation unit 1303 can use an orthogonal transformation of two dimensions or more instead of using a one-dimensional orthogonal transformation. For example, the transformation unit 1303 maps the prediction residual into a two-dimensional arrangement in a certain scan order and applies a two-dimensional orthogonal transformation to the obtained two-dimensional arrangement. Also, the transformation unit 1303 can select the orthogonal transformation method to be used from multiple orthogonal transformation methods. In this case, the three-dimensional data encoding device 1300 attaches information indicating which orthogonal transformation method is used to the bitstream. And it can be that the transformation unit 1303 selects the orthogonal transformation method to be used from multiple orthogonal transformation methods with different dimensions. In this case, the three-dimensional data encoding device 1300 attaches information indicating which dimension of the orthogonal transformation method is used to the bitstream.
[0536] For example, the transformation unit 1303 aligns the scan order of the prediction residual with the scan order (such as breadth-first or depth-first) in the octree within the volume. Accordingly, since there is no need to attach information indicating the scan order of the prediction residual to the bitstream, the overhead can be reduced. Also, the transformation unit 1303 may apply a scan order different from the scan order of the octree. In this case, the three-dimensional data encoding device 1300 attaches information indicating the scan order of the prediction residual to the bitstream. Accordingly, the three-dimensional data encoding device 1300 can efficiently encode the prediction residual. Also, it may be that the three-dimensional data encoding device 1300 attaches information (such as a flag) indicating whether the scan order of the octree is applied to the bitstream, and in the case where the scan order of the octree is not applied, attaches information indicating the scan order of the prediction residual to the bitstream.
[0537] The transformation unit 1303 can transform not only the prediction residual of the color information but also other attribute information possessed by the voxels. For example, it may be that the transformation unit 1303 transforms and encodes information such as reflectance obtained when acquiring point clouds through LiDAR or the like.
[0538] When the space does not have attribute information such as color information, the transformation unit 1303 can skip the processing. Also, the three-dimensional data encoding device 1300 can attach information (a flag) indicating whether to skip the processing of the transformation unit 1303 to the bitstream.
[0539] The quantization unit 1304 quantizes the frequency components of the prediction residual generated in the transformation unit 1303 using quantization control parameters to generate quantization coefficients. Thereby, the amount of information is reduced. The generated quantization coefficients are output to the entropy encoding unit 1313. The quantization unit 1304 can control the quantization control parameters in world space units, space units, or volume units. At this time, the three-dimensional data encoding device 1300 attaches the quantization control parameters to respective header information and the like. Also, the quantization unit 1304 can change the weights for quantization control according to each frequency component of the prediction residual. For example, the quantization unit 1304 can perform fine quantization on low-frequency components and rough quantization on high-frequency components. In this case, the three-dimensional data encoding device 1300 can attach parameters representing the weights of the respective frequency components to the header.
[0540] When the space does not have attribute information such as color information, the quantization unit 1304 can skip the processing. Also, the three-dimensional data encoding device 1300 can attach information (a flag) indicating whether to skip the processing of the quantization unit 1304 to the bitstream.
[0541] The inverse quantization unit 1305 performs inverse quantization on the quantization coefficients generated by the quantization unit 1304 by using quantization control parameters, and thereby generates inverse quantization coefficients of the prediction residual, and outputs the generated inverse quantization coefficients to the inverse transform unit 1306.
[0542] The inverse transform unit 1306 applies an inverse transform to the inverse quantization coefficients generated in the inverse quantization unit 1305, thereby generating a prediction residual after the inverse transform is applied. Since the prediction residual after the inverse transform is applied is the prediction residual generated after quantization, it may not be exactly the same as the prediction residual output by the transform unit 1303.
[0543] The addition unit 1307 adds the prediction volume generated by the inverse transform unit 1306 after the inverse transform is applied and the prediction volume used in the generation of the prediction residual before quantization and generated by intra-frame prediction or inter-frame prediction described later, to generate a reconstructed volume. The reconstructed volume is stored in the reference volume memory 1308 or the reference space memory 1310.
[0544] The intra-frame prediction unit 1309 generates a prediction volume of the volume to be encoded by using the attribute information of adjacent volumes stored in the reference volume memory 1308. The attribute information includes voxel color information or reflectivity. The intra-frame prediction unit 1309 generates a predicted value of the color information or reflectivity of the volume to be encoded.
[0545] Figure 44 is a diagram for explaining the operation of the intra-frame prediction unit 1309. For example, Figure 44 As shown, the intra-frame prediction unit 1309 generates a prediction volume of the volume to be encoded (volume idx = 3) based on an adjacent volume (volume idx = 0). Here, the volume idx is identifier information attached to volumes in the space, and different values are assigned to each volume. The order of assignment of the volume idx may be the same as the encoding order or different from the encoding order. For example, as Figure 44 the predicted value of the color information of the volume to be encoded shown, the intra-frame prediction unit 1309 uses the average value of the color information of the voxels included in the adjacent volume with volume idx = 0 as the adjacent volume. In this case, by subtracting the predicted value of the color information from the color information of each voxel included in the volume to be encoded, a prediction residual is generated. The processing after the transform unit 1303 is performed on the prediction residual. And, in this case, the three-dimensional data encoding device 1300 attaches adjacent volume information and prediction mode information to the bitstream. Here, the adjacent volume information is information showing the adjacent volume used in the prediction, for example, showing the volume idx of the adjacent volume used in the prediction. And, the prediction mode information shows the mode used in the generation of the prediction volume. The mode is, for example, an average value mode in which a predicted value is generated based on the average value of the voxels in the adjacent volume, or a median value mode in which a predicted value is generated based on the median value of the voxels in the adjacent volume, etc.
[0546] The intra prediction unit 1309 may also generate a prediction volume based on a plurality of adjacent volumes. For example, in the Figure 44 configuration shown, the intra prediction unit 1309 generates a prediction volume 0 based on the volume with volume idx = 0, and generates a prediction volume 1 based on the volume with volume idx = 1. Then, the intra prediction unit 1309 generates the average of the prediction volume 0 and the prediction volume 1 as the final prediction volume. In this case, the three-dimensional data encoding device 1300 may also attach the plurality of volume idxs of the plurality of volumes used in the generation of the prediction volume to the bitstream.
[0547] Figure 45 The inter prediction process according to the present embodiment is shown in terms of a mode. The inter prediction unit 1311 performs encoding (inter prediction) on the space (SPC) at a certain time T_Cur using the encoded spaces at different times T_LX. In this case, the inter prediction unit 1311 applies rotation and translation processing to the encoded spaces at different times T_LX to perform the encoding process.
[0548] Furthermore, the three-dimensional data encoding device 1300 attaches RT information related to the rotation and translation processing of the spaces at different times T_LX to the bitstream. The different times T_LX are, for example, the time T_L0 before the certain time T_Cur. At this time, the three-dimensional data encoding device 1300 may also attach the RT information RT_L0 related to the rotation and translation processing of the space at the time T_L0 to the bitstream.
[0549] Alternatively, the different times T_LX are, for example, the time T_L1 after the certain time T_Cur. At this time, the three-dimensional data encoding device 1300 may attach the RT information RT_L1 related to the rotation and translation processing of the space at the time T_L1 to the bitstream.
[0550] Alternatively, the inter prediction unit 1311 performs encoding (dual prediction) by referring to the spaces at both different times T_L0 and time T_L1. In this case, the three-dimensional data encoding device 1300 may attach both the RT information RT_L0 and RT_L1 related to the rotation and translation applied to the spaces respectively to the bitstream.
[0551] In addition, although T_L0 is set as the time before T_Cur and T_L1 is set as the time after T_Cur above, it is not limited thereto. For example, both T_L0 and T_L1 may be the times before T_Cur. Or, both T_L0 and T_L1 may be the times after T_Cur.
[0552] Alternatively, when the three-dimensional data encoding device 1300 performs encoding with reference to spaces at multiple different times, RT information related to the rotation and translation applicable to each space is attached to the bitstream. For example, the three-dimensional data encoding device 1300 manages the multiple encoded spaces to be referred to through two reference lists (L0 list and L1 list). When the first reference space in the L0 list is set as L0R0, the second reference space in the L0 list is set as L0R1, the first reference space in the L1 list is set as L1R0, and the second reference space in the L1 list is set as L1R1, the three-dimensional data encoding device 1300 attaches the RT information RT_L0R0 of L0R0, the RT information RT_L0R1 of L0R1, the RT information RT_L1R0 of L1R0, and the RT information RT_L1R1 of L1R1 to the bitstream. For example, the three-dimensional data encoding device 1300 attaches this RT information to the header of the bitstream or the like.
[0553] Alternatively, when the three-dimensional data encoding device 1300 performs encoding with reference to reference spaces at multiple different times, it determines whether rotation and translation are applied for each reference space. At this time, the three-dimensional data encoding device 1300 can attach information (such as RT application flag) indicating whether rotation and translation are applied for each reference space to the header information of the bitstream or the like. For example, the three-dimensional data encoding device 1300 calculates the RT information and the ICP error value for each reference space to be referred to according to the encoding target space using the ICP (Interactive Closest Point) algorithm. When the ICP error value is equal to or less than a predetermined fixed value, the three-dimensional data encoding device 1300 determines that rotation and translation are not required and sets the RT application flag to OFF (invalid). In addition, when the ICP error value is greater than the above-mentioned fixed value, the three-dimensional data encoding device 1300 sets the RT application flag to ON (valid) and attaches the RT information to the bitstream.
[0554] Figure 46 A syntax example of attaching the RT information and the RT application flag to the header is shown. In addition, the number of bits allocated to each syntax can be determined according to the range that the syntax can take. For example, when the number of reference spaces included in the reference list L0 is 8, 3 bits can be allocated to MaxRefSpc_l0. The number of allocated bits can be changed according to the values that each syntax can take, or the number of allocated bits can be fixed regardless of the values that can be taken. When the number of allocated bits is fixed, the three-dimensional data encoding device 1300 can attach this fixed number of bits to other header information.
[0555] Here, Figure 46The shown MaxRefSpc_l0 shows the number of reference spaces included in the reference list L0. RT_flag_l0[i] is the RT application flag for the reference space i in the reference list L0. When RT_flag_l0[i] is 1, rotation and translation are applied to the reference space i. When RT_flag_l0[i] is 0, rotation and translation are not applied to the reference space i.
[0556] R_l0[i] and T_l0[i] are the RT information of the reference space i in the reference list L0. R_l0[i] is the rotation information of the reference space i in the reference list L0. The rotation information shows the content of the applied rotation process, such as a rotation matrix or a quaternion, etc. T_l0[i] is the translation information of the reference space i in the reference list L0. The translation information shows the content of the applied translation process, such as a translation vector, etc.
[0557] MaxRefSpc_l1 shows the number of reference spaces included in the reference list L1. RT_flag_l1[i] is the RT application flag for the reference space i in the reference list L1. When RT_flag_l1[i] is 1, rotation and translation are applied to the reference space i. When RT_flag_l1[i] is 0, rotation and translation are not applied to the reference space i.
[0558] R_l1[i] and T_l1[i] are the RT information of the reference space i in the reference list L1. R_l1[i] is the rotation information of the reference space i in the reference list L1. The rotation information shows the content of the applied rotation process, such as a rotation matrix or a quaternion, etc. T_l1[i] is the translation information of the reference space i in the reference list L1. The translation information shows the content of the applied translation process, such as a translation vector, etc.
[0559] The inter-frame prediction unit 1311 generates a predicted volume of the coding target volume by using the information of the encoded reference spaces stored in the reference space memory 1310. As described above, before generating the predicted volume of the coding target volume, the inter-frame prediction unit 1311 uses the ICP (Interactive Closest Point) algorithm in the coding target space and the reference space to find the RT information in order to make the positional relationship between the coding target space and the whole reference space closer. Then, the inter-frame prediction unit 1311 applies rotation and translation processing to the reference space by using the obtained RT information, thereby obtaining the reference space B. After that, the inter-frame prediction unit 1311 generates a predicted volume of the coding target volume in the coding target space by using the information in the reference space B. Here, the three-dimensional data coding device 1300 attaches the RT information used to obtain the reference space B to the header information of the coding target space, etc.
[0560] In this way, the inter-frame prediction unit 1311 applies rotation and translation processing to the reference space, so that after making the positional relationship between the coding target space and the overall reference space closer, the prediction volume is generated using the information of the reference space. In this way, the accuracy of the prediction volume can be improved. Moreover, since the prediction residual can be suppressed, the coding amount can be reduced. In addition, although an example of performing ICP using the coding target space and the reference space is shown here, it is not limited thereto. For example, in order to reduce the processing amount, the inter-frame prediction unit 1311 may also perform ICP using at least one of the coding target space with the voxel or point cloud number extracted and the reference space with the voxel or point cloud number extracted, so as to obtain the RT information.
[0561] Moreover, when the ICP error value obtained from the result of ICP is smaller than a preset first threshold value, that is, for example, when the positional relationship between the coding target space and the reference space is close, the inter-frame prediction unit 1311 may determine that rotation and translation processing are not required and not perform rotation and translation. In this case, the three-dimensional data coding device 1300 may not attach the RT information to the bitstream, thereby being able to suppress the overhead.
[0562] Moreover, when the ICP error value is larger than a preset second threshold value, the inter-frame prediction unit 1311 determines that the shape change in space is large, and intra-frame prediction can be applied to all volumes of the coding target space. Hereinafter, the space to which intra-frame prediction is applied is referred to as the intra-frame space. Moreover, the second threshold value is a value larger than the above-mentioned first threshold value. And it is not limited to ICP. As long as it is a method for obtaining the RT information from two voxel sets or two point cloud sets, any method can be applied.
[0563] Moreover, when attribute information such as shape or color is included in the three-dimensional data, as the prediction volume of the coding target volume in the coding target space, the inter-frame prediction unit 1311 searches, for example, for the volume in the reference space that is closest to the shape or color attribute information of the coding target volume. And this reference space is, for example, the reference space after the above-mentioned rotation and translation processing. The inter-frame prediction unit 1311 generates a prediction volume based on the volume (reference volume) obtained through the search. Figure 47 It is a diagram for explaining the generation operation of the prediction volume. The inter-frame prediction unit 1311 is for Figure 47In the case of encoding the shown encoded object volume (volume idx = 0) using inter-frame prediction, while sequentially scanning the reference volumes in the reference space, the volume with the smallest difference (i.e., prediction residual) between the encoded object volume and the reference volume is searched for. The inter-frame prediction unit 1311 selects the volume with the smallest prediction residual as the prediction volume. The prediction residual between the encoded object volume and the prediction volume is encoded by the processing after the transform unit 1303. Here, the prediction residual refers to the difference between the attribute information of the encoded object volume and the attribute information of the prediction volume. And the three-dimensional data encoding device 1300 attaches the volume idx of the reference volume in the reference space used as the prediction volume to the head of the bitstream, etc.
[0564] In Figure 47 In the shown example, the reference volume with volume idx = 4 in the reference space L0R0 is selected as the prediction volume of the encoded object volume. Then, the prediction residual between the encoded object volume and the reference volume and the reference volume idx = 4 are encoded and attached to the bitstream.
[0565] In addition, although an example of generating a prediction volume of attribute information has been described here, the same processing can be performed for the prediction volume of position information.
[0566] The prediction control unit 1312 controls which of intra-frame prediction and inter-frame prediction is used to encode the encoded object volume. Here, the mode including intra-frame prediction and inter-frame prediction is called the prediction mode. For example, the prediction control unit 1312 calculates the prediction residual in the case where the encoded object volume is predicted by intra-frame prediction and the prediction residual in the case where it is predicted by inter-frame prediction as evaluation values, and selects the prediction mode with the smaller evaluation value. Alternatively, the prediction control unit 1312 may apply orthogonal transformation, quantization, and entropy coding to the prediction residual of intra-frame prediction and the prediction residual of inter-frame prediction respectively to calculate the actual coding amount, and use the calculated coding amount as the evaluation value to select the prediction mode. And additional overhead information (such as reference volume idx information) other than the prediction residual can also be added to the evaluation value. And the prediction control unit 1312 may usually select intra-frame prediction even when the encoded object space is predetermined to be encoded in the intra-frame space.
[0567] The entropy coding unit 1313 generates an encoded signal (encoded bitstream) by performing variable-length coding on the input from the quantization unit 1304, i.e., the quantization coefficients. Specifically, the entropy coding unit 1313, for example, binarizes the quantization coefficients and performs arithmetic coding on the obtained binary signal.
[0568] Next, a three-dimensional data decoding device that decodes the encoded signal generated by the three-dimensional data encoding device 1300 will be described. Figure 48It is a block diagram of the three-dimensional data decoding device 1400 according to this embodiment. The three-dimensional data decoding device 1400 includes: an entropy decoding unit 1401, an inverse quantization unit 1402, an inverse transformation unit 1403, an addition unit 1404, a reference volume memory 1405, an intra prediction unit 1406, a reference space memory 1407, an inter prediction unit 1408, and a prediction control unit 1409.
[0569] The entropy decoding unit 1401 performs variable-length decoding on the encoded signal (encoded bitstream). For example, the entropy decoding unit 1401 performs arithmetic decoding on the encoded signal to generate a binary signal, and generates quantization coefficients based on the generated binary signal.
[0570] The inverse quantization unit 1402 performs inverse quantization on the quantization coefficients input from the entropy decoding unit 1401 using quantization parameters attached to the bitstream or the like, thereby generating inverse quantization coefficients.
[0571] The inverse transformation unit 1403 performs inverse transformation on the inverse quantization coefficients input from the inverse quantization unit 1402, thereby generating a prediction residual. For example, the inverse transformation unit 1403 performs inverse orthogonal transformation on the inverse quantization coefficients according to the information attached to the bitstream, thereby generating a prediction residual.
[0572] The addition unit 1404 adds the prediction residual generated by the inverse transformation unit 1403 and the prediction volume generated by intra prediction or inter prediction to generate a reconstructed volume. This reconstructed volume is output as decoded three-dimensional data and is stored in the reference volume memory 1405 or the reference space memory 1407.
[0573] The intra prediction unit 1406 generates a prediction volume by intra prediction using the reference volume in the reference volume memory 1405 and the information attached to the bitstream. Specifically, the intra prediction unit 1406 obtains prediction mode information and adjacent volume information (such as volume idx) attached to the bitstream, and uses the adjacent volume indicated by the adjacent volume information to generate a prediction volume in the mode indicated by the prediction mode information. In addition, the details of these processes are the same as those of the intra prediction unit 1309 described above except for using the information attached to the bitstream.
[0574] The inter-frame prediction unit 1408 generates a prediction volume through inter-frame prediction by using the reference spaces in the reference space memory 1407 and the information appended to the bitstream. Specifically, the inter-frame prediction unit 1408 applies rotation and translation processing to the reference spaces by using the RT information of each reference space appended to the bitstream, and generates a prediction volume by using the reference spaces after the application. In addition, when the RT application flag for each reference space exists in the bitstream, the inter-frame prediction unit 1408 applies rotation and translation processing to the reference spaces according to the RT application flag. In addition, regarding the details of the above processing, except for using the information appended to the bitstream, it is the same as the processing of the above inter-frame prediction unit 1311.
[0575] Regarding whether to decode the volume to be decoded by intra-frame prediction or inter-frame prediction, it will be controlled by the prediction control unit 1409. For example, the prediction control unit 1409 selects intra-frame prediction or inter-frame prediction according to the information appended to the bitstream and indicating the prediction mode to be used. In addition, the prediction control unit 1409 may generally select intra-frame prediction when it is predetermined that the space to be decoded is decoded by the intra-frame space.
[0576] A modification example of the present embodiment will be described below. In the present embodiment, although rotation and translation are applied in units of space as an example, rotation and translation may also be applied in finer units. For example, the three-dimensional data encoding device 1300 may divide the space into sub-spaces and apply rotation and translation in units of sub-spaces. In this case, the three-dimensional data encoding device 1300 generates RT information for each sub-space and appends the generated RT information to the head of the bitstream or the like. And, the three-dimensional data encoding device 1300 may apply rotation and translation in units of volume units as the encoding unit. In this case, the three-dimensional data encoding device 1300 generates RT information in units of encoding volumes and appends the generated RT information to the head of the bitstream or the like. Moreover, the above can be combined. That is, the three-dimensional data encoding device 1300 may apply rotation and translation in a large unit first, and then apply rotation and translation in a finer unit. For example, it may be that the three-dimensional data encoding device 1300 applies rotation and translation in units of space, and applies different rotations and translations to each of the multiple volumes included in the obtained space.
[0577] Also, although in this embodiment, rotation and translation are applied to the reference space as an example, it is not limited thereto. For example, it may be that the three-dimensional data encoding device 1300 applies a scaling process to change the size of the three-dimensional data. Also, the three-dimensional data encoding device 1300 may apply any one or two of rotation, translation, and scaling. Also, as described above, when processing is applied in multiple stages with different units, the types of processing applied in each unit may be different. For example, it may be that rotation and translation are applied in the space unit, and translation is applied in the volume unit.
[0578] In addition, regarding these modification examples, the same can be applied to the three-dimensional data decoding device 1400.
[0579] As described above, the three-dimensional data encoding device 1300 according to this embodiment performs the following processing. Figure 48 It is a flowchart of the inter-frame prediction process performed by the three-dimensional data encoding device 1300.
[0580] First, the three-dimensional data encoding device 1300 uses the position information of the three-dimensional points included in the target three-dimensional data (e.g., the encoding target space) and the reference three-dimensional data (e.g., the reference space) at different times to generate prediction position information (e.g., the prediction volume) (S1301). Specifically, the three-dimensional data encoding device 1300 generates prediction position information by applying rotation and translation processing to the position information of the three-dimensional points included in the reference three-dimensional data.
[0581] In addition, the three-dimensional data encoding device 1300 performs rotation and translation processing in the first unit (e.g., space), and generates prediction position information in the second unit (e.g., volume) that is finer than the first unit. For example, it may be that the three-dimensional data encoding device 1300 searches for the volume in the multiple volumes included in the reference space after rotation and translation processing, where the difference between the encoded object volume included in the encoded object space and the position information is the smallest, and uses the obtained volume as the prediction volume. In addition, the three-dimensional data encoding device 1300 may also perform rotation and translation processing and generation of prediction position information in the same unit.
[0582] And it may be that the three-dimensional data encoding device 1300 applies the first rotation and translation processing to the position information of the three-dimensional points included in the reference three-dimensional data in the first unit (e.g., space), and applies the second rotation and translation processing to the position information of the three-dimensional points obtained by the first rotation and translation processing in the second unit (e.g., volume) that is finer than the first unit, thereby generating prediction position information.
[0583] Here, the position information of the three-dimensional points and the prediction position information are as Figure 41As shown, it is represented in an octree structure. For example, the position information and predicted position information of three-dimensional points are represented in a width-first scan order in terms of depth and width in the octree structure. Alternatively, the position information and predicted position information of three-dimensional points are represented in a depth-first scan order in terms of depth and width in the octree structure.
[0584] And, as Figure 46 shown, the three-dimensional data encoding device 1300 encodes an RT application flag indicating whether rotation and translation processing are applied to the position information of the three-dimensional points included in the reference three-dimensional data. That is, the three-dimensional data encoding device 1300 generates an encoded signal (encoded bitstream) including the RT application flag. And, the three-dimensional data encoding device 1300 encodes RT information indicating the content of the rotation and translation processing. That is, the three-dimensional data encoding device 1300 generates an encoded signal (encoded bitstream) including the RT information. Additionally, it may be that the three-dimensional data encoding device 1300 encodes the RT information when the RT application flag indicates that rotation and translation processing are applied, and does not encode the RT information when the RT application flag indicates that rotation and translation processing are not applied.
[0585] And, the three-dimensional data includes, for example, the position information of three-dimensional points and the attribute information (such as color information) of each three-dimensional point. The three-dimensional data encoding device 1300 generates predicted attribute information (S1302) using the attribute information of the three-dimensional points included in the reference three-dimensional data.
[0586] Next, the three-dimensional data encoding device 1300 encodes the position information of the three-dimensional points included in the target three-dimensional data using the predicted position information. For example, as Figure 38 shown, the three-dimensional data encoding device 1300 calculates the difference between the position information of the three-dimensional points included in the target three-dimensional data and the predicted position information, that is, the differential position information (S1303).
[0587] And, the three-dimensional data encoding device 1300 encodes the attribute information of the three-dimensional points included in the target three-dimensional data using the predicted attribute information. For example, the three-dimensional data encoding device 1300 calculates the difference between the attribute information of the three-dimensional points included in the target three-dimensional data and the predicted attribute information, that is, the differential attribute information (S1304). Next, the three-dimensional data encoding device 1300 performs transformation and quantization on the calculated differential attribute information (S1305).
[0588] Finally, the three-dimensional data encoding device 1300 encodes the differential position information and the quantized differential attribute information (for example, entropy encoding) (S1306). That is, the three-dimensional data encoding device 1300 generates an encoded signal (encoded bitstream) including the differential position information and the differential attribute information.
[0589] In addition, when the attribute information is not included in the three-dimensional data, the three-dimensional data encoding apparatus 1300 may also not perform steps S1302, S1304, and S1305. Moreover, the three-dimensional data encoding apparatus 1300 may also perform only one of the encoding of the position information of the three-dimensional points and the encoding of the attribute information of the three-dimensional points.
[0590] Moreover, Figure 49 The order of the processes shown is merely an example and is not limited thereto. For example, since the processes for the position information (S1301, S1303) and the processes for the attribute information (S1302, S1304, S1305) are independent of each other, they can be executed in any order, or a part of them can be processed in parallel.
[0591] As described above, in the present embodiment, the three-dimensional data encoding apparatus 1300 uses the position information of the three-dimensional points included in the target three-dimensional data and the reference three-dimensional data at different times to generate predicted position information, and encodes the difference between the position information of the three-dimensional points included in the target three-dimensional data and the predicted position information, that is, the differential position information. Accordingly, since the data amount of the encoded signal can be reduced, the encoding efficiency can be improved.
[0592] Moreover, in the present embodiment, the three-dimensional data encoding apparatus 1300 uses the attribute information of the three-dimensional points included in the reference three-dimensional data to generate predicted attribute information, and encodes the difference between the attribute information of the three-dimensional points included in the target three-dimensional data and the predicted attribute information, that is, the differential attribute information. Accordingly, since the data amount of the encoded signal can be reduced, the encoding efficiency can be improved.
[0593] For example, the three-dimensional data encoding apparatus 1300 includes a processor and a memory, and the processor uses the memory to perform the above processes.
[0594] Figure 48 is a flowchart of the inter-frame prediction process performed by the three-dimensional data decoding apparatus 1400.
[0595] First, the three-dimensional data decoding apparatus 1400 decodes (for example, entropy decodes) the differential position information and the differential attribute information according to the encoded signal (encoded bitstream) (S1401).
[0596] Further, the three-dimensional data decoding device 1400 decodes an RT application flag indicating whether rotation and translation processing are applicable to the position information of the three-dimensional points included in the reference three-dimensional data, based on the encoded signal. Further, the three-dimensional data decoding device 1400 decodes RT information indicating the details of the rotation and translation processing. In addition, the three-dimensional data decoding device 1400 decodes the RT information when the RT application flag indicates that rotation and translation processing are applicable, and does not decode the RT information when the RT application flag indicates that rotation and translation processing are not applicable.
[0597] Next, the three-dimensional data decoding device 1400 performs inverse quantization and inverse transformation on the decoded differential attribute information (S1402).
[0598] Next, the three-dimensional data decoding device 1400 generates predicted position information (e.g., predicted volume) by using the position information of the three-dimensional points included in the object three-dimensional data (e.g., decoding target space) and the reference three-dimensional data (e.g., reference space) at different times (S1403). Specifically, the three-dimensional data decoding device 1400 generates the predicted position information by applying rotation and translation processing to the position information of the three-dimensional points included in the reference three-dimensional data.
[0599] More specifically, when the RT application flag indicates that rotation and translation processing are applicable, the three-dimensional data decoding device 1400 applies rotation and translation processing to the position information of the three-dimensional points included in the reference three-dimensional data indicated by the RT information. When the RT application flag indicates that rotation and translation processing are not applicable, the three-dimensional data decoding device 1400 does not apply rotation and translation processing to the position information of the three-dimensional points included in the reference three-dimensional data.
[0600] In addition, the three-dimensional data decoding device 1400 may perform rotation and translation processing in a first unit (e.g., space), and may generate predicted position information in a second unit (e.g., volume) that is finer than the first unit. In addition, the three-dimensional data decoding device 1400 may also perform rotation and translation processing and generation of predicted position information in the same unit.
[0601] It may be that the three-dimensional data decoding device 1400 applies first rotation and translation processing to the position information of the three-dimensional points included in the reference three-dimensional data in a first unit (e.g., space), and applies second rotation and translation processing to the position information of the three-dimensional points obtained by the first rotation and translation processing in a second unit (e.g., volume) that is finer than the first unit, thereby generating predicted position information.
[0602] Here, the position information of the three-dimensional points and the predicted position information are, for example Figure 41As shown, it is represented in an octree structure. For example, the position information and predicted position information of three-dimensional points are represented in the scanning order with priority given to the width among the depth and width in the octree structure. Alternatively, the position information and predicted position information of three-dimensional points are represented in the scanning order with priority given to the depth among the depth and width in the octree structure.
[0603] The three-dimensional data decoding device 1400 generates predicted attribute information by using the attribute information of the three-dimensional points included in the reference three-dimensional data (S1404).
[0604] Next, the three-dimensional data decoding device 1400 decodes the encoded position information included in the encoded signal by using the predicted position information, thereby restoring the position information of the three-dimensional points included in the target three-dimensional data. Here, the encoded position information is, for example, differential position information, and the three-dimensional data decoding device 1400 restores the position information of the three-dimensional points included in the target three-dimensional data by adding the differential position information and the predicted position information (S1405).
[0605] Furthermore, the three-dimensional data decoding device 1400 decodes the encoded attribute information included in the encoded signal by using the predicted attribute information, thereby restoring the attribute information of the three-dimensional points included in the target three-dimensional data. Here, the encoded attribute information is, for example, differential attribute information, and the three-dimensional data decoding device 1400 restores the attribute information of the three-dimensional points included in the target three-dimensional data by adding the differential attribute information and the predicted attribute information (S1406).
[0606] Alternatively, in the case where the attribute information is not included in the three-dimensional data, the three-dimensional data decoding device 1400 may not execute steps S1402, S1404, and S1406. Also, the three-dimensional data decoding device 1400 may perform only one of the decoding of the position information of the three-dimensional points and the decoding of the attribute information of the three-dimensional points.
[0607] And Figure 50 The order of the processes shown is an example and is not limited thereto. For example, since the processes for the position information (S1403, S1405) and the processes for the attribute information (S1402, S1404, S1406) are independent of each other, they can be performed in any order, and a part of them can also be processed in parallel.
[0608] (Embodiment 8)
[0609] In this embodiment, a method for representing three-dimensional points (point cloud) in the encoding of three-dimensional data will be described.
[0610] Figure 51 is a block diagram showing the configuration of the three-dimensional data distribution system according to this embodiment.Figure 51 The distribution system shown includes a server 1501 and a plurality of clients 1502.
[0611] The server 1501 includes a storage unit 1511 and a control unit 1512. The storage unit 1511 stores the encoded three-dimensional data, i.e., the encoded three-dimensional map 1513.
[0612] Figure 52 A configuration example of the bitstream of the encoded three-dimensional map 1513 is shown. The three-dimensional map is divided into a plurality of sub-maps, and each sub-map is encoded. A random access header (RA) including sub-coordinate information is attached to each sub-map. The sub-coordinate information is used to improve the encoding efficiency of the sub-map. The sub-coordinate information shows the sub-coordinates of the sub-map. The sub-coordinates are the coordinates of the sub-map with respect to a reference coordinate. In addition, the three-dimensional map including a plurality of sub-maps is called the entire map. And, in the entire map, the coordinate that serves as a reference (e.g., the origin) is called the reference coordinate. That is, the sub-coordinates are the coordinates of the sub-map in the coordinate system of the entire map. In other words, the sub-coordinates show the deviation between the coordinate system of the entire map and the coordinate system of the sub-map. And, the coordinates in the coordinate system of the entire map with respect to the reference coordinate are called the global coordinates. The coordinates in the coordinate system of the sub-map with respect to the sub-coordinates are called the differential coordinates.
[0613] The client 1502 sends a message to the server 1501. The message includes the location information of the client 1502. The control unit 1512 included in the server 1501 obtains the bitstream of the sub-map at the location closest to the location of the client 1502 based on the location information included in the received message. The bitstream of the sub-map includes sub-coordinate information and is sent to the client 1502. The decoder 1521 included in the client 1502 uses the sub-coordinate information to obtain the global coordinates of the sub-map with respect to the reference coordinate. The application program 1522 included in the client 1502 executes an application program related to its own location using the obtained global coordinates of the sub-map.
[0614] Moreover, the sub-map shows a partial area of the entire map. The sub-coordinates are the coordinates of the position where the sub-map is located in the reference coordinate space of the entire map. For example, it is assumed that there is a sub-map A of AA and a sub-map B of AB in the entire map of A. When the vehicle wants to refer to the map of AA, it starts decoding from the sub-map A, and when it wants to refer to the map of AB, it starts decoding from the sub-map B. Here, the sub-map is a random access point. Specifically, A is Osaka Prefecture, AA is Osaka City, AB is Takatsuki City, etc.
[0615] Each sub-map is sent to the client together with sub-coordinate information. The sub-coordinate information is included in the header information of each sub-map, or in the transmitted data packet, etc.
[0616] The reference coordinate, which is the coordinate serving as the basis for the sub-coordinate information of each sub-map, can also be attached to the header information of a space higher than the sub-map, such as the header information of the entire map.
[0617] A sub-map can be composed of one space (SPC). Also, a sub-map can be composed of multiple SPCs.
[0618] Also, a sub-map can include a GOS (Group of Space). Also, a sub-map can be composed of world spaces. For example, in the case where there are multiple objects in a sub-map, if the multiple objects are assigned to different SPCs, the sub-map is composed of multiple SPCs. And if the multiple objects are assigned to one SPC, the sub-map is composed of one SPC.
[0619] Next, the improvement effect on the coding efficiency in the case of adopting sub-coordinate information will be described. Figure 53 This is a diagram for explaining this effect. For example, to encode the three-dimensional point A at a position far from the reference coordinate as shown, Figure 53 more bits are required. Here, the distance between the sub-coordinate and the three-dimensional point A is shorter than the distance between the reference coordinate and the three-dimensional point A. Therefore, compared with the case of encoding the coordinates of the three-dimensional point A based on the reference coordinate, the coding efficiency can be improved in the case of encoding the coordinates of the three-dimensional point A based on the sub-coordinate. And the bitstream of the sub-map includes sub-coordinate information. By sending the bitstream of the sub-map and the reference coordinate to the decoding side (client), the entire coordinates of the sub-map can be restored on the decoding side.
[0620] Figure 54 This is a flowchart of the processing performed by the server 1501, which is the sending side of the sub-map.
[0621] First, the server 1501 receives a message (S1501) including the position information of the client 1502 from the client 1502. The control unit 1512 obtains the encoded bitstream of the sub-map based on the position information of the client from the storage unit 1511 (S1502). Then, the server 1501 sends the encoded bitstream of the sub-map and the reference coordinate to the client 1502 (S1503).
[0622] Figure 55 This is a flowchart of the processing performed by the client 1502, which is the receiving side of the sub-map.
[0623] First, the client 1502 receives the encoded bitstream of the sub-map and the reference coordinates sent from the server 1501 (S1511). Next, the client 1502 decodes the encoded bitstream to obtain the sub-map and the sub-coordinate information (S1512). Next, the client 1502 uses the reference coordinates and the sub-coordinates to restore the differential coordinates in the sub-map to the overall coordinates (S1513).
[0624] Next, a syntactic example of the information related to the sub-map will be described. In the encoding of the sub-map, the three-dimensional data encoding device calculates the differential coordinates by subtracting the sub-coordinates from the coordinates of each point cloud (three-dimensional point). Then, the three-dimensional data encoding device encodes the differential coordinates as a bitstream as the value of each point cloud. And, the encoding device encodes the sub-coordinate information indicating the sub-coordinates as the header information of the bitstream. Accordingly, the three-dimensional data decoding device can obtain the overall coordinates of each point cloud. For example, the three-dimensional data encoding device is included in the server 1501, and the three-dimensional data decoding device is included in the client 1502.
[0625] Figure 56 A syntactic example of the sub-map is shown. Figure 56 The shown NumOfPoint represents the number of point clouds included in the sub-map. sub_coordinate_x, sub_coordinate_y, and sub_coordinate_z are the sub-coordinate information. sub_coordinate_x represents the x coordinate of the sub-coordinates. sub_coordinate_y represents the y coordinate of the sub-coordinates. sub_coordinate_z represents the z coordinate of the sub-coordinates.
[0626] And, diff_x[i], diff_y[i], and diff_z[i] are the differential coordinates of the i-th point cloud in the sub-map. diff_x[i] represents the difference value between the x coordinate of the i-th point cloud in the sub-map and the x coordinate of the sub-coordinates. diff_y[i] represents the difference value between the y coordinate of the i-th point cloud in the sub-map and the y coordinate of the sub-coordinates. diff_z[i] represents the difference value between the z coordinate of the i-th point cloud in the sub-map and the z coordinate of the sub-coordinates.
[0627] The three-dimensional data decoding device decodes the point_cloud[i]_x, point_cloud[i]_y, and point_cloud[i]_z, which are the overall coordinates of the i-th point cloud, using the following formula. point_cloud[i]_x is the x coordinate of the overall coordinates of the i-th point cloud. point_cloud[i]_y is the y coordinate of the overall coordinates of the i-th point cloud. point_cloud[i]_z is the z coordinate of the overall coordinates of the i-th point cloud.
[0628] point_cloud[i]_x = sub_coordinate_x + diff_x[i]
[0629] point_cloud[i]_y = sub_coordinate_y + diff_y[i]
[0630] point_cloud[i]_z = sub_coordinate_z + diff_z[i]
[0631] Next, the applicable switching process of octree encoding will be described. When the three-dimensional data encoding device performs sub-map encoding, it either selects the octree representation to encode each point cloud (hereinafter referred to as octree encoding (octree coding)), or selects to encode the difference value from the sub-coordinates (hereinafter referred to as non-octree encoding (non-octree coding)). Figure 57 This operation is shown in the mode. For example, when the number of point clouds in the sub-map is equal to or greater than a pre-specified threshold, the three-dimensional data encoding device applies octree encoding to the sub-map. When the number of point clouds in the sub-map is smaller than the above threshold, the three-dimensional data encoding device applies non-octree encoding to the sub-map. Accordingly, the three-dimensional data encoding device appropriately selects whether to use octree encoding or non-octree encoding according to the shape and density of the object included in the sub-map, and thus the encoding efficiency can be improved.
[0632] In addition, the three-dimensional data encoding device attaches information indicating which of octree encoding and non-octree encoding is applicable to the sub-map (hereinafter referred to as octree encoding applicability information) to the head of the sub-map or the like. Accordingly, the three-dimensional data decoding device can determine whether the bitstream is a bitstream obtained by octree encoding the sub-map or a bitstream obtained by non-octree encoding the sub-map.
[0633] In addition, the three-dimensional data encoding device can calculate the encoding efficiency when octree encoding and non-octree encoding are respectively applied to the same point cloud, and apply the encoding method with higher encoding efficiency to the sub-map.
[0634] Figure 58 A syntax example of the sub-map in the case of performing such a switch is shown. Figure 58 The shown coding_type is information indicating the coding type and is the above-mentioned octree encoding applicability information. coding_type = 00 indicates that octree encoding is applied. coding_type = 01 indicates that non-octree encoding is applied. coding_type = 10 or 11 indicates that other coding methods than the above are applied, etc.
[0635] When the coding type is non - octree, the sub - map includes NumOfPoint and sub - coordinate information (sub_coordinate_x, sub_coordinate_y, and sub_coordinate_z).
[0636] When the coding type is octree, the sub - map includes octree_info. Octree_info is the information required in octree coding, for example, it includes depth information, etc.
[0637] When the coding type is non - octree, the sub - map includes differential coordinates (diff_x[i], diff_y[i], and diff_z[i]).
[0638] When the coding type is octree, the sub - map includes coding data related to octree coding, that is, octree_data.
[0639] In addition, here, although an example of using the xyz coordinate system is shown as the coordinate system of the point cloud, the polar coordinate system can also be used.
[0640] Figure 59 It is a flowchart of the 3D data coding process performed by the 3D data coding device. First, the 3D data coding device calculates the number of point clouds in the object sub - map, which is the sub - map to be processed (S1521). Then, the 3D data coding device determines whether the calculated number of point clouds is above a pre - specified threshold (S1522).
[0641] When the number of point clouds is above the threshold (Yes in S1522), the 3D data coding device applies octree coding to the object sub - map (S1523). And the 3D point data coding device attaches octree coding application information indicating that octree coding has been applied to the object sub - map to the head of the bitstream (S1525).
[0642] In addition, when the number of point clouds is below the threshold (No in S1522), the 3D data coding device applies non - octree coding to the object sub - map (S1524). And the 3D point data coding device attaches octree coding application information indicating that non - octree coding has been applied to the object sub - map to the head of the bitstream (S1525).
[0643] Figure 60It is a flowchart of the 3D data decoding process performed by the 3D data decoding device. First, the 3D data decoding device decodes the octree coding applicable information from the head of the bitstream (S1531). Next, the 3D data decoding device determines whether the coding type applied to the target sub-map is octree coding based on the decoded octree coding applicable information (S1532).
[0644] When the coding type shown in the octree coding applicable information is octree coding (Yes in S1532), the 3D data decoding device decodes the target sub-map using octree decoding (S1533). Additionally, when the coding type shown in the octree coding applicable information is non-octree coding (No in S1532), the 3D data decoding device decodes the target sub-map using non-octree decoding (S1534).
[0645] A modification example of the present embodiment will be described below. Figures 61 to 63 The operation of a modification example of the coding type switching process is shown in the mode.
[0646] As Figure 61 shown, the 3D data encoding device can select whether to apply octree coding or non-octree coding for each space. In this case, the 3D data encoding device attaches the octree coding applicable information to the head of the space. Accordingly, the 3D data decoding device can determine whether octree coding is applicable for each space. And in this case, the 3D data encoding device sets sub-coordinates for each space and encodes the difference value obtained by subtracting the sub-coordinate value from the coordinates of each point cloud within the space.
[0647] Accordingly, since the 3D data encoding device can appropriately switch whether to apply octree coding according to the shape or the number of point clouds of the object within the space, the encoding efficiency can be improved.
[0648] And, as Figure 62 shown, the 3D data encoding device can select whether to apply octree coding or non-octree coding for each volume. In this case, the 3D data encoding device attaches the octree coding applicable information to the head of the volume. Accordingly, the 3D data decoding device can determine whether octree coding is applicable for each volume. And in this case, the 3D data encoding device sets sub-coordinates for each volume and encodes the difference value obtained by subtracting the sub-coordinate value from the coordinates of each point cloud within the volume.
[0649] Accordingly, since the 3D data encoding device can appropriately switch whether to apply octree coding according to the shape or the number of point clouds of the object within the volume, the encoding efficiency can be improved.
[0650] Also, in the above description, as a non-octree encoding, an example of encoding the difference obtained by subtracting the sub-coordinates from the coordinates of each point cloud is shown. However, it is not limited thereto, and any encoding method other than octree encoding can be used for encoding. For example Figure 63 As shown, the three-dimensional data encoding device may not use the difference from the sub-coordinates, but may use a method of encoding the values of the point clouds themselves within the sub-map, space, or volume (hereinafter referred to as original coordinate encoding) as the non-octree encoding.
[0651] In this case, the three-dimensional data encoding device stores information indicating that the original coordinate encoding is applied to the target space (sub-map, space, or volume) in the header. Accordingly, the three-dimensional data decoding device can determine whether the original coordinate encoding is applied to the target space.
[0652] Also, in the case where the original coordinate encoding is applied, the three-dimensional data encoding device can perform encoding without applying quantization and arithmetic coding to the original coordinates. Also, the three-dimensional data encoding device can encode the original coordinates with a predetermined fixed bit length. Accordingly, the three-dimensional data encoding device can generate a stream having a certain bit length at a certain timing.
[0653] Also, in the above description, although an example of encoding the difference obtained by subtracting the sub-coordinates from the coordinates of each point cloud is shown as the non-octree encoding, it is not limited thereto.
[0654] For example, the three-dimensional data encoding device can sequentially encode the difference values between the coordinates of each point cloud. Figure 64 is a diagram for explaining the operation in this case. For example, in Figure 64 the example shown, when encoding the point cloud PA, the three-dimensional data encoding device uses the sub-coordinates as the prediction coordinates and encodes the difference value between the coordinates of the point cloud PA and the prediction coordinates. Also, when encoding the point cloud PB, the three-dimensional data encoding device uses the coordinates of the point cloud PA as the prediction coordinates and encodes the difference value between the point cloud PB and the prediction coordinates. Also, when encoding the point cloud PC, the three-dimensional data encoding device uses the point cloud PB as the prediction coordinates and encodes the difference value between the point cloud PB and the prediction coordinates. In this way, the three-dimensional data encoding device can set a scanning order for a plurality of point clouds and encode the coordinates of the target point cloud to be processed and the difference value between the coordinates of the point cloud that is previous to the target point cloud in the scanning order.
[0655] Also, in the above description, although the sub-coordinates are the coordinates of the lower left front corner of the sub-map, the position of the sub-coordinates is not limited thereto. Figures 65 to 67Other examples of the positions of the sub - coordinates are shown. Regarding the set position of the sub - coordinates, they can be set to any coordinates within the object space (sub - map, space, or volume). That is, as described above, the sub - coordinates can be the coordinates of the lower - left front corner. As Figure 65 shown, the sub - coordinates can also be the coordinates of the center of the object space. As Figure 66 shown, the sub - coordinates can also be the coordinates of the upper - right rear corner of the object space. And the sub - coordinates are not limited to the coordinates of the lower - left front or upper - right rear corners of the object space, and can be the coordinates of any corner in the object space.
[0656] And the set position of the sub - coordinates can also be the same as the coordinates of a certain point cloud within the object space (sub - map, space, or volume). For example, in the Figure 67 shown example, the coordinates of the sub - coordinates are the same as the coordinates of the point cloud PD.
[0657] And although an example of switching between applying octree coding and non - octree coding is shown in this embodiment, it is not limited thereto. For example, the three - dimensional data encoding device can also switch between applying other tree structures other than the octree and non - tree structures outside of the tree structure. For example, other tree structures refer to kd - trees and the like that are divided using a plane perpendicular to one of the coordinate axes. Additionally, any method can be adopted as other tree structures.
[0658] And although an example of encoding the coordinate information possessed by the point cloud is shown in this embodiment, it is not limited thereto. The three - dimensional data encoding device can, for example, also encode color information, three - dimensional feature quantities, or feature quantities of visible light in the same way as the coordinate information. For example, the three - dimensional data encoding device can also set the average value of the color information possessed by each point cloud within the sub - map as sub - color information, and encode the difference between the color information of each point cloud and the sub - color information.
[0659] And although an example of selecting a highly efficient encoding method (octree coding or non - octree coding) according to the number of point clouds, etc. is shown in this embodiment, it is not limited thereto. For example, as a three - dimensional data encoding device on the server side, it can pre - store the bitstreams of point clouds encoded by octree coding, the bitstreams of point clouds encoded by non - octree coding, and the bitstreams of point clouds encoded by both methods, and switch the bitstreams sent to the three - dimensional data decoding device according to the communication environment or the processing ability of the three - dimensional data decoding device.
[0660] Figure 68 An example of the syntax of the volume in the case of switching the application of octree coding is shown. Figure 68 The shown syntax is the same asFigure 58 The syntax shown is basically the same, except for the information where each piece of information is in volume units. Specifically, NumOfPoint shows the number of point clouds contained in the volume. sub_coordinate_x, sub_coordinate_y, and sub_coordinate_z are the sub - coordinate information of the volume.
[0661] Moreover, diff_x[i], diff_y[i], and diff_z[i] are the differential coordinates of the i - th point cloud within the volume. diff_x[i] represents the difference value between the x - coordinate of the i - th point cloud within the volume and the x - coordinate of the sub - coordinate. diff_y[i] represents the difference value between the y - coordinate of the i - th point cloud within the volume and the y - coordinate of the sub - coordinate. diff_z[i] represents the difference value between the z - coordinate of the i - th point cloud within the volume and the z - coordinate of the sub - coordinate.
[0662] In addition, when the relative position of the volume in space can be calculated, the three - dimensional data encoding device may not include the sub - coordinate information in the header of the volume. That is, the three - dimensional data encoding device can calculate the relative position of the volume in space without including the sub - coordinate information in the header, and use the calculated position as the sub - coordinate of each volume.
[0663] As described above, the three - dimensional data encoding device according to the present embodiment determines whether to encode an object space unit (e.g., Figure 59 in S1522) among a plurality of spatial units (such as sub - maps, spaces, or volumes) included in the three - dimensional data in an octree structure. For example, when the number of three - dimensional points included in the object space unit is more than a preset threshold, the three - dimensional data encoding device determines to encode the object space unit in an octree structure. And when the number of three - dimensional points included in the object space unit is below the above - mentioned threshold, the three - dimensional data encoding device determines not to encode the object space unit in an octree structure.
[0664] When it is determined to encode the object space unit in an octree structure ( "Yes" in S1522), the three - dimensional data encoding device encodes the object space unit in an octree structure (S1523). And when it is determined not to encode the object space unit in an octree structure ( "No" in S1522), the three - dimensional data encoding device encodes the object space unit in a manner different from the octree structure (S1524). For example, as a different manner, the three - dimensional data encoding device encodes the coordinates of the three - dimensional points included in the object space unit. Specifically, as a different manner, the three - dimensional data encoding device encodes the reference coordinates of the object space unit and the difference between the coordinates of the three - dimensional points included in the object space unit.
[0665] Next, the three-dimensional data encoding device attaches information indicating whether to encode the object space unit in an octree structure to the bitstream (S1525).
[0666] Accordingly, the three-dimensional data encoding device can reduce the data amount of the encoded signal, thereby improving the encoding efficiency.
[0667] For example, the three-dimensional data encoding device includes a processor and a memory, and the processor uses the memory to perform the above processing.
[0668] In addition, the three-dimensional data decoding device according to the present embodiment decodes, from the bitstream, information indicating whether to decode the object space unit (e.g., sub-map, space, or volume) included in the three-dimensional data in an octree structure (e.g., Figure 60 S1531). When the above information indicates that the object space unit is to be decoded in an octree structure (Yes in S1532), the three-dimensional data decoding device decodes the object space unit in an octree structure (S1533).
[0669] When the above information indicates that the object space unit is not to be decoded in an octree structure (No in S1532), the three-dimensional data decoding device decodes the object space unit in a manner different from the octree structure (S1534). For example, the three-dimensional data decoding device decodes the coordinates of the three-dimensional points included in the object space unit in a different manner. Specifically, the three-dimensional data decoding device decodes the difference between the reference coordinates of the object space unit and the coordinates of the three-dimensional points included in the object space unit in a different manner.
[0670] Thereby, the three-dimensional data decoding device can reduce the data amount of the encoded signal, and thus can improve the encoding efficiency.
[0671] For example, the three-dimensional data decoding device includes a processor and a memory, and the processor uses the memory to perform the above processing.
[0672] (Embodiment 9)
[0673] In the present embodiment, an encoding method of a tree structure such as an octree structure will be described.
[0674] By identifying the important area and preferentially decoding the three-dimensional data of the important area, the efficiency can be improved.
[0675] Figure 69This is a diagram showing examples of important regions in a three-dimensional map. An important region is, for example, a region that contains three-dimensional points with large feature quantity values among the three-dimensional points in the three-dimensional map above a certain number. Or, an important region can also be, for example, a region that contains the three-dimensional points required for a client such as a vehicle to estimate its own position above a certain number. Or, an important region can also be the region of the face in a three-dimensional model of a person. In this way, an important region can be defined for each application program, and the important region can also be switched according to the application program.
[0676] In this embodiment, as a way to represent an octree structure or the like, occupancy coding and location coding are used. In addition, the bit string obtained by occupancy coding is called an occupancy code. The bit string obtained by location coding is called a location code.
[0677] Figure 70 This is a diagram showing an example of an occupancy code. Figure 70 An example of an occupancy code representing a quadtree structure. In Figure 70 each node is assigned an occupancy code. Each occupancy code indicates whether a three-dimensional point is included in the child nodes or leaf nodes of each node. For example, in the case of a quadtree, the information indicating whether each of the four child nodes or leaf nodes of each node contains a three-dimensional point is represented by a 4-bit occupancy code. In addition, in the case of an octree, the information indicating whether each of the 8 child nodes or leaf nodes of each node contains a three-dimensional point is represented by an 8-bit occupancy code. In addition, here, for the sake of simplicity of explanation, a quadtree structure is taken as an example for explanation, but it can also be similarly applied to an octree structure. For example, as shown in Figure 70 the occupancy code is a bit example in which the nodes and leaf nodes are scanned in breadth-first order as described in Figure 40 etc. In the occupancy code, since the information of multiple three-dimensional points is decoded in a fixed order, the information of an arbitrary three-dimensional point cannot be preferentially decoded. In addition, the occupancy code can also be a bit string in which the nodes and leaf nodes are scanned in depth-first order as described in Figure 40 etc.
[0678] Hereinafter, location coding will be described. By using a location code, it is possible to directly decode important parts in an octree structure. In addition, it is possible to efficiently encode important three-dimensional points in the deep layer.
[0679] Figure 71 This is a diagram for explaining location coding and is a diagram showing an example of a quadtree structure. In Figure 71[[ENDIn the example shown, the three-dimensional points A to I are represented by a quadtree structure. In addition, the three-dimensional points A and C are important three-dimensional points included in the important area.
[0680] represents a diagram showing the occupancy rate code and the position code representing the important three-dimensional points A and C in the quadtree structure shown.
[0681] In position encoding, in the tree structure, the indexes of the nodes and the index of the leaf node existing in the path to the leaf node to which the object three-dimensional point to be encoded belongs are encoded. Here, the index is a numerical value assigned to each node and leaf node. In other words, the index is an identifier for identifying multiple child nodes of the object node. As shown, in the case of a quadtree, the index represents any one of 0 to 3.
[0682] For example, in the quadtree structure shown, when the leaf node A is the object three-dimensional point, the leaf node A is represented as 0→2→1→0→1→2→1. Here, when the maximum value of each index is as shown in the right figure, it is 4 (which can be represented by 2 bits), so the number of bits required for the position code of the leaf node A is 7×2 bits = 14 bits. When the leaf node C is the encoding object, the same number of bits required is 14 bits. In addition, in the case of an octree, since the maximum value of each index is 8 (which can be represented by 3 bits), the number of bits required can be calculated by 3 bits × the depth of the leaf node. In addition, the three-dimensional data encoding device can also perform entropy encoding after binarizing each index to reduce the data amount.
[0683] In addition, as shown, in the occupancy rate code, in order to decode the leaf nodes A and C, it is necessary to decode all the nodes above them. On the other hand, in the position code, it is possible to decode only the data of the leaf nodes A and C. Thus, as shown, by using the position code, the number of bits can be reduced compared to the occupancy rate code.
[0684] In addition, as shown, by performing dictionary-based compression such as LZ77 on a part or all of the position code, the code amount can be further reduced.
[0685] Next, an example of applying position encoding to three-dimensional points (point cloud) obtained by LiDAR is described. This is a diagram showing an example of three-dimensional points obtained by LiDAR. The three-dimensional points obtained by LiDAR are sparse. That is, when representing these three-dimensional points with occupancy codes, the number of zeros increases. Additionally, high three-dimensional accuracy is required for these three-dimensional points. That is, the depth of the octree structure becomes deeper.
[0686] This is a diagram showing an example of such a sparse deep octree structure. The occupancy code of the octree structure shown is 136 bits (= 8 bits × 17 nodes). Additionally, since the depth is 6 and there are 6 three-dimensional points, the position code is 3 bits × 6 × 6 = 108 bits. That is, the position code can reduce the code amount by 20% relative to the occupancy code. Thus, by applying position encoding to the sparse deep octree structure, the code amount can be reduced.
[0687] Hereinafter, the code amounts of the occupancy code and the position code will be described. When the depth of the octree structure is 10, the maximum number of three-dimensional points is 8 10 = 1073741824. Additionally, the number of bits L o of the occupancy code of the octree structure is represented as follows.
[0688] L o = 8 + 8 2 + … + 8 10 = 127133512 bits
[0689] Therefore, the number of bits per three-dimensional point is 1.143 bits. Additionally, in the occupancy code, this number of bits does not change even when the number of three-dimensional points included in the octree structure changes.
[0690] On the other hand, in the position code, the number of bits per three-dimensional point directly affects the depth of the octree structure. Specifically, the number of bits of the position code per three-dimensional point is 3 bits × depth 10 = 30 bits.
[0691] Therefore, the number of bits L l of the position code of the octree structure is represented as follows.
[0692] L l = 30 × N
[0693] Here, N is the number of three-dimensional points included in the octree structure.
[0694] Therefore, when N < L o / 30 = 40904450.4, that is, when the number of three-dimensional points is less than 40904450, the code amount of the position code becomes less than the code amount of the occupancy code (L l < L o ).
[0695] In this way, when the number of three-dimensional points is small, the code amount of the position code is smaller than that of the occupancy rate code, and when the number of three-dimensional points is large, the code amount of the position code is larger than that of the occupancy rate code.
[0696] Therefore, the three-dimensional data encoding device can also switch which one of the position encoding and the occupancy rate encoding to use according to the number of the input three-dimensional points. In this case, the three-dimensional data encoding device can also attach information indicating which one of the position encoding and the occupancy rate encoding has been used to the header information of the bitstream or the like.
[0697] Hereinafter, a hybrid encoding combining the position encoding and the occupancy rate encoding will be described. The hybrid encoding combining the position encoding and the occupancy rate encoding is effective when encoding a dense important area. It is a diagram showing this example. In In the example shown, important three-dimensional points are densely arranged. In this case, the three-dimensional data encoding device performs position encoding on the upper layer with a shallow depth and uses occupancy rate encoding for the lower layer. Specifically, position encoding is used until the deepest common node, and occupancy rate encoding is used at positions deeper than the deepest common node. Here, the deepest common node refers to the deepest node among the nodes that are the common ancestors of multiple important three-dimensional points.
[0698] Next, a hybrid encoding that prioritizes compression efficiency will be described. The three-dimensional data encoding device can also switch the position encoding and the occupancy rate encoding according to rules predefined in the encoding of the octree.
[0699] It is a diagram showing an example of this rule. First, the three-dimensional data encoding device confirms the ratio of the nodes including three-dimensional points at each level (depth). When this ratio is higher than a predefined threshold, the three-dimensional data encoding device performs occupancy rate encoding on several nodes in the upper layer of the target level. For example, the three-dimensional data encoding device applies occupancy rate encoding to the levels from the target level to the deepest common node.
[0700] For example, in In the example shown, the ratio of the nodes including three-dimensional points in the third level is higher than the threshold. Therefore, the three-dimensional data encoding device applies occupancy rate encoding to the second and third levels from this third level to the deepest common node, and applies position encoding to the first and fourth levels other than these.
[0701] A description is given of the method for calculating the above-mentioned threshold value. In one layer of the octree structure, there is a root node and eight child nodes. Therefore, in occupancy encoding, eight bits are required to encode one layer of the octree structure. On the other hand, in position encoding, three bits are required for each child node containing a three-dimensional point. Therefore, when the number of nodes containing three-dimensional points is greater than two, occupancy encoding is more efficient than position encoding. That is, in this case, the threshold value is two.
[0702] Next, a description is given of a configuration example of a bitstream generated by the above-mentioned position encoding, occupancy encoding, or hybrid encoding.
[0703] FIG. is an example of a bitstream generated by position encoding. As shown, the bitstream generated by position encoding includes a header and a plurality of position codes. Each position code corresponds to one three-dimensional point.
[0704] With this configuration, the three-dimensional data decoding device can decode a plurality of three-dimensional points with high precision respectively. In addition, FIG. shows an example of a bitstream in the case of a quadtree structure. In the case of an octree structure, each index can take a value from 0 to 7.
[0705] In addition, the three-dimensional data encoding device may perform entropy encoding after binarizing a column (string) of indexes representing one three-dimensional point. For example, when the column of indexes is 0121, the three-dimensional data encoding device may binarize 0121 into 00011001 and perform arithmetic encoding on this bit string.
[0706] FIG. is an example of a bitstream generated by hybrid encoding in the case of including important three-dimensional points. As shown, the position codes of the upper layer, the occupancy codes of the important three-dimensional points of the lower layer, and the occupancy codes of the non-important three-dimensional points other than the important three-dimensional points of the lower layer are arranged in sequence. In addition, the position code length shown indicates the amount of code of the subsequent position code. In addition, the occupancy code amount indicates the amount of code of the subsequent occupancy code.
[0707] With this configuration, the three-dimensional data decoding device can select different decoding plans according to the application program.
[0708] In addition, the encoded data of the important three-dimensional points is stored near the beginning of the bitstream, and the encoded data of the non-important three-dimensional points not included in the important area is stored after the encoded data of the important three-dimensional points.
[0709] FIG. is shown by Diagram of the tree structure of the occupancy rate code representation of the important three-dimensional points shown. is a diagram representing the tree structure of the occupancy rate code representation of the unimportant three-dimensional points shown by . As shown by , in the occupancy rate code of the important three-dimensional points, information related to the unimportant three-dimensional points is excluded. Specifically, in nodes 0 and 3 at depth 5, there are no important three-dimensional points, so the values 0 indicating no three-dimensional points are assigned to nodes 0 and 3.
[0710] On the other hand, as shown by , in the occupancy rate code of the unimportant three-dimensional points, information related to the important three-dimensional points is excluded. Specifically, in node 1 at depth 5, there are no unimportant three-dimensional points, so the value 0 indicating no three-dimensional points is assigned to node 1.
[0711] In this way, the three-dimensional data encoding device divides the original tree structure into a first tree structure containing important three-dimensional points and a second tree structure containing unimportant three-dimensional points, and performs occupancy rate encoding on the first tree structure and the second tree structure independently. As a result, the three-dimensional data decoding device can preferentially decode important three-dimensional points.
[0712] Next, a configuration example of the bitstream generated by the efficiency-oriented hybrid encoding will be described. is a diagram showing a configuration example of the bitstream generated by the efficiency-oriented hybrid encoding. As shown by , for each subtree, the position of the subtree root node, the amount of occupancy rate code, and the occupancy rate code are sequentially arranged. The subtree position shown is the position code of the root node of the subtree.
[0713] In the above configuration, when only one of position encoding and occupancy rate encoding is applied to the octree structure, the following holds.
[0714] When the length of the position encoding of the root node of the subtree is equal to the depth of the octree structure, the subtree has no child nodes. That is, position encoding is applied to the entire tree structure.
[0715] When the root node of the subtree is equal to the root node of the octree structure, occupancy rate encoding is applied to the entire tree structure.
[0716] For example, based on the above rules, the three-dimensional data decoding device can determine whether the bitstream contains a position code or an occupancy rate encoding.
[0717] In addition, the bitstream may also include encoding mode information indicating which of position encoding, occupancy rate encoding, and hybrid encoding is used. is a diagram showing an example of the bitstream in this case. For example, as shown by As shown, 2-bit coding mode information representing the coding mode is appended to the bitstream.
[0718] In addition, (1) the "number of 3D points" in position coding represents the number of subsequent 3D points. In addition, (2) the "amount of occupancy code" in occupancy rate coding represents the amount of the subsequent occupancy rate code. In addition, (3) the "number of important subtrees" in hybrid coding (important 3D points) represents the number of subtrees containing important 3D points. Furthermore, (4) the "number of occupancy rate subtrees" in hybrid coding (efficiency emphasis) represents the number of subtrees after occupancy rate coding.
[0719] Next, a syntax example used to switch the application of occupancy rate coding and position coding will be described. It is a figure showing this syntax example.
[0720] The isleaf shown represents a flag indicating whether the object node is a leaf node. isleaf = 1 indicates that the object node is a leaf node, and isleaf = 0 indicates that the object node is not a leaf node but a node.
[0721] When the object node is a leaf node, point_flag is appended to the bitstream. point_flag is a flag indicating whether the object node (leaf node) contains 3D points. point_flag = 1 indicates that the object node contains 3D points, and point_flag = 0 indicates that the object node does not contain 3D points.
[0722] When the object node is not a leaf node, coding_type is appended to the bitstream. coding_type is coding type information indicating the applicable coding type. coding_type = 00 indicates that position coding is applicable, coding_type = 01 indicates that occupancy rate coding is applicable, and coding_type = 10 or 11 indicates that other coding methods are applicable, etc.
[0723] When the coding type is position coding, numPoint, num_idx[i], and idx[i][j] are appended to the bitstream.
[0724] numPoint represents the number of 3D points for which position coding is performed. num_idx[i] represents the number (depth) of the index from the object node to 3D point i. When all the 3D points for which position coding is performed are at the same depth, num_idx[i] are all the same value. Therefore, it can also be that, before the for statement shown (for(i = 0; i < numPoint; i++){), num_idx is defined as a common value.
[0725] idx[i][j] represents the value of the j-th index in the index from the object node to the three-dimensional point i. In the case of an octree, the number of bits of idx[i][j] is 3 bits.
[0726] In addition, as described above, an index is an identifier for identifying a plurality of child nodes of an object node. In the case of an octree, idx[i][j] represents any one of 0 to 7. In addition, in the case of an octree, there are 8 child nodes, and each child node corresponds to each of the 8 sub-blocks obtained by spatially dividing the object block corresponding to the object node into 8 parts. Therefore, idx[i][j] can also be information representing the three-dimensional position of the sub-block corresponding to the child node. For example, idx[i][j] can also be 3-bit information in total including 1 bit of information representing the position of each of x, y, and z of the sub-block.
[0727] In the case where the coding type is occupancy coding, an occupancy_code is appended to the bitstream. The occupancy_code is the occupancy rate code of the object node. In the case of an octree, the occupancy_code is an 8-bit bitstring such as the bitstring "00101000", for example.
[0728] When the value of the (i + 1)-th bit of the occupancy_code is 1, the process proceeds to the child node. That is, the child node is set as the next object node, and a bitstring is generated recursively.
[0729] In the present embodiment, an example is shown in which the end of the octree is represented by appending leaf node information (isleaf, point_flag) to the bitstream, but it is not necessarily limited to this. For example, the three-dimensional data encoding device can append the maximum depth (depth) from the start node (root node) of the occupancy rate code to the end (leaf node) where the three-dimensional point exists to the head of the start node. Then, the three-dimensional data encoding device can also bitstringize the information of the child nodes recursively while increasing the depth from the start node, and determine that it has reached the leaf node when the depth becomes the maximum depth. In addition, the three-dimensional data encoding device can append the information representing the maximum depth to the first node where the coding_type becomes occupancy coding, or can also append it to the start node (root node) of the octree.
[0730] As described above, the three-dimensional data encoding device can also append information for switching between occupancy coding and position coding to the bitstream as the header information of each node.
[0731] In addition, the 3D data encoding device can also perform entropy encoding on the coding_type, numPoint, num_idx, idx, and occupancy_code of each node generated by the above method. For example, the 3D data encoding device performs arithmetic encoding after binarizing each value.
[0732] In addition, in the above syntax, the case where a depth-first bit string of an octree structure is used as the occupancy code is illustrated, but it is not necessarily limited to this. The 3D data encoding device can also use a breadth-first bit string of the octree structure as the occupancy code. When the 3D data encoding device uses a breadth-first bit string, it can also attach information for switching between occupancy encoding and position encoding to the bitstream as the header information of each node.
[0733] In the present embodiment, the octree structure is used as an example for illustration, but it is not necessarily limited to this. The above method can also be applied to N-ary trees (N is an integer of 2 or more) such as quadtrees and hexadecimal trees, or other tree structures.
[0734] Hereinafter, an example of the encoding process flow for switching the application between occupancy encoding and position encoding will be described. It is a flowchart of the encoding process of the present embodiment.
[0735] First, the 3D data encoding device represents a plurality of 3D points included in the 3D data using an octree structure (S1601). Next, the 3D data encoding device sets the root node in the octree structure as the object node (S1602). Next, the 3D data encoding device generates a bit string of the octree structure by performing node encoding processing on the object node (S1603). Next, the 3D data encoding device generates a bitstream by performing entropy encoding on the generated bit string (S1604).
[0736] It is a flowchart of the node encoding process (S1603). First, the 3D data encoding device determines whether the object node is a leaf node (S1611). When the object node is not a leaf node (the "no" in S1611), the 3D data encoding device sets the leaf node flag (isleaf) to 0 and attaches the leaf node flag to the bit string (S1612).
[0737] Next, the 3D data encoding device determines whether the number of child nodes containing 3D points is larger than a pre-specified threshold (S1613). In addition, the 3D data encoding device can also attach the threshold to the bit string.
[0738] When the number of child nodes containing three-dimensional points is greater than a predefined threshold (Yes in S1613), the three-dimensional data encoding device sets the coding type (coding_type) to occupancy encoding and attaches this coding type to the bit string (S1614).
[0739] Next, the three-dimensional data encoding device sets occupancy encoding information and attaches this occupancy encoding information to the bit string. Specifically, the three-dimensional data encoding device generates an occupancy code for the object node and attaches this occupancy code to the bit string (S1615).
[0740] Next, the three-dimensional data encoding device sets the next object node according to the occupancy code (S1616). Specifically, the three-dimensional data encoding device sets an unprocessed child node with an occupancy code of "1" as the next object node.
[0741] Next, the three-dimensional data encoding device performs node encoding processing on the newly set object node (S1617). That is, it performs the processing shown in the figure.
[0742] If the processing of all child nodes is not completed (No in S1618), the processing after step S1616 is performed again. On the other hand, if the processing of all child nodes is completed (Yes in S1618), the three-dimensional data encoding device ends the node encoding processing.
[0743] In addition, in step S1613, when the number of child nodes containing three-dimensional points is less than or equal to the predefined threshold (No in S1613), the three-dimensional data encoding device sets the coding type to position encoding and attaches this coding type to the bit string (S1619).
[0744] Next, the three-dimensional data encoding device sets position encoding information and attaches this position encoding information to the bit string. Specifically, the three-dimensional data encoding device generates a position code and attaches this position encoding to the bit string (S1620). The position code includes numPoint, num_idx, and idx.
[0745] In addition, in step S1611, when the object node is a leaf node (Yes in S1611), the three-dimensional data encoding device sets the leaf node flag to 1 and attaches this leaf node flag to the bit string (S1621). In addition, the three-dimensional data encoding device sets information indicating whether the leaf node contains three-dimensional points, that is, the point flag (point_flag), and attaches this point flag to the bit string (S1622).
[0746] Next, an example of the decoding process for switching the application of occupancy encoding and position encoding will be described. This is a flowchart of the decoding process of this embodiment.
[0747] The 3D data decoding device generates a bit string by performing entropy decoding on the bitstream (S1631). Next, the 3D data decoding device restores the octree structure by performing node decoding processing on the obtained bit string (S1632). Next, the 3D data decoding device generates 3D points based on the restored octree structure (S1633).
[0748] This is a flowchart of the node decoding process (S1632). First, the 3D data decoding device obtains (decodes) the leaf node flag (isleaf) from the bit string (S1641). Next, the 3D data decoding device determines whether the object node is a leaf node based on the leaf node flag (S1642).
[0749] When the object node is not a leaf node (the "No" in S1642), the 3D data decoding device obtains the coding type (coding_type) from the bit string (S1643). The 3D data decoding device determines whether the coding type is occupancy coding (S1644).
[0750] When the coding type is occupancy coding (the "Yes" in S1644), the 3D data decoding device obtains occupancy coding information from the bit string. Specifically, the 3D data decoding device obtains the occupancy code from the bit string (S1645).
[0751] Next, the 3D data decoding device sets the next object node according to the occupancy coding. Specifically, the 3D data decoding device sets the unprocessed child node with the occupancy code of "1" as the next object node.
[0752] Next, the 3D data decoding device performs node decoding processing on the newly set object node (S1647). That is, perform the processing shown in shown.
[0753] When the processing of all child nodes is not completed (the "No" in S1648), the processing after step S1646 is performed again. On the other hand, when the processing of all child nodes is completed (the "Yes" in S1648), the 3D data decoding device ends the node decoding process.
[0754] In addition, when the coding type is position coding in step S1644 (the "No" in S1644), the 3D data decoding device obtains position coding information from the bit string. Specifically, the 3D data decoding device obtains the position code from the bit string (S1649). The position code includes numPoint, num_idx, and idx.
[0755] In addition, when the object node is a leaf node in step S1642 ("Yes" in S1642), the three-dimensional data decoding device obtains, from the bit string, information indicating whether the leaf node includes three-dimensional points, that is, a point flag (S1650).
[0756] In addition, in the present embodiment, an example in which the encoding type is switched for each node is shown, but it is not necessarily limited thereto. The encoding type may also be fixed in units of volume, space, or world space. In this case, the three-dimensional data encoding device may also attach the encoding type information to the header information of the volume, space, or world space.
[0757] As described above, the three-dimensional data encoding device of the present embodiment generates first information in the form of an N-ary tree structure (where N is an integer of 2 or more) representing a plurality of three-dimensional points included in the three-dimensional data in the first method (position encoding), and generates a bit stream including the first information. The first information includes three-dimensional point information (position code) corresponding to each of the plurality of three-dimensional points. Each three-dimensional point information includes an index (idx) corresponding to each of the plurality of layers in the N-ary tree structure. Each index indicates the sub-block to which the corresponding three-dimensional point belongs among the N sub-blocks belonging to the corresponding layer.
[0758] In other words, each three-dimensional point information represents a path in the N-ary tree structure up to the corresponding three-dimensional point. Each index indicates the child node included in the above path among the N child nodes belonging to the corresponding layer (node).
[0759] Thus, this three-dimensional data encoding method can generate a bit stream that can selectively decode three-dimensional points.
[0760] For example, the three-dimensional point information (position code) includes information (num_idx) indicating the number of indexes included in the three-dimensional point information. In other words, this information represents the depth (number of layers) in the N-ary tree structure up to the corresponding three-dimensional point.
[0761] For example, the first information includes information (numPoint) indicating the number of three-dimensional point information included in the first information. In other words, this information represents the number of three-dimensional points included in the N-ary tree structure.
[0762] For example, N is 8 and the index is 3 bits.
[0763] For example, a three-dimensional data encoding device has: a first encoding mode that generates first information; and a second encoding mode that generates second information (occupancy code) representing an N-ary tree structure in a second manner (occupancy encoding), and generates a bitstream including the second information. The second information includes a 1-bit information corresponding to each of a plurality of sub-blocks belonging to a plurality of layers in the N-ary tree structure, and represents whether there are three-dimensional points in the corresponding sub-block.
[0764] For example, when the number of a plurality of three-dimensional points is equal to or less than a preset threshold, the three-dimensional data encoding device uses the first encoding mode, and when the number of the plurality of three-dimensional points is greater than the threshold, the three-dimensional data encoding device uses the second encoding mode. Thereby, the three-dimensional data encoding device can reduce the code amount of the bitstream.
[0765] For example, the first information and the second information include information (encoding mode information) indicating whether the information represents the N-ary tree structure in the first manner or in the second manner.
[0766] For example, as shown, etc., the three-dimensional data encoding device uses the first encoding mode in a part of the N-ary tree structure and uses the second encoding mode in another part of the N-ary tree structure.
[0767] For example, the three-dimensional data encoding device includes a processor and a memory, and the processor uses the memory to perform the above processing.
[0768] In addition, the three-dimensional data decoding device of the present embodiment obtains, from a bitstream, first information (position code) representing an N-ary tree structure (where N is an integer greater than or equal to 2) of a plurality of three-dimensional points included in three-dimensional data in a first manner (position encoding). The first information includes three-dimensional point information (position code) corresponding to each of the plurality of three-dimensional points. Each three-dimensional point information includes an index (idx) corresponding to each of a plurality of layers in the N-ary tree structure. Each index represents a sub-block to which the corresponding three-dimensional point belongs among N sub-blocks belonging to the corresponding layer.
[0769] In other words, each three-dimensional point information represents a path to the corresponding three-dimensional point in the N-ary tree structure. Each index represents a sub-node included in the above path among N sub-nodes belonging to the corresponding layer (node).
[0770] The three-dimensional data decoding device further uses the three-dimensional point information to restore the three-dimensional point corresponding to the three-dimensional point information.
[0771] Thereby, the three-dimensional data decoding device can selectively decode three-dimensional points from the bitstream.
[0772] For example, the three-dimensional point information (position code) includes information (num_idx) indicating the number of indices included in the three-dimensional point information. In other words, this information represents the depth (number of layers) up to the corresponding three-dimensional point in the N-ary tree structure.
[0773] For example, the first information includes information (numPoint) indicating the number of three-dimensional point information included in the first information. In other words, this information represents the number of three-dimensional points included in the N-ary tree structure.
[0774] For example, N is 8 and the index is 3 bits.
[0775] For example, the three-dimensional data decoding device further obtains, from the bitstream, second information (occupancy code) representing the N-ary tree structure in the second manner (occupancy encoding). The three-dimensional data decoding device uses the second information to reconstruct a plurality of three-dimensional points. The second information includes a plurality of 1-bit information corresponding to each of the plurality of sub-blocks belonging to the plurality of layers in the N-ary tree structure and indicating whether there is a three-dimensional point in the corresponding sub-block.
[0776] For example, the first information and the second information include information (encoding mode information) indicating whether the information represents the N-ary tree structure in the first manner or the second manner.
[0777] For example, as shown, etc., a part of the N-ary tree structure is represented in the first manner, and another part of the N-ary tree structure is represented in the second manner.
[0778] For example, the three-dimensional data decoding device includes a processor and a memory, and the processor uses the memory to perform the above processing.
[0779] (Embodiment 10)
[0780] In this embodiment, another example of an encoding method for a tree structure such as an octree structure will be described. is a diagram showing an example of the tree structure related to this embodiment. In addition, shows an example of a quadtree structure.
[0781] A leaf node containing a three-dimensional point is called an effective leaf node, and a leaf node not containing a three-dimensional point is called an invalid leaf node. A branch where the number of effective leaf nodes is equal to or greater than a threshold is called a dense branch. A branch where the number of effective leaf nodes is less than the threshold is called a sparse branch.
[0782] The three-dimensional data encoding device calculates, in a certain layer of the tree structure, the number of three-dimensional points included in each branch (i.e., the number of effective leaf nodes). An example is shown where the threshold is 5. In this example, there are two branches in layer 1. Since the left branch contains 7 three-dimensional points, the left branch is determined to be a dense branch. Since the right branch contains two three-dimensional points, the right branch is determined to be a sparse branch.
[0783] For example, it is a diagram showing the number of valid leaf nodes (3D points) of each branch in layer 5. The horizontal axis represents the identification number (index) of the branches in layer 5. As Figure 89 shown, in a specific branch, there are significantly more three-dimensional points than in other branches. In such a dense branch, occupancy encoding is more effective than in a sparse branch.
[0784] Hereinafter, the application methods of occupancy encoding and position encoding will be described. Figure 90 It is a diagram showing the relationship between the number of three-dimensional points (number of valid leaf nodes) contained in each branch in layer 5 and the encoding method applied. As Figure 90 shown, the three-dimensional data encoding device applies occupancy encoding to dense branches and position encoding to sparse branches. Thereby, the encoding efficiency can be improved.
[0785] Figure 91 It is a diagram showing an example of a dense branch region in LiDAR data. As Figure 91 shown, depending on the region, the density of the three-dimensional points calculated based on the number of three-dimensional points contained in each branch is different.
[0786] In addition, by separating dense three-dimensional points (branches) from sparse three-dimensional points (branches), there are the following advantages. The closer to the LiDAR sensor, the higher the density of the three-dimensional points. Therefore, by separating the branches according to density, zoning in the distance direction can be performed. Such zoning is effective in specific applications. In addition, for sparse branches, it is effective to use methods other than occupancy encoding.
[0787] In the present embodiment, the three-dimensional data encoding device separates the input three-dimensional point cloud into two or more sub-three-dimensional point clouds, and applies different encoding methods to each sub-three-dimensional point cloud.
[0788] For example, the three-dimensional data encoding device separates the input three-dimensional point cloud into a sub-three-dimensional point cloud A (dense three-dimensional point cloud: dense cloud) including dense branches and a sub-three-dimensional point cloud B (sparse three-dimensional point cloud: sparse cloud) including sparse branches. Figure 92 It is a diagram showing an example of a sub-three-dimensional point cloud A (dense three-dimensional point cloud) including dense branches separated from the Figure 88 tree structure shown.Figure 93 This is a diagram showing an example of a sub-three-dimensional point group B (sparse three-dimensional point group) including sparse branches separated from the tree structure shown in Figure 88 Figure 88 .
[0789] Next, the three-dimensional data encoding device encodes the sub-three-dimensional point group A by occupancy encoding and encodes the sub-three-dimensional point group B by position encoding.
[0790] In addition, an example is shown here in which different encoding methods (occupancy encoding and position encoding) are applied as different encoding methods. However, for example, the three-dimensional data encoding device can also use the same encoding method for the sub-three-dimensional point group A and the sub-three-dimensional point group B, and make the parameters used in the encoding different between the sub-three-dimensional point group A and the sub-three-dimensional point group B.
[0791] Hereinafter, the process of three-dimensional data encoding processing performed by the three-dimensional data encoding device will be described. Figure 94 This is a flowchart of three-dimensional data encoding processing performed by the three-dimensional data encoding device according to the present embodiment.
[0792] First, the three-dimensional data encoding device separates the input three-dimensional point group into sub-three-dimensional point groups (S1701). The three-dimensional data encoding device can perform this separation either automatically or based on information input by the user. For example, the user can also specify the range of the sub-three-dimensional point group, etc. In addition, as an example of automatic execution, for example, when the input data is LiDAR data, the three-dimensional data encoding device uses the distance information to each point group for separation. Specifically, the three-dimensional data encoding device separates the point group within a certain range from the measurement location and the point group outside the range. In addition, the three-dimensional data encoding device can also use the information of important areas and unimportant areas for separation.
[0793] Next, the three-dimensional data encoding device encodes the sub-three-dimensional point group A by method A to generate encoded data (encoded bitstream) (S1702). In addition, the three-dimensional data encoding device encodes the sub-three-dimensional point group B by method B to generate encoded data (S1703). In addition, the three-dimensional data encoding device can also encode the sub-three-dimensional point group B by method A. In this case, the three-dimensional data encoding device encodes the sub-three-dimensional point group B using encoding parameters different from those used in the encoding of the sub-three-dimensional point group A. For example, this parameter can also be a quantization parameter. For example, the three-dimensional data encoding device encodes the sub-three-dimensional point group B using a quantization parameter larger than the quantization parameter used in the encoding of the sub-three-dimensional point group A. In this case, the three-dimensional data encoding device can also attach information indicating the quantization parameter used in the encoding of the sub-three-dimensional point group to the header of the encoded data of each sub-three-dimensional point group.
[0794] Next, the three-dimensional data encoding device generates a bitstream (S1704) by combining the encoded data obtained in step S1702 and the encoded data obtained in step S1703.
[0795] In addition, the three-dimensional data encoding device may also encode, as header information of the bitstream, information for decoding each sub-three-dimensional point group. For example, the three-dimensional data encoding device may encode the following information.
[0796] The header information may also include information indicating the number of encoded sub-three-dimensional points. In this example, this information indicates 2.
[0797] The header information may also include information indicating the number of three-dimensional points included in each sub-three-dimensional point group and the encoding method. In this example, this information indicates the number of three-dimensional points included in sub-three-dimensional point group A, the encoding method (technique A) applied to sub-three-dimensional point group A, the number of three-dimensional points included in sub-three-dimensional point group B, and the encoding method (technique B) applied to sub-three-dimensional point group B.
[0798] The header information may also include information for identifying the start position or end position of the encoded data of each sub-three-dimensional point group.
[0799] In addition, the three-dimensional data encoding device may also encode sub-three-dimensional point group A and sub-three-dimensional point group B in parallel. Or, the three-dimensional data encoding device may also encode sub-three-dimensional point group A and sub-three-dimensional point group B sequentially.
[0800] In addition, the method of separating into sub-three-dimensional point groups is not limited to the above. For example, the three-dimensional data encoding device changes the separation method, encodes using each of a plurality of separation methods, and calculates the encoding efficiency of the encoded data obtained using each separation method. And, the three-dimensional data encoding device selects the separation method with the highest encoding efficiency. For example, the three-dimensional data encoding device may also separate the three-dimensional point group in each of a plurality of layers, calculate the encoding efficiency in each case, select the separation method (i.e., the layer for separation) with the highest encoding efficiency, and generate sub-three-dimensional point groups using the selected separation method for encoding.
[0801] In addition, when combining the encoded data, the three-dimensional data encoding device may arrange the encoding information of a more important sub-three-dimensional point group closer to the beginning of the bitstream. Thus, the three-dimensional data decoding device can obtain important information only by decoding the beginning bitstream, so it can obtain important information earlier.
[0802] Next, the process of three-dimensional data decoding processing performed by the three-dimensional data decoding device will be described. Figure 95 It is a flowchart of three-dimensional data decoding processing performed by the three-dimensional data decoding device according to this embodiment.
[0803] First, a three-dimensional data decoding device, for example, obtains a bitstream generated by the above-described three-dimensional data encoding device. Next, the three-dimensional data decoding device separates the encoded data of the sub-three-dimensional point group A and the encoded data of the sub-three-dimensional point group B from the obtained bitstream (S1711). Specifically, the three-dimensional data decoding device decodes the information used to decode each sub-three-dimensional point group from the header information of the bitstream, and uses this information to separate the encoded data of each sub-three-dimensional point group.
[0804] Next, the three-dimensional data decoding device obtains the sub-three-dimensional point group A by decoding the encoded data of the sub-three-dimensional point group A using method A (S1712). In addition, the three-dimensional data decoding device obtains the sub-three-dimensional point group B by decoding the encoded data of the sub-three-dimensional point group B using method B (S1713). Next, the three-dimensional data decoding device combines the sub-three-dimensional point group A and the sub-three-dimensional point group B (S1714).
[0805] In addition, the three-dimensional data decoding device may decode the sub-three-dimensional point group A and the sub-three-dimensional point group B in parallel. Or, the three-dimensional data decoding device may decode the sub-three-dimensional point group A and the sub-three-dimensional point group B sequentially.
[0806] Furthermore, the three-dimensional data decoding device may also decode the required sub-three-dimensional point group. For example, the three-dimensional data decoding device may decode the sub-three-dimensional point group A without decoding the sub-three-dimensional point group B. For example, when the sub-three-dimensional point group A is a three-dimensional point group included in an important area of LiDAR data, the three-dimensional data decoding device decodes the three-dimensional point group of this important area. The three-dimensional point group of this important area is used for self-position estimation of a vehicle or the like.
[0807] Next, a specific example of the encoding process according to this embodiment will be described. Figure 96 It is a flowchart of the three-dimensional data encoding process performed by the three-dimensional data encoding device according to this embodiment.
[0808] First, the three-dimensional data encoding device separates the input three-dimensional points into a sparse three-dimensional point group and a dense three-dimensional point group (S1721). Specifically, the three-dimensional data encoding device counts the number of valid leaf nodes of the branches of a certain layer of the octree structure. The three-dimensional data encoding device sets each branch as a dense branch or a sparse branch according to the number of valid leaf nodes of each branch. And, the three-dimensional data encoding device generates a sub-three-dimensional point group (dense three-dimensional point group) that aggregates the dense branches and a sub-three-dimensional point group (sparse three-dimensional point group) that aggregates the sparse branches.
[0809] Next, the three-dimensional data encoding device generates encoded data by encoding the sparse three-dimensional point group (S1722). For example, the three-dimensional data encoding device encodes the sparse three-dimensional point group using position encoding.
[0810] In addition, the three-dimensional data encoding device generates encoded data by encoding a dense three-dimensional point group (S1723). For example, the three-dimensional data encoding device encodes the dense three-dimensional point group using occupancy encoding.
[0811] Next, the three-dimensional data encoding device generates a bitstream by combining the encoded data of the sparse three-dimensional point group obtained in step S1722 and the encoded data of the dense three-dimensional point group obtained in step S1723 (S1724).
[0812] In addition, the three-dimensional data encoding device may also encode information used to decode the sparse three-dimensional point group and the dense three-dimensional point group as header information of the bitstream. For example, the three-dimensional data encoding device may encode information as follows.
[0813] The header information may also include information indicating the number of encoded sub-three-dimensional point groups. In this example, this information indicates 2.
[0814] The header information may also include information indicating the number of three-dimensional points included in each sub-three-dimensional point group and the encoding method. In this example, this information indicates the number of three-dimensional points included in the sparse three-dimensional point group, the encoding method applied to the sparse three-dimensional point group (position encoding), the number of three-dimensional points included in the dense three-dimensional point group, and the encoding method applied to the dense three-dimensional point group (occupancy encoding).
[0815] The header information may also include information used to identify the start position or end position of the encoded data of each sub-three-dimensional point group. In this example, this information indicates at least one of the start position and end position of the encoded data of the sparse three-dimensional point group and the start position and end position of the encoded data of the dense three-dimensional point group.
[0816] In addition, the three-dimensional data encoding device may also encode the sparse three-dimensional point group and the dense three-dimensional point group in parallel. Alternatively, the three-dimensional data encoding device may also encode the sparse three-dimensional point group and the dense three-dimensional point group sequentially.
[0817] Next, a specific example of the three-dimensional data decoding process will be described. Figure 97 It is a flowchart of the three-dimensional data decoding process performed by the three-dimensional data decoding device according to this embodiment.
[0818] First, a three-dimensional data decoding device, for example, obtains a bitstream generated by the above-described three-dimensional data encoding device. Next, the three-dimensional data decoding device separates the obtained bitstream into encoded data of a sparse three-dimensional point cloud and encoded data of a dense three-dimensional point cloud (S1731). Specifically, the three-dimensional data decoding device decodes information used to decode each sub three-dimensional point cloud from the header information of the bitstream, and separates the encoded data of each sub three-dimensional point cloud using this information. In this example, the three-dimensional data decoding device separates the encoded data of the sparse three-dimensional point cloud and the dense three-dimensional point cloud from the bitstream using the header information.
[0819] Next, the three-dimensional data decoding device obtains a sparse three-dimensional point cloud by decoding the encoded data of the sparse three-dimensional point cloud (S1732). For example, the three-dimensional data decoding device decodes the sparse three-dimensional point cloud using position decoding used to decode the encoded data encoded by position.
[0820] In addition, the three-dimensional data decoding device obtains a dense three-dimensional point cloud by decoding the encoded data of the dense three-dimensional point cloud (S1733). For example, the three-dimensional data decoding device decodes the dense three-dimensional point cloud using occupancy decoding used to decode the encoded data encoded by occupancy rate.
[0821] Next, the three-dimensional data decoding device combines the sparse three-dimensional point cloud obtained in step S1732 with the dense three-dimensional point cloud obtained in step S1733 (S1734).
[0822] In addition, the three-dimensional data decoding device may decode the sparse three-dimensional point cloud and the dense three-dimensional point cloud in parallel. Or, the three-dimensional data decoding device may decode the sparse three-dimensional point cloud and the dense three-dimensional point cloud sequentially.
[0823] In addition, the three-dimensional data decoding device may also decode a part of the required sub three-dimensional point clouds. For example, the three-dimensional data decoding device may decode the dense three-dimensional point cloud without decoding the sparse three-dimensional data. For example, in the case where the dense three-dimensional point cloud is a three-dimensional point cloud included in an important area of LiDAR data, the three-dimensional data decoding device decodes the three-dimensional point cloud of this important area. The three-dimensional point cloud of this important area is used for self-position estimation of a vehicle or the like.
[0824] Figure 98 This is a flowchart of the encoding process according to this embodiment. First, the three-dimensional data encoding device generates a sparse three-dimensional point cloud and a dense three-dimensional point cloud by separating the input three-dimensional point cloud into a sparse three-dimensional point cloud and a dense three-dimensional point cloud (S1741).
[0825] Next, the three-dimensional data encoding device generates encoded data by encoding a dense three-dimensional point group (S1742). In addition, the three-dimensional data encoding device generates encoded data by encoding a sparse three-dimensional point group (S1743). Finally, the three-dimensional data encoding device generates a bitstream by combining the encoded data of the sparse three-dimensional point group obtained in step S1742 and the encoded data of the dense three-dimensional point group obtained in step S1743 (S1744).
[0826] Figure 99 is a flowchart of the decoding process according to this embodiment. First, the three-dimensional data decoding device extracts the encoded data of the dense three-dimensional point group and the encoded data of the sparse three-dimensional point group from the bitstream (S1751). Next, the three-dimensional data decoding device obtains the decoded data of the dense three-dimensional point group by decoding the encoded data of the dense three-dimensional point group (S1752). In addition, the three-dimensional data decoding device obtains the decoded data of the sparse three-dimensional point group by decoding the encoded data of the sparse three-dimensional point group (S1753). Next, the three-dimensional data decoding device generates a three-dimensional point group by combining the decoded data of the dense three-dimensional point group obtained in step S1752 and the decoded data of the sparse three-dimensional point group obtained in step S1753 (S1754).
[0827] In addition, the three-dimensional data encoding device and the three-dimensional data decoding device can encode or decode either the dense three-dimensional point group or the sparse three-dimensional point group first. In addition, the encoding process or the decoding process can be performed in parallel by multiple processors or the like.
[0828] In addition, the three-dimensional data encoding device can also encode one of the dense three-dimensional point group and the sparse three-dimensional point group. For example, when important information is included in the dense three-dimensional point group, the three-dimensional data encoding device extracts the dense three-dimensional point group and the sparse three-dimensional point group from the input three-dimensional point group, encodes the dense three-dimensional point group, and does not encode the sparse three-dimensional point group. Thus, the three-dimensional data encoding device can suppress the amount of bits while attaching important information to the stream. For example, between a server and a client, when the client requests the server to send three-dimensional point group information around the client, the server encodes the important information around the client as a dense three-dimensional point group and sends it to the client. Thus, the server can send the information requested by the client while suppressing the network bandwidth.
[0829] In addition, the three-dimensional data decoding device can also decode one of the dense three-dimensional point group and the sparse three-dimensional point group. For example, when important information is included in the dense three-dimensional point group, the three-dimensional data decoding device decodes the dense three-dimensional point group and does not decode the sparse three-dimensional point group. Thus, the three-dimensional data decoding device can obtain the required information while suppressing the processing load of the decoding process.
[0830] Figure 100 Yes Figure 98 Figure 98 is a flowchart of the separation process (S1741) of three-dimensional points shown. First, the three-dimensional data encoding device sets layer L and threshold TH (S1761). Additionally, the three-dimensional data encoding device may also attach information representing the set layer L and threshold TH to the bitstream. That is, the three-dimensional data encoding device may also generate a bitstream containing information representing the set layer L and threshold TH.
[0831] Next, the three-dimensional data encoding device moves the position of the object to be processed from the root node of the octree to the beginning branch of layer L. That is, the three-dimensional data encoding device selects the beginning branch of layer L as the branch of the object to be processed (S1762).
[0832] Next, the three-dimensional data encoding device counts the number of valid leaf nodes of the branch of the object to be processed in layer L (S1763). When the number of valid leaf nodes of the branch of the object to be processed is more than the threshold TH (Yes in S1764), the three-dimensional data encoding device registers the branch of the object to be processed as a dense branch in the dense three-dimensional point group (S1765). On the other hand, when the number of valid leaf nodes of the branch of the object to be processed is less than or equal to the threshold TH (No in S1764), the three-dimensional data encoding device registers the branch of the object to be processed as a sparse branch in the sparse three-dimensional point group (S1766).
[0833] When the processing of all branches in layer L is not completed (No in S1767), the three-dimensional data encoding device moves the position of the object to be processed to the next branch in layer L. That is, the three-dimensional data encoding device selects the next branch in layer L as the branch of the object to be processed (S1768). And the three-dimensional data encoding device performs the processing after step S1763 on the selected next branch of the object to be processed.
[0834] The above processing is repeated until the processing of all branches in layer L is completed (Yes in S1767).
[0835] In addition, in the above description, the layer L and the threshold TH are preset, but they are not necessarily limited thereto. For example, the three-dimensional data encoding device sets multiple styles of the group of the layer L and the threshold TH, generates a dense three-dimensional point cloud and a sparse three-dimensional point cloud using each group, and encodes them separately. The three-dimensional data encoding device finally encodes the dense three-dimensional point cloud and the sparse three-dimensional point cloud using the group of the layer L and the threshold TH with the highest encoding efficiency among the multiple groups. Thereby, the encoding efficiency can be improved. In addition, the three-dimensional data encoding device can also calculate the layer L and the threshold TH, for example. For example, the three-dimensional data encoding device can set a value that is half of the maximum value of the layers included in the tree structure for the layer L. In addition, the three-dimensional data encoding device can set a value that is half of the total number of the multiple three-dimensional points included in the tree structure as the threshold TH.
[0836] In addition, in the above description, an example in which the input three-dimensional point cloud is classified into two types, a dense three-dimensional point cloud and a sparse three-dimensional point cloud, is described, but the three-dimensional data encoding device can also classify the inpu...
Claims
1. A three-dimensional data encoding method, wherein, a bit string of an N-ary tree structure representing a plurality of three-dimensional points included in three-dimensional data is entropy encoded using the first to Nth contexts, N being an integer of 2 or more, the bit string includes information of the first to Nth bits for each node in the N-ary tree structure, the information of the first to Nth bits includes N pieces of 1-bit information indicating whether three-dimensional points exist in respective N child nodes corresponding to the node, in the entropy encoding, the information of the first to Nth bits is entropy encoded using the first to Nth contexts respectively corresponding to the first to Nth bits.
2. The three-dimensional data encoding method according to claim 1, wherein, in the entropy encoding, based on whether three-dimensional points exist in respective a plurality of adjacent nodes adjacent to an object node, the first to Nth contexts respectively corresponding to the first to Nth bits of the object node are determined.
3. The three-dimensional data encoding method according to claim 2, wherein, in the entropy encoding, based on a configuration pattern indicating a configuration position of adjacent nodes having three-dimensional points among the plurality of adjacent nodes, the first to Nth contexts respectively corresponding to the first to Nth bits of the object node are determined.
4. The three-dimensional data encoding method according to claim 3, wherein, for configuration patterns in the configuration pattern that become the same by rotation, the same first to Nth contexts are determined.
5. The three-dimensional data encoding method according to claim 1, wherein, in the entropy encoding, based on a layer to which an object node belongs, the first to Nth contexts respectively corresponding to the first to Nth bits of the object node are determined.
6. The three-dimensional data encoding method according to claim 1, wherein, in the entropy encoding, based on a normal vector of an object node, the first to Nth contexts respectively corresponding to the first to Nth bits of the object node are determined.
7. The three-dimensional data encoding method according to claim 1, wherein, in the entropy encoding, the first to Nth contexts respectively corresponding to the first to Nth bits of an object node are determined.
8. The three-dimensional data encoding method according to any one of claims 1 to 7, wherein, N is 8.
9. A three-dimensional data decoding method, wherein, a bit string of an N-ary tree structure representing a plurality of three-dimensional points included in three-dimensional data is entropy decoded using the first to Nth contexts, N being an integer of 2 or more, the bit string includes information of the first to Nth bits for each node in the N-ary tree structure, the information of the first to Nth bits includes N pieces of 1-bit information indicating whether three-dimensional points exist in respective N child nodes corresponding to the node, in the entropy decoding, the information of the first to Nth bits is entropy decoded using the first to Nth contexts respectively corresponding to the first to Nth bits.
10. The three-dimensional data decoding method according to claim 9, wherein, In the entropy decoding, based on whether there are three-dimensional points in each of a plurality of adjacent nodes adjacent to an object node, the first to Nth contexts corresponding to the first to Nth bits of the object node are determined.
11. The three-dimensional data decoding method according to claim 10, wherein, in the entropy decoding, based on a configuration pattern representing the configuration positions of the adjacent nodes having three-dimensional points among the plurality of adjacent nodes, the first to Nth contexts corresponding to the first to Nth bits of the object node are determined.
12. The three-dimensional data decoding method according to claim 11, wherein, for the configuration patterns in the configuration pattern that become the same through rotation, the same first to Nth contexts are determined.
13. The three-dimensional data decoding method according to claim 9, wherein, in the entropy decoding, based on the layer to which the object node belongs, the first to Nth contexts corresponding to the first to Nth bits of the object node are determined.
14. The three-dimensional data decoding method according to claim 9, wherein, in the entropy decoding, based on the normal vector of the object node, the first to Nth contexts corresponding to the first to Nth bits of the object node are determined.
15. The three-dimensional data decoding method according to claim 9, wherein, in the entropy decoding, the first to Nth contexts corresponding to the first to Nth bits of the object node are determined.
16. The three-dimensional data decoding method according to any one of claims 9 to 15, wherein, N is 8.
17. A three-dimensional data encoding apparatus, wherein, comprises: a processor; and a memory, the processor uses the memory, entropy-encodes a bit string of an N-ary tree structure representing a plurality of three-dimensional points included in three-dimensional data using first to Nth contexts, N being an integer of 2 or more, the bit string includes information of first to Nth bits for each node in the N-ary tree structure, the information of the first to Nth bits includes N pieces of 1-bit information indicating whether there are three-dimensional points in respective N child nodes of the corresponding node, in the entropy encoding, the information of the first to Nth bits is entropy-encoded using the first to Nth contexts corresponding to the first to Nth bits respectively.
18. A three-dimensional data decoding apparatus, wherein, comprises: a processor; and a memory, the processor uses the memory, entropy-decodes a bit string of an N-ary tree structure representing a plurality of three-dimensional points included in three-dimensional data using first to Nth contexts, N being an integer of 2 or more, the bit string includes information of first to Nth bits for each node in the N-ary tree structure, the information of the first to Nth bits includes N pieces of 1-bit information indicating whether there are three-dimensional points in respective N child nodes of the corresponding node, in the entropy decoding, the information of the first to Nth bits is entropy-decoded using the first to Nth contexts corresponding to the first to Nth bits respectively.
Citation Information
Patent Citations
Map display device
WO2014020663A1