Three-dimensional data encoding method, three-dimensional data decoding method, three-dimensional data encoding device, and three-dimensional data decoding device
By referencing only the shared information of spatially adjacent parent nodes in 3D data encoding, the problems of low encoding efficiency and large processing volume are solved, achieving more efficient 3D data compression.
Patent Information
- Application Number
- CN201980023416.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2018-02-08
- Filing Date
- 2019-02-07
- Publication Date
- 2026-01-02
- Estimated Expiration
- 2039-02-07
AI Technical Summary
Existing technologies have low encoding efficiency and large processing volume in 3D data encoding, making it difficult to effectively compress point cloud data.
The encoding is performed using an N-ary tree structure for multiple 3D points in the 3D data. Only the information of the same parent node that is spatially adjacent to the object node is referenced, and reference to the information of different parent nodes is prohibited.
It improves encoding efficiency, reduces processing volume, and achieves more efficient 3D data compression.
Smart Images

Figure CN112041889B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present disclosure relates to a three-dimensional data encoding method, a three-dimensional data decoding method, a three-dimensional data encoding apparatus, and a three-dimensional data decoding apparatus. BACKGROUND
[0002] In the future, devices or services that flexibly use three-dimensional data will become widespread in the fields of computer vision map information monitoring infrastructure inspection or video distribution for autonomous work of an automobile or a robot. Three-dimensional data is obtained by various methods such as a distance sensor, a stereo camera, or a combination of a plurality of monocular cameras.
[0003] One of the methods of expressing three-dimensional data is a method called point cloud, which expresses the shape of a three-dimensional structure by a point group in a three-dimensional space (for example, refer to Non-Patent Literature 1). The position and color of the point group are stored in the point cloud. Although it is expected that the point cloud will become mainstream as a method of expressing three-dimensional data, the data amount of the point group is very large. Therefore, in accumulation or transmission of three-dimensional data, as with two-dimensional moving images (as one example, MPEG-4 AVC or HEVC standardized by MPEG, or the like), compression of the data amount by encoding is required.
[0004] Furthermore, regarding compression of the point cloud, a part is supported by a publicly known library such as Point Cloud Library (PCL) that performs processing of associating point clouds.
[0005] Furthermore, there is known a technique of searching for facilities in the vicinity of a vehicle using three-dimensional map data and displaying the searched facilities (for example, refer to Patent Literature 1).
[0006] PRIOR ART DOCUMENTS
[0007] PATENT LITERATURE
[0008] Patent Literature 1 International Publication No. 2014 / 020663 SUMMARY
[0009] PROBLEMS TO BE SOLVED BY THE INVENTION
[0010] It is desirable to improve the encoding efficiency and reduce the processing amount in encoding of three-dimensional data.
[0011] An object of the present disclosure is to provide a three-dimensional data encoding method, a three-dimensional data decoding method, a three-dimensional data encoding apparatus, or a three-dimensional data decoding apparatus that can improve the encoding efficiency and reduce the processing amount.
[0012] MEANS FOR SOLVING THE PROBLEMS
[0013] The three-dimensional data encoding method of one aspect of the present disclosure encodes information of an object node included in an N (N is an integer of 2 or more) -ary tree structure of a plurality of three-dimensional points included in three-dimensional data, in which, in the encoding, reference to information of a first node that is the same as a parent node of the object node among a plurality of neighboring nodes that are spatially adjacent to the object node is permitted, and reference to information of a second node that is different from the parent node of the object node is prohibited.
[0014] The three-dimensional data decoding method of one aspect of the present disclosure decodes information of an object node included in an N (N is an integer of 2 or more) -ary tree structure of a plurality of three-dimensional points included in three-dimensional data, in which, in the decoding, reference to information of a first node that is the same as a parent node of the object node among a plurality of neighboring nodes that are spatially adjacent to the object node is permitted, and reference to information of a second node that is different from the parent node of the object node is prohibited.
[0015] Effects of Invention
[0016] The present disclosure can provide a three-dimensional data encoding method, a three-dimensional data decoding method, a three-dimensional data encoding device, or a three-dimensional data decoding device that can improve encoding efficiency and reduce processing amount. BRIEF DESCRIPTION OF DRAWINGS
[0017] Figure 1 A configuration of encoding three-dimensional data according to Embodiment 1 is shown.
[0018] Figure 2 An example of a prediction structure between SPCs belonging to the lowest layer of GOS according to Embodiment 1 is shown.
[0019] Figure 3 An example of a prediction structure between layers according to Embodiment 1 is shown.
[0020] Figure 4 An example of an encoding order of GOS according to Embodiment 1 is shown.
[0021] Figure 5 An example of an encoding order of GOS according to Embodiment 1 is shown.
[0022] Figure 6 A block diagram of a three-dimensional data encoding device according to Embodiment 1 is shown.
[0023] Figure 7 A flowchart of an encoding process according to Embodiment 1 is shown.
[0024] Figure 8 A block diagram of a three-dimensional data decoding device according to Embodiment 1 is shown.
[0025] Figure 9 is a flowchart of the decoding process according to Embodiment 1.
[0026] Figure 10 An example of meta information according to Embodiment 1 is shown.
[0027] Figure 11 A configuration example of SWLD according to Embodiment 2 is shown.
[0028] Figure 12 An operation example of the server and the client according to Embodiment 2 is shown.
[0029] Figure 13 An operation example of the server and the client according to Embodiment 2 is shown.
[0030] Figure 14 An operation example of the server and the client according to Embodiment 2 is shown.
[0031] Figure 15 An operation example of the server and the client according to Embodiment 2 is shown.
[0032] Figure 16 is a block diagram of a three-dimensional data encoding apparatus according to Embodiment 2.
[0033] Figure 17 is a flowchart of the encoding process according to Embodiment 2.
[0034] Figure 18 is a block diagram of a three-dimensional data decoding apparatus according to Embodiment 2.
[0035] Figure 19 is a flowchart of the decoding process according to Embodiment 2.
[0036] Figure 20 A configuration example of WLD according to Embodiment 2 is shown.
[0037] Figure 21 An example of octree structure of WLD according to Embodiment 2 is shown.
[0038] Figure 22 A configuration example of SWLD according to Embodiment 2 is shown.
[0039] Figure 23 An example of octree structure of SWLD according to Embodiment 2 is shown.
[0040] Figure 24 is a block diagram of a three-dimensional data producing apparatus according to Embodiment 3.
[0041] Figure 25is a block diagram of a three-dimensional data transmitting apparatus according to Embodiment 3.
[0042] Figure 26 is a block diagram of a three-dimensional information processing apparatus according to Embodiment 4.
[0043] Figure 27 is a block diagram of a three-dimensional data creating apparatus according to Embodiment 5.
[0044] Figure 28 shows a configuration of a system according to Embodiment 6.
[0045] Figure 29 is a block diagram of a client apparatus according to Embodiment 6.
[0046] Figure 30 is a block diagram of a server according to Embodiment 6.
[0047] Figure 31 is a flowchart of a three-dimensional data creating process by a client apparatus according to Embodiment 6.
[0048] Figure 32 is a flowchart of a sensor information transmitting process by a client apparatus according to Embodiment 6.
[0049] Figure 33 is a flowchart of a three-dimensional data creating process by a server according to Embodiment 6.
[0050] Figure 34 is a flowchart of a three-dimensional map transmitting process by a server according to Embodiment 6.
[0051] Figure 35 shows a configuration of a modification example of a system according to Embodiment 6.
[0052] Figure 36 shows a configuration of a server and a client apparatus according to Embodiment 6.
[0053] Figure 37 is a block diagram of a three-dimensional data encoding apparatus according to Embodiment 7.
[0054] Figure 38 shows an example of a prediction residual according to Embodiment 7.
[0055] Figure 39 shows an example of a volume according to Embodiment 7.
[0056] Figure 40 shows an example of an octree representation of a volume according to Embodiment 7.
[0057] Figure 41 An example of a bit string of a volume involved in Embodiment 7 is shown.
[0058] Figure 42 An example of an octree representation of a volume involved in Embodiment 7 is shown.
[0059] Figure 43 An example of a volume involved in Embodiment 7 is shown.
[0060] Figure 44 is a diagram for explaining intra prediction processing involved in Embodiment 7.
[0061] Figure 45 is a diagram for explaining rotation and translation processing involved in Embodiment 7.
[0062] Figure 46 An example of syntax of RT applicability flag and RT information involved in Embodiment 7 is shown.
[0063] Figure 47 is a diagram for explaining inter prediction processing involved in Embodiment 7.
[0064] Figure 48 is a block diagram of a three-dimensional data decoding apparatus involved in Embodiment 7.
[0065] Figure 49 is a flowchart of three-dimensional data encoding processing performed by a three-dimensional data encoding apparatus involved in Embodiment 7.
[0066] Figure 50 is a flowchart of three-dimensional data decoding processing performed by a three-dimensional data decoding apparatus involved in Embodiment 7.
[0067] Figure 51 A configuration of a distribution system involved in Embodiment 8 is shown.
[0068] Figure 52 An example of a configuration of a bit stream of an encoded three-dimensional map involved in Embodiment 8 is shown.
[0069] Figure 53 is a diagram for explaining an improvement effect of encoding efficiency involved in Embodiment 8.
[0070] Figure 54 is a flowchart of processing performed by a server involved in Embodiment 8.
[0071] Figure 55 is a flowchart of processing performed by a client involved in Embodiment 8.
[0072] Figure 56 An example of syntax of a submap involved in Embodiment 8 is shown.
[0073] Figure 57 The operation of a modification of the switching process of the encoding type involved in Embodiment 8 is shown in a mode.
[0074] Figure 58 A syntax example of the submap involved in Embodiment 8 is shown.
[0075] Figure 59 is a flowchart of the three-dimensional data encoding process involved in Embodiment 8.
[0076] Figure 60 is a flowchart of the three-dimensional data decoding process involved in Embodiment 8.
[0077] Figure 61 The operation of a modification of the switching process of the encoding type involved in Embodiment 8 is shown in a mode.
[0078] Figure 62 The operation of a modification of the switching process of the encoding type involved in Embodiment 8 is shown in a mode.
[0079] Figure 63 The operation of a modification of the switching process of the encoding type involved in Embodiment 8 is shown in a mode.
[0080] Figure 64 The operation of a modification of the difference value calculation process involved in Embodiment 8 is shown in a mode.
[0081] Figure 65 The operation of a modification of the difference value calculation process involved in Embodiment 8 is shown in a mode.
[0082] Figure 66 The operation of a modification of the difference value calculation process involved in Embodiment 8 is shown in a mode.
[0083] Figure 67 The operation of a modification of the difference value calculation process involved in Embodiment 8 is shown in a mode.
[0084] Figure 68 A syntax example of the volume involved in Embodiment 8 is shown.
[0085] Figure 69 is a diagram showing an example of the important region involved in Embodiment 9.
[0086] Figure 70 is a diagram showing an example of the occupancy code involved in Embodiment 9.
[0087] Figure 71 is a diagram showing an example of the quadtree structure involved in Embodiment 9.
[0088] Figure 72 is a diagram showing one example of occupancy code and position code related to Embodiment 9.
[0089] Figure 73 is a diagram showing an example of three-dimensional points obtained in the LiDAR related to Embodiment 9.
[0090] Figure 74 is a diagram showing an example of octree structure related to Embodiment 9.
[0091] Figure 75 is a diagram showing an example of hybrid coding related to Embodiment 9.
[0092] Figure 76 is a diagram for explaining a switching method of position coding and occupancy coding related to Embodiment 9.
[0093] Figure 77 is a diagram showing one example of bitstream of position coding related to Embodiment 9.
[0094] Figure 78 is a diagram showing one example of bitstream of hybrid coding related to Embodiment 9.
[0095] Figure 79 is a diagram showing a tree structure of occupancy code of important three-dimensional points related to Embodiment 9.
[0096] Figure 80 is a diagram showing a tree structure of occupancy code of unimportant three-dimensional points related to Embodiment 9.
[0097] Figure 81 is a diagram showing one example of bitstream of hybrid coding related to Embodiment 9.
[0098] Figure 82 is a diagram showing one example of bitstream including coding mode information related to Embodiment 9.
[0099] Figure 83 is a diagram showing a syntax example related to Embodiment 9.
[0100] Figure 84 is a flowchart of an encoding process related to Embodiment 9.
[0101] Figure 85 is a flowchart of a node encoding process related to Embodiment 9.
[0102] Figure 86 is a flowchart of a decoding process related to Embodiment 9.
[0103] Figure 87 is a flowchart of the node decoding process according to Embodiment 9.
[0104] Figure 88 is a diagram showing an example of a tree structure according to Embodiment 10.
[0105] Figure 89 is a diagram showing an example of the number of valid leaf nodes possessed by each branch according to Embodiment 10.
[0106] Figure 90 is a diagram showing an example of the application of the encoding method according to Embodiment 10.
[0107] Figure 91 is a diagram showing an example of a dense branch region according to Embodiment 10.
[0108] Figure 92 is a diagram showing an example of a dense three-dimensional point group according to Embodiment 10.
[0109] Figure 93 is a diagram showing an example of a sparse three-dimensional point group according to Embodiment 10.
[0110] Figure 94 is a flowchart of the encoding process according to Embodiment 10.
[0111] Figure 95 is a flowchart of the decoding process according to Embodiment 10.
[0112] Figure 96 is a flowchart of the encoding process according to Embodiment 10.
[0113] Figure 97 is a flowchart of the decoding process according to Embodiment 10.
[0114] Figure 98 is a flowchart of the encoding process according to Embodiment 10.
[0115] Figure 99 is a flowchart of the decoding process according to Embodiment 10.
[0116] Figure 100 is a flowchart showing the separation process of a three-dimensional point according to Embodiment 10.
[0117] Figure 101 is a diagram showing a syntax example according to Embodiment 10.
[0118] Figure 102 is a diagram showing an example of a dense branch according to Embodiment 10.
[0119] Figure 103is a diagram showing an example of a sparse branch involved in Embodiment 10.
[0120] Figure 104 is a flowchart of an encoding process of a modification example involved in Embodiment 10.
[0121] Figure 105 is a flowchart of a decoding process of a modification example involved in Embodiment 10.
[0122] Figure 106 is a flowchart of a separation process of a three-dimensional point of a modification example involved in Embodiment 10.
[0123] Figure 107 is a diagram showing a syntax example of a modification example involved in Embodiment 10.
[0124] Figure 108 is a flowchart of an encoding process involved in Embodiment 10.
[0125] Figure 109 is a flowchart of a decoding process involved in Embodiment 10.
[0126] Figure 110 is a diagram showing an example of a tree structure involved in Embodiment 11.
[0127] Figure 111 is a diagram showing an example of an occupancy rate code involved in Embodiment 11.
[0128] Figure 112 is a diagram schematically showing an action of a three-dimensional data encoding apparatus involved in Embodiment 11.
[0129] Figure 113 is a diagram showing an example of geometry information involved in Embodiment 11.
[0130] Figure 114 is a diagram showing a selection example of an encoding table using geometry information involved in Embodiment 11.
[0131] Figure 115 is a diagram showing a selection example of an encoding table using structure information involved in Embodiment 11.
[0132] Figure 116 is a diagram showing a selection example of an encoding table using attribute information involved in Embodiment 11.
[0133] Figure 117 is a diagram showing a selection example of an encoding table using attribute information involved in Embodiment 11.
[0134] Figure 118 is a diagram showing a structure example of a bitstream involved in Embodiment 11.
[0135] Figure 119 FIG. 11B is a diagram illustrating an example of an encoding table related to Embodiment 11.
[0136] Figure 120 FIG. 11B is a diagram illustrating an example of an encoding table related to Embodiment 11.
[0137] Figure 121 FIG. 11B is a diagram illustrating an example of an encoding table related to Embodiment 11.
[0138] Figure 122 FIG. 11B is a diagram illustrating an example of an encoding table related to Embodiment 11.
[0139] Figure 123 FIG. 11B is a diagram illustrating an example of an encoding table related to Embodiment 11.
[0140] Figure 124 FIG. 11B is a diagram illustrating an example of an encoding table related to Embodiment 11.
[0141] Figure 125 FIG. 11B is a diagram illustrating an example of an encoding table related to Embodiment 11.
[0142] Figure 126 FIG. 11B is a diagram illustrating an example of an encoding table related to Embodiment 11.
[0143] Figure 127 FIG. 11B is a diagram illustrating an example of an encoding table related to Embodiment 11.
[0144] Figure 128 FIG. 11B is a diagram illustrating an example of an encoding table related to Embodiment 11.
[0145] Figure 129 FIG. 11B is a diagram illustrating an example of an encoding table related to Embodiment 11.
[0146] Figure 130 FIG. 11B is a diagram illustrating an example of an encoding table related to Embodiment 11.
[0147] Figure 131 FIG. 11B is a diagram illustrating an example of an encoding table related to Embodiment 11.
[0148] Figure 132 FIG. 11B is a diagram illustrating an example of an encoding table related to Embodiment 11.
[0149] Figure 133 FIG. 11B is a diagram illustrating an example of an encoding table related to Embodiment 11.
[0150] Figure 134is a block diagram of a three-dimensional data encoding apparatus according to Embodiment 11.
[0151] Figure 135 is a block diagram of a three-dimensional data decoding apparatus according to Embodiment 11.
[0152] Figure 136 is a diagram showing a reference relationship in an octree structure according to Embodiment 12.
[0153] Figure 137 is a diagram showing a reference relationship in a spatial region according to Embodiment 12.
[0154] Figure 138 is a diagram showing an example of adjacent reference nodes according to Embodiment 12.
[0155] Figure 139 is a diagram showing a relationship between a parent node and a node according to Embodiment 12.
[0156] Figure 140 is a diagram showing an example of an occupancy code of a parent node according to Embodiment 12.
[0157] Figure 141 is a block diagram of a three-dimensional data encoding apparatus according to Embodiment 12.
[0158] Figure 142 is a block diagram of a three-dimensional data decoding apparatus according to Embodiment 12.
[0159] Figure 143 is a flowchart showing a three-dimensional data encoding process according to Embodiment 12.
[0160] Figure 144 is a flowchart showing a three-dimensional data decoding process according to Embodiment 12.
[0161] Figure 145 is a diagram showing a switching example of an encoding table according to Embodiment 12.
[0162] Figure 146 is a diagram showing a reference relationship in a spatial region according to Modified Example 1 of Embodiment 12.
[0163] Figure 147 is a diagram showing a syntax example of header information according to Modified Example 1 of Embodiment 12.
[0164] Figure 148 is a diagram showing a syntax example of header information according to Modified Example 1 of Embodiment 12.
[0165] Figure 149is a diagram representing an example of the adjacent reference nodes involved in Embodiment 12, Modification 2.
[0166] Figure 150 is a diagram representing an example of the object node and the adjacent nodes involved in Embodiment 12, Modification 2.
[0167] Figure 151 is a diagram representing the reference relationship in the octree structure involved in Embodiment 12, Modification 3.
[0168] Figure 152 is a diagram representing the reference relationship in the spatial region involved in Embodiment 12, Modification 3. DETAILED DESCRIPTION
[0169] The three-dimensional data encoding method of one aspect of the present disclosure encodes information of an object node included in an N (N is an integer of 2 or more) octree structure of a plurality of three-dimensional points included in three-dimensional data, in the encoding, permitting reference to information of a first node that is the same as a parent node of the object node among a plurality of adjacent nodes that are spatially adjacent to the object node, and prohibiting reference to information of a second node that is different from the parent node of the object node.
[0170] Thus, the three-dimensional data encoding method can improve the encoding efficiency by referring to information of the first node that is the same as the parent node of the object node among the plurality of adjacent nodes that are spatially adjacent to the object node. In addition, the three-dimensional data encoding method can reduce the processing amount because it does not refer to information of the second node that is different from the parent node of the object node among the plurality of adjacent nodes. Thus, the three-dimensional data encoding method can improve the encoding efficiency and can reduce the processing amount.
[0171] For example, the three-dimensional data encoding method further decides whether to prohibit reference to the information of the second node, in the encoding, based on a result of the decision, it switches whether to prohibit or to permit reference to the information of the second node, and the three-dimensional data encoding method further generates a bitstream containing prohibition switching information that is the result of the decision and indicates whether to prohibit reference to the information of the second node.
[0172] Thus, the three-dimensional data encoding method can switch whether to prohibit reference to the information of the second node. In addition, the three-dimensional data decoding device can appropriately perform the decoding process using the prohibition switching information.
[0173] For example, it can also be that the information of the object node is information indicating whether or not a three-dimensional point exists in each of the child nodes belonging to the object node, the information of the first node is information indicating whether or not a three-dimensional point exists in the first node, and the information of the second node is information indicating whether or not a three-dimensional point exists in the second node.
[0174] For example, it can also be that, in the encoding, an encoding table is selected based on whether or not a three-dimensional point exists in the first node, and the information of the object node is entropy encoded using the selected encoding table.
[0175] For example, it can also be that, in the encoding, information of a child node of the first node among the plurality of neighboring nodes is referred to.
[0176] Thus, the three-dimensional data encoding method can refer to more detailed information of neighboring nodes, and thus can improve encoding efficiency.
[0177] For example, it can also be that, in the encoding, the neighboring node to be referred to among the plurality of neighboring nodes is switched according to a spatial position within a parent node of the object node.
[0178] Thus, the three-dimensional data encoding method can refer to an appropriate neighboring node according to a spatial position within a parent node of an object node.
[0179] A three-dimensional data decoding method of one aspect of the present disclosure decodes information of an object node included in an N (N is an integer of 2 or more) tree structure of a plurality of three-dimensional points included in three-dimensional data, and, in the decoding, information of a first node whose parent node is the same as a parent node of the object node among a plurality of neighboring nodes spatially adjacent to the object node is referred to, and information of a second node whose parent node is different from the parent node of the object node is not referred to.
[0180] Thus, the three-dimensional data decoding method can improve encoding efficiency by referring to information of a first node whose parent node is the same as a parent node of an object node among a plurality of neighboring nodes spatially adjacent to the object node. In addition, the three-dimensional data decoding method can reduce processing amount because information of a second node whose parent node is different from the parent node of the object node is not referred to. Thus, the three-dimensional data decoding method can improve encoding efficiency and can reduce processing amount.
[0181] For example, it can also be that the three-dimensional data decoding method further obtains, from a bitstream, prohibition switching information indicating whether or not information of the second node is prohibited from being referred to, and, in the decoding, whether or not information of the second node is referred to is switched based on the prohibition switching information.
[0182] Thus, the three-dimensional data decoding method can appropriately perform decoding processing using the prohibited switching information.
[0183] For example, the information of the object node can be information indicating whether or not a three-dimensional point exists in each of the child nodes belonging to the object node, the information of the first node can be information indicating whether or not a three-dimensional point exists in the first node, and the information of the second node can be information indicating whether or not a three-dimensional point exists in the second node.
[0184] For example, in the decoding, an encoding table can be selected based on whether or not a three-dimensional point exists in the first node, and the information of the object node can be entropy-decoded using the selected encoding table.
[0185] For example, in the decoding, information of a child node of the first node among the plurality of adjacent nodes can be permitted to be referred to.
[0186] Thus, the three-dimensional data decoding method can refer to more detailed information of adjacent nodes, and thus can improve the reference encoding efficiency.
[0187] For example, in the decoding, the adjacent node to be referred to among the plurality of adjacent nodes can be switched according to a spatial position within a parent node of the object node.
[0188] Thus, the three-dimensional data decoding method can refer to an appropriate adjacent node according to a spatial position within a parent node of an object node.
[0189] In addition, a three-dimensional data encoding device according to one aspect of the present disclosure includes a processor and a memory. The processor encodes information of an object node included in an N (N is an integer of 2 or more) tree structure of a plurality of three-dimensional points included in three-dimensional data using the memory. In the encoding, information of a first node whose parent node is the same as a parent node of the object node among a plurality of adjacent nodes adjacent to the object node in space is permitted to be referred to, and information of a second node whose parent node is different from the parent node of the object node is prohibited from being referred to.
[0190] Thus, the three-dimensional data encoding device can improve the encoding efficiency by referring to information of a first node whose parent node is the same as a parent node of an object node among a plurality of adjacent nodes adjacent to the object node in space. In addition, the three-dimensional data encoding device can reduce the processing amount because information of a second node whose parent node is different from the parent node of the object node among the plurality of adjacent nodes is not referred to. Thus, the three-dimensional data encoding device can improve the encoding efficiency and can reduce the processing amount.
[0191] Further, a three-dimensional data decoding device according to one aspect of the present disclosure includes a processor and a memory, and the processor decodes information of an object node included in an N (N is an integer of 2 or more)-ary tree structure of a plurality of three-dimensional points included in three-dimensional data using the memory. In the decoding, information of a first node, which is one of a plurality of neighboring nodes spatially adjacent to the object node and has a parent node identical to a parent node of the object node, is permitted to be referred to, and information of a second node, which is one of the plurality of neighboring nodes and has a parent node different from the parent node of the object node, is prohibited from being referred to.
[0192] Thus, the three-dimensional data decoding device can improve encoding efficiency by referring to the information of the first node, which is one of the plurality of neighboring nodes spatially adjacent to the object node and has the parent node identical to the parent node of the object node. Further, the three-dimensional data decoding device can reduce processing amount because the information of the second node, which is one of the plurality of neighboring nodes and has the parent node different from the parent node of the object node, is not referred to. Thus, the three-dimensional data decoding device can improve encoding efficiency and can reduce processing amount.
[0193] Further, these general or specific aspects can be implemented by system, method, integrated circuit, computer program, or recording medium such as a CD-ROM, and any combination of the system, method, integrated circuit, computer program, and recording medium.
[0194] Embodiments will be specifically described below with reference to the drawings. Further, the embodiments to be described below are each one specific example illustrating the present disclosure. Numerical values, shapes, materials, component elements, arrangement positions of component elements, connection forms, steps, orders of steps, and the like shown in the following embodiments are each one example, and the gist of the present disclosure is not limited to them. Also, among component elements of the following embodiments, component elements not described in the technical solution showing the most general concept are described as arbitrary component elements.
[0195] (Embodiment 1)
[0196] First, a data structure of encoded three-dimensional data (hereinafter also referred to as encoded data) to which the present embodiment is applied will be described. Figure 1 The configuration of the encoded three-dimensional data to which the present embodiment is applied is shown.
[0197] In the present embodiment, a three-dimensional space is divided into spaces (SPC) corresponding to pictures in encoding of a moving image, and the three-dimensional data is encoded in units of space. The space is further divided into volumes (VLM) corresponding to macroblocks and the like in encoding of a moving image, and prediction and conversion are performed in units of VLM. The volume includes a plurality of voxels (VXL) as the minimum unit corresponding to position coordinates. In addition, prediction means that, similarly to prediction performed in a two-dimensional image, prediction three-dimensional data similar to a processing unit of an object of processing is generated with reference to other processing units, and a difference between the prediction three-dimensional data and the processing unit of the object of processing is encoded. Furthermore, the prediction includes not only spatial prediction with reference to other prediction units at the same time, but also temporal prediction with reference to prediction units at different times.
[0198] For example, a three-dimensional data encoding apparatus (hereinafter also referred to as an encoding apparatus) encodes a three-dimensional space represented by point cloud data and the like by encoding each point of a point cloud or a plurality of points included in a voxel in units of the size of the voxel. If the voxel is subdivided, a three-dimensional shape of the point cloud can be represented with high accuracy, and if the size of the voxel is increased, the three-dimensional shape of the point cloud can be represented roughly.
[0199] In addition, although the following description is given taking the case of three-dimensional data as point cloud data as an example, the three-dimensional data is not limited to point cloud data, and can be three-dimensional data in any form.
[0200] Furthermore, a voxel of a hierarchical structure can be used. In this case, in the n-th hierarchical layer, whether or not a sampling point exists in a hierarchical layer lower than the n-th hierarchical layer (a lower layer of the n-th hierarchical layer) can be sequentially shown. For example, when a sampling point exists in a hierarchical layer lower than the n-th hierarchical layer, it can be considered that a sampling point exists at the center of a voxel of the n-th hierarchical layer, and decoding can be performed when only the n-th hierarchical layer is decoded.
[0201] Furthermore, the encoding apparatus obtains point cloud data by a distance sensor, a stereo camera, a monocular camera, a gyroscope, or an inertial sensor.
[0202] With respect to a space, similarly to encoding of a moving image, at least any one of the following three prediction structures is classified: an intra-space (I-SPC) that can be decoded independently, a predictive space (P-SPC) that can be referred to only in one direction, and a bidirectional space (B-SPC) that can be referred to in both directions. Furthermore, the space has both time information of a decoding time and a display time.
[0203] Furthermore, as Figure 1As a processing unit including a plurality of spaces, there is a GOS (Group Of Space) as a random access unit. Further, as a processing unit including a plurality of GOSs, there is a WLD (World).
[0204] A space region occupied by a WLD corresponds to an absolute position on the earth by GPS or latitude and longitude information. This position information is stored as meta information. The meta information can be included in encoded data or can be transmitted separately from the encoded data.
[0205] Further, within a GOS, all SPCs can be contiguous in three dimensions, or there can be SPCs that are not contiguous in three dimensions with other SPCs.
[0206] Further, encoding, decoding, or reference processing and the like of encoded data corresponding to three-dimensional data included in a processing unit such as a GOS, SPC, or VLM will be simply referred to as encoding, decoding, or reference processing and the like of the processing unit. Further, three-dimensional data included in a processing unit includes, for example, at least one set of spatial positions such as three-dimensional coordinates and characteristic values such as color information.
[0207] Next, the prediction structure of SPCs in a GOS will be described. A plurality of SPCs within the same GOS or a plurality of VLMs within the same SPC hold the same time information (decoding time and display time) although they occupy different spaces from each other.
[0208] Further, in a GOS, an SPC at the beginning in decoding order is an I-SPC. Further, there are two types of GOSs, a closed GOS and an open GOS. A closed GOS is a GOS in which all SPCs within the GOS can be decoded when decoding starts from the beginning I-SPC. In an open GOS, a part of SPCs within the GOS refer to a different GOS earlier than the display time of the beginning I-SPC, and can be decoded only in the GOS.
[0209] Further, in encoded data of map information or the like, there is a case where a WLD is decoded from a direction opposite to the encoding order, and if there is a dependency between GOSs, it is difficult to reproduce in the reverse direction. Therefore, in this case, a closed GOS is basically adopted.
[0210] Further, a GOS has a layer structure in the height direction, and SPCs of a lower layer are sequentially encoded or decoded.
[0211] Figure 2 An example of a prediction structure between SPCs belonging to the lowest layer of a GOS is shown. Figure 3 An example of a prediction structure between layers is shown.
[0212] There is one or more I-SPC within the GOS. This is effective when small-sized objects are encoded as I-SPC, although there are objects such as people, animals, cars, bicycles, signal lights, or buildings that become landmarks in three-dimensional space. For example, a three-dimensional data decoding device (hereinafter also referred to as a decoding device) decodes only the I-SPC within the GOS when decoding the GOS at low processing capacity or at high speed.
[0213] Further, the encoding device can switch the encoding interval or the occurrence frequency of the I-SPC according to the density of the objects within the WLD.
[0214] Further, in the configuration shown in Figure 3 In the configuration shown in the drawing, the encoding device or the decoding device encodes or decodes the multiple layers in order from the lower layer (Layer 1). Accordingly, for example, for an automatically walking vehicle or the like, the priority of the data near the ground, which has a large amount of information, can be improved.
[0215] In addition, in the encoded data used for a drone or the like, the SPCs of the layers from the upper layer in the height direction within the GOS can be encoded or decoded in order.
[0216] Further, the encoding device or the decoding device can encode or decode the multiple layers in a manner in which the decoding device roughly grasps the GOS and can gradually improve the resolution. For example, the encoding device or the decoding device can encode or decode in the order of Layer 3, 8, 1, 9,....
[0217] Next, the corresponding method of static objects and dynamic objects will be described.
[0218] In three-dimensional space, there are static objects or scenes (hereinafter collectively referred to as static objects) such as buildings or roads, and dynamic objects (hereinafter referred to as dynamic objects) such as vehicles or people. The detection of the objects can be performed separately by extracting feature points from the data of point clouds or captured images of a stereo camera or the like. Here, an example of the encoding method of dynamic objects will be described.
[0219] The first method is a method of encoding without distinguishing between static objects and dynamic objects. The second method is a method of distinguishing between static objects and dynamic objects by recognition information.
[0220] For example, the GOS is used as a unit of recognition. In this case, the GOS including the SPCs constituting the static objects and the GOS including the SPCs constituting the dynamic objects are distinguished within the encoded data or by recognition information stored separately from the encoded data.
[0221] Alternatively, the SPC is used as the identification unit. In this case, the SPC including the VLM constituting the static object and the SPC including the VLM constituting the dynamic object are distinguished by the above-mentioned identification information.
[0222] Alternatively, the VLM or the VXL can be used as the identification unit. In this case, the VLM or the VXL including the static object and the VLM or the VXL including the dynamic object are distinguished by the above-mentioned identification information.
[0223] Further, the encoding device can encode the dynamic object as one or more VLMs or SPCs, and encode the VLM or SPC including the static object and the SPC including the dynamic object as different GOSs from each other. Further, the encoding device stores the size of the GOS as meta information separately in the case where the size of the GOS becomes variable depending on the size of the dynamic object.
[0224] Further, the encoding device encodes the static object and the dynamic object independently from each other, and can superimpose the dynamic object on the world space constituted by the static object. At this time, the dynamic object is constituted by one or more SPCs, and each SPC corresponds to one or more SPCs constituting the static object on which the SPC is superimposed. Alternatively, the dynamic object can not be represented by the SPC, and can be represented by one or more VLMs or VXLs.
[0225] Further, the encoding device can encode the static object and the dynamic object as different streams from each other.
[0226] Further, the encoding device can generate a GOS including one or more SPCs constituting the dynamic object. Further, the encoding device can set the GOS including the dynamic object (GOS_M) and the GOS of the static object corresponding to the spatial region of the GOS_M to the same size (occupy the same spatial region). In this way, it is possible to perform superimposition processing in units of GOS.
[0227] The P-SPC or the B-SPC constituting the dynamic object can also refer to the SPC included in a different GOS that has been encoded. The position of the dynamic object changes with time, and in the case where the same dynamic object is encoded as different GOSs at different times, the reference across the GOSs is effective from the viewpoint of compression rate.
[0228] Further, the above-mentioned first method and the second method can be switched depending on the use of the encoded data. For example, in the case where three-dimensional data is encoded as a map, it is desirable to separate the dynamic object, and therefore the encoding device adopts the second method. Alternatively, in the case where three-dimensional data of an event such as a concert or sports is encoded, the first method is adopted if there is no need to separate the dynamic object.
[0229] Also, the decoding time and the display time of the GOS or the SPC can be stored in the encoded data or as meta information. Also, the time information of the static object can be all the same. In this case, the actual decoding time and the display time can be determined by the decoding device. Or, as the decoding time, different values can be assigned to each GOS or SPC, and as the display time, the same value can be assigned to all. Also, as shown in the decoder model in the dynamic image encoding such as the HRD (Hypothetical Reference Decoder) of the HEVC, the decoder has a buffer of a prescribed size, and as long as the bit stream is read at a prescribed bit rate according to the decoding time, a model can be introduced that is not broken and guarantees decodability.
[0230] Next, the configuration of the GOS in the world space is explained. The coordinates of the three-dimensional space in the world space are expressed by three coordinate axes (x-axis, y-axis, z-axis) that are orthogonal to each other. By setting a prescribed rule in the encoding order of the GOS, GOSs that are spatially adjacent can be continuously encoded in the encoded data. For example, in the example shown in FIG. 8, the GOSs in the xz plane are continuously encoded. After the encoding of all the GOSs in one xz plane is completed, the value of the y-axis is updated. That is, as the encoding proceeds, the world space extends in the direction of the y-axis. Also, the index number of the GOS is set to the encoding order. Figure 4
[0231] Here, the three-dimensional space of the world space corresponds one-to-one to the absolute coordinates in geography such as the GPS or the latitude and longitude. Or, the three-dimensional space can be expressed by the relative position with respect to a reference position set in advance. The directions of the x-axis, the y-axis, and the z-axis of the three-dimensional space are expressed as direction vectors determined based on the latitude and longitude or the like, and the direction vectors are stored as meta information together with the encoded data.
[0232] Also, the size of the GOS is set to be fixed, and the encoding device stores the size as meta information. Also, the size of the GOS can be switched, for example, according to whether it is in a city or indoors or outdoors or the like. That is, the size of the GOS can be switched according to the amount or the properties of the objects that have value as information. Or, the encoding device can appropriately switch the size of the GOS or the interval of the I-SPC in the GOS according to the density of the objects or the like in the same world space. For example, the encoding device sets the size of the GOS to be small and the interval of the I-SPC in the GOS to be short as the density of the objects is higher.
[0233] In the case of the example shown in FIG. 8, the GOSs in the xz plane are continuously encoded. After the encoding of all the GOSs in one xz plane is completed, the value of the y-axis is updated. That is, as the encoding proceeds, the world space extends in the direction of the y-axis. Also, the index number of the GOS is set to the encoding order. Figure 5 In the example, in the region from the 3rd to the 10th GOS, due to the high density of objects, the GOS is subdivided to achieve fine-grained random access. Furthermore, the 7th to 10th GOS are located on the back sides of the 3rd to 6th GOS, respectively.
[0234] Next, the structure and operation flow of the three-dimensional data encoding device involved in this embodiment will be explained. Figure 6 This is a block diagram of the three-dimensional data encoding device 100 according to this embodiment. Figure 7 This is a flowchart illustrating an example of the operation of the three-dimensional data encoding device 100.
[0235] Figure 6 The 3D data encoding apparatus 100 shown generates encoded 3D data 112 by encoding 3D data 111. The 3D data encoding apparatus 100 includes: an acquisition unit 101, an encoding region determination unit 102, a segmentation unit 103, and an encoding unit 104.
[0236] like Figure 7 As shown, firstly, the acquisition unit 101 acquires three-dimensional data 111 as point group data (S101).
[0237] Next, the encoding region determination unit 102 determines the region of the encoding object from the spatial region corresponding to the obtained point group data (S102). For example, the encoding region determination unit 102 determines the spatial region surrounding the user or vehicle's location as the region of the encoding object.
[0238] Next, the segmentation unit 103 divides the point group data contained in the region of the encoding object into individual processing units. Here, the processing units are the aforementioned GOS and SPC, etc. Furthermore, the region of the encoding object corresponds, for example, to the aforementioned world space. Specifically, the segmentation unit 103 divides the point group data into processing units based on a pre-set GOS size, the presence or size of dynamic objects (S103). Furthermore, the segmentation unit 103 determines the starting position of the SPC that will be the first in the encoding sequence within each GOS.
[0239] Next, the encoding unit 104 generates encoded three-dimensional data 112 by sequentially encoding multiple SPCs within each GOS (S104).
[0240] Furthermore, although an example of encoding each GOS is shown here after dividing the region of the encoded object into GOS and SPC, the processing order is not limited to the above. For example, a GOS can be encoded after its structure is determined, and then the order of GOS structure can be determined afterward.
[0241] Thus, the three-dimensional data encoding apparatus 100 generates the encoded three-dimensional data 112 by encoding the three-dimensional data 111. Specifically, the three-dimensional data encoding apparatus 100 divides the three-dimensional data into random access units, i.e., first processing units (GOS) each corresponding to a three-dimensional coordinate, divides each of the first processing units (GOS) into a plurality of second processing units (SPC), and divides each of the second processing units (SPC) into a plurality of third processing units (VLM). Further, the third processing unit (VLM) includes one or more voxels (VXL) which are the smallest units corresponding to positional information.
[0242] Next, the three-dimensional data encoding apparatus 100 generates the encoded three-dimensional data 112 by encoding each of the plurality of first processing units (GOS). Specifically, the three-dimensional data encoding apparatus 100 encodes each of the plurality of second processing units (SPC) in each of the first processing units (GOS). Further, the three-dimensional data encoding apparatus 100 encodes each of the plurality of third processing units (VLM) in each of the second processing units (SPC).
[0243] For example, in a case where the first processing unit (GOS) of the processing target is a closed GOS, the three-dimensional data encoding apparatus 100 encodes the second processing unit (SPC) of the processing target included in the first processing unit (GOS) of the processing target with reference to other second processing units (SPC) included in the first processing unit (GOS) of the processing target. That is, the three-dimensional data encoding apparatus 100 does not refer to the second processing units (SPC) included in the first processing units (GOS) other than the first processing unit (GOS) of the processing target.
[0244] Further, in a case where the first processing unit (GOS) of the processing target is an open GOS, the three-dimensional data encoding apparatus 100 encodes the second processing unit (SPC) of the processing target included in the first processing unit (GOS) of the processing target with reference to other second processing units (SPC) included in the first processing unit (GOS) of the processing target or the second processing units (SPC) included in the first processing units (GOS) other than the first processing unit (GOS) of the processing target.
[0245] Further, the three-dimensional data encoding apparatus 100 selects one of a first type (I-SPC) which does not refer to other second processing units (SPC), a second type (P-SPC) which refers to one other second processing unit (SPC), and a third type which refers to two other second processing units (SPC) as the type of the second processing unit (SPC) of the processing target, and encodes the second processing unit (SPC) of the processing target in accordance with the selected type.
[0246] Next, the configuration of the three-dimensional data decoding apparatus and the flow of the operation thereof according to the present embodiment will be described. Figure 8 Fig. 2 is a block diagram of a three-dimensional data decoding apparatus 200 according to the present embodiment. Figure 9 Fig. 3 is a flowchart showing an example of the operation of the three-dimensional data decoding apparatus 200.
[0247] Figure 8 The three-dimensional data decoding apparatus 200 shown in Fig. 2 generates decoded three-dimensional data 212 by decoding encoded three-dimensional data 211. Here, the encoded three-dimensional data 211 is, for example, the encoded three-dimensional data 112 generated by the three-dimensional data encoding apparatus 100. The three-dimensional data decoding apparatus 200 includes an acquisition unit 201, a decoding start GOS decision unit 202, a decoding SPC decision unit 203, and a decoding unit 204.
[0248] First, the acquisition unit 201 acquires the encoded three-dimensional data 211 (S201). Next, the decoding start GOS decision unit 202 decides the GOS to be decoded (S202). Specifically, the decoding start GOS decision unit 202 decides the GOS to be decoded by referring to the GOS including the spatial position, the object, or the SPC corresponding to the time at which the decoding is started, which is included in the encoded three-dimensional data 211 or the meta information stored separately from the encoded three-dimensional data.
[0249] Next, the decoding SPC decision unit 203 decides the type (I, P, B) of the SPC to be decoded in the GOS (S203). For example, the decoding SPC decision unit 203 decides (1) whether to decode only the I-SPC, (2) whether to decode the I-SPC and the P-SPC, or (3) whether to decode all types. In addition, in a case where the type of the SPC to be decoded is predetermined, such as decoding all SPCs, the present step can not be performed.
[0250] Next, the decoding unit 204 acquires the SPC at the beginning in the decoding order (the same as the encoding order) in the GOS, acquires the address position at which the decoding is started in the encoded three-dimensional data 211, acquires the encoded data of the SPC at the beginning from the address position, and decodes each SPC in order from the SPC at the beginning (S204). In addition, the address position is stored in the meta information or the like.
[0251] Thus, the three-dimensional data decoding apparatus 200 decodes the decoded three-dimensional data 212. Specifically, the three-dimensional data decoding apparatus 200 generates the decoded three-dimensional data 212 of the first processing unit (GOS) as a random access unit by decoding each of the encoded three-dimensional data 211 of the first processing unit (GOS) corresponding to the three-dimensional coordinates. More specifically, the three-dimensional data decoding apparatus 200 decodes each of the plurality of second processing units (SPC) in each of the first processing units (GOS). Further, the three-dimensional data decoding apparatus 200 decodes each of the plurality of third processing units (VLM) in each of the second processing units (SPC).
[0252] The meta information for random access is described below. The meta information is generated by the three-dimensional data encoding apparatus 100 and included in the encoded three-dimensional data 112 (211).
[0253] In the conventional random access of a two-dimensional dynamic image, decoding is started from the beginning frame of the random access unit in the vicinity of the specified time. However, in the world space, random access is assumed for (coordinates or objects, etc.) in addition to the time.
[0254] Therefore, in order to achieve random access of at least three elements of coordinates, objects, and time, a table in which each element is associated with an index number of a GOS is prepared. Further, the index number of the GOS is associated with the address of an I-SPC that is the beginning of the GOS. Figure 10 An example of the table included in the meta information is shown. In addition, all of the tables shown in FIG. 8 need not be used. At least one table can be used. Figure 10
[0255] Random access starting from coordinates is described below as an example. When access is performed for coordinates (x2, y2, z2), first, the coordinate-GOS table is referred to, and it can be known that the place of coordinates (x2, y2, z2) is included in the second GOS. Next, the GOS address table is referred to, and since it can be known that the address of the I-SPC at the beginning of the second GOS is addr(2), the decoding section 204 obtains data from the address and starts decoding.
[0256] In addition, the address can be an address in a logical format or a physical address of an HDD or a memory. Further, information for specifying a file segment can be used instead of the address. For example, the file segment is a unit in which one or more GOSs are segmented.
[0257] Also, in a case where the object is a plurality of GOSs, the GOSs to which the plurality of objects belong can be shown in the object GOS table. If the plurality of GOSs are closed GOSs, the encoding apparatus and the decoding apparatus can perform encoding or decoding in parallel. Also, if the plurality of GOSs are open GOSs, the compression efficiency can be further improved by the plurality of GOSs referring to each other.
[0258] Examples of the object are a person, an animal, a car, a bicycle, a signal light, or a building that becomes a landmark on land. For example, the three-dimensional data encoding apparatus 100 extracts a feature point unique to the object from a three-dimensional point cloud or the like at the time of encoding in the world space, detects the object based on the feature point, and can set the detected object as a random access point.
[0259] Thus, the three-dimensional data encoding apparatus 100 generates first information showing a plurality of first processing units (GOSs) and a three-dimensional coordinate corresponding to each of the plurality of first processing units (GOSs). Also, the encoded three-dimensional data 112 (211) includes the first information. Also, the first information further shows at least one of an object, a time, and a data storage destination corresponding to each of the plurality of first processing units (GOSs).
[0260] The three-dimensional data decoding apparatus 200 obtains the first information from the encoded three-dimensional data 211, determines the encoded three-dimensional data 211 of the first processing unit corresponding to the specified three-dimensional coordinate, object, or time using the first information, and decodes the encoded three-dimensional data 211.
[0261] Examples of other meta information will be described below. The three-dimensional data encoding apparatus 100 can generate and store the following meta information in addition to the meta information for random access. Also, the three-dimensional data decoding apparatus 200 can use the meta information at the time of decoding.
[0262] In a case where the three-dimensional data is used as map information or the like, a profile is specified according to the use, and information showing the profile can be included in the meta information. For example, a profile for an urban area or a suburban area is specified, or a profile for a flying object is specified, and the maximum or minimum size of the world space, the SPC, or the VLM is defined, for example. For example, in the profile for the urban area, more detailed information than the suburban area is required, and thus the minimum size of the VLM is set to be small.
[0263] The meta-information can also include a label value showing the kind of the object. The label value corresponds to the VLM, SPC, or GOS that constitutes the object. The label value can be set in accordance with the kind of the object, and so on, for example, a label value "0" indicates "person", a label value "1" indicates "car", and a label value "2" indicates "traffic light". Alternatively, in a case where the kind of the object is difficult to determine or is not required to be determined, a label value showing the size, or the property of being a dynamic object or a static object, and so on can be used.
[0264] Further, the meta-information can also include information showing the range of the spatial region occupied by the world space.
[0265] Further, the meta-information can also store the size of the SPC or VXL as header information common to the entire stream of the encoded data, or to a plurality of SPCs within the GOS, and so on.
[0266] Further, the meta-information can also include identification information of a distance sensor or a camera, and so on, used in the generation of the point cloud, or information showing the positional accuracy of the point group within the point cloud.
[0267] Further, the meta-information can include information showing whether the world space is constituted only of static objects or contains dynamic objects.
[0268] A modification of the present embodiment will be described below.
[0269] The encoding device or the decoding device can encode or decode two or more SPCs or GOSs that are different from each other in parallel. The GOSs that are encoded or decoded in parallel can be determined in accordance with the meta-information showing the spatial position of the GOS, and so on.
[0270] In a case where three-dimensional data is used as a spatial map at the time of movement of a vehicle or a flying object, and so on, or such a spatial map is generated, the encoding device or the decoding device can encode or decode the GOS or SPC contained in the space determined on the basis of GPS, path information, or a zoom ratio, and so on.
[0271] Further, the decoding device can start decoding from the space close to the own position or the walking path in order. The encoding device or the decoding device can also encode or decode the space far from the own position or the walking path with a lower priority than the space close thereto. Here, lowering the priority means lowering the processing order, lowering the resolution (post-filtering), or lowering the image quality (improving the encoding efficiency. For example, increasing the quantization step size), and so on.
[0272] Further, the decoding device can decode only the lower hierarchy when decoding the encoded data encoded in a hierarchical manner within the space.
[0273] Also, the decoding device can decode from the lower hierarchy first according to the scale or use of the map.
[0274] Also, in uses such as self-position estimation or object recognition during automatic travel of a car or a robot, the encoding device or the decoding device can reduce the resolution of regions outside the region (recognition target region) within a prescribed height from the road surface and encode or decode.
[0275] Also, the encoding device can encode the point clouds representing the shapes of the indoor and outdoor spaces separately. For example, by separating the GOS representing the indoor space (indoor GOS) from the GOS representing the outdoor space (outdoor GOS), the decoding device can select the GOS to be decoded according to the viewpoint position when using the encoded data.
[0276] Also, the encoding device can encode the indoor GOS and the outdoor GOS with close coordinates to be adjacent in the encoded stream. For example, the encoding device stores information showing the identifiers corresponding to the identifiers of the indoor GOS and the outdoor GOS in the encoded stream or in meta information stored separately. Accordingly, the decoding device can identify the indoor GOS and the outdoor GOS with close coordinates by referring to the information in the meta information.
[0277] Also, the encoding device can switch the size of the GOS or the SPC between the indoor GOS and the outdoor GOS. For example, the encoding device can set the size of the GOS to be smaller in the indoor space than in the outdoor space. Also, the encoding device can change the accuracy when extracting feature points from the point cloud or the accuracy of object detection between the indoor GOS and the outdoor GOS.
[0278] Also, the encoding device can attach information for the decoding device to display a dynamic object and a static object separately to the encoded data. Accordingly, the decoding device can represent the dynamic object in combination with a red frame or an explanatory text or the like. Alternatively, the decoding device can represent only the red frame or the explanatory text instead of the dynamic object. Also, the decoding device can represent a more detailed object category. For example, a car can be represented by a red frame and a person can be represented by a yellow frame.
[0279] Also, the encoding device or the decoding device can determine whether to encode or decode a dynamic object and a static object as different SPCs or GOSs according to the frequency of appearance of the dynamic object or the ratio of the static object to the dynamic object or the like. For example, in a case where the frequency of appearance of the dynamic object or the ratio exceeds a threshold value, an SPC or a GOS in which the dynamic object and the static object are mixed is allowed, and in a case where the frequency of appearance of the dynamic object or the ratio does not exceed the threshold value, an SPC or a GOS in which the dynamic object and the static object are mixed is not allowed.
[0280] When the dynamic object is detected from two-dimensional image information of a camera instead of a point cloud, the encoding apparatus can separately obtain information (a frame or a character, etc.) for identifying a detection result and an object position, and encode the information as a part of three-dimensional encoding data. In this case, the decoding apparatus superimposes and displays auxiliary information (a frame or a character) indicating the dynamic object on a decoding result of the static object.
[0281] Also, the encoding apparatus can change the density of the VXL or the VLM according to the complexity of the shape of the static object, etc. For example, the encoding apparatus sets the VXL or the VLM to be denser as the shape of the static object is more complex. Also, the encoding apparatus can determine a quantization step size, etc. when quantizing the spatial position or the color information according to the density of the VXL or the VLM. For example, the encoding apparatus sets the quantization step size to be smaller as the VXL or the VLM is denser.
[0282] As described above, the encoding apparatus or the decoding apparatus according to the present embodiment encodes or decodes a space in units of space having coordinate information.
[0283] Also, the encoding apparatus and the decoding apparatus encode or decode in units of volume in a space. The volume includes a voxel which is a minimum unit corresponding to position information.
[0284] Also, the encoding apparatus and the decoding apparatus establish a correspondence between arbitrary elements by encoding or decoding a table in which each element of spatial information including coordinates, objects, and time, etc. is associated with a GOP, or a table in which each element is associated with each other. Also, the decoding apparatus determines a volume, a voxel, or a space by using a value of a selected element, and decodes a space including the volume or the voxel, or a determined space according to the coordinates.
[0285] Also, the encoding apparatus determines a volume, a voxel, or a space that can be selected by an element by feature point extraction or object recognition, and encodes as a volume, a voxel, or a space that can be randomly accessed.
[0286] A space is classified into three types, i.e., an I-SPC which can be encoded or decoded in a unit of the space, a P-SPC which is encoded or decoded by referring to any one of processed spaces, and a B-SPC which is encoded or decoded by referring to any two of processed spaces.
[0287] One or more volumes correspond to a static object or a dynamic object. A space including a static object and a space including a dynamic object are encoded or decoded as different GOSs. That is, an SPC including a static object and an SPC including a dynamic object are allocated to different GOSs.
[0288] The dynamic objects are encoded or decoded per object, corresponding to one or more spaces containing only static objects. That is, the plurality of dynamic objects are encoded separately, and the resulting encoded data of the plurality of dynamic objects correspond to SPCs containing only static objects.
[0289] The encoding apparatus and the decoding apparatus increase the priority of I-SPC within GOS to perform encoding or decoding. For example, the encoding apparatus performs encoding in such a manner that degradation of I-SPC is reduced (after decoding, the original three-dimensional data can be reproduced more faithfully). Also, the decoding apparatus, for example, decodes only I-SPC.
[0290] The encoding apparatus can change the frequency of use of I-SPC according to the density or the number of objects within the world space to perform encoding. That is, the encoding apparatus changes the frequency of selection of I-SPC according to the number or the density of objects contained in the three-dimensional data. For example, the encoding apparatus increases the frequency of use of I space as the density of objects within the world space is greater.
[0291] Also, the encoding apparatus sets random access points in units of GOS, and stores information showing the spatial region corresponding to GOS to the header information.
[0292] The encoding apparatus, for example, adopts a default value as the spatial size of GOS. Alternatively, the encoding apparatus can change the size of GOS according to the number or the density of objects or dynamic objects. For example, the encoding apparatus sets the spatial size of GOS to be small as the objects or dynamic objects are denser or the number is greater.
[0293] Also, the space or volume includes a group of feature points derived using information obtained by a depth sensor, a gyroscope, or a camera, or the like. The coordinates of the feature points are set to the center positions of the voxels. Also, by the subdivision of the voxels, high accuracy of position information can be achieved.
[0294] The group of feature points is derived using a plurality of pictures. The plurality of pictures have at least two kinds of time information, i.e., actual time information, and the same time information in the plurality of pictures corresponding to a space (for example, encoding time for rate control, etc.).
[0295] Also, encoding or decoding is performed in units of GOS including one or more spaces.
[0296] The encoding apparatus and the decoding apparatus refer to the space within the GOS that has been processed to predict the P space or the B space within the GOS that is the processing target.
[0297] Alternatively, the encoding apparatus and the decoding apparatus do not refer to different GOSs, and predict the P space or the B space within the GOS of the processing target using the processed space within the GOS of the processing target.
[0298] Also, the encoding apparatus and the decoding apparatus transmit or receive the encoded stream in units of a world space including one or more GOSs.
[0299] Also, the GOS has a hierarchical structure in at least one direction in the world space, and the encoding apparatus and the decoding apparatus encode or decode from the lower layer. For example, a GOS that can be randomly accessed belongs to the lowest layer. A GOS belonging to an upper layer refers only to a GOS belonging to a layer lower than the same layer. That is, the GOS is spatially divided in a predetermined direction, and includes a plurality of layers each having one or more SPCs. The encoding apparatus and the decoding apparatus encode or decode each SPC by referring to an SPC included in the same layer or a layer lower than the SPC.
[0300] Also, the encoding apparatus and the decoding apparatus sequentially encode or decode GOSs in a world space unit including a plurality of GOSs. The encoding apparatus and the decoding apparatus write or read information showing the order (direction) of encoding or decoding as metadata. That is, the encoded data includes information showing the encoding order of the plurality of GOSs.
[0301] Also, the encoding apparatus and the decoding apparatus encode or decode two or more spaces or GOSs different from each other in parallel.
[0302] Also, the encoding apparatus and the decoding apparatus encode or decode spatial information (coordinates, size, etc.) of a space or a GOS.
[0303] Also, the encoding apparatus and the decoding apparatus encode or decode a space or a GOS included in a specific space determined based on external information such as GPS, path information, or a scale related to the own position or / and the area size.
[0304] The encoding apparatus or the decoding apparatus encodes or decodes a space farther from the own position with a lower priority than a space closer to the own position.
[0305] The encoding apparatus sets one direction in the world space according to the scale or the use, and encodes a GOS having a hierarchical structure in the direction. Also, the decoding apparatus preferentially decodes from the lower layer a GOS having a hierarchical structure in one direction in the world space set according to the scale or the use.
[0306] The encoding device changes the accuracy of feature point extraction, object recognition, or the size of a spatial region included in an indoor and outdoor space. However, the encoding device and the decoding device encode or decode the indoor GOS and the outdoor GOS, which are adjacent in the world space, and also encode or decode the identifiers corresponding to each other.
[0307] (Embodiment 2)
[0308] In encoding data of a point cloud for use in an actual device or service, in order to suppress network bandwidth, it is desirable to transmit and receive only necessary information according to the use. However, such a function is not present in the encoding structure of three-dimensional data so far, and thus there is no encoding method corresponding to this.
[0309] In the present embodiment, a three-dimensional data encoding method and a three-dimensional data encoding device for providing a function of transmitting and receiving only necessary information according to the use in encoding data of a point cloud in three dimensions, and a three-dimensional data decoding method and a three-dimensional data decoding device for decoding the encoding data will be described.
[0310] A voxel (VXL) having a feature amount of one or more is defined as a feature voxel (FVXL), and a world space (WLD) constituted by FVXLs is defined as a sparse world space (SWLD). Figure 11 A configuration example of the sparse world space and the world space is shown. In the SWLD, FGOS constituted by FVXLs, FSPC constituted by FVXLs, and FVLM constituted by FVXLs are included. The data structure and the prediction structure of FGOS, FSPC, and FVLM can be the same as those of GOS, SPC, and VLM.
[0311] The feature amount refers to a feature amount representing three-dimensional position information of a VXL or visible light information of a VXL position, and in particular, a feature amount that can be detected more from a corner or an edge of a solid object. Specifically, the feature amount is a three-dimensional feature amount or a visible light feature amount described below, but can be any feature amount as long as it is a feature amount representing a position, brightness, or color information of a VXL.
[0312] As the three-dimensional feature amount, a SHOT feature amount (Signature of Histograms of OrienTations), a PFH feature amount (Point Feature Histograms), or a PPF feature amount (Point Pair Feature) is used.
[0313] The SHOT feature quantity is obtained by segmenting the periphery of the VXL, and calculating the inner product of the normal vector of the reference point and the segmented region, and histogramizing. This SHOT feature quantity has a high dimension and a high feature expression.
[0314] The PFH feature quantity is obtained by selecting a plurality of 2-point groups in the vicinity of the VXL, calculating the normal vector and the like from the 2 points, and histogramizing. This PFH feature quantity has a high feature expression because it is a histogram feature, and has a high robustness against a small amount of interference.
[0315] The PPF feature quantity is a feature quantity calculated using the normal vector and the like from the VXL of 2 points. In this PPF feature quantity, since all of the VXLs are used, it has a high robustness against occlusion.
[0316] Also, as the feature quantity of the visible light, SIFT (Scale-Invariant Feature Transform), SURF (Speeded Up Robust Features), or HOG (Histogram of Oriented Gradients), which uses information such as luminance gradient information of an image, can be used.
[0317] The SWLD is generated by calculating the above feature quantity from each VXL of the WLD, and extracting the FVXL. Here, the SWLD can be updated each time the WLD is updated, or can be updated periodically after a certain time elapses, regardless of the timing of the update of the WLD.
[0318] The SWLD can be generated for each feature quantity. For example, as shown in SWLD 1 based on the SHOT feature quantity and SWLD 2 based on the SIFT feature quantity, the SWLD can be generated for each feature quantity, respectively, and the SWLDs can be distinguished in use according to the use. Also, the feature quantity of each FVXL calculated can be held as feature quantity information to each FVXL.
[0319] Next, the method of using the sparse world space (SWLD) will be described. Since the SWLD contains only the feature voxels (FVXLs), the data size is generally small compared to the WLD which includes all of the VXLs.
[0320] In an application that utilizes the feature amount for a certain purpose, by utilizing the information of the SWLD instead of the WLD, the readout time from the hard disk can be suppressed, and the frequency band and the transmission time at the time of network transmission can be suppressed. For example, as map information, the WLD and the SWLD are held to the server in advance, by switching the transmitted map information to the WLD or the SWLD according to the demand from the client, the network band and the transmission time can be suppressed. A specific example is shown below.
[0321] Figure 12 and Figure 13 An example of utilization of the SWLD and the WLD is shown. As shown in Figure 12 , in a case where the client 1 that is a vehicle-mounted device needs map information for a purpose of self-position estimation, the client 1 transmits a demand for acquisition of map data for self-position estimation to the server (S301). The server transmits the SWLD to the client 1 according to the demand (S302). The client 1 performs self-position estimation using the received SWLD (S303). At this time, the client 1 acquires VXL information of the periphery of the client 1 by various methods such as a distance sensor such as a range finder, a stereo camera, or a combination of a plurality of monocular cameras, and estimates self-position information from the obtained VXL information and the SWLD. Here, the self-position information includes three-dimensional position information and an orientation of the client 1 and the like.
[0322] As shown in Figure 13 , in a case where the client 2 that is a vehicle-mounted device needs map information for a purpose of map drawing such as a three-dimensional map, the client 2 transmits a demand for acquisition of map data for map drawing to the server (S311). The server transmits the WLD to the client 2 according to the demand (S312). The client 2 performs map drawing using the received WLD (S313). At this time, the client 2 creates a concept image using an image captured by a visible light camera or the like and the WLD acquired from the server, and draws the created image to a screen of a car navigation or the like.
[0323] As shown above, the server transmits the SWLD to the client in a case where a feature amount of each VXL such as self-position estimation is mainly needed, and transmits the WLD to the client in a case where detailed VXL information is needed like map drawing. Thereby, the map data can be efficiently transmitted and received.
[0324] In addition, the client can judge which one of the SWLD and the WLD is needed by the client, and request transmission of the SWLD or the WLD to the server. Also, the server can judge which one of the SWLD and the WLD should be transmitted according to the state of the client or the network.
[0325] Next, a method of switching the transmission and reception of a sparse world space (SWLD) and a world space (WLD) will be described.
[0326] The reception of the WLD or the SWLD can be switched according to the network bandwidth. Figure 14 An example of the operation in this case will be shown. For example, in a case where a low-speed network using a network bandwidth of an LTE (Long Term Evolution) environment or the like is used, when the client accesses the server via the low-speed network (S321), the SWLD is acquired from the server as map information (S322). On the other hand, in a case where a high-speed network using a network bandwidth of a WiFi environment or the like is used, when the client accesses the server via the high-speed network (S323), the WLD is acquired from the server (S324). Accordingly, the client can acquire appropriate map information according to the network bandwidth of the client.
[0327] Specifically, the client receives the SWLD via the LTE outdoors, and acquires the WLD via the WiFi when entering a facility or the like indoors. Accordingly, the client can acquire more detailed map information indoors.
[0328] In this way, the client can request the WLD or the SWLD from the server according to the frequency band of the network used by the client. Alternatively, the client can transmit information showing the frequency band of the network used by the client to the server, and the server can transmit appropriate data (WLD or SWLD) to the client according to the information. Alternatively, the server can determine the network bandwidth of the client, and transmit appropriate data (WLD or SWLD) to the client.
[0329] Also, the reception of the WLD or the SWLD can be switched according to the moving speed. Figure 15 An example of the operation in this case will be shown. For example, in a case where the client moves at a high speed (S331), the client receives the SWLD from the server (S332). On the other hand, in a case where the client moves at a low speed (S333), the client receives the WLD from the server (S334). Accordingly, the client can suppress the network bandwidth and acquire map information according to the speed. Specifically, the client can update the map information at an appropriate speed by receiving the SWLD having a small amount of data in highway driving. On the other hand, the client can acquire more detailed map information by receiving the WLD in general road driving.
[0330] Thus, the client can request the WLD or the SWLD from the server according to the moving speed of the client. Alternatively, the client can transmit information showing the moving speed of the client to the server, and the server can transmit appropriate data (WLD or SWLD) to the client according to the information. Alternatively, the server can determine the moving speed of the client and transmit appropriate data (WLD or SWLD) to the client.
[0331] Also, the client can first acquire the SWLD from the server and then acquire the WLD of the important region in the SWLD. For example, the client acquires the outline map information in the SWLD when acquiring the map data, and extracts the region in which the features such as buildings, signs, or persons appear more frequently, and then acquires the WLD of the extracted region. Thus, the client can suppress the amount of data received from the server and acquire detailed information of the desired region.
[0332] Also, the server can create the SWLD for each object from the WLD, and the client can receive the SWLD according to the use. Thus, the network bandwidth can be suppressed. For example, the server creates the SWLD of a person and the SWLD of a car by previously recognizing the person or the car from the WLD. The client receives the SWLD of the person when the client wants to acquire information of the person around the client and receives the SWLD of the car when the client wants to acquire information of the car. Also, the types of the SWLD can be distinguished according to the information (sign or type, etc.) attached to the head or the like.
[0333] Next, the configuration of the three-dimensional data encoding apparatus (for example, the server) according to the present embodiment and the flow of the operation will be described. Figure 16 is a block diagram of the three-dimensional data encoding apparatus 400 according to the present embodiment. Figure 17 is a flowchart of the three-dimensional data encoding process performed by the three-dimensional data encoding apparatus 400.
[0334] Figure 16 The three-dimensional data encoding apparatus 400 shown in the figure generates the encoded three-dimensional data 413 and 414 as an encoded stream by encoding the input three-dimensional data 411. Here, the encoded three-dimensional data 413 is the encoded three-dimensional data corresponding to the WLD, and the encoded three-dimensional data 414 is the encoded three-dimensional data corresponding to the SWLD. The three-dimensional data encoding apparatus 400 includes an obtaining section 401, an encoding region determining section 402, a SWLD extracting section 403, a WLD encoding section 404, and a SWLD encoding section 405.
[0335] As shown in Figure 17 First, the obtaining section 401 obtains the input three-dimensional data 411 as the point group data in the three-dimensional space (S401).
[0336] Next, the encoding region decision section 402 decides a spatial region of the encoding target from a spatial region in which the point group data exists (S402).
[0337] Next, the SWLD extraction section 403 defines the spatial region of the encoding target as a WLD, and calculates a feature value from each VXL included in the WLD. Also, the SWLD extraction section 403 extracts a VXL whose feature value is equal to or greater than a predetermined threshold value, defines the extracted VXL as an FVXL, and generates the extraction three-dimensional data 412 by adding the FVXL to the SWLD (S403). That is, the extraction three-dimensional data 412 whose feature value is equal to or greater than the threshold value is extracted from the input three-dimensional data 411.
[0338] Next, the WLD encoding section 404 generates the encoding three-dimensional data 413 corresponding to the WLD by encoding the input three-dimensional data 411 corresponding to the WLD (S404). At this time, the WLD encoding section 404 adds information for distinguishing that the encoding three-dimensional data 413 is a stream including the WLD to a header of the encoding three-dimensional data 413.
[0339] Also, the SWLD encoding section 405 generates the encoding three-dimensional data 414 corresponding to the SWLD by encoding the extraction three-dimensional data 412 corresponding to the SWLD (S405). At this time, the SWLD encoding section 405 adds information for distinguishing that the encoding three-dimensional data 414 is a stream including the SWLD to a header of the encoding three-dimensional data 414.
[0340] Also, the processing order of the processing of generating the encoding three-dimensional data 413 and the processing of generating the encoding three-dimensional data 414 can be reversed from the above. Also, a part or all of the above processing can be performed in parallel.
[0341] The information given to the header of the encoding three-dimensional data 413 and 414 is defined as a parameter such as "world_type", for example. In the case of world_type = 0, it indicates that the stream includes the WLD, and in the case of world_type = 1, it indicates that the stream includes the SWLD. In the case of defining other more categories, the assigned value can be increased like world_type = 2. Also, a specific flag can be included in one of the encoding three-dimensional data 413 and 414. For example, the encoding three-dimensional data 414 can be given a flag showing that the stream includes the SWLD. In this case, the decoding apparatus can discriminate whether it is a stream including the WLD or a stream including the SWLD from the presence or absence of the flag.
[0342] Also, the encoding method used by the WLD encoding section 404 when encoding the WLD can be different from the encoding method used by the SWLD encoding section 405 when encoding the SWLD.
[0343] For example, since the SWLD data is thinned out, the correlation with the surrounding data can be lower than that of the WLD. Therefore, in the encoding method for the SWLD, the inter prediction among the intra prediction and the inter prediction is prioritized compared to the encoding method for the WLD.
[0344] Also, the expression method of the three-dimensional position can be different between the encoding method for the SWLD and the encoding method for the WLD. For example, it can be that the three-dimensional position of the FVXL is expressed by three-dimensional coordinates in the FWLD, the three-dimensional position is expressed by the octree described later in the WLD, and it can be the reverse.
[0345] Also, the SWLD encoding section 405 encodes in such a manner that the data size of the encoded three-dimensional data 414 of the SWLD is smaller than the data size of the encoded three-dimensional data 413 of the WLD. For example, as described above, the correlation between the data can be lower in the SWLD than in the WLD. Accordingly, the encoding efficiency decreases, and the data size of the encoded three-dimensional data 414 can be larger than the data size of the encoded three-dimensional data 413 of the WLD. Therefore, the SWLD encoding section 405, in a case where the data size of the obtained encoded three-dimensional data 414 is larger than the data size of the encoded three-dimensional data 413 of the WLD, generates the encoded three-dimensional data 414 again by performing re-encoding so as to reduce the data size.
[0346] For example, the SWLD extraction section 403 generates the extracted three-dimensional data 412 again in which the number of extracted feature points is reduced, and the SWLD encoding section 405 encodes the extracted three-dimensional data 412. Alternatively, the degree of quantization in the SWLD encoding section 405 can be made coarse. For example, in the octree structure described later, the degree of quantization can be made coarse by rounding the data of the lowest layer.
[0347] Also, the SWLD encoding section 405, in a case where the data size of the encoded three-dimensional data 414 of the SWLD cannot be made smaller than the data size of the encoded three-dimensional data 413 of the WLD, can not generate the encoded three-dimensional data 414 of the SWLD. Alternatively, the encoded three-dimensional data 413 of the WLD can be copied to the encoded three-dimensional data 414 of the SWLD. That is, the encoded three-dimensional data 413 of the WLD can be directly used as the encoded three-dimensional data 414 of the SWLD.
[0348] Next, the configuration of the three-dimensional data decoding apparatus (for example, a client) according to the present embodiment and the flow of the operation will be described. Figure 18 is a block diagram of the three-dimensional data decoding apparatus 500 according to the present embodiment. Figure 19 is a flowchart of the three-dimensional data decoding process performed by the three-dimensional data decoding apparatus 500.
[0349] Figure 18 The illustrated 3D data decoding apparatus 500 generates decoded 3D data 512 or 513 by decoding encoded 3D data 511. Here, encoded 3D data 511 is, for example, encoded 3D data 413 or 414 generated by the 3D data encoding apparatus 400.
[0350] The 3D data decoding device 500 includes: an acquisition unit 501, a head parsing unit 502, a WLD decoding unit 503, and a SWLD decoding unit 504.
[0351] like Figure 19 As shown, firstly, the acquisition unit 501 acquires the encoded three-dimensional data 511 (S501). Next, the header parsing unit 502 parses the header of the encoded three-dimensional data 511 and determines whether the encoded three-dimensional data 511 is a stream containing WLD or a stream containing SWLD (S502). For example, the determination is made by referring to the world_type parameter mentioned above.
[0352] If the encoded 3D data 511 is a stream containing WLD (S503 "Yes"), the WLD decoding unit 503 decodes the encoded 3D data 511 to generate decoded 3D data 512 of WLD (S504). Alternatively, if the encoded 3D data 511 is a stream containing SWLD (S503 "No"), the SWLD decoding unit 504 decodes the encoded 3D data 511 to generate decoded 3D data 513 of SWLD (S505).
[0353] Furthermore, similar to the encoding apparatus, the decoding method used by the WLD decoding unit 503 when decoding the WLD can be different from the decoding method used by the SWLD decoding unit 504 when decoding the SWLD. For example, in the decoding method for the SWLD, compared with the decoding method for the WLD, inter-frame prediction in intra-frame prediction and inter-frame prediction can be prioritized.
[0354] Furthermore, the methods used to represent the three-dimensional position can differ between the decoding methods used for SWLD and WLD. For example, SWLD can represent the three-dimensional position of FVXL using three-dimensional coordinates, while WLD can represent the three-dimensional position using an octree (described later), and vice versa.
[0355] Next, the octree representation as a method of representing three-dimensional location will be explained. The VXL data contained in the three-dimensional data is converted into an octree structure and then encoded. Figure 20 An example of VXL for WLD is shown. Figure 21 It shows Figure 20 The octree structure of WLD is shown.Figure 20 In the example shown, there are three VXL1 to VXL3 that constitute the VXL (hereinafter, valid VXL) of the point group. Figure 21 As shown, the octree structure consists of nodes and leaf nodes. Each node has a maximum of 8 nodes or leaf nodes. Each leaf node contains VXL information. Here, Figure 21 Among the leaf nodes shown, leaf nodes 1, 2, and 3 respectively represent Figure 20 VXL1, VXL2, and VXL3 are shown.
[0356] Specifically, each node and leaf node corresponds to a 3D position. Node 1 and... Figure 20 All the blocks shown correspond to each other. The block corresponding to node 1 is divided into 8 blocks. Among these 8 blocks, the block with valid VXL is set as a node, and the rest are set as leaf nodes. The block corresponding to a node is further divided into 8 nodes or leaf nodes, and this process is repeated the same number of times as the level in the tree structure. Furthermore, all the blocks at the bottom level are set as leaf nodes.
[0357] and, Figure 22 It shows from Figure 20 The example shown is a SWLD generated from a WLD. Figure 20 The feature extraction results of VXL1 and VXL2 shown are identified as FVXL1 and FVXL2 and added to SWLD. VXL3, however, is not identified as FVXL and therefore is not included in SWLD. Figure 23 It shows Figure 22 The octree structure of SWLD is shown. Figure 23 In the octree structure shown, Figure 21 Leaf node 3, which corresponds to VXL3, is deleted. Accordingly, Figure 21 Node 3 shown does not have a valid VXL and has been changed to a leaf node. Thus, generally speaking, the number of leaf nodes in SWLD is less than that in WLD, and the encoded 3D data of SWLD is also smaller than that of WLD.
[0358] The following describes variations of this embodiment.
[0359] For example, when a vehicle-mounted device or other client is estimating its own position, it receives a SWLD from the server, uses the SWLD to estimate its own position, and performs obstacle detection. Then, it uses various methods such as distance sensors such as rangefinders, stereo cameras, or combinations of multiple monocular cameras to perform obstacle detection based on the three-dimensional information of the surrounding environment it has obtained.
[0360] Also, in general, it is difficult to include VXL data of a flat region in the SWLD. For this reason, the server holds a subsampled world space (SubWLD) that is subsampled from the WLD for detection of stationary obstacles, and can transmit the SWLD and the SubWLD to the client. Thereby, it is possible to suppress the network bandwidth, and to perform self-position estimation and obstacle detection on the client side.
[0361] Also, when the client rapidly draws three-dimensional map data, it is convenient if the map information is in a grid structure. Therefore, the server can generate a grid from the WLD in advance, and hold it as a grid world space (MWLD). For example, the MWLD is transmitted when the client needs to perform rough three-dimensional drawing, and the WLD is transmitted when the client needs to perform detailed three-dimensional drawing. Thereby, it is possible to suppress the network bandwidth.
[0362] Also, although the server sets VXLs in which the feature amount is equal to or greater than a threshold as FVXLs from each VXL, the FVXLs can be calculated by a different method. For example, the server can include VXLs, VLMs, SPCs, or GOSs that constitute a signal or a cross point, etc., as FVXLs, FVLMs, FSPCs, or FGOSs in the SWLD when they are determined to be needed in self-position estimation, driving assistance, or autonomous driving, etc. Also, the determination can be performed manually. In addition, the FVXLs, etc., obtained by the above method can be added to the FVXLs, etc., set based on the feature amount. That is, the SWLD extraction unit 403 can further extract data corresponding to an object having a predetermined attribute as the extracted three-dimensional data 412 from the input three-dimensional data 411.
[0363] Also, a label different from the feature amount can be given to a situation in which it is needed for these uses. The server can hold FVXLs needed in self-position estimation, driving assistance, or autonomous driving, etc., of a signal or a cross point, etc., as a higher layer (for example, a lane world space) of the SWLD.
[0364] Also, the server can attach an attribute to a VXL within the WLD in a random access unit or a predetermined unit. The attribute includes, for example, information showing whether it is needed or not needed in self-position estimation, or information showing whether traffic information such as a signal or a cross point is important, etc. Also, the attribute can include a correspondence relationship with a Feature (a cross point or a road, etc.) in lane information (GDF: Geographic Data Files, etc.).
[0365] Also, as a method of updating the WLD or the SWLD, the following method can be employed.
[0366] Update information such as a change in a person, construction, or a street tree (facing a track) is loaded to the server as a point group or metadata. The server updates the WLD according to the load, and after that, the SWLD is updated with the updated WLD.
[0367] Also, in a case where a mismatch between the three-dimensional information generated by itself and the three-dimensional information received from the server is detected at the time of self-position estimation by the client, the three-dimensional information generated by itself can be transmitted to the server together with the update notification. In this case, the server updates the SWLD with the WLD. In a case where the SWLD is not updated, the server judges that the WLD itself is old.
[0368] Also, as the header information of the encoded stream, although information for distinguishing whether it is the WLD or the SWLD is added, in a case where a plurality of world spaces such as a mesh world space or a lane world space exist, information for distinguishing them can be added to the header information. Also, in a case where a plurality of SWLDs exist which differ in the feature amount, information for distinguishing them respectively can be added to the header information.
[0369] Also, although the SWLD is constituted by the FVXL, it can include a VXL which is not judged as the FVXL. For example, the SWLD can include an adjoining VXL used when the feature amount of the FVXL is calculated. According to this, even in a case where the feature amount information is not added to each FVXL of the SWLD, the client can calculate the feature amount of the FVXL at the time of receiving the SWLD. Also at this time, the SWLD can include information for distinguishing whether each VXL is the FVXL or the VXL.
[0370] As described above, the three-dimensional data encoding device 400 extracts the extracted three-dimensional data 412 (2nd three-dimensional data) in which the feature amount is equal to or greater than the threshold value from the input three-dimensional data 411 (1st three-dimensional data), and generates the encoded three-dimensional data 414 (1st encoded three-dimensional data) by encoding the extracted three-dimensional data 412.
[0371] According to this, the three-dimensional data encoding device 400 generates the encoded three-dimensional data 414 obtained by encoding the data in which the feature amount is equal to or greater than the threshold value. Thus, compared to a case where the input three-dimensional data 411 is directly encoded, it is possible to reduce the data amount. Therefore, the three-dimensional data encoding device 400 can reduce the data amount at the time of transmission.
[0372] Also, the three-dimensional data encoding device 400 further generates the encoded three-dimensional data 413 (2nd encoded three-dimensional data) by encoding the input three-dimensional data 411.
[0373] Accordingly, the three-dimensional data encoding device 400 can selectively transmit the encoded three-dimensional data 413 and the encoded three-dimensional data 414, for example, in accordance with a use or the like.
[0374] Further, the extracted three-dimensional data 412 is encoded by the first encoding method, and the input three-dimensional data 411 is encoded by a second encoding method different from the first encoding method.
[0375] Accordingly, the three-dimensional data encoding device 400 can employ appropriate encoding methods for the input three-dimensional data 411 and the extracted three-dimensional data 412.
[0376] Further, in the first encoding method, inter prediction among intra prediction and inter prediction is prioritized compared to the second encoding method.
[0377] Accordingly, the three-dimensional data encoding device 400 can increase the priority of inter prediction for the extracted three-dimensional data 412 for which the correlation between adjacent data easily becomes low.
[0378] Further, in the first encoding method and the second encoding method, the expression method of the three-dimensional position is different. For example, in the second encoding method, the three-dimensional position is expressed by an octree, and in the first encoding method, the three-dimensional position is expressed by three-dimensional coordinates.
[0379] Accordingly, the three-dimensional data encoding device 400 can employ a more appropriate expression method of the three-dimensional position for three-dimensional data in which the number of data (the number of VXL or FVXL) is different.
[0380] Further, at least one of the encoded three-dimensional data 413 and 414 includes an identifier that shows whether the encoded three-dimensional data is encoded three-dimensional data obtained by encoding the input three-dimensional data 411 or encoded three-dimensional data obtained by encoding a part of the input three-dimensional data 411. That is, the identifier shows whether the encoded three-dimensional data is the encoded three-dimensional data 413 of the WLD or the encoded three-dimensional data 414 of the SWLD.
[0381] Accordingly, the decoding device can easily determine whether the acquired encoded three-dimensional data is the encoded three-dimensional data 413 or the encoded three-dimensional data 414.
[0382] Further, the three-dimensional data encoding device 400 encodes the extracted three-dimensional data 412 in such a manner that the amount of data of the encoded three-dimensional data 414 is less than the amount of data of the encoded three-dimensional data 413.
[0383] Accordingly, the three-dimensional data encoding device 400 can make the amount of data of the encoded three-dimensional data 414 less than the amount of data of the encoded three-dimensional data 413.
[0384] Further, the three-dimensional data encoding device 400 further extracts, from the input three-dimensional data 411, data corresponding to an object having a predetermined attribute, as the extracted three-dimensional data 412. The object having the predetermined attribute is, for example, an object required in self-position estimation, driving assistance, or autonomous driving, or a signal or a crossing.
[0385] Accordingly, the three-dimensional data encoding device 400 can generate the encoded three-dimensional data 414 including data required by the decoding device.
[0386] Further, the three-dimensional data encoding device 400 (server) further transmits one of the encoded three-dimensional data 413 and 414 to the client in accordance with a state of the client.
[0387] Accordingly, the three-dimensional data encoding device 400 can transmit appropriate data in accordance with the state of the client.
[0388] Further, the state of the client includes a communication condition (for example, a network bandwidth) of the client or a moving speed of the client.
[0389] Further, the three-dimensional data encoding device 400 further transmits one of the encoded three-dimensional data 413 and 414 to the client in accordance with a request of the client.
[0390] Accordingly, the three-dimensional data encoding device 400 can transmit appropriate data in accordance with the request of the client.
[0391] Further, the three-dimensional data decoding device 500 according to the present embodiment decodes the encoded three-dimensional data 413 or 414 generated by the three-dimensional data encoding device 400 described above.
[0392] That is, the three-dimensional data decoding device 500 decodes the encoded three-dimensional data 414 in which the extracted three-dimensional data 412 in which the feature amount extracted from the input three-dimensional data 411 is equal to or greater than a threshold value is encoded, by a first decoding method. Further, the three-dimensional data decoding device 500 decodes the encoded three-dimensional data 413 in which the input three-dimensional data 411 is encoded, by a second decoding method different from the first decoding method.
[0393] Accordingly, the three-dimensional data decoding device 500 can selectively receive the encoded three-dimensional data 414 in which data having a feature amount equal to or greater than a threshold value is encoded and the encoded three-dimensional data 413, for example, in accordance with a use or the like. Accordingly, the three-dimensional data decoding device 500 can reduce the amount of data at the time of transmission. Further, the three-dimensional data decoding device 500 can employ appropriate decoding methods with respect to the input three-dimensional data 411 and the extracted three-dimensional data 412, respectively.
[0394] Also, in the first decoding method, inter prediction among intra prediction and inter prediction is prioritized compared to the second decoding method.
[0395] Accordingly, the three-dimensional data decoding apparatus 500 can improve the priority of inter prediction for three-dimensional data for which the correlation between adjacent data easily becomes low.
[0396] Also, in the first decoding method and the second decoding method, the method of expressing three-dimensional positions is different. For example, in the second decoding method, three-dimensional positions are expressed by an octree, and in the first decoding method, three-dimensional positions are expressed by three-dimensional coordinates.
[0397] Accordingly, the three-dimensional data decoding apparatus 500 can adopt a more appropriate method of expressing three-dimensional positions for three-dimensional data for which the number of data (the number of VXL or FVXL) is different.
[0398] Also, at least one of the encoded three-dimensional data 413 and 414 includes an identifier that shows whether the encoded three-dimensional data is encoded three-dimensional data obtained by encoding the input three-dimensional data 411 or encoded three-dimensional data obtained by encoding a part of the input three-dimensional data 411. The three-dimensional data decoding apparatus 500 refers to the identifier to identify the encoded three-dimensional data 413 and 414.
[0399] Accordingly, the three-dimensional data decoding apparatus 500 can easily determine whether the obtained encoded three-dimensional data is the encoded three-dimensional data 413 or the encoded three-dimensional data 414.
[0400] Also, the three-dimensional data decoding apparatus 500 further notifies the server of the state of the client (the three-dimensional data decoding apparatus 500). The three-dimensional data decoding apparatus 500 receives one of the encoded three-dimensional data 413 and 414 transmitted from the server in accordance with the state of the client.
[0401] Accordingly, the three-dimensional data decoding apparatus 500 can receive appropriate data in accordance with the state of the client.
[0402] Also, the state of the client includes the communication status (for example, network bandwidth) of the client or the moving speed of the client.
[0403] Also, the three-dimensional data decoding apparatus 500 further requests one of the encoded three-dimensional data 413 and 414 from the server, and receives one of the encoded three-dimensional data 413 and 414 transmitted from the server in accordance with the request.
[0404] Accordingly, the three-dimensional data decoding apparatus 500 can receive appropriate data corresponding to the use.
[0405] (Embodiment 3)
[0406] In the present embodiment, a method of transmitting and receiving three-dimensional data between vehicles is described. For example, three-dimensional data is transmitted and received between a host vehicle and surrounding vehicles.
[0407] Figure 24 is a block diagram of a three-dimensional data creating apparatus 620 according to the present embodiment. The three-dimensional data creating apparatus 620 is included in, for example, a host vehicle, and creates denser third three-dimensional data 636 by synthesizing the received second three-dimensional data 635 with first three-dimensional data 632 created by the three-dimensional data creating apparatus 620.
[0408] The three-dimensional data creating apparatus 620 includes a three-dimensional data creating section 621, a request range determining section 622, a searching section 623, a receiving section 624, a decoding section 625, and a synthesizing section 626.
[0409] First, the three-dimensional data creating section 621 creates the first three-dimensional data 632 using sensor information 631 detected by a sensor included in the host vehicle. Next, the request range determining section 622 determines a request range, which is a range of three-dimensional space in which data is insufficient in the created first three-dimensional data 632.
[0410] Next, the searching section 623 searches for surrounding vehicles that hold three-dimensional data of the request range, and transmits request range information 633 indicating the request range to the surrounding vehicles determined by the search. Next, the receiving section 624 receives encoded three-dimensional data 634, which is an encoded stream of the request range, from the surrounding vehicles (S624). Alternatively, the searching section 623 can issue a request to all vehicles existing in the determined range without discrimination, and receive the encoded three-dimensional data 634 from the responding counterpart. Further, the searching section 623 can issue a request to an object such as a signal or a sign, and receive the encoded three-dimensional data 634 from the object, instead of a vehicle.
[0411] Next, the received encoded three-dimensional data 634 is decoded by the decoding section 625 to obtain second three-dimensional data 635. Next, the first three-dimensional data 632 and the second three-dimensional data 635 are synthesized by the synthesizing section 626 to create denser third three-dimensional data 636.
[0412] Next, the configuration and operation of a three-dimensional data transmitting apparatus 640 according to the present embodiment are described. Figure 25 is a block diagram of the three-dimensional data transmitting apparatus 640.
[0413] The three-dimensional data transmitting device 640, for example, included in the surrounding vehicle described above, processes the fifth three-dimensional data 652 made by the surrounding vehicle into the sixth three-dimensional data 654 requested by the own vehicle, and generates the encoded three-dimensional data 634 by encoding the sixth three-dimensional data 654, and transmits the encoded three-dimensional data 634 to the own vehicle.
[0414] The three-dimensional data transmitting device 640 includes a three-dimensional data making section 641, a receiving section 642, an extracting section 643, an encoding section 644, and a transmitting section 645.
[0415] First, the three-dimensional data making section 641 makes the fifth three-dimensional data 652 using the sensor information 651 detected by the sensor included in the surrounding vehicle. Next, the receiving section 642 receives the request range information 633 transmitted from the own vehicle.
[0416] Next, the extracting section 643 extracts the three-dimensional data of the request range indicated by the request range information 633 from the fifth three-dimensional data 652, and processes the fifth three-dimensional data 652 into the sixth three-dimensional data 654. Next, the encoding section 644 encodes the sixth three-dimensional data 654, thereby generating the encoded three-dimensional data 634 as an encoded stream. Then, the transmitting section 645 transmits the encoded three-dimensional data 634 to the own vehicle.
[0417] In addition, here, although the example in which the own vehicle includes the three-dimensional data making device 620 and the surrounding vehicle includes the three-dimensional data transmitting device 640 is described, each vehicle can have the functions of the three-dimensional data making device 620 and the three-dimensional data transmitting device 640.
[0418] (Embodiment 4)
[0419] In this embodiment, the operation regarding the abnormal situation in the own position estimation based on the three-dimensional map is described.
[0420] The use of mobile bodies such as an autonomous vehicle, a robot, or a flying object such as a drone will expand in the future. As one example of a method of realizing such autonomous movement, there is a method in which a mobile body estimates its position in a three-dimensional map (own position estimation) while traveling according to the map.
[0421] The own position estimation is realized by matching a three-dimensional map and three-dimensional information of the surroundings of the own vehicle (hereinafter referred to as own vehicle detection three-dimensional data) obtained by a sensor such as a range finder (LIDAR, etc.) or a stereo camera mounted on the own vehicle, and estimating the position of the own vehicle in the three-dimensional map.
[0422] A three-dimensional map such as an HD map proposed by the company HERE is not only a three-dimensional point cloud but also can include two-dimensional map data of shapes of roads and intersections, or information that changes in real time such as congestion and accidents. A three-dimensional map is constituted of a plurality of levels of three-dimensional data, two-dimensional data, and metadata that changes in real time, and a device can obtain only necessary data or can refer to necessary data.
[0423] Data of a point cloud can be SWLD described above or can include point group data that is not a feature point. Also, reception and transmission of data of a point cloud is basically performed in one or a plurality of random access units.
[0424] As a matching method of a three-dimensional map and three-dimensional data detected by a self vehicle, the following method can be adopted. For example, a device compares shapes of point groups in point clouds of each other and determines a portion in which similarity between feature points is high as the same position. Also, in a case where a three-dimensional map is constituted of SWLD, a device compares and matches feature points constituting SWLD and three-dimensional feature points extracted from three-dimensional data detected by a self vehicle.
[0425] Here, in order to perform self position estimation with high accuracy, (A) that three-dimensional map and three-dimensional data detected by a self vehicle have been obtained and (B) that their accuracies satisfy a predetermined criterion need to be satisfied. However, in the following abnormal cases, (A) or (B) cannot be satisfied.
[0426] (1) A three-dimensional map cannot be obtained through a communication path.
[0427] (2) A three-dimensional map does not exist or an obtained three-dimensional map is damaged.
[0428] (3) A sensor of a self vehicle malfunctions or generation accuracy of three-dimensional data detected by a self vehicle is not sufficient due to bad weather.
[0429] The following describes operations for coping with these abnormal cases. The following describes operations of a vehicle as an example, and the following method can be applied to all moving objects that perform autonomous movement such as robots and drones.
[0430] The following describes a configuration and operations of a three-dimensional information processing device according to the present embodiment for coping with abnormal cases in a three-dimensional map or three-dimensional data detected by a self vehicle. Figure 26 is a block diagram showing a configuration example of a three-dimensional information processing device 700 according to the present embodiment.
[0431] The three-dimensional information processing device 700 is mounted on a moving object such as a motor vehicle, for example. As shown in Figure 26As shown, the three-dimensional information processing apparatus 700 includes a three-dimensional map obtaining section 701, a self-vehicle detection data obtaining section 702, an abnormal situation judging section 703, a countermeasure work deciding section 704, and a work control section 705.
[0432] In addition, the three-dimensional information processing apparatus 700 can include a camera that obtains a two-dimensional image, or can include a two-dimensional or one-dimensional sensor that detects a structure object or a moving object around the self-vehicle, such as a sensor that uses one-dimensional data of an ultrasonic wave or a laser, which is not shown. Further, the three-dimensional information processing apparatus 700 can include a communication section (not shown) that obtains a three-dimensional map through a mobile communication network such as 4G or 5G, or a communication between vehicles or a communication between a road and a vehicle.
[0433] The three-dimensional map obtaining section 701 obtains a three-dimensional map 711 around a travel route. For example, the three-dimensional map obtaining section 701 obtains the three-dimensional map 711 through a mobile communication network, or a communication between vehicles or a communication between a road and a vehicle.
[0434] Next, the self-vehicle detection data obtaining section 702 obtains self-vehicle detection three-dimensional data 712 based on sensor information. For example, the self-vehicle detection data obtaining section 702 generates the self-vehicle detection three-dimensional data 712 based on sensor information obtained by a sensor included in the self-vehicle.
[0435] Next, the abnormal situation judging section 703 detects an abnormal situation by performing a predetermined check on at least one of the obtained three-dimensional map 711 and the self-vehicle detection three-dimensional data 712. That is, the abnormal situation judging section 703 judges whether at least one of the obtained three-dimensional map 711 and the self-vehicle detection three-dimensional data 712 is abnormal.
[0436] In a case where an abnormal situation is detected, the countermeasure work deciding section 704 decides a countermeasure work for the abnormal situation. Next, the work control section 705 controls the work of each processing section required in the implementation of the countermeasure work, such as the three-dimensional map obtaining section 701.
[0437] In addition, in a case where an abnormal situation is not detected, the three-dimensional information processing apparatus 700 ends the processing.
[0438] Further, the three-dimensional information processing apparatus 700 estimates a self-position of a vehicle including the three-dimensional information processing apparatus 700 using the three-dimensional map 711 and the self-vehicle detection three-dimensional data 712. Next, the three-dimensional information processing apparatus 700 causes the vehicle to perform automatic driving using a result of the self-position estimation.
[0439] Accordingly, the three-dimensional information processing apparatus 700 obtains map data (three-dimensional map 711) including first three-dimensional position information via a channel. For example, the first three-dimensional position information is encoded in units of partial spaces having coordinate information in three dimensions, and the first three-dimensional position information includes a plurality of random access units each of which is a collection of one or more partial spaces and can be independently decoded. For example, the first three-dimensional position information is data (SWLD) in which a feature point in which a feature amount in three dimensions becomes equal to or greater than a predetermined threshold is encoded.
[0440] Also, the three-dimensional information processing apparatus 700 generates second three-dimensional position information (self-vehicle detection three-dimensional data 712) based on information detected by a sensor. Next, the three-dimensional information processing apparatus 700 determines whether the first three-dimensional position information or the second three-dimensional position information is abnormal by performing abnormality determination processing on the first three-dimensional position information or the second three-dimensional position information.
[0441] The three-dimensional information processing apparatus 700 determines a countermeasure against the abnormality when the first three-dimensional position information or the second three-dimensional position information is determined to be abnormal. Next, the three-dimensional information processing apparatus 700 performs control required when the countermeasure is implemented.
[0442] Accordingly, the three-dimensional information processing apparatus 700 can detect an abnormality in the first three-dimensional position information or the second three-dimensional position information and can perform a countermeasure.
[0443] (Embodiment 5)
[0444] In this embodiment, a three-dimensional data transmission method and the like for a rear vehicle is described.
[0445] Figure 27 is a block diagram showing a configuration example of a three-dimensional data production apparatus 810 according to the present embodiment. The three-dimensional data production apparatus 810 is mounted on a vehicle, for example. The three-dimensional data production apparatus 810 performs transmission and reception of three-dimensional data with an external traffic cloud monitor, a front vehicle, or a rear vehicle, and produces and accumulates three-dimensional data.
[0446] The three-dimensional data production apparatus 810 includes a data reception section 811, a communication section 812, a reception control section 813, a format conversion section 814, a plurality of sensors 815, a three-dimensional data production section 816, a three-dimensional data synthesis section 817, a three-dimensional data accumulation section 818, a communication section 819, a transmission control section 820, a format conversion section 821, and a data transmission section 822.
[0447] The data reception section 811 receives three-dimensional data 831 from the traffic cloud monitor or the preceding vehicle. The three-dimensional data 831 includes, for example, a point cloud including information on an area that the sensor 815 of the own vehicle cannot detect, a visible light image, depth information, sensor position information, or speed information.
[0448] The communication section 812 communicates with the traffic cloud monitor or the preceding vehicle, and transmits a data transmission request or the like to the traffic cloud monitor or the preceding vehicle.
[0449] The reception control section 813 establishes communication with the communication partner by exchanging corresponding format information or the like with the communication partner via the communication section 812.
[0450] The format conversion section 814 generates three-dimensional data 832 by performing format conversion or the like on the three-dimensional data 831 received by the data reception section 811. Also, the format conversion section 814 performs decompression or decoding processing in the case where the three-dimensional data 831 is compressed or encoded.
[0451] The plurality of sensors 815 are a group of sensors that obtain information on the outside of the vehicle, such as a LiDAR, a visible light camera, or an infrared camera, and generate sensor information 833. For example, in the case where the sensor 815 is a laser sensor such as a LiDAR, the sensor information 833 is three-dimensional data such as a point cloud (point group data). Also, the sensor 815 can not be plural.
[0452] The three-dimensional data production section 816 generates three-dimensional data 834 from the sensor information 833. The three-dimensional data 834 includes, for example, information on a point cloud, a visible light image, depth information, sensor position information, or speed information.
[0453] The three-dimensional data synthesis section 817 synthesizes the three-dimensional data 832 produced by the traffic cloud monitor or the preceding vehicle or the like into the three-dimensional data 834 produced from the sensor information 833 of the own vehicle, and thereby can construct three-dimensional data 835 that includes a space in front of the preceding vehicle of the own vehicle that cannot be detected by the sensor 815 of the own vehicle.
[0454] The three-dimensional data accumulation section 818 accumulates the generated three-dimensional data 835 or the like.
[0455] The communication section 819 communicates with the traffic cloud monitor or the preceding vehicle, and transmits a data transmission request or the like to the traffic cloud monitor or the preceding vehicle.
[0456] The transmission control unit 820 exchanges corresponding format and other information with the communication counterpart via the communication unit 819 to establish communication with the communication counterpart. Furthermore, the transmission control unit 820 determines the transmission area of the spatial dimension of the dimension data to be transmitted based on the dimension data construction information of the dimension data 832 generated in the dimension data synthesis unit 817 and the data transmission request from the communication counterpart.
[0457] Specifically, the transmission control unit 820 determines the transmission area, including the space in front of its own vehicle that cannot be detected by the sensors of the rear vehicle, based on data transmission requests from traffic cloud monitoring or rear vehicles. Furthermore, the transmission control unit 820 determines the transmission area by judging whether there are updates to the space that can be transmitted or the space that has already been transmitted, based on three-dimensional data construction information. For example, the transmission control unit 820 determines the transmission area as the area that is both specified by the data transmission request and where the corresponding three-dimensional data 835 exists. The transmission control unit 820 also notifies the format conversion unit 821 of the format corresponding to the communication counterpart and the transmission area.
[0458] The format conversion unit 821 generates three-dimensional data 837 by converting the three-dimensional data 836 of the transmission area stored in the three-dimensional data 835 of the three-dimensional data storage unit 818 into a format corresponding to the receiving side. Alternatively, the format conversion unit 821 can compress or encode the three-dimensional data 837 to reduce the data volume.
[0459] The data transmission unit 822 sends the three-dimensional data 837 to the traffic cloud monitoring or the vehicle behind. The three-dimensional data 837 may include, for example, point clouds of the front of the vehicle containing information about areas that would become blind spots for the vehicle behind, visible light images, depth information, or sensor position information.
[0460] Furthermore, although the format conversion units 814 and 821 are used as examples for format conversion, it is also possible to omit format conversion.
[0461] With this configuration, the 3D data generation apparatus 810 obtains 3D data 831 from an external source in an area that cannot be detected by the vehicle's own sensors 815, and generates 3D data 835 by combining the 3D data 831 with 3D data 834 based on sensor information 833 detected by the vehicle's own sensors 815. Accordingly, the 3D data generation apparatus 810 is capable of generating 3D data for a range that cannot be detected by the vehicle's own sensors 815.
[0462] Furthermore, the 3D data generation device 810 can send 3D data, including the space in front of its own vehicle that cannot be detected by the sensors of the rear vehicle, to the traffic cloud monitoring or rear vehicle according to data transmission requests from traffic cloud monitoring or rear vehicles.
[0463] (Embodiment 6)
[0464] An example to be described in Embodiment 5 is that a client device such as a vehicle transmits three-dimensional data to another vehicle or a server such as a traffic cloud monitor. In the present embodiment, the client device transmits sensor information obtained by a sensor to a server or another client device.
[0465] First, the configuration of a system to which the present embodiment is applied will be described. Figure 28 The configuration of a three-dimensional map and a system for transmitting and receiving sensor information to which the present embodiment is applied is shown. The system includes a server 901, client devices 902A and 902B. In addition, in the case where the client devices 902A and 902B are not specially distinguished, it is also written as a client device 902.
[0466] The client device 902 is, for example, an in-vehicle device mounted on a mobile body such as a vehicle. The server 901 is, for example, a traffic cloud monitor or the like, and can communicate with a plurality of client devices 902.
[0467] The server 901 transmits a three-dimensional map constituted by a point cloud to the client device 902. In addition, the configuration of the three-dimensional map is not limited to a point cloud, and can be expressed by other three-dimensional data such as a mesh structure.
[0468] The client device 902 transmits sensor information obtained by the client device 902 to the server 901. The sensor information includes, for example, at least one of LiDAR obtained information, a visible light image, an infrared image, a depth image, sensor position information, and speed information.
[0469] As for data transmitted and received between the server 901 and the client device 902, it can be compressed in the case where it is desired to reduce data, and can not be compressed in the case where it is desired to maintain the accuracy of data. In the case where data is compressed, for example, a three-dimensional compression method based on an octree can be employed in a point cloud. Also, a two-dimensional image compression method can be employed in a visible light image, an infrared image, and a depth image. The two-dimensional image compression method is, for example, MPEG-4 AVC or HEVC standardized by MPEG or the like.
[0470] Also, the server 901 transmits the three-dimensional map managed by the server 901 to the client device 902 in accordance with a transmission request of the three-dimensional map from the client device 902. In addition, the server 901 can transmit the three-dimensional map without waiting for the transmission request of the three-dimensional map from the client device 902. For example, the server 901 can broadcast the three-dimensional map to one or more client devices 902 in a predetermined space. Also, the server 901 can transmit the three-dimensional map suitable for the position of the client device 902 to the client device 902 at a predetermined time interval after the client device 902 has made a transmission request once. Also, the server 901 can transmit the three-dimensional map to the client device 902 every time the three-dimensional map managed by the server 901 is updated.
[0471] The client device 902 makes a transmission request of the three-dimensional map to the server 901. For example, in a case where the client device 902 wants to perform self-position estimation while traveling, the client device 902 makes a transmission request of the three-dimensional map to the server 901.
[0472] In addition, the client device 902 can make a transmission request of the three-dimensional map to the server 901 in the following cases. The client device 902 can make a transmission request of the three-dimensional map to the server 901 in a case where the three-dimensional map held by the client device 902 is old. For example, in a case where the client device 902 has obtained the three-dimensional map a predetermined period of time ago, the client device 902 can make a transmission request of the three-dimensional map to the server 901.
[0473] Also, the client device 902 can make a transmission request of the three-dimensional map to the server 901 before a predetermined time at which the client device 902 is to exit a space shown by the three-dimensional map held by the client device 902. For example, the client device 902 can make a transmission request of the three-dimensional map to the server 901 in a case where the client device 902 is present within a predetermined distance from a boundary of the space shown by the three-dimensional map held by the client device 902. Also, in a case where the movement path and the movement speed of the client device 902 are grasped, the time at which the client device 902 exits the space shown by the three-dimensional map held by the client device 902 can be predicted from the grasped movement path and movement speed.
[0474] The client device 902 can make a transmission request of the three-dimensional map to the server 901 in a case where an error in position matching between the three-dimensional data produced by the client device 902 from sensor information and the three-dimensional map is equal to or greater than a predetermined range.
[0475] The client device 902 transmits the sensor information to the server 901 in accordance with a transmission request of the sensor information from the server 901. Alternatively, the client device 902 can transmit the sensor information to the server 901 without waiting for the transmission request of the sensor information from the server 901. For example, the client device 902 can periodically transmit the sensor information to the server 901 within a certain period after receiving the transmission request of the sensor information from the server 901. Alternatively, when the error in the position of the three-dimensional data created by the client device 902 based on the sensor information with respect to the three-dimensional map obtained from the server 901 is equal to or greater than a certain range, the client device 902 can determine that there is a possibility that the three-dimensional map around the client device 902 has changed, and transmit the determination result and the sensor information to the server 901.
[0476] The server 901 transmits the transmission request of the sensor information to the client device 902. For example, the server 901 receives the position information of the client device 902 such as GPS from the client device 902. The server 901 transmits the transmission request of the sensor information to the client device 902 in order to regenerate the three-dimensional map when the server 901 determines that the client device 902 approaches a space in which the information in the three-dimensional map managed by the server 901 is small, based on the position information of the client device 902. Alternatively, the server 901 can transmit the transmission request of the sensor information when the server 901 wants to update the three-dimensional map, when the server 901 wants to confirm the road conditions at the time of snowfall or disaster, or the like, or when the server 901 wants to confirm the congestion conditions or the event accident conditions, or the like.
[0477] Alternatively, the client device 902 can set the data amount of the sensor information transmitted to the server 901 in accordance with the communication state or the frequency band at the time of receiving the transmission request of the sensor information from the server 901. The setting of the data amount of the sensor information transmitted to the server 901 refers to, for example, increasing or decreasing the data itself, or selecting an appropriate compression method.
[0478] Figure 29 is a block diagram illustrating an example of the configuration of the client device 902. The client device 902 receives the three-dimensional map configured by the point cloud or the like from the server 901, and estimates the own position of the client device 902 based on the three-dimensional data created based on the sensor information of the client device 902. The client device 902 transmits the obtained sensor information to the server 901.
[0479] The client device 902 is provided with a data reception section 1011, a communication section 1012, a reception control section 1013, a format conversion section 1014, a plurality of sensors 1015, a three-dimensional data production section 1016, a three-dimensional image processing section 1017, a three-dimensional data accumulation section 1018, a format conversion section 1019, a communication section 1020, a transmission control section 1021, and a data transmission section 1022.
[0480] The data reception section 1011 receives a three-dimensional map 1031 from the server 901. The three-dimensional map 1031 is data including a point cloud of a WLD or a SWLD or the like. The three-dimensional map 1031 can include either compressed data or uncompressed data.
[0481] The communication section 1012 communicates with the server 901 and transmits a data transmission request (for example, a three-dimensional map transmission request) or the like to the server 901.
[0482] The reception control section 1013 establishes communication with a communication partner by exchanging information on a corresponding format or the like with the communication partner via the communication section 1012.
[0483] The format conversion section 1014 generates a three-dimensional map 1032 by performing format conversion or the like on the three-dimensional map 1031 received by the data reception section 1011. Also, the format conversion section 1014 performs decompression or decoding processing in the case where the three-dimensional map 1031 is compressed or encoded. In addition, the format conversion section 1014 does not perform decompression or decoding processing in the case where the three-dimensional map 1031 is uncompressed data.
[0484] The plurality of sensors 1015 are a group of sensors mounted on the client device 902 for obtaining information on the outside of the vehicle, such as a LiDAR, a visible light camera, an infrared camera, or a depth sensor, and generate sensor information 1033. For example, in the case where the sensor 1015 is a laser sensor such as a LiDAR, the sensor information 1033 is three-dimensional data such as a point cloud (point group data). In addition, the sensor 1015 can not be plural.
[0485] The three-dimensional data production section 1016 produces three-dimensional data 1034 of the surroundings of the own vehicle on the basis of the sensor information 1033. For example, the three-dimensional data production section 1016 produces point cloud data having color information of the surroundings of the own vehicle using information obtained by a LiDAR and a visible light image obtained by a visible light camera.
[0486] The three-dimensional image processing section 1017 performs self-position estimation processing and the like of the own vehicle using the received three-dimensional map 1032 and three-dimensional data 1034 of the surroundings of the own vehicle generated from the sensor information 1033. In addition, the three-dimensional image processing section 1017 can synthesize the three-dimensional map 1032 and the three-dimensional data 1034 to create three-dimensional data 1035 of the surroundings of the own vehicle, and perform self-position estimation processing using the created three-dimensional data 1035.
[0487] The three-dimensional data accumulation section 1018 accumulates the three-dimensional map 1032, the three-dimensional data 1034, and the three-dimensional data 1035.
[0488] The format conversion section 1019 generates sensor information 1037 by converting the sensor information 1033 into a format corresponding to the receiving side. In addition, the format conversion section 1019 can reduce the data amount by compressing or encoding the sensor information 1037. Furthermore, the format conversion section 1019 can omit the processing in the case where format conversion is not necessary. Furthermore, the format conversion section 1019 can control the data amount to be transmitted in accordance with the designation of the transmission range.
[0489] The communication section 1020 communicates with the server 901, and receives a data transmission request (a sensor information transmission request) and the like from the server 901.
[0490] The transmission control section 1021 establishes communication by exchanging information on the corresponding format and the like with the communication counterpart via the communication section 1020.
[0491] The data transmission section 1022 transmits the sensor information 1037 to the server 901. The sensor information 1037 includes, for example, information obtained by the LiDAR, a brightness image (a visible light image) obtained by a visible light camera, an infrared image obtained by an infrared camera, a depth image obtained by a depth sensor, sensor position information, speed information, and the like obtained by the plurality of sensors 1015.
[0492] Next, the configuration of the server 901 will be described. Figure 30 is a block diagram showing a configuration example of the server 901. The server 901 receives sensor information transmitted from the client device 902, and creates three-dimensional data from the received sensor information. The server 901 updates a three-dimensional map managed by the server 901 using the created three-dimensional data. Furthermore, the server 901 transmits the updated three-dimensional map to the client device 902 in accordance with a transmission request of the three-dimensional map from the client device 902.
[0493] The server 901 includes a data reception unit 1111, a communication unit 1112, a reception control unit 1113, a format conversion unit 1114, a three-dimensional data creation unit 1116, a three-dimensional data synthesis unit 1117, a three-dimensional data accumulation unit 1118, a format conversion unit 1119, a communication unit 1120, a transmission control unit 1121, and a data transmission unit 1122.
[0494] The data reception unit 1111 receives the sensor information 1037 from the client device 902. The sensor information 1037 includes, for example, information obtained by a LiDAR, a brightness image (visible light image) obtained by a visible light camera, an infrared image obtained by an infrared camera, a depth image obtained by a depth sensor, sensor position information, speed information, and the like.
[0495] The communication unit 1112 communicates with the client device 902 and transmits a data transmission request (for example, a sensor information transmission request) and the like to the client device 902.
[0496] The reception control unit 1113 establishes communication with a communication partner by exchanging information on a corresponding format and the like via the communication unit 1112.
[0497] The format conversion unit 1114 generates the sensor information 1132 by performing decompression or decoding processing in the case where the received sensor information 1037 is compressed or encoded. In the case where the sensor information 1037 is uncompressed data, the format conversion unit 1114 does not perform decompression or decoding processing.
[0498] The three-dimensional data creation unit 1116 creates three-dimensional data 1134 of the surroundings of the client device 902 on the basis of the sensor information 1132. For example, the three-dimensional data creation unit 1116 creates point cloud data in which the surroundings of the client device 902 have color information, using information obtained by a LiDAR and a visible light image obtained by a visible light camera.
[0499] The three-dimensional data synthesis unit 1117 synthesizes the three-dimensional data 1134 created on the basis of the sensor information 1132 with a three-dimensional map 1135 managed by the server 901, and updates the three-dimensional map 1135 on the basis of the synthesis.
[0500] The three-dimensional data accumulation unit 1118 accumulates the three-dimensional map 1135 and the like.
[0501] The format conversion section 1119 generates the three-dimensional map 1031 by converting the three-dimensional map 1135 into a format corresponding to the receiving side. In addition, the format conversion section 1119 can also reduce the data amount by compressing or encoding the three-dimensional map 1135. Also, the format conversion section 1119 can omit the processing in the case where the format conversion is not necessary. Also, the format conversion section 1119 can control the data amount to be transmitted in accordance with the designation of the transmission range.
[0502] The communication section 1120 communicates with the client device 902, and receives a data transmission request (a three-dimensional map transmission request) and the like from the client device 902.
[0503] The transmission control section 1121 establishes communication by exchanging information on the corresponding format and the like with the communication counterpart via the communication section 1120.
[0504] The data transmission section 1122 transmits the three-dimensional map 1031 to the client device 902. The three-dimensional map 1031 is data including a point cloud of a WLD or a SWLD or the like. Either of compressed data and uncompressed data can be included in the three-dimensional map 1031.
[0505] Next, the workflow of the client device 902 will be described. Figure 31 is a flowchart showing the workflow of the client device 902 at the time of obtaining a three-dimensional map.
[0506] First, the client device 902 requests the server 901 for transmission of a three-dimensional map (point cloud or the like) (S1001). At this time, the client device 902 also transmits the position information of the client device 902 obtained by a GPS or the like, and on the basis of this, can request the server 901 for transmission of a three-dimensional map related to the position information.
[0507] Next, the client device 902 receives a three-dimensional map from the server 901 (S1002). If the received three-dimensional map is compressed data, the client device 902 decodes the received three-dimensional map, and generates an uncompressed three-dimensional map (S1003).
[0508] Next, the client device 902 creates three-dimensional data 1034 of the surroundings of the client device 902 on the basis of the sensor information 1033 obtained from the plurality of sensors 1015 (S1004). Next, the client device 902 estimates the own position of the client device 902 using the three-dimensional map 1032 received from the server 901 and the three-dimensional data 1034 created on the basis of the sensor information 1033 (S1005).
[0509] Figure 32is a flowchart showing the operation of the client device 902 at the time of transmission of sensor information. First, the client device 902 receives a request for transmission of sensor information from the server 901 (S1011). The client device 902 that has received the request for transmission transmits the sensor information 1037 to the server 901 (S1012). In addition, in a case where the sensor information 1033 includes a plurality of information obtained by a plurality of sensors 1015, the client device 902 compresses each information in a compression method suitable for the information, thereby generating the sensor information 1037.
[0510] Next, the operation flow of the server 901 will be described. Figure 33 is a flowchart showing the operation of the server 901 at the time of acquisition of sensor information. First, the server 901 requests the client device 902 for transmission of sensor information (S1021). Next, the server 901 receives the sensor information 1037 transmitted from the client device 902 in response to the request (S1022). Next, the server 901 creates three-dimensional data 1134 using the received sensor information 1037 (S1023). Next, the server 901 reflects the created three-dimensional data 1134 to a three-dimensional map 1135 (S1024).
[0511] Figure 34 is a flowchart showing the operation of the server 901 at the time of transmission of a three-dimensional map. First, the server 901 receives a request for transmission of a three-dimensional map from the client device 902 (S1031). The server 901 that has received the request for transmission of a three-dimensional map transmits the three-dimensional map 1031 to the client device 902 (S1032). At this time, the server 901 can extract a three-dimensional map in the vicinity thereof in correspondence with the position information of the client device 902, and transmit the extracted three-dimensional map. Also, it can be that the server 901 compresses a three-dimensional map constituted by a point cloud, for example, using a compression method by an octree, and transmits the compressed three-dimensional map.
[0512] Next, a modification of the present embodiment will be described.
[0513] The server 901 creates three-dimensional data 1134 in the vicinity of the position of the client device 902 using the sensor information 1037 received from the client device 902. Then, the server 901 matches the created three-dimensional data 1134 with a three-dimensional map 1135 of the same area managed by the server 901, and calculates a difference between the three-dimensional data 1134 and the three-dimensional map 1135. The server 901 determines that some abnormality has occurred in the vicinity of the client device 902 when the difference is equal to or greater than a threshold value determined in advance. For example, when a subsidence of the ground surface has occurred due to a natural disaster such as an earthquake, a large difference is expected to occur between the three-dimensional map 1135 managed by the server 901 and the three-dimensional data 1134 created based on the sensor information 1037.
[0514] The sensor information 1037 can include at least one of the kind of the sensor, the performance of the sensor, and the model of the sensor. The sensor information 1037 can also include a category ID or the like corresponding to the performance of the sensor. For example, when the sensor information 1037 is information obtained by a LiDAR, it is possible to assign an identifier for the performance of the sensor, such as category 1 for a sensor capable of obtaining information with an accuracy of several mm, category 2 for a sensor capable of obtaining information with an accuracy of several cm, and category 3 for a sensor capable of obtaining information with an accuracy of several m. The server 901 can also estimate the performance information of the sensor or the like from the model of the client device 902. For example, when the client device 902 is mounted on a vehicle, the server 901 can determine the specification information of the sensor from the model of the vehicle. In this case, the server 901 can obtain the information of the model of the vehicle in advance, or can include the information in the sensor information. The server 901 can also switch the degree of correction for the three-dimensional data 1134 created using the sensor information 1037 using the obtained sensor information 1037. For example, when the performance of the sensor is high accuracy (category 1), the server 901 does not perform correction for the three-dimensional data 1134. When the performance of the sensor is low accuracy (category 3), the server 901 applies correction appropriate for the accuracy of the sensor to the three-dimensional data 1134. For example, the server 901 can enhance the degree (intensity) of correction as the accuracy of the sensor decreases.
[0515] The server 901 can also simultaneously issue a request for transmission of sensor information to a plurality of client devices 902 present in a certain space. In a case where the server 901 receives a plurality of pieces of sensor information from the plurality of client devices 902, it is not necessary to utilize all of the pieces of sensor information in the creation of the three-dimensional data 1134, and, for example, the pieces of sensor information to be utilized can be selected in accordance with the performance of the sensors. For example, in a case where the three-dimensional map 1135 is updated, the server 901 can select high-accuracy sensor information (category 1) from among the plurality of pieces of sensor information received, and create the three-dimensional data 1134 using the selected piece of sensor information.
[0516] The server 901 is not limited to a server such as a traffic cloud monitoring server, and can also be another client device (onboard). Figure 35 The system configuration in this case is shown.
[0517] For example, the client device 902C issues a request for transmission of sensor information to the client device 902A present in the vicinity, and obtains sensor information from the client device 902A. The client device 902C then creates three-dimensional data using the obtained sensor information of the client device 902A, and updates the three-dimensional map of the client device 902C. In this way, the client device 902C can utilize the performance of the client device 902C to generate a three-dimensional map of the space that can be obtained from the client device 902A. For example, this can be considered to occur in a case where the performance of the client device 902C is high.
[0518] Also, in this case, the client device 902A that provided the sensor information is given the right to obtain the high-accuracy three-dimensional map generated by the client device 902C. The client device 902A receives the high-accuracy three-dimensional map from the client device 902C in accordance with this right.
[0519] Also, the client device 902C can issue a request for transmission of sensor information to a plurality of client devices 902 (the client device 902A and the client device 902B) present in the vicinity. In a case where the sensors of the client device 902A or the client device 902B are high-performance, the client device 902C can create three-dimensional data using sensor information obtained by the high-performance sensors.
[0520] Figure 36 is a block diagram showing the functional configuration of the server 901 and the client device 902. The server 901 includes, for example, a three-dimensional map compression / decoding processing section 1201 that compresses and decodes a three-dimensional map, and a sensor information compression / decoding processing section 1202 that compresses and decodes sensor information.
[0521] The client device 902 has a three-dimensional map decoding processing section 1211 and a sensor information compression processing section 1212. The three-dimensional map decoding processing section 1211 receives encoded data of a compressed three-dimensional map, decodes the encoded data, and acquires a three-dimensional map. The sensor information compression processing section 1212 compresses sensor information itself, rather than three-dimensional data created from the acquired sensor information, and transmits encoded data of the compressed sensor information to the server 901. According to this configuration, the client device 902 can hold a processing section (device or LSI) for decoding processing of a three-dimensional map (point cloud or the like) inside, without holding a processing section for compression processing of three-dimensional data of a three-dimensional map (point cloud or the like) inside. In this way, it is possible to suppress the cost and power consumption of the client device 902 and the like.
[0522] As described above, the client device 902 according to the present embodiment is mounted on a mobile body, and creates three-dimensional data 1034 of the surroundings of the mobile body from sensor information 1033 showing the surroundings of the mobile body, which is obtained by the sensor 1015 mounted on the mobile body. The client device 902 estimates the own position of the mobile body using the created three-dimensional data 1034. The client device 902 transmits the acquired sensor information 1033 to the server 901 or another mobile body 902.
[0523] Accordingly, the client device 902 transmits the sensor information 1033 to the server 901 or the like. In this way, it is possible to reduce the amount of data to be transmitted, compared to the case where three-dimensional data is transmitted. Further, since it is not necessary to perform processing such as compression or encoding of three-dimensional data at the client device 902, it is possible to reduce the amount of processing at the client device 902. Therefore, the client device 902 can achieve reduction of the amount of data to be transmitted or simplification of the configuration of the device.
[0524] Further, the client device 902 further transmits a transmission request of a three-dimensional map to the server 901, and receives a three-dimensional map 1031 from the server 901. The client device 902 estimates the own position using the three-dimensional data 1034 and the three-dimensional map 1032 in the estimation of the own position.
[0525] Further, the sensor information 1033 includes at least one of information obtained by a laser sensor, a brightness image (visible light image), an infrared image, a depth image, position information of the sensor, and speed information of the sensor.
[0526] Further, the sensor information 1033 includes information showing the performance of the sensor.
[0527] Also, the client device 902 encodes or compresses the sensor information 1033, and transmits the encoded or compressed sensor information 1037 to the server 901 or other mobile body 902 in the transmission of the sensor information. Thereby, the client device 902 can reduce the amount of data to be transmitted.
[0528] For example, the client device 902 is provided with a processor and a memory, and the processor performs the above-described processing using the memory.
[0529] Also, the server 901 according to the present embodiment can communicate with the client device 902 mounted on a mobile body, and receives sensor information 1037 showing the surrounding situation of the mobile body from the client device 902, which is obtained by the sensor 1015 mounted on the mobile body. The server 901 creates three-dimensional data 1134 of the surroundings of the mobile body based on the received sensor information 1037.
[0530] Thereby, the server 901 creates the three-dimensional data 1134 using the sensor information 1037 transmitted from the client device 902. In this way, compared to the case where the client device 902 transmits the three-dimensional data, it is possible to reduce the amount of data to be transmitted. Also, since it is not necessary to perform processing such as compression or encoding of the three-dimensional data at the client device 902, it is possible to reduce the processing amount of the client device 902. In this way, the server 901 can achieve reduction of the amount of data to be transmitted, or simplification of the configuration of the device.
[0531] Also, the server 901 further transmits a transmission request of the sensor information to the client device 902.
[0532] Also, the server 901 further updates the three-dimensional map 1135 using the created three-dimensional data 1134, and transmits the three-dimensional map 1135 to the client device 902 in accordance with a transmission request of the three-dimensional map 1135 from the client device 902.
[0533] Also, the sensor information 1037 includes at least one of information obtained by a laser sensor, a brightness image (visible light image), an infrared image, a depth image, position information of the sensor, and speed information of the sensor.
[0534] Also, the sensor information 1037 includes information showing the performance of the sensor.
[0535] Also, the server 901 further corrects the three-dimensional data in accordance with the performance of the sensor. Thereby, the three-dimensional data creation method can improve the quality of the three-dimensional data.
[0536] Also, the server 901 receives a plurality of pieces of sensor information 1037 from a plurality of client apparatuses 902 in the reception of the sensor information, and selects the sensor information 1037 used in the production of the three-dimensional data 1134, based on a plurality of pieces of information included in the plurality of pieces of sensor information 1037, which show the performance of the sensor. Thereby, the server 901 can improve the quality of the three-dimensional data 1134.
[0537] Also, the server 901 decodes or decompresses the received sensor information 1037, and produces the three-dimensional data 1134 based on the decoded or decompressed sensor information 1132. Thereby, the server 901 can reduce the amount of data to be transmitted.
[0538] For example, the server 901 has a processor and a memory, and the processor performs the above-described processing using the memory.
[0539] (Embodiment 7)
[0540] In the present embodiment, an encoding method and a decoding method of three-dimensional data using inter prediction processing are described.
[0541] Figure 37 is a block diagram of a three-dimensional data encoding apparatus 1300 according to the present embodiment. The three-dimensional data encoding apparatus 1300 generates an encoded bitstream (hereinafter, simply referred to as a bitstream) as an encoded signal by encoding three-dimensional data. As shown in Figure 37 The three-dimensional data encoding apparatus 1300 has a division unit 1301, a subtraction unit 1302, a transform unit 1303, a quantization unit 1304, an inverse quantization unit 1305, an inverse transform unit 1306, an addition unit 1307, a reference volume memory 1308, an intra prediction unit 1309, a reference space memory 1310, an inter prediction unit 1311, a prediction control unit 1312, and an entropy encoding unit 1313.
[0542] The division unit 1301 divides each space (SPC) included in the three-dimensional data into a plurality of volumes (VLM) as an encoding unit. Also, the division unit 1301 octree-encodes the voxels within each volume. Alternatively, the division unit 1301 can octree-encode the spaces so that the spaces and the volumes have the same size. Also, the division unit 1301 can attach information required for octree-encoding (depth information, etc.) to the header of the bitstream or the like.
[0543] The subtraction unit 1302 calculates the difference between the volume (encoding target volume) output from the division unit 1301 and the prediction volume generated by the intra prediction or the inter prediction described later, and outputs the calculated difference as a prediction residual to the transform unit 1303. Figure 38An example of calculating the prediction residual is shown. Additionally, the bit strings of the encoded object volume and the prediction volume shown here, for example, indicate the location information of the three-dimensional points (e.g., point clouds) contained within the volume.
[0544] The following explains the octree representation and the voxel scanning order. The volume is transformed into an octree structure (octreeification) and then encoded. The octree structure consists of nodes and leaf nodes. Each node has 8 nodes or leaf nodes, and each leaf node contains voxel (VXL) information. Figure 39 An example of the composition of a volume including multiple voxels is shown. Figure 40 It shows that Figure 39 The volume transformation shown is an example of an octree structure. Here, Figure 40 Leaf nodes 1, 2, and 3 in the leaf nodes shown represent, respectively Figure 39 The voxels VXL1, VXL2, and VXL3 shown represent VXL including point groups (hereinafter referred to as effective VXL).
[0545] An octree can be represented as a binary sequence of 0s and 1s. For example, when a node or valid VXL is set to a value of 1, and all others are set to a value of 0, the nodes and leaf nodes are assigned values. Figure 40 The binary sequence is shown. Therefore, this binary sequence is scanned according to either width-first or depth-first scanning order. For example, in the case of a width-first scan, we obtain... Figure 41 The binary sequence A is shown. After a depth-first scan, the following is obtained: Figure 41 The binary sequence shown in B. The binary sequence obtained through this scan is encoded by entropy coding, thereby reducing the amount of information.
[0546] Next, we will explain the depth information in the octree representation. The depth in the octree representation controls the granularity at which the point cloud information contained within a volume is preserved. Setting a large depth allows for the reproduction of point cloud information at a finer level, but this increases the amount of data used to represent nodes and leaf nodes. Conversely, setting a small depth reduces the amount of data, but multiple point cloud information points at different locations and with different colors will be treated as the same location and color, thus losing the original information inherent in the point cloud.
[0547] For example, Figure 42 It shows that Figure 40 The example shown is an octree with a depth of 2, represented as an octree with a depth of 1. Figure 42 The octree shown is compared to Figure 40 The octree shown has a small amount of data. That is, Figure 42 The octree shown Figure 42The octree shown above has fewer bits than the binary sequence. In this case, Figure 40 Leaf node 1 and leaf node 2 shown above are expressed as Figure 41 Leaf node 1 shown above. That is, the information that leaf node 1 and leaf node 2 are different positions is lost. Figure 40 Leaf node 1 and leaf node 2 shown above are expressed as
[0548] Figure 43 The volume corresponding to the octree shown above is shown. Figure 42 VXL1 and VXL2 shown above correspond to VXL12. In this case, the three-dimensional data encoding apparatus 1300 generates color information of VXL12 from color information of VXL1 and VXL2 shown above. For example, the three-dimensional data encoding apparatus 1300 calculates the average, the median, or the weighted average of the color information of VXL1 and VXL2 as the color information of VXL12. In this way, the three-dimensional data encoding apparatus 1300 can control the reduction of the data amount by changing the depth of the octree. Figure 39 Figure 43 Figure 39 Figure 43
[0549] The three-dimensional data encoding apparatus 1300 can set the depth information of the octree using any one of the world space unit, the space unit, and the volume unit. Also at this time, the three-dimensional data encoding apparatus 1300 can attach the depth information to the header information of the world space, the header information of the space, or the header information of the volume. Also, the same value can be used as the depth information in all the world spaces, spaces, and volumes that are different in time. In this case, the three-dimensional data encoding apparatus 1300 can attach the depth information to the header information that manages all the world spaces in time.
[0550] In the case where the color information is included in the voxels, the transform unit 1303 applies a frequency transform such as an orthogonal transform to the prediction residual of the color information of the voxels in the volume. For example, the transform unit 1303 scans the prediction residual in a certain scan order to make a one-dimensional arrangement. After that, the transform unit 1303 transforms the one-dimensional arrangement into a frequency domain by applying a one-dimensional orthogonal transform to the made one-dimensional arrangement. Accordingly, in the case where the values of the prediction residual in the volume are close, the values of the frequency components of the low frequency band become large and the values of the frequency components of the high frequency band become small. Therefore, the quantization unit 1304 can more effectively reduce the amount of encoding.
[0551] Furthermore, the transformation unit 1303 can utilize orthogonal transformations of two dimensions or higher, instead of one-dimensional orthogonal transformations. For example, the transformation unit 1303 maps the prediction residuals to a two-dimensional arrangement in a certain scanning order, and applies a two-dimensional orthogonal transformation to the resulting two-dimensional arrangement. The transformation unit 1303 can also select the orthogonal transformation method to be used from multiple orthogonal transformation methods. In this case, the three-dimensional data encoding device 1300 appends information indicating which orthogonal transformation method was used to the bitstream. Alternatively, the transformation unit 1303 can select the orthogonal transformation method to be used from multiple orthogonal transformation methods with different dimensions. In this case, the three-dimensional data encoding device 1300 appends information indicating which dimension of orthogonal transformation method was used to the bitstream.
[0552] For example, the transformation unit 1303 matches the scan order of the predicted residual with the scan order (width-first or depth-first, etc.) in the octree within the volume. Therefore, since it is not necessary to append information indicating the scan order of the predicted residual to the bitstream, additional overhead can be reduced. Furthermore, the transformation unit 1303 can also apply a scan order different from the octree scan order. In this case, the three-dimensional data encoding device 1300 appends information indicating the scan order of the predicted residual to the bitstream. Therefore, the three-dimensional data encoding device 1300 can efficiently encode the predicted residual. Alternatively, the three-dimensional data encoding device 1300 may append information indicating whether the octree scan order is applicable (a flag, etc.) to the bitstream; if the octree scan order is not applicable, it may append information indicating the scan order of the predicted residual to the bitstream.
[0553] The transformation unit 1303 can transform not only the prediction residual of color information, but also other attribute information of voxels. For example, the transformation unit 1303 can transform and encode information such as reflectance obtained when acquiring point clouds through LiDAR or the like.
[0554] If the transformation unit 1303 does not have attribute information such as color information in the space, it can skip processing. Furthermore, the three-dimensional data encoding device 1300 can append information (flags) indicating whether to skip the processing of the transformation unit 1303 to the bit stream.
[0555] The quantization section 1304 quantizes the frequency components of the prediction residual generated by the transform section 1303 using a quantization control parameter, thereby generating quantization coefficients. The information amount is reduced accordingly. The generated quantization coefficients are output to the entropy coding section 1313. The quantization section 1304 can control the quantization control parameter in terms of world space units, space units, or volume units. In this case, the three-dimensional data encoding apparatus 1300 attaches the quantization control parameter to the respective header information and the like. Also, the quantization section 1304 can change the weight to perform quantization control for each frequency component of the prediction residual. For example, the quantization section 1304 can perform fine quantization for low frequency components and perform coarse quantization for high frequency components. In this case, the three-dimensional data encoding apparatus 1300 can attach a parameter indicating the weight of each frequency component to the header.
[0556] The quantization section 1304 can skip the processing in the case where the attribute information such as color information is not present in the space. Also, the three-dimensional data encoding apparatus 1300 can attach information (flag) indicating whether the processing of the quantization section 1304 is skipped to the bitstream.
[0557] The inverse quantization section 1305 inversely quantizes the quantization coefficients generated by the quantization section 1304 using the quantization control parameter, thereby generating inverse quantization coefficients of the prediction residual. The generated inverse quantization coefficients are output to the inverse transform section 1306.
[0558] The inverse transform section 1306 applies inverse transform to the inverse quantization coefficients generated by the inverse quantization section 1305, thereby generating inverse transform-applied prediction residuals. Since the inverse transform-applied prediction residuals are generated after quantization, they can not be identical to the prediction residuals output from the transform section 1303.
[0559] The addition section 1307 adds the inverse transform-applied prediction residuals generated by the inverse transform section 1306 to the prediction volume generated by the intra prediction or the inter prediction described later and used in the generation of the prediction residual before quantization, thereby generating a reconstruction volume. The reconstruction volume is stored in the reference volume storage 1308 or the reference space storage 1310.
[0560] The intra prediction section 1309 generates a prediction volume of the encoding target volume using the attribute information of the neighboring volumes stored in the reference volume storage 1308. The attribute information includes color information or reflectance of the voxels. The intra prediction section 1309 generates a predicted value of the color information or the reflectance of the encoding target volume.
[0561] Figure 44 is a diagram for explaining the operation of the intra prediction section 1309. For example, Figure 44As shown, the intra-frame prediction unit 1309 generates a predicted volume for the encoded object volume (volume idx = 3) based on adjacent volumes (volume idx = 0). Here, volume idx is identifier information added to volumes within the space, and different values are assigned to each volume. The order in which volume idx is assigned can be the same as or different from the encoding order. For example, as... Figure 44 The intra-frame prediction unit 1309 uses the average value of the color information of the voxels contained in the adjacent volume idx = 0 as the predicted value of the color information of the encoded object volume. In this case, a prediction residual is generated by subtracting the predicted value of the color information from the color information of each voxel contained in the encoded object volume. The transformation unit 1303 and subsequent processing are performed on this prediction residual. In this case, the three-dimensional data encoding apparatus 1300 appends adjacent volume information and prediction mode information to the bitstream. Here, the adjacent volume information shows the information of the adjacent volumes used in the prediction, such as the volume idx of the adjacent volumes used in the prediction. The prediction mode information shows the mode used in the generation of the prediction volume. The mode is, for example, an average value mode that generates the prediction value based on the average value of the voxels in the adjacent volumes, or an intermediate value mode that generates the prediction value based on the median value of the voxels in the adjacent volumes, etc.
[0562] The intra-frame prediction unit 1309 can also generate a prediction volume based on multiple adjacent volumes. For example, in Figure 44 In the configuration shown, the intra-frame prediction unit 1309 generates prediction volume 0 based on the volume where volume idx = 0, and generates prediction volume 1 based on the volume where volume idx = 1. Then, the intra-frame prediction unit 1309 generates the final prediction volume by averaging prediction volume 0 and prediction volume 1. In this case, the 3D data encoding apparatus 1300 can also append multiple volume idx values of the multiple volumes used in generating the prediction volume to the bitstream.
[0563] Figure 45 The inter-frame prediction process involved in this embodiment is illustrated in the diagram. The inter-frame prediction unit 1311 performs encoding (inter-frame prediction) on the space (SPC) of a certain time T_Cur using the encoded space of different times T_LX. In this case, the inter-frame prediction unit 1311 performs encoding processing by applying rotation and translation processing to the encoded space of different times T_LX.
[0564] Also, the three-dimensional data encoding apparatus 1300 adds RT information related to rotation and translation processing of a space applied at a different time T_LX to the bitstream. The different time T_LX is, for example, a time T_L0 before the certain time T_Cur. At this time, the three-dimensional data encoding apparatus 1300 can also add RT information RT_L0 related to rotation and translation processing of a space applied at the time T_L0 to the bitstream.
[0565] Alternatively, the different time T_LX is, for example, a time T_L1 after the certain time T_Cur. At this time, the three-dimensional data encoding apparatus 1300 can add RT information RT_L1 related to rotation and translation processing of a space applied at the time T_L1 to the bitstream.
[0566] Alternatively, the inter prediction section 1311 performs encoding with reference to both a space at the time T_L0 and a space at the time T_L1 (bi-prediction). In this case, the three-dimensional data encoding apparatus 1300 can add both RT information RT_L0 and RT_L1 related to rotation and translation of the respective spaces to the bitstream.
[0567] Note that, although T_L0 is set to a time before T_Cur and T_L1 is set to a time after T_Cur above, this is not limiting. For example, T_L0 and T_L1 can both be times before T_Cur. Alternatively, T_L0 and T_L1 can both be times after T_Cur.
[0568] Also, in a case where the three-dimensional data encoding apparatus 1300 performs encoding with reference to a plurality of spaces at different times, the three-dimensional data encoding apparatus 1300 can add RT information related to rotation and translation of each of the spaces to the bitstream. For example, the three-dimensional data encoding apparatus 1300 manages a plurality of encoded spaces to be referred to by two reference lists (an L0 list and an L1 list). In a case where a first reference space in the L0 list is set to L0R0, a second reference space in the L0 list is set to L0R1, a first reference space in the L1 list is set to L1R0, and a second reference space in the L1 list is set to L1R1, the three-dimensional data encoding apparatus 1300 adds RT information RT_L0R0 of L0R0, RT information RT_L0R1 of L0R1, RT information RT_L1R0 of L1R0, and RT information RT_L1R1 of L1R1 to the bitstream. For example, the three-dimensional data encoding apparatus 1300 adds these RT information to a header or the like of the bitstream.
[0569] Also, in a case where the three-dimensional data encoding apparatus 1300 encodes with reference to a plurality of reference spaces at different times, the three-dimensional data encoding apparatus 1300 can determine whether to apply rotation and translation for each reference space. In this case, the three-dimensional data encoding apparatus 1300 can attach information (RT application flag or the like) indicating whether rotation and translation are applied for each reference space to the header information or the like of the bitstream. For example, the three-dimensional data encoding apparatus 1300 calculates RT information and an ICP (Interactive Closest Point) error value for each reference space to be referred to using the ICP algorithm according to the encoding target space. The three-dimensional data encoding apparatus 1300 determines that rotation and translation are not necessary when the ICP error value is equal to or less than a predetermined value, and sets the RT application flag to OFF (invalid). In addition, the three-dimensional data encoding apparatus 1300 sets the RT application flag to ON (valid) and attaches the RT information to the bitstream when the ICP error value is greater than the predetermined value.
[0570] Figure 46 A syntax example in which the RT information and the RT application flag are attached to the header is shown. In addition, the number of bits allocated to each syntax can be determined according to the range that can be taken by the syntax. For example, in a case where the number of reference spaces included in the reference list L0 is 8, 3 bits can be allocated to MaxRefSpc_l0. The number of bits allocated can be changed according to the value that can be taken by each syntax, or can be fixed regardless of the value that can be taken. In a case where the number of bits allocated is fixed, the three-dimensional data encoding apparatus 1300 can attach the fixed number of bits to other header information.
[0571] Here, Figure 46 MaxRefSpc_l0 shown indicates the number of reference spaces included in the reference list L0. RT_flag_l0[i] is the RT application flag for the reference space i in the reference list L0. In a case where RT_flag_l0[i] is 1, rotation and translation are applied to the reference space i. In a case where RT_flag_l0[i] is 0, rotation and translation are not applied to the reference space i.
[0572] R_l0[i] and T_l0[i] are the RT information for the reference space i in the reference list L0. R_l0[i] is the rotation information for the reference space i in the reference list L0. The rotation information indicates the content of the rotation process to be applied, such as a rotation matrix or a quaternion or the like. T_l0[i] is the translation information for the reference space i in the reference list L0. The translation information indicates the content of the translation process to be applied, such as a translation vector or the like.
[0573] MaxRefSpc_l1 shows the number of reference spaces included in the reference list L1. RT_flag_l1[i] is an RT application flag of the reference space i in the reference list L1. In the case where RT_flag_l1[i] is 1, rotation and translation are applied to the reference space i. In the case where RT_flag_l1[i] is 0, rotation and translation are not applied to the reference space i.
[0574] R_l1[i] and T_l1[i] are RT information of the reference space i in the reference list L1. R_l1[i] is rotation information of the reference space i in the reference list L1. The rotation information shows the content of the applied rotation process, such as a rotation matrix or a quaternion, and the like. T_l1[i] is translation information of the reference space i in the reference list L1. The translation information shows the content of the applied translation process, such as a translation vector, and the like.
[0575] The inter prediction section 1311 generates a prediction volume of the encoding target volume using the information of the encoded reference space stored in the reference space storage 1310. As described above, the inter prediction section 1311 calculates the RT information using the ICP (Interactive Closest Point) algorithm for the encoding target space and the reference space in order to approximate the positional relationship of the entire encoding target space and the reference space before generating the prediction volume of the encoding target volume. Then, the inter prediction section 1311 applies the rotation and translation processes to the reference space using the calculated RT information, thereby obtaining the reference space B. After that, the inter prediction section 1311 generates the prediction volume of the encoding target volume in the encoding target space using the information in the reference space B. Here, the three-dimensional data encoding apparatus 1300 attaches the RT information used to obtain the reference space B to the header information or the like of the encoding target space.
[0576] In this way, the inter prediction section 1311 generates the prediction volume using the information of the reference space after approximating the positional relationship of the entire encoding target space and the reference space by applying the rotation and translation processes to the reference space, so that the accuracy of the prediction volume can be improved. Also, since the prediction residual can be suppressed, the amount of encoding can be reduced. Note that, although the example in which the ICP is performed using the encoding target space and the reference space is shown here, the present application is not limited to this. For example, the inter prediction section 1311 can perform the ICP using at least one of the encoding target space from which the number of voxels or point clouds is extracted and the reference space from which the number of voxels or point clouds is extracted in order to reduce the amount of processing, thereby calculating the RT information.
[0577] Further, the inter-frame prediction unit 1311 can determine that rotation and translation processing is not needed and does not perform the rotation and translation, in a case where the ICP error value obtained from the result of the ICP is smaller than a first threshold value that is prescribed in advance, i.e., in a case where the positional relationship between the encoding target space and the reference space is close, for example. In this case, the three-dimensional data encoding apparatus 1300 can not attach the RT information to the bitstream, and thus can suppress the overhead.
[0578] Further, the inter-frame prediction unit 1311 determines that the shape change in the space is large, in a case where the ICP error value is larger than a second threshold value that is prescribed in advance, and can apply the intra-frame prediction to all of the volumes of the encoding target space. Hereinafter, the space to which the intra-frame prediction is applied is referred to as an intra-frame space. Further, the second threshold value is a value that is larger than the first threshold value described above. Further, any method can be applied as long as the method is a method of obtaining the RT information from two sets of volumes or two sets of point clouds, and is not limited to the ICP.
[0579] Further, in a case where the three-dimensional data contains attribute information such as shape or color, the inter-frame prediction unit 1311 searches for a volume that is closest to the shape or color attribute information of the encoding target volume, for example, in the reference space, as a prediction volume of the encoding target volume in the encoding target space. Further, the reference space is the reference space after the rotation and translation processing described above, for example. The inter-frame prediction unit 1311 generates the prediction volume based on the volume (reference volume) obtained by the search. Figure 47 is a diagram for explaining the generation of the prediction volume. The inter-frame prediction unit 1311 encodes the encoding target volume (volume idx = 0) shown in Figure 47 by using the inter-frame prediction. The inter-frame prediction unit 1311 selects the volume in which the prediction residual between the encoding target volume and the reference volume is the smallest, as the prediction volume. The prediction residual between the encoding target volume and the prediction volume is encoded by the processing after the transform unit 1303. Here, the prediction residual refers to the difference between the attribute information of the encoding target volume and the attribute information of the prediction volume. Further, the three-dimensional data encoding apparatus 1300 attaches the volume idx of the reference volume in the reference space referred to as the prediction volume to the header or the like of the bitstream.
[0580] In the example shown in Figure 47 , the reference volume of volume idx = 4 of the reference space L0R0 is selected as the prediction volume of the encoding target volume. Then, the prediction residual between the encoding target volume and the reference volume and the reference volume idx = 4 are encoded and attached to the bitstream.
[0581] Further, although the prediction volume of the attribute information is described as an example, the same processing can be performed for the prediction volume of the position information.
[0582] The prediction control section 1312 controls which of the intra prediction and the inter prediction is used to encode the encoding target volume. Here, the mode including the intra prediction and the inter prediction is referred to as a prediction mode. For example, the prediction control section 1312 calculates the prediction residual in the case where the encoding target volume is predicted by the intra prediction and the prediction residual in the case where the encoding target volume is predicted by the inter prediction as evaluation values, and selects the prediction mode of the smaller one of the evaluation values. Alternatively, the prediction control section 1312 can apply orthogonal transformation, quantization, and entropy coding to the prediction residual of the intra prediction and the prediction residual of the inter prediction, respectively, to calculate the actual amount of encoding, and select the prediction mode using the calculated amount of encoding as an evaluation value. Further, additional overhead information (reference volume idx information and the like) other than the prediction residual can be added to the evaluation value. Furthermore, the prediction control section 1312 can normally select the intra prediction in the case where the encoding target space is predetermined to be encoded in the intra space.
[0583] The entropy coding section 1313 generates an encoded signal (encoded bit stream) by variable length coding the input from the quantization section 1304, that is, the quantization coefficient. Specifically, the entropy coding section 1313, for example, binarizes the quantization coefficient, and arithmetically encodes the obtained binary signal.
[0584] Next, a three-dimensional data decoding apparatus that decodes the encoded signal generated by the three-dimensional data encoding apparatus 1300 will be described. Figure 48 is a block diagram of a three-dimensional data decoding apparatus 1400 according to the present embodiment. The three-dimensional data decoding apparatus 1400 includes an entropy decoding section 1401, an inverse quantization section 1402, an inverse transformation section 1403, an addition section 1404, a reference volume storage 1405, an intra prediction section 1406, a reference space storage 1407, an inter prediction section 1408, and a prediction control section 1409.
[0585] The entropy decoding section 1401 variable length decodes the encoded signal (encoded bit stream). For example, the entropy decoding section 1401 arithmetically decodes the encoded signal to generate a binary signal, and generates a quantization coefficient from the generated binary signal.
[0586] The inverse quantization section 1402 inversely quantizes the quantization coefficient input from the entropy decoding section 1401 using the quantization parameter attached to the bit stream or the like, thereby generating an inverse quantization coefficient.
[0587] The inverse transform unit 1403 performs inverse transform on the inverse quantization coefficients input from the inverse quantization unit 1402, thereby generating a prediction residual. For example, the inverse transform unit 1403 performs inverse orthogonal transform on the inverse quantization coefficients in accordance with the information attached to the bitstream, thereby generating a prediction residual.
[0588] The addition unit 1404 adds the prediction residual generated by the inverse transform unit 1403 to the prediction volume generated by the intra prediction or the inter prediction, thereby generating a reconstructed volume. The reconstructed volume is output as decoded three-dimensional data, and is stored in the reference volume storage 1405 or the reference space storage 1407.
[0589] The intra prediction unit 1406 generates a prediction volume by intra prediction using the reference volume in the reference volume storage 1405 and the information attached to the bitstream. Specifically, the intra prediction unit 1406 obtains prediction mode information and neighboring volume information (e.g., volume idx) attached to the bitstream, generates a prediction volume by a mode shown by the prediction mode information using the neighboring volume shown by the neighboring volume information. In addition, details of these processes are the same as those of the above-described intra prediction unit 1309 except that the information attached to the bitstream is used.
[0590] The inter prediction unit 1408 generates a prediction volume by inter prediction using the reference space in the reference space storage 1407 and the information attached to the bitstream. Specifically, the inter prediction unit 1408 applies rotation and translation processing to the reference space using the RT information of each reference space attached to the bitstream, generates a prediction volume using the reference space to which the processing is applied. In addition, in a case where the RT application flag of each reference space is present in the bitstream, the inter prediction unit 1408 applies rotation and translation processing to the reference space in accordance with the RT application flag. In addition, details of the above-described processing are the same as those of the above-described inter prediction unit 1311 except that the information attached to the bitstream is used.
[0591] Whether to decode the decoding target volume by intra prediction or by inter prediction is controlled by the prediction control unit 1409. For example, the prediction control unit 1409 selects intra prediction or inter prediction in accordance with the information attached to the bitstream and showing the prediction mode to be used. In addition, the prediction control unit 1409 can also normally select intra prediction in a case where it is decided in advance that the decoding target space is decoded by intra space.
[0592] The following describes modifications of the present embodiment. In the present embodiment, although the rotation and the translation are described as being applied in units of spaces, the rotation and the translation can be applied in finer units. For example, the three-dimensional data encoding apparatus 1300 can divide a space into subspaces and apply the rotation and the translation in units of subspaces. In this case, the three-dimensional data encoding apparatus 1300 generates the RT information for each subspace and attaches the generated RT information to a header or the like of a bitstream. Also, the three-dimensional data encoding apparatus 1300 can apply the rotation and the translation in units of volumes as an encoding unit. In this case, the three-dimensional data encoding apparatus 1300 generates the RT information in units of encoding volumes and attaches the generated RT information to a header or the like of a bitstream. Furthermore, the above can be combined. That is, the three-dimensional data encoding apparatus 1300 can apply the rotation and the translation in finer units after applying the rotation and the translation in larger units. For example, the three-dimensional data encoding apparatus 1300 can apply the rotation and the translation in units of spaces and apply different rotation and translation for each of a plurality of volumes included in the resulting space.
[0593] Also, although the rotation and the translation are described as being applied to the reference space in the present embodiment, the present embodiment is not limited thereto. For example, the three-dimensional data encoding apparatus 1300 can apply a scaling process to change the size of the three-dimensional data. Also, the three-dimensional data encoding apparatus 1300 can apply any one or two of the rotation, the translation, and the scaling. Also, as described above, in the case where the processes are applied in different units in multiple stages, the types of processes applied in the respective units can be different. For example, the rotation and the translation can be applied in units of spaces and the translation can be applied in units of volumes.
[0594] In addition, the above modifications can be similarly applied to the three-dimensional data decoding apparatus 1400.
[0595] As described above, the three-dimensional data encoding apparatus 1300 according to the present embodiment performs the following processes. Figure 48 is a flowchart of an inter-frame prediction process performed by the three-dimensional data encoding apparatus 1300.
[0596] First, the three-dimensional data encoding apparatus 1300 generates prediction position information (e.g., a prediction volume) using position information of three-dimensional points included in the object three-dimensional data (e.g., an encoding target space) and the reference three-dimensional data (e.g., a reference space) at different times (S1301). Specifically, the three-dimensional data encoding apparatus 1300 generates the prediction position information by applying the rotation and the translation to the position information of the three-dimensional points included in the reference three-dimensional data.
[0597] Further, the three-dimensional data encoding apparatus 1300 performs the rotation and translation process in the first unit (e.g., a space), and generates the prediction position information in a second unit (e.g., a volume) that is finer than the first unit. For example, the three-dimensional data encoding apparatus 1300 can search for a volume in which the difference between the position information of the encoding target volume included in the encoding target space and the position information is the smallest, from among a plurality of volumes included in the reference space after the rotation and translation process, and use the obtained volume as the prediction volume. Further, the three-dimensional data encoding apparatus 1300 can perform the rotation and translation process and the generation of the prediction position information in the same unit.
[0598] Further, the three-dimensional data encoding apparatus 1300 can apply the first rotation and translation process to the position information of the three-dimensional point included in the reference three-dimensional data in the first unit (e.g., a space), apply the second rotation and translation process to the position information of the three-dimensional point obtained by the first rotation and translation process in a second unit (e.g., a volume) that is finer than the first unit, and thereby generate the prediction position information.
[0599] Here, the position information of the three-dimensional point and the prediction position information are expressed in an octree structure as shown in FIG. 8, for example. For example, the position information of the three-dimensional point and the prediction position information are expressed in a scan order in which the depth in the octree structure is prioritized over the width. Alternatively, the position information of the three-dimensional point and the prediction position information are expressed in a scan order in which the depth in the octree structure is prioritized over the width. Figure 41
[0600] Further, as shown in FIG. 9, the three-dimensional data encoding apparatus 1300 encodes an RT application flag that indicates whether or not the rotation and translation process is applied to the position information of the three-dimensional point included in the reference three-dimensional data. That is, the three-dimensional data encoding apparatus 1300 generates an encoded signal (encoded bitstream) including the RT application flag. Further, the three-dimensional data encoding apparatus 1300 encodes RT information that indicates the content of the rotation and translation process. That is, the three-dimensional data encoding apparatus 1300 generates an encoded signal (encoded bitstream) including the RT information. Further, the three-dimensional data encoding apparatus 1300 can encode the RT information when the rotation and translation process is indicated by the RT application flag, and not encode the RT information when the rotation and translation process is not indicated by the RT application flag. Figure 46 Further, the three-dimensional data includes, for example, the position information of the three-dimensional point, and attribute information (color information, etc.) of each three-dimensional point. The three-dimensional data encoding apparatus 1300 generates prediction attribute information using the attribute information of the three-dimensional point included in the reference three-dimensional data (S1302).
[0601]
[0602] Next, the three-dimensional data encoding apparatus 1300 encodes the position information of the three-dimensional point included in the object three-dimensional data using the predicted position information. For example, as shown in FIG. 13B, the three-dimensional data encoding apparatus 1300 calculates the difference between the position information of the three-dimensional point included in the object three-dimensional data and the predicted position information, that is, the difference position information (S1303). Figure 38
[0603] Also, the three-dimensional data encoding apparatus 1300 encodes the attribute information of the three-dimensional point included in the object three-dimensional data using the predicted attribute information. For example, the three-dimensional data encoding apparatus 1300 calculates the difference between the attribute information of the three-dimensional point included in the object three-dimensional data and the predicted attribute information, that is, the difference attribute information (S1304). Next, the three-dimensional data encoding apparatus 1300 performs transformation and quantization with respect to the calculated difference attribute information (S1305).
[0604] Finally, the three-dimensional data encoding apparatus 1300 encodes the difference position information and the quantized difference attribute information (for example, entropy encoding) (S1306). That is, the three-dimensional data encoding apparatus 1300 generates an encoded signal (encoded bit stream) including the difference position information and the difference attribute information.
[0605] In addition, in a case where the attribute information is not included in the three-dimensional data, the three-dimensional data encoding apparatus 1300 can not perform steps S1302, S1304, and S1305. Also, the three-dimensional data encoding apparatus 1300 can perform only one of the encoding of the position information of the three-dimensional point and the encoding of the attribute information of the three-dimensional point.
[0606] Also, Figure 49 The order of the processes shown in FIG. 13B is merely one example, and is not limited thereto. For example, since the processes with respect to the position information (S1301, S1303) and the processes with respect to the attribute information (S1302, S1304, S1305) are independent of each other, they can be executed in any order, or a part of them can be processed in parallel.
[0607] As described above, in the present embodiment, the three-dimensional data encoding apparatus 1300 generates predicted position information using the position information of the three-dimensional point included in the object three-dimensional data and the reference three-dimensional data at different times, and encodes the difference between the position information of the three-dimensional point included in the object three-dimensional data and the predicted position information, that is, the difference position information. Thereby, since the data amount of the encoded signal can be reduced, the encoding efficiency can be improved.
[0608] Further, in the present embodiment, the three-dimensional data encoding apparatus 1300 generates predicted attribute information using attribute information of three-dimensional points included in the reference three-dimensional data, and encodes difference attribute information which is a difference between attribute information of three-dimensional points included in the object three-dimensional data and the predicted attribute information. Thereby, since it is possible to reduce the data amount of the encoded signal, it is possible to improve the encoding efficiency.
[0609] For example, the three-dimensional data encoding apparatus 1300 has a processor and a memory, and the processor performs the above-described processing using the memory.
[0610] Figure 48 is a flowchart of inter prediction processing performed by the three-dimensional data decoding apparatus 1400.
[0611] First, the three-dimensional data decoding apparatus 1400 decodes (for example, entropy decodes) the difference position information and the difference attribute information from the encoded signal (encoded bitstream) (S1401).
[0612] Further, the three-dimensional data decoding apparatus 1400 decodes the RT application flag which shows whether or not to apply the rotation and the translation processing to the position information of the three-dimensional points included in the reference three-dimensional data from the encoded signal. Further, the three-dimensional data decoding apparatus 1400 decodes the RT information which shows the content of the rotation and the translation processing. In addition, the three-dimensional data decoding apparatus 1400 decodes the RT information in a case where the RT application flag shows that the rotation and the translation processing are applied, and does not decode the RT information in a case where the RT application flag shows that the rotation and the translation processing are not applied.
[0613] Next, the three-dimensional data decoding apparatus 1400 inverse quantizes and inverse transforms the decoded difference attribute information (S1402).
[0614] Next, the three-dimensional data decoding apparatus 1400 generates predicted position information (for example, a predicted volume) using the position information of the three-dimensional points included in the object three-dimensional data (for example, a decoded object space) and the reference three-dimensional data (for example, a reference space) of different time points (S1403). Specifically, the three-dimensional data decoding apparatus 1400 generates the predicted position information by applying the rotation and the translation processing to the position information of the three-dimensional points included in the reference three-dimensional data.
[0615] More specifically, the three-dimensional data decoding apparatus 1400 applies the rotation and the translation processing to the position information of the three-dimensional points included in the reference three-dimensional data shown by the RT information in a case where the RT application flag shows that the rotation and the translation processing are applied. Further, the three-dimensional data decoding apparatus 1400 does not apply the rotation and the translation processing to the position information of the three-dimensional points included in the reference three-dimensional data in a case where the RT application flag shows that the rotation and the translation processing are not applied.
[0616] Further, the three-dimensional data decoding apparatus 1400 can perform the rotation and the translation process in the first unit (e.g., space), and can perform the generation of the predicted position information in a second unit (e.g., volume) finer than the first unit. Further, the three-dimensional data decoding apparatus 1400 can perform the rotation and the translation process and the generation of the predicted position information in the same unit.
[0617] Further, the three-dimensional data decoding apparatus 1400 can perform the rotation and the translation process in the first unit (e.g., space), and can perform the generation of the predicted position information in a second unit (e.g., volume) finer than the first unit. Further, the three-dimensional data decoding apparatus 1400 can perform the rotation and the translation process and the generation of the predicted position information in the same unit.
[0618] Here, the position information of the three-dimensional point and the predicted position information are expressed, for example, by an octree structure as shown in FIG. 8. For example, the position information of the three-dimensional point and the predicted position information are expressed in a scan order in which the width is prioritized among the depth and the width in the octree structure. Alternatively, the position information of the three-dimensional point and the predicted position information are expressed in a scan order in which the depth is prioritized among the depth and the width in the octree structure. Figure 41
[0619] The three-dimensional data decoding apparatus 1400 generates the predicted attribute information by referring to the attribute information of the three-dimensional point included in the three-dimensional data (S1404).
[0620] Next, the three-dimensional data decoding apparatus 1400 restores the position information of the three-dimensional point included in the object three-dimensional data by decoding the encoded position information included in the encoded signal using the predicted position information. Here, the encoded position information is, for example, the differential position information, and the three-dimensional data decoding apparatus 1400 restores the position information of the three-dimensional point included in the object three-dimensional data by adding the differential position information and the predicted position information (S1405).
[0621] Further, the three-dimensional data decoding apparatus 1400 restores the attribute information of the three-dimensional point included in the object three-dimensional data by decoding the encoded attribute information included in the encoded signal using the predicted attribute information. Here, the encoded attribute information is, for example, the differential attribute information, and the three-dimensional data decoding apparatus 1400 restores the attribute information of the three-dimensional point included in the object three-dimensional data by adding the differential attribute information and the predicted attribute information (S1406).
[0622] In addition, in a case where the attribute information is not included in the three-dimensional data, the three-dimensional data decoding device 1400 can not perform the steps S1402, S1404, and S1406. Also, the three-dimensional data decoding device 1400 can perform only one of decoding of the position information of the three-dimensional points and decoding of the attribute information of the three-dimensional points.
[0623] Also, Figure 50 The order of the processes illustrated is one example, and is not limited thereto. For example, since the process for the position information (S1403, S1405) and the process for the attribute information (S1402, S1404, S1406) are independent of each other, they can be performed in any order, and a part of them can be processed in parallel.
[0624] (Embodiment 8)
[0625] In this embodiment, a method of expressing three-dimensional points (point cloud) in encoding of three-dimensional data is described.
[0626] Figure 51 is a block diagram illustrating a configuration of a distribution system of three-dimensional data according to the present embodiment. Figure 51 The distribution system illustrated includes a server 1501 and a plurality of clients 1502.
[0627] The server 1501 includes a storage 1511 and a control unit 1512. The storage 1511 stores encoded three-dimensional data, i.e., an encoded three-dimensional map 1513.
[0628] Figure 52 An example of a configuration of a bitstream of the encoded three-dimensional map 1513 is illustrated. The three-dimensional map is divided into a plurality of sub-maps, and each of the sub-maps is encoded. A random access header (RA) including sub-coordinate information is attached to each of the sub-maps. The sub-coordinate information is used to improve the encoding efficiency of the sub-maps. The sub-coordinate information shows a sub-coordinate of the sub-map. The sub-coordinate is a coordinate of the sub-map with reference to a reference coordinate. In addition, a three-dimensional map including a plurality of sub-maps is referred to as an entire map. Also, in the entire map, a coordinate that becomes a reference (e.g., an origin) is referred to as a reference coordinate. That is, the sub-coordinate is a coordinate of the sub-map in a coordinate system of the entire map. In other words, the sub-coordinate shows a deviation of the coordinate system of the entire map from the coordinate system of the sub-map. Also, a coordinate in the coordinate system of the entire map with reference to the reference coordinate is referred to as an entire coordinate. A coordinate in the coordinate system of the sub-map with reference to the sub-coordinate is referred to as a differential coordinate.
[0629] The client 1502 transmits a message to the server 1501. The message includes position information of the client 1502. The control section 1512 included in the server 1501 obtains a bit stream of a sub map of a position closest to the position of the client 1502, based on the position information included in the received message. The bit stream of the sub map includes sub coordinate information, which is transmitted to the client 1502. The decoder 1521 included in the client 1502 obtains the entire coordinates of the sub map with reference to the reference coordinates, using the sub coordinate information. The application program 1522 included in the client 1502 executes an application program related to the own position, using the obtained entire coordinates of the sub map.
[0630] Also, the sub map shows a part of the entire map. The sub coordinate is a coordinate of a position where the sub map is located in the reference coordinate space of the entire map. For example, consider that there are a sub map A of AA and a sub map B of AB in the entire map of A. The vehicle starts decoding from the sub map A in a case where the map of AA is intended to be referred to, and starts decoding from the sub map B in a case where the map of AB is intended to be referred to. Here, the sub map is a random access point. Specifically, A is Osaka Prefecture, AA is Osaka City, and AB is Takatsuki City or the like.
[0631] Each sub map is transmitted to the client together with the sub coordinate information. The sub coordinate information is included in the header information of each sub map, or in a transmission packet or the like.
[0632] The reference coordinate, which is a reference of the sub coordinate information of each sub map, can also be attached to the header information of the entire map or the header information of a space higher than the sub map.
[0633] The sub map can be constituted by one space (SPC). Also, the sub map can be constituted by a plurality of SPCs.
[0634] Also, the sub map can include a group of spaces (GOS). Also, the sub map can be constituted by a world space. For example, in a case where there are a plurality of objects in the sub map, if the plurality of objects are allocated to different SPCs, the sub map is constituted by a plurality of SPCs. Also, if the plurality of objects are allocated to one SPC, the sub map is constituted by one SPC.
[0635] Next, the effect of improvement of encoding efficiency in a case where the sub coordinate information is employed will be described. Figure 53 is a diagram for explaining the effect. For example, consider that it is intended to encode the entire map of A shown in FIG. 10. Figure 53The three-dimensional point A that is shown at a position far from the reference coordinate is encoded, and thus a large number of bits is required. Here, the sub-coordinate is shorter than the distance from the reference coordinate to the three-dimensional point A. Therefore, in the case where the coordinates of the three-dimensional point A are encoded on the basis of the sub-coordinate, the coding efficiency can be improved as compared with the case where the coordinates of the three-dimensional point A are encoded on the basis of the reference coordinate. Furthermore, the bit stream of the sub-map includes the sub-coordinate information. By sending the bit stream of the sub-map and the reference coordinate to the decoding side (the client), the entire coordinates of the sub-map can be restored on the decoding side.
[0636] Figure 54 is a flowchart of the process performed by the transmitting side, i.e., the server 1501, of the sub-map.
[0637] First, the server 1501 receives a message including the position information of the client 1502 from the client 1502 (S1501). The control section 1512 obtains the encoded bit stream of the sub-map based on the position information of the client from the storage section 1511 (S1502). Then, the server 1501 transmits the encoded bit stream of the sub-map and the reference coordinate to the client 1502 (S1503).
[0638] Figure 55 is a flowchart of the process performed by the receiving side, i.e., the client 1502, of the sub-map.
[0639] First, the client 1502 receives the encoded bit stream of the sub-map and the reference coordinate transmitted from the server 1501 (S1511). Next, the client 1502 obtains the sub-map and the sub-coordinate information by decoding the encoded bit stream (S1512). Next, the client 1502 restores the differential coordinates within the sub-map to the entire coordinates using the reference coordinate and the sub-coordinate (S1513).
[0640] Next, a syntax example of information related to the sub-map will be described. In the encoding of the sub-map, the three-dimensional data encoding apparatus calculates the differential coordinates by subtracting the sub-coordinate from the coordinates of each point cloud (three-dimensional point). Then, the three-dimensional data encoding apparatus encodes the differential coordinates as a bit stream as the value of each point cloud. Furthermore, the encoding apparatus encodes the sub-coordinate information showing the sub-coordinate as the header information of the bit stream. Accordingly, the three-dimensional data decoding apparatus can obtain the entire coordinates of each point cloud. For example, the three-dimensional data encoding apparatus is included in the server 1501, and the three-dimensional data decoding apparatus is included in the client 1502.
[0641] Figure 56 A syntax example of the sub-map is shown. Figure 56NumOfPoint indicates the number of point clouds included in the sub map. sub_coordinate_x, sub_coordinate_y, and sub_coordinate_z are sub coordinate information. sub_coordinate_x indicates the x coordinate of the sub coordinate. sub_coordinate_y indicates the y coordinate of the sub coordinate. sub_coordinate_z indicates the z coordinate of the sub coordinate.
[0642] Also, diff_x[i], diff_y[i], and diff_z[i] are the differential coordinates of the i-th point cloud within the sub map. diff_x[i] indicates the differential value of the x coordinate of the i-th point cloud within the sub map from the x coordinate of the sub coordinate. diff_y[i] indicates the differential value of the y coordinate of the i-th point cloud within the sub map from the y coordinate of the sub coordinate. diff_z[i] indicates the differential value of the z coordinate of the i-th point cloud within the sub map from the z coordinate of the sub coordinate.
[0643] The three-dimensional data decoding device decodes point_cloud[i]_x, point_cloud[i]_y, and point_cloud[i]_z, which are the overall coordinates of the i-th point cloud, using the following equations. point_cloud[i]_x is the x coordinate of the overall coordinates of the i-th point cloud. point_cloud[i]_y is the y coordinate of the overall coordinates of the i-th point cloud. point_cloud[i]_z is the z coordinate of the overall coordinates of the i-th point cloud.
[0644] point_cloud[i]_x = sub_coordinate_x + diff_x[i]
[0645] point_cloud[i]_y = sub_coordinate_y + diff_y[i]
[0646] point_cloud[i]_z = sub_coordinate_z + diff_z[i]
[0647] Next, the adaptive switching process for octree encoding is described. The three-dimensional data encoding device, when performing sub map encoding, either encodes each point cloud using octree representation (hereinafter referred to as octree encoding) or encodes the differential values from the sub coordinate (hereinafter referred to as non-octree encoding). Figure 57This operation is shown in a schema. For example, the three-dimensional data encoding apparatus applies octree encoding to a submap when the number of point clouds in the submap is equal to or greater than a predetermined threshold value. The three-dimensional data encoding apparatus applies non-octree encoding to a submap when the number of point clouds in the submap is less than the threshold value. Accordingly, the three-dimensional data encoding apparatus appropriately selects whether to use octree encoding or non-octree encoding in accordance with the shape and density of the object included in the submap, and thus the encoding efficiency can be improved.
[0648] Also, the three-dimensional data encoding apparatus attaches information showing which of octree encoding and non-octree encoding is applied to the submap (hereinafter referred to as octree encoding application information) to the header of the submap or the like. Accordingly, the three-dimensional data decoding apparatus can determine whether the bitstream is a bitstream obtained by octree encoding of the submap or a bitstream obtained by non-octree encoding of the submap.
[0649] Also, the three-dimensional data encoding apparatus can calculate the encoding efficiency when octree encoding and non-octree encoding are applied to the same point cloud, respectively, and apply the encoding method with higher encoding efficiency to the submap.
[0650] Figure 58 A syntax example of a submap in a case where such switching is performed is shown. Figure 58 The coding_type shown is information showing the encoding type, and is the octree encoding application information described above. coding_type = 00 indicates that octree encoding is applied. coding_type = 01 indicates that non-octree encoding is applied. coding_type = 10 or 11 indicates that another encoding method or the like other than the above is applied.
[0651] In a case where the encoding type is non-octree encoding (non_octree), the submap includes NumOfPoint and sub-coordinate information (sub_coordinate_x, sub_coordinate_y, and sub_coordinate_z).
[0652] In a case where the encoding type is octree encoding (octree), the submap includes octree_info. The octree_info is information required in octree encoding, and includes, for example, depth information or the like.
[0653] In a case where the encoding type is non-octree encoding (non_octree), the submap includes difference coordinates (diff_x[i], diff_y[i], and diff_z[i]).
[0654] In a case where the encoding type is octree encoding, the submap includes encoding data related to the octree encoding, i.e., octree_data.
[0655] In addition, although an example in which an xyz coordinate system is employed is shown as the coordinate system of the point cloud here, a polar coordinate system can be employed.
[0656] Figure 59 is a flowchart of a three-dimensional data encoding process performed by a three-dimensional data encoding apparatus. First, the three-dimensional data encoding apparatus calculates the number of point clouds in a processing target submap, i.e., a target submap (S1521). Next, the three-dimensional data encoding apparatus determines whether the calculated number of point clouds is equal to or greater than a threshold value that is specified in advance (S1522).
[0657] In a case where the number of point clouds is equal to or greater than the threshold value (YES in S1522), the three-dimensional data encoding apparatus applies octree encoding to the target submap (S1523). Also, the three-dimensional point data encoding apparatus attaches octree encoding application information indicating that octree encoding is applied to the target submap to a header of a bitstream (S1525).
[0658] In addition, in a case where the number of point clouds is less than the threshold value (NO in S1522), the three-dimensional data encoding apparatus applies non-octree encoding to the target submap (S1524). Also, the three-dimensional point data encoding apparatus attaches octree encoding application information indicating that non-octree encoding is applied to the target submap to the header of the bitstream (S1525).
[0659] Figure 60 is a flowchart of a three-dimensional data decoding process performed by a three-dimensional data decoding apparatus. First, the three-dimensional data decoding apparatus decodes octree encoding application information from a header of a bitstream (S1531). Next, the three-dimensional data decoding apparatus determines whether the encoding type applied to a target submap is octree encoding based on the decoded octree encoding application information (S1532).
[0660] In a case where the encoding type indicated by the octree encoding application information is octree encoding (YES in S1532), the three-dimensional data decoding apparatus decodes the target submap using octree decoding (S1533). In a case where the encoding type indicated by the octree encoding application information is non-octree encoding (NO in S1532), the three-dimensional data decoding apparatus decodes the target submap using non-octree decoding (S1534).
[0661] A modification of the present embodiment will be described below. Figures 61 to 63 An operation of a modification of the switching process of the encoding type is shown in a flowchart.
[0662] As Figure 61As shown, the 3D data encoding device can select whether to use octree encoding or non-octree encoding for each space. In this case, the 3D data encoding device appends octree encoding application information to the header of the space. Accordingly, the 3D data decoding device can determine whether octree encoding is applicable for each space. Furthermore, in this case, the 3D data encoding device sets sub-coordinates for each space and encodes the difference value obtained by subtracting the sub-coordinate value from the coordinates of each point cloud in the space.
[0663] Therefore, since the 3D data encoding device can appropriately switch between octree encoding and other methods based on the shape of the object or the number of point clouds in space, the encoding efficiency can be improved.
[0664] And, as Figure 62 As shown, the 3D data encoding device can select whether to use octree encoding or non-octree encoding for each volume. In this case, the 3D data encoding device appends octree encoding application information to the header of the volume. Accordingly, the 3D data decoding device can determine whether octree encoding is applicable for each volume. Furthermore, in this case, the 3D data encoding device sets sub-coordinates for each volume and encodes the difference value obtained by subtracting the sub-coordinates from the coordinates of each point cloud within the volume.
[0665] Therefore, since the 3D data encoding device can appropriately switch between octree encoding and other methods based on the shape of the object or the number of point clouds within the volume, the encoding efficiency can be improved.
[0666] Furthermore, the above description, as a non-octree coding example, shows encoding the difference after subtracting the sub-coordinates from the coordinates of each point cloud. However, this is not a limitation, and any coding method other than octree coding can be used. For example... Figure 63 As shown, the 3D data encoding device can also use a method of encoding the values of the point cloud within the sub-map, space, or volume (hereinafter referred to as original coordinate encoding) instead of sub-coordinate difference, as a non-octree encoding.
[0667] In this scenario, the 3D data encoding device stores information showing that original coordinate encoding has been applied to the object space (submap, space, or volume) in the header. Based on this, the 3D data decoding device can determine whether original coordinate encoding has been applied to the object space.
[0668] Furthermore, when using original coordinate encoding, the 3D data encoding device can encode the original coordinates without applying quantization and arithmetic encoding. Moreover, the 3D data encoding device can encode the original coordinates with a pre-defined fixed bit length. Accordingly, the 3D data encoding device can generate a stream with a certain bit length at a specific timing.
[0669] Also, in the above description, although an example of encoding a difference from a sub-coordinate subtracted from a coordinate of each point cloud is shown as non-octree encoding, the present application is not limited to this.
[0670] For example, the three-dimensional data encoding apparatus can sequentially encode difference values between coordinates of each point cloud. Figure 64 is a diagram for explaining the operation in this case. For example, in the example shown in Figures 65 to 67 In the example shown in the example shown in
[0671] Also, in the above description, the sub-coordinate is the coordinate of the left lower front corner of the sub-map, but the position of the sub-coordinate is not limited to this. Figure 65 Another example of the position of the sub-coordinate is shown. As for the set position of the sub-coordinate, it can be set to an arbitrary coordinate within the object space (sub-map, space, or volume). That is, as described above, the sub-coordinate can be the coordinate of the left lower front corner. As shown in Figure 66 As shown in Figure 67 As shown in
[0672] Also, the set position of the sub-coordinate can be the same as the coordinate of a certain point cloud within the object space (sub-map, space, or volume). For example, in the example shown in Figure 68 In the example shown in
[0673] Also, in the present embodiment, although a case where switching between application of octree encoding and application of non-octree encoding is shown, this is not limiting. For example, the three-dimensional data encoding apparatus can also switch between application of another tree structure other than an octree and application of a non-tree structure other than the tree structure. For example, another tree structure refers to a kd-tree or the like that is partitioned using a plane perpendicular to one of the coordinate axes. Also, as another tree structure, any method can be used.
[0674] Also, in the present embodiment, although a case where coordinate information possessed by a point cloud is encoded is shown, this is not limiting. For example, the three-dimensional data encoding apparatus can also encode color information, three-dimensional feature amounts, or feature amounts of visible light, or the like, in the same manner as the coordinate information. For example, the three-dimensional data encoding apparatus can also set an average value of color information possessed by each point cloud within a submap as sub-color information, and encode a difference between the color information of each point cloud and the sub-color information.
[0675] Also, in the present embodiment, although a case where an encoding method with high encoding efficiency (octree encoding or non-octree encoding) is selected in accordance with the number of point clouds or the like is shown, this is not limiting. For example, as the three-dimensional data encoding apparatus on the server side, a bitstream of a point cloud encoded by octree encoding, a bitstream of a point cloud encoded by non-octree encoding, and a bitstream of a point cloud encoded by both methods can be held in advance, and a bitstream transmitted to the three-dimensional data decoding apparatus can be switched in accordance with the communication environment or the processing capacity of the three-dimensional data decoding apparatus.
[0676] Figure 68 A syntax example of a volume in a case where switching of application of octree encoding is shown. Figure 58 The syntax shown is basically the same as the syntax shown in Figure 59 The syntax shown is basically the same as the syntax shown in
[0677] Also, diff_x[i], diff_y[i], and diff_z[i] are difference coordinates of the i-th point cloud within the volume. diff_x[i] indicates a difference value between the x coordinate of the i-th point cloud within the volume and the x coordinate of the sub-coordinate. diff_y[i] indicates a difference value between the y coordinate of the i-th point cloud within the volume and the y coordinate of the sub-coordinate. diff_z[i] indicates a difference value between the z coordinate of the i-th point cloud within the volume and the z coordinate of the sub-coordinate.
[0678] In addition, in a case where the relative positions of the volumes in the space can be calculated, the three-dimensional data encoding apparatus can not include the sub-coordinate information in the header of the volume. That is, the three-dimensional data encoding apparatus can not include the sub-coordinate information in the header, and calculate the relative positions of the volumes in the space, and use the calculated positions as the sub-coordinates of the respective volumes.
[0679] As described above, the three-dimensional data encoding apparatus according to the present embodiment determines whether to encode an object spatial unit among the plurality of spatial units (e.g., sub-maps, spaces, or volumes) included in the three-dimensional data in the octree structure (e.g., Figure 60 For example, the three-dimensional data encoding apparatus determines to encode the object spatial unit in the octree structure in a case where the number of three-dimensional points included in the object spatial unit is more than a threshold value prescribed in advance. Also, the three-dimensional data encoding apparatus determines not to encode the object spatial unit in the octree structure in a case where the number of three-dimensional points included in the object spatial unit is equal to or less than the threshold value.
[0680] In a case where it is determined to encode the object spatial unit in the octree structure (YES in S1522), the three-dimensional data encoding apparatus encodes the object spatial unit in the octree structure (S1523). Also, in a case where it is determined not to encode the object spatial unit in the octree structure (NO in S1522), the three-dimensional data encoding apparatus encodes the object spatial unit in a manner different from the octree structure (S1524). For example, as the different manner, the three-dimensional data encoding apparatus encodes the coordinates of the three-dimensional points included in the object spatial unit. Specifically, as the different manner, the three-dimensional data encoding apparatus encodes the difference between the reference coordinates of the object spatial unit and the coordinates of the three-dimensional points included in the object spatial unit.
[0681] Next, the three-dimensional data encoding apparatus attaches information showing whether to encode the object spatial unit in the octree structure to the bitstream (S1525).
[0682] According to this, the three-dimensional data encoding apparatus can reduce the data amount of the encoded signal, and thus can improve the encoding efficiency.
[0683] For example, the three-dimensional data encoding apparatus includes a processor and a memory, and the processor performs the above-described processing using the memory.
[0684] Also, the three-dimensional data decoding apparatus according to the present embodiment decodes, from the bitstream, information showing whether to decode an object spatial unit among a plurality of object spatial units (e.g., sub-maps, spaces, or volumes) included in the three-dimensional data in the octree structure (e.g., Figure 69The three-dimensional data decoding device decodes the object space unit in the octree structure (S1533) if the information indicates that the object space unit is decoded in the octree structure (YES in S1532).
[0685] The three-dimensional data decoding device decodes the object space unit in a manner different from the octree structure (S1534) if the information indicates that the object space unit is not decoded in the octree structure (NO in S1532). For example, the three-dimensional data decoding device decodes coordinates of the three-dimensional points included in the object space unit in the different manner. Specifically, the three-dimensional data decoding device decodes a difference between a reference coordinate of the object space unit and the coordinates of the three-dimensional points included in the object space unit in the different manner.
[0686] Thus, the three-dimensional data decoding device can reduce the amount of data of the encoded signal, and thus can improve the encoding efficiency.
[0687] For example, the three-dimensional data decoding device includes a processor and a memory, and the processor performs the above-described processing using the memory.
[0688] (Embodiment 9)
[0689] In this embodiment, an encoding method of a tree structure such as an octree structure is described.
[0690] By recognizing an important area, three-dimensional data of the important area is preferentially decoded, and thus the efficiency can be improved.
[0691] Figure 70 This is a diagram illustrating an example of an important area in a three-dimensional map. The important area is, for example, an area including three-dimensional points having values of a feature amount in the three-dimensional points in the three-dimensional map that are greater than or equal to a certain number. Alternatively, the important area can be, for example, an area including three-dimensional points required in a case where a client such as a vehicle performs self-position estimation. Alternatively, the important area can be an area of a face in a three-dimensional model of a person. In this way, the important area can be defined for each application, and the important area can be switched according to the application.
[0692] In this embodiment, as a manner of expressing an octree structure and the like, occupancy coding and location coding are used. In addition, a bit string obtained by occupancy coding is referred to as an occupancy code. A bit string obtained by location coding is referred to as a location code.
[0693] Figure 70is a diagram showing an example of the occupancy code. Figure 70 An example of the occupancy code representing a quadtree structure. In Figure 70 , the occupancy code is assigned to each node. Each occupancy code represents whether or not a three-dimensional point is contained in a child node or a leaf node of each node. For example, in the case of a quadtree, information representing whether or not four child nodes or leaf nodes possessed by each node respectively contain a three-dimensional point is represented by an occupancy code of 4 bits. In addition, in the case of an octree, information representing whether or not eight child nodes or leaf nodes possessed by each node respectively contain a three-dimensional point is represented by an occupancy code of 8 bits. In addition, here, a quadtree structure is described for the sake of simplicity of explanation, but the same can be applied to an octree structure. For example, as shown in Figure 40 , the occupancy code is a bit string in which nodes and leaf nodes are scanned in a breadth-first manner as described in Figure 40 and the like. In the occupancy code, since information of a plurality of three-dimensional points is decoded in a fixed order, any three-dimensional point information cannot be decoded preferentially. In addition, the occupancy code can be a bit string in which nodes and leaf nodes are scanned in a depth-first manner as described in Figure 71 and the like.
[0694] Next, position encoding is described. By using the position code, an important part in an octree structure can be decoded directly. In addition, an important three-dimensional point in a deep layer can be encoded efficiently.
[0695] Figure 71 is a diagram for explaining position encoding, and is a diagram showing an example of a quadtree structure. In Figure 72 , three-dimensional points A to I are represented by a quadtree structure. In addition, three-dimensional points A and C are important three-dimensional points contained in an important region.
[0696] Figure 71 is a diagram showing an occupancy code and a position code representing important three-dimensional points A and C in the quadtree structure shown in Figure 71
[0697] In position encoding, in a tree structure, an index of a node and an index of a leaf node existing in a path up to a leaf node to which an object three-dimensional point belonging to an object three-dimensional point to be encoded belongs are encoded. Here, the index is a numerical value assigned to each node and leaf node. In other words, the index refers to an identifier for identifying a plurality of child nodes of an object node. As shown in Figure 71 , in the case of a quadtree, the index represents any one of 0 to 3.
[0698] For example, in Figure 72 In the quadtree structure shown, in the case where the leaf node A is a three-dimensional point of an object, the leaf node A appears as 0→2→1→0→1→2→1. Here, in the case where the maximum value of each index is 4 (which can be expressed in 2 bits) in the right figure, 7 x 2 bits = 14 bits are required for the bit number of the position code of the leaf node A. In the case where the leaf node C is an encoding object, the same number of bits is required. In the case of an octree, since the maximum value of each index is 8 (which can be expressed in 3 bits), the required bit number can be calculated in 3 bits x the depth of the leaf node. In addition, the three-dimensional data encoding apparatus can also perform entropy encoding after binarizing each index to reduce the data amount.
[0699] In addition, as shown in Figure 72 In the occupancy code, in order to decode the leaf nodes A and C, all nodes above them need to be decoded. On the other hand, in the position code, only the data of the leaf nodes A and C can be decoded. Thus, as shown in Figure 72 By using the position code, the bit number can be reduced compared to the occupancy code.
[0700] In addition, as shown in Figure 73 By performing dictionary compression such as LZ77 on part or all of the position code, the code amount can be further reduced.
[0701] Next, an example of applying position encoding to a three-dimensional point (point cloud) obtained by a LiDAR will be described. Figure 74 is a diagram showing an example of a three-dimensional point obtained by a LiDAR. The three-dimensional point obtained by the LiDAR is sparse. That is, in the case where the three-dimensional point is expressed in the occupancy code, the number of zeros increases. In addition, high three-dimensional precision is requested for the three-dimensional point. That is, the hierarchy of the octree structure becomes deeper.
[0702] Figure 74 is a diagram showing an example of such a sparse deep octree structure. Figure 75 The occupancy code of the octree structure shown is 136 bits (= 8 bits x 17 nodes). In addition, since the depth is 6 and there are 6 three-dimensional points, the position code is 3 bits x 6 x 6 = 108 bits. That is, the position code can reduce the code amount by 20% with respect to the occupancy code. In this way, by applying position encoding to a sparse deep octree structure, the code amount can be reduced.
[0703] The code amount of the occupancy code and the position code will be described below. In the case where the depth of the octree structure is 10, the maximum number of three-dimensional points is 8 10 = 1073741824. In addition, the bit number L o of the occupancy code of the octree structure is represented by the following.
[0704] L o = 8 + 8 2 +... + 8 10 = 127133512 bits
[0705] Therefore, the number of bits per three-dimensional point is 1.143 bits. In addition, in the occupancy code, the number of bits does not change even if the number of three-dimensional points included in the octree structure changes.
[0706] On the other hand, in the position code, the number of bits per three-dimensional point directly affects the depth of the octree structure. Specifically, the number of bits of the position code of each three-dimensional point is 3 bits x depth 10 = 30 bits.
[0707] Therefore, the number of bits L l of the position code of the octree structure is represented by the following expression.
[0708] L l = 30 x N
[0709] Here, N is the number of three-dimensional points included in the octree structure.
[0710] Therefore, in the case of N < L o / 30 = 40904450.4, that is, in the case of the number of three-dimensional points being less than 40904450, the code amount of the position code becomes less than that of the occupancy code (L l < L o ).
[0711] Thus, in the case of a small number of three-dimensional points, the code amount of the position code is less than that of the occupancy code, and in the case of a large number of three-dimensional points, the code amount of the position code is more than that of the occupancy code.
[0712] Therefore, the three-dimensional data encoding apparatus can also switch which one of the position encoding and the occupancy encoding to use depending on the number of input three-dimensional points. In this case, the three-dimensional data encoding apparatus can also attach information indicating which one of the position encoding and the occupancy encoding is encoded in the header information or the like of the bit stream.
[0713] Next, a hybrid encoding combining the position encoding and the occupancy encoding will be described. The hybrid encoding combining the position encoding and the occupancy encoding is effective in the case of encoding an important region that is dense. Figure 75 is a diagram illustrating this example. In Figure 76In the example shown, the important three-dimensional points are densely arranged. In this case, the three-dimensional data encoding device encodes the upper layer with a position code and the lower layer with an occupancy code. Specifically, the position code is used up to the deepest common node, and the occupancy code is used at a position deeper than the deepest common node. Here, the deepest common node refers to the deepest node among the nodes that become the ancestor of the common of a plurality of important three-dimensional points.
[0714] Next, a hybrid encoding that prioritizes compression efficiency is described. The three-dimensional data encoding device can switch the position code and the occupancy code in accordance with a rule that is predetermined in the encoding of the octree.
[0715] Figure 76 is a diagram showing one example of the rule. First, the three-dimensional data encoding device confirms the proportion of the nodes that contain the three-dimensional points in each level (depth). In the case where the proportion is higher than a predetermined threshold, the three-dimensional data encoding device encodes several nodes of the upper layer of the target level with the occupancy code. For example, the three-dimensional data encoding device applies the occupancy code to the levels from the target level to the deepest common node.
[0716] For example, in the case where the proportion of the nodes that contain the three-dimensional points in the third level is higher than the threshold, the three-dimensional data encoding device applies the occupancy code to the second and third levels from the third level to the deepest common node and applies the position code to the first and fourth levels other than this. Figure 77 In the example shown, the proportion of the nodes that contain the three-dimensional points in the third level is higher than the threshold. Therefore, the three-dimensional data encoding device applies the occupancy code to the second and third levels from the third level to the deepest common node and applies the position code to the first and fourth levels other than this.
[0717] The method of calculating the above threshold is described. In one layer of the octree structure, there is one root node and eight child nodes. Therefore, in the occupancy code, eight bits are required to encode one layer of the octree structure. On the other hand, in the position code, three bits are required for each child node that contains a three-dimensional point. Therefore, in the case where the number of nodes that contain a three-dimensional point is greater than two, the occupancy code is more efficient than the position code. That is, in this case, the threshold is two.
[0718] Next, an example of the configuration of the bit stream generated by the above position code, occupancy code, or hybrid encoding is described.
[0719] Figure 77 is a diagram showing one example of the bit stream generated by the position code. As shown in Figure 77 , the bit stream generated by the position code contains a header and a plurality of position codes. Each position code is for one three-dimensional point.
[0720] With this configuration, the three-dimensional data decoding device can decode a plurality of three-dimensional points with high precision respectively. In addition, Figure 78An example of a bitstream in the case of a quadtree structure. In the case of an octree structure, each index can take a value of 0 to 7.
[0721] In addition, the three-dimensional data encoding device can perform entropy encoding after binarizing the column (string) of indices indicating one three-dimensional point. For example, in the case where the column of indices is 0121, the three-dimensional data encoding device can binarize 0121 into 00011001, and perform arithmetic encoding on the bit string.
[0722] Figure 78 is a diagram indicating one example of a bitstream generated by hybrid encoding including important three-dimensional points. As shown in Figure 78 , a position code of an upper layer, an occupancy code of important three-dimensional points of a lower layer, and an occupancy code of non-important three-dimensional points other than the important three-dimensional points of the lower layer are sequentially arranged. In addition, Figure 79 , the position code length indicates the code amount of the position code following. In addition, the occupancy code amount indicates the code amount of the occupancy code following.
[0723] With this configuration, the three-dimensional data decoding device can select different decoding plans according to the application program.
[0724] In addition, the encoding data of the important three-dimensional points is stored near the beginning of the bitstream, and the encoding data of the non-important three-dimensional points not included in the important region is stored after the encoding data of the important three-dimensional points.
[0725] Figure 78 is a diagram indicating a tree structure represented by Figure 80 the occupancy code of the important three-dimensional points. Figure 78 is a diagram indicating a tree structure represented by the occupancy code of the non-important three-dimensional points. As shown in Figure 79 , in the occupancy code of the important three-dimensional points, information related to the non-important three-dimensional points is excluded. Specifically, the nodes 0 and 3 at depth 5 do not include important three-dimensional points, and therefore, the nodes 0 and 3 are assigned a value 0 indicating that no three-dimensional points are included.
[0726] On the other hand, as shown in Figure 80 , in the occupancy code of the non-important three-dimensional points, information related to the important three-dimensional points is excluded. Specifically, the node 1 at depth 5 does not include non-important three-dimensional points, and therefore, the node 1 is assigned a value 0 indicating that no three-dimensional points are included.
[0727] Thus, the three-dimensional data encoding apparatus divides the original tree structure into a first tree structure including important three-dimensional points and a second tree structure including unimportant three-dimensional points, and independently encodes the first tree structure and the second tree structure by occupancy coding. Thereby, the three-dimensional data decoding apparatus can preferentially decode the important three-dimensional points.
[0728] Next, a configuration example of a bitstream generated by the hybrid coding that emphasizes efficiency will be described. Figure 81 is a diagram showing a configuration example of a bitstream generated by the hybrid coding that emphasizes efficiency. As shown in Figure 81 , for each sub-tree, a sub-tree root node position, an occupancy code amount, and an occupancy code are sequentially arranged. Figure 81 The sub-tree position shown in
[0729] In the above configuration, in a case where only one of the position coding and the occupancy coding is applied to the octree structure, the following holds.
[0730] In a case where the length of the position coding of the root node of the sub-tree is equal to the depth of the octree structure, the sub-tree does not have a child node. That is, the position coding is applied to the entire tree structure.
[0731] In a case where the root node of the sub-tree is equal to the root node of the octree structure, the occupancy coding is applied to the entire tree structure.
[0732] For example, based on the above rule, the three-dimensional data decoding apparatus can determine whether the position code or the occupancy code is included in the bitstream.
[0733] In addition, the bitstream can include coding mode information indicating which one of the position coding, the occupancy coding, and the hybrid coding is used. Figure 82 is a diagram showing an example of a bitstream in this case. For example, as shown in Figure 82 , 2-bit coding mode information indicating the coding mode is attached to the bitstream.
[0734] In addition, (1) the "number of three-dimensional points" in the position coding indicates the number of three-dimensional points that follow. In addition, (2) the "occupancy code amount" in the occupancy coding indicates the code amount of the occupancy code that follows. In addition, (3) the "number of important sub-trees" in the hybrid coding (important three-dimensional points) indicates the number of sub-trees including important three-dimensional points. Further, (4) the "number of occupancy sub-trees" in the hybrid coding (emphasizing efficiency) indicates the number of sub-trees after occupancy coding.
[0735] Next, a syntax example used for switching the application of the occupancy coding and the position coding will be described. Figure 83 is a diagram showing this syntax example.
[0736] Figure 83 isleaf is a flag indicating whether the object node is a leaf node. isleaf = 1 indicates that the object node is a leaf node, and isleaf = 0 indicates that the object node is not a leaf node but a node.
[0737] In the case where the object node is a leaf node, point_flag is appended to the bitstream. point_flag is a flag indicating whether the object node (leaf node) contains a three-dimensional point. point_flag = 1 indicates that the object node contains a three-dimensional point, and point_flag = 0 indicates that the object node does not contain a three-dimensional point.
[0738] In the case where the object node is not a leaf node, coding_type is appended to the bitstream. coding_type is coding type information indicating the coding type applied. coding_type = 00 indicates that position coding is applied, coding_type = 01 indicates that occupancy coding is applied, coding_type = 10 or 11 indicates that another coding mode is applied, and the like.
[0739] In the case where the coding type is position coding, numPoint, num_idx[i], and idx[i][j] are appended to the bitstream.
[0740] numPoint indicates the number of three-dimensional points subjected to position coding. num_idx[i] indicates the number (depth) of indices from the object node to three-dimensional point i. In the case where all the three-dimensional points subjected to position coding are located at the same depth, num_idx[i] are all the same value. Therefore, it is also possible that, in the case where numPoint is 1, Figure 83 num_idx is defined as a common value before the for statement (for (i = 0; i < numPoint; i++) {}) shown above.
[0741] idx[i][j] indicates the value of the j-th index from the object node to the index of three-dimensional point i. In the case of an octree, the number of bits of idx[i][j] is 3 bits.
[0742] Further, as described above, an index refers to an identifier for identifying a plurality of child nodes of an object node. In the case of an octree, idx[i][j] indicates any one of 0 to 7. Further, in the case of an octree, there are eight child nodes, each of which corresponds to each of eight sub-blocks obtained by spatially dividing an object block corresponding to the object node by eight. Therefore, idx[i][j] can also be information indicating a three-dimensional position of a sub-block corresponding to a child node. For example, idx[i][j] can also be 3-bit information including each 1 bit of information indicating the position of x, y, and z of a sub-block.
[0743] In a case where the coding type is the occupancy coding, the occupancy_code is appended to the bitstream. The occupancy_code is an occupancy code of the object node. In a case of the octree, the occupancy_code is a bit string of 8 bits such as a bit string "00101000" or the like.
[0744] In a case where the value of the (i+1)th bit of the occupancy_code is 1, the processing of the child node is shifted. That is, the child node is set as the next object node, and the bit string is generated recursively.
[0745] In the present embodiment, an example in which the end of the octree is represented by appending the leaf node information (isleaf, point_flag) to the bitstream is shown, but is not necessarily limited thereto. For example, the three-dimensional data encoding apparatus can append the maximum depth (depth) from the start node (root node) to the end (leaf node) where the three-dimensional point exists to the header of the start node from the occupancy code. Then, the three-dimensional data encoding apparatus can also bit string the information of the child node while increasing the depth from the start node recursively, and determine that the leaf node is reached at the point in time when the depth becomes the maximum depth. Further, the three-dimensional data encoding apparatus can append the information representing the maximum depth to the initial node where the coding_type becomes the occupancy coding, or can append it to the start node (root node) of the octree.
[0746] As described above, the three-dimensional data encoding apparatus can also append the information for switching the occupancy coding and the position coding as the header information of each node in the bitstream.
[0747] Further, the three-dimensional data encoding apparatus can also entropy-encode the coding_type, the numPoint, the num_idx, the idx, and the occupancy_code of each node generated by the above-described method. For example, the three-dimensional data encoding apparatus performs arithmetic encoding after binarizing each value.
[0748] In addition, in the above-described syntax, a case where the depth-first bit string using the octree structure is used as the occupancy code is exemplified, but is not necessarily limited thereto. The three-dimensional data encoding apparatus can also use the width-first bit string using the octree structure as the occupancy code. The three-dimensional data encoding apparatus can also append the information for switching the occupancy coding and the position coding as the header information of each node in the bitstream in a case where the width-first bit string is used.
[0749] In the present embodiment, the representation is performed using the octree structure as an example, but is not necessarily limited thereto, and the above-described method can also be applied to the N-ary tree (N is an integer of 2 or more) such as the quadtree and the hexoctree or other tree structures.
[0750] Next, a flow of the encoding process that switches the application of the occupancy coding and the position coding will be described. Figure 84 is a flowchart of the encoding process of the present embodiment.
[0751] First, the three-dimensional data encoding apparatus expresses a plurality of three-dimensional points included in the three-dimensional data using an octree structure (S1601). Next, the three-dimensional data encoding apparatus sets a root node in the octree structure as an object node (S1602). Next, the three-dimensional data encoding apparatus generates a bit string of the octree structure by performing a node encoding process on the object node (S1603). Next, the three-dimensional data encoding apparatus generates a bit stream by performing entropy coding on the generated bit string (S1604).
[0752] Figure 85 is a flowchart of the node encoding process (S1603). First, the three-dimensional data encoding apparatus determines whether the object node is a leaf node (S1611). In the case where the object node is not a leaf node (NO in S1611), the three-dimensional data encoding apparatus sets a leaf node flag (isleaf) to 0 and appends the leaf node flag to the bit string (S1612).
[0753] Next, the three-dimensional data encoding apparatus determines whether the number of child nodes that include three-dimensional points is greater than a threshold value that is specified in advance (S1613). In addition, the three-dimensional data encoding apparatus can also append the threshold value to the bit string.
[0754] In the case where the number of child nodes that include three-dimensional points is greater than the threshold value that is specified in advance (YES in S1613), the three-dimensional data encoding apparatus sets a coding type (coding_type) to occupancy coding and appends the coding type to the bit string (S1614).
[0755] Next, the three-dimensional data encoding apparatus sets occupancy coding information and appends the occupancy coding information to the bit string. Specifically, the three-dimensional data encoding apparatus generates an occupancy code of the object node and appends the occupancy code to the bit string (S1615).
[0756] Next, the three-dimensional data encoding apparatus sets a next object node in accordance with the occupancy code (S1616). Specifically, the three-dimensional data encoding apparatus sets an unprocessed child node for which the occupancy code is "1" as the next object node.
[0757] Next, the three-dimensional data encoding apparatus performs a node encoding process on the newly set object node (S1617). That is, the three-dimensional data encoding apparatus performs the processing illustrated in Figure 85
[0758] In a case where the processing of all child nodes is not completed (NO in S1618), the processing after S1616 is performed again. On the other hand, in a case where the processing of all child nodes is completed (YES in S1618), the three-dimensional data encoding apparatus ends the node encoding processing.
[0759] Further, in a case where the number of child nodes including the three-dimensional point is equal to or less than a predetermined threshold in S1613 (NO in S1613), the three-dimensional data encoding apparatus sets the encoding type to position encoding, and attaches the encoding type to the bit string (S1619).
[0760] Next, the three-dimensional data encoding apparatus sets position encoding information, and attaches the position encoding information to the bit string. Specifically, the three-dimensional data encoding apparatus generates a position code, and attaches the position code to the bit string (S1620). The position code includes numPoint, num_idx, and idx.
[0761] Further, in a case where the object node is a leaf node in S1611 (YES in S1611), the three-dimensional data encoding apparatus sets a leaf node flag to 1, and attaches the leaf node flag to the bit string (S1621). Further, the three-dimensional data encoding apparatus sets information indicating whether or not the leaf node includes a three-dimensional point, that is, a point flag (point_flag), and attaches the point flag to the bit string (S1622).
[0762] Next, a flow example of the decoding processing that switches the application of the occupancy encoding and the position encoding will be described. Figure 85 is a flowchart of the decoding processing of the present embodiment.
[0763] The three-dimensional data decoding apparatus generates a bit string by entropy decoding the bit stream (S1631). Next, the three-dimensional data decoding apparatus restores an octree structure by performing node decoding processing on the obtained bit string (S1632). Next, the three-dimensional data decoding apparatus generates three-dimensional points from the restored octree structure (S1633).
[0764] Figure 87 is a flowchart of the node decoding processing (S1632). First, the three-dimensional data decoding apparatus obtains (decodes) a leaf node flag (isleaf) from the bit string (S1641). Next, the three-dimensional data decoding apparatus determines whether or not the object node is a leaf node based on the leaf node flag (S1642).
[0765] In a case where the object node is not a leaf node (NO in S1642), the three-dimensional data decoding apparatus obtains an encoding type (coding_type) from the bit string (S1643). The three-dimensional data decoding apparatus determines whether or not the encoding type is occupancy encoding (S1644).
[0766] In a case where the encoding type is the occupancy encoding (YES in S1644), the three-dimensional data decoding apparatus obtains the occupancy encoding information from the bit string. Specifically, the three-dimensional data decoding apparatus obtains the occupancy code from the bit string (S1645).
[0767] Next, the three-dimensional data decoding apparatus sets the next object node in accordance with the occupancy encoding (S1646). Specifically, the three-dimensional data decoding apparatus sets an unprocessed child node for which the occupancy code is "1" as the next object node.
[0768] Next, the three-dimensional data decoding apparatus performs the node decoding process on the newly set object node (S1647). That is, the three-dimensional data decoding apparatus performs the node decoding process on the newly set object node in accordance with the processing shown in FIG. 17. Figure 87
[0769] In a case where the processing of all the child nodes is not completed (NO in S1648), the processing after S1646 is performed again. On the other hand, in a case where the processing of all the child nodes is completed (YES in S1648), the three-dimensional data decoding apparatus ends the node decoding process.
[0770] Further, in a case where the encoding type is the position encoding in S1644 (NO in S1644), the three-dimensional data decoding apparatus obtains the position encoding information from the bit string. Specifically, the three-dimensional data decoding apparatus obtains the position code from the bit string (S1649). The position code includes numPoint, num_idx, and idx.
[0771] Further, in a case where the object node is the leaf node in S1642 (YES in S1642), the three-dimensional data decoding apparatus obtains information indicating whether or not the leaf node includes the three-dimensional point, that is, the point flag (point_flag), from the bit string (S1650).
[0772] Further, in the present embodiment, an example in which the encoding type is switched per node is shown, but it is not necessarily limited thereto. The encoding type can be fixed in a volume, space, or world space unit. In this case, the three-dimensional data encoding apparatus can also attach the encoding type information to the header information of the volume, space, or world space.
[0773] As described above, the three-dimensional data encoding apparatus of the present embodiment generates first information that represents an N (N is an integer of 2 or more) -ary tree structure of a plurality of three-dimensional points included in three-dimensional data in a first manner (position encoding), and generates a bitstream that includes the first information. The first information includes three-dimensional point information (position code) corresponding to each of the plurality of three-dimensional points. Each of the three-dimensional point information includes an index (idx) corresponding to each of a plurality of layers in the N-ary tree structure. Each of the indices represents a sub-block to which the corresponding three-dimensional point belongs among N sub-blocks belonging to the corresponding layer.
[0774] In other words, each of the three-dimensional point information represents a path in the N-ary tree structure to the corresponding three-dimensional point. Each of the indices represents a sub-node included in the path among N sub-nodes belonging to the corresponding layer (node).
[0775] Thus, the three-dimensional data encoding method can generate a bitstream that can selectively decode three-dimensional points.
[0776] For example, the three-dimensional point information (position code) includes information (num_idx) that represents a number of indices included in the three-dimensional point information. In other words, the information represents a depth (number of layers) in the N-ary tree structure to the corresponding three-dimensional point.
[0777] For example, the first information includes information (numPoint) that represents a number of three-dimensional point information included in the first information. In other words, the information represents a number of three-dimensional points included in the N-ary tree structure.
[0778] For example, N is 8, and the index is 3 bits.
[0779] For example, the three-dimensional data encoding apparatus has a first encoding mode that generates the first information, and a second encoding mode that generates second information (occupancy code) that represents the N-ary tree structure in a second manner (occupancy encoding), and generates a bitstream that includes the second information. The second information includes a plurality of 1-bit information that corresponds to each of a plurality of sub-blocks belonging to a plurality of layers in the N-ary tree structure, and represents whether or not a three-dimensional point exists in the corresponding sub-block.
[0780] For example, the three-dimensional data encoding apparatus uses the first encoding mode when a number of the plurality of three-dimensional points is equal to or less than a predetermined threshold value, and uses the second encoding mode when the number of the plurality of three-dimensional points is more than the threshold value. Thus, the three-dimensional data encoding apparatus can reduce a code amount of the bitstream.
[0781] For example, the first information and the second information include information (encoding mode information) that represents whether the information represents the N-ary tree structure in the first manner or the N-ary tree structure in the second manner.
[0782] For example, as described above, the three-dimensional data encoding apparatus generates the first information that represents the N-ary tree structure in the first manner when a number of the plurality of three-dimensional points is equal to or less than a predetermined threshold value, and generates the second information that represents the N-ary tree structure in the second manner when the number of the plurality of three-dimensional points is more than the threshold value. Figure 75As illustrated in FIG. 1, the three-dimensional data encoding apparatus uses the first encoding mode in a part of the N-ary tree structure and uses the second encoding mode in another part of the N-ary tree structure.
[0783] For example, the three-dimensional data encoding apparatus includes a processor and a memory, and the processor performs the above-described processing using the memory.
[0784] In addition, the three-dimensional data decoding apparatus of the present embodiment obtains first information (position code) indicating an N (N is an integer of 2 or more)-ary tree structure of a plurality of three-dimensional points included in the three-dimensional data in a first manner (position encoding) from the bitstream. The first information includes three-dimensional point information (position code) corresponding to each of the plurality of three-dimensional points. Each three-dimensional point information includes an index (idx) corresponding to each of a plurality of layers in the N-ary tree structure. Each index indicates a sub-block to which the corresponding three-dimensional point belongs among N sub-blocks belonging to the corresponding layer.
[0785] In other words, each three-dimensional point information indicates a path to the corresponding three-dimensional point in the N-ary tree structure. Each index indicates a sub-node included in the above-described path among N sub-nodes belonging to the corresponding layer (node).
[0786] The three-dimensional data decoding apparatus further restores the three-dimensional point corresponding to the three-dimensional point information using the three-dimensional point information.
[0787] Thus, the three-dimensional data decoding apparatus can selectively decode the three-dimensional points from the bitstream.
[0788] For example, the three-dimensional point information (position code) includes information (num_idx) indicating the number of indexes included in the three-dimensional point information. In other words, the information indicates the depth (number of layers) to the corresponding three-dimensional point in the N-ary tree structure.
[0789] For example, the first information includes information (numPoint) indicating the number of three-dimensional point information included in the first information. In other words, the information indicates the number of three-dimensional points included in the N-ary tree structure.
[0790] For example, N is 8, and the index is 3 bits.
[0791] For example, the three-dimensional data decoding apparatus further obtains second information (occupancy code) indicating the N-ary tree structure in a second manner (occupancy encoding) from the bitstream. The three-dimensional data decoding apparatus restores the plurality of three-dimensional points using the second information. The second information includes a plurality of 1-bit information corresponding to each of a plurality of sub-blocks belonging to a plurality of layers in the N-ary tree structure and indicates whether or not there is a three-dimensional point in the corresponding sub-block.
[0792] For example, the first information and the second information include information (coding mode information) indicating whether the information indicates the N-ary tree structure in the first manner or in the second manner.
[0793] For example, as shown in FIG. 1, a part of the N-ary tree structure is indicated in the first manner, and another part of the N-ary tree structure is indicated in the second manner. Figure 75
[0794] For example, the three-dimensional data decoding apparatus includes a processor and a memory, and the processor performs the above-described processing using the memory.
[0795] (Embodiment 10)
[0796] In this embodiment, another example of a coding method of a tree structure such as an octree structure will be described. Figure 88 is a diagram indicating an example of a tree structure according to this embodiment. In addition, Figure 88 indicates an example of a quadtree structure.
[0797] A leaf node including a three-dimensional point is referred to as an effective leaf node, and a leaf node not including a three-dimensional point is referred to as an ineffective leaf node. A branch in which the number of effective leaf nodes is equal to or greater than a threshold value is referred to as a dense branch. A branch in which the number of effective leaf nodes is less than the threshold value is referred to as a sparse branch.
[0798] The three-dimensional data encoding apparatus calculates the number of three-dimensional points (i.e., the number of effective leaf nodes) included in each branch in a certain layer of the tree structure. Figure 88 indicates an example in which the threshold value is 5. In this example, there are two branches in layer 1. Since the left branch includes seven three-dimensional points, the left branch is determined to be a dense branch. Since the right branch includes two three-dimensional points, the right branch is determined to be a sparse branch.
[0799] Figure 89 For example, is a diagram indicating an example of the number of effective leaf nodes (3D points) possessed by each branch of layer 5. Figure 89 The horizontal axis of indicates the identification number (index) of the branch of layer 5. As shown in Figure 89 In a certain branch, a significantly larger number of three-dimensional points are included than in other branches. In such a dense branch, the occupancy coding is more effective than in a sparse branch.
[0800] Next, a method of applying the occupancy coding and the position coding will be described. Figure 90 is a diagram indicating the relationship between the number of three-dimensional points (the number of effective leaf nodes) included in each branch of layer 5 and the coding method applied. As shown in Figure 90 As shown, the 3D data encoding device applies occupancy-based encoding to dense branches and position-based encoding to sparse branches. This improves encoding efficiency.
[0801] Figure 91 This is a diagram illustrating an example of a densely branched region in LiDAR data. For example... Figure 91 As shown, the density of 3D points calculated based on the number of 3D points contained in each branch varies depending on the region.
[0802] Furthermore, separating dense 3D points (branches) from sparse 3D points (branches) offers the following advantages: The closer to the LiDAR sensor, the higher the density of 3D points. Therefore, by separating the branches according to their density, it is possible to perform range-direction partitioning. Such partitioning is effective in certain applications. Moreover, for sparse branches, methods other than occupancy rate coding are effective.
[0803] In this embodiment, the three-dimensional data encoding device separates the input three-dimensional point group into two or more sub-three-dimensional point groups and applies different encoding methods to each sub-three-dimensional point group.
[0804] For example, a 3D data encoding device can separate an input 3D point group into a sub-3D point group A (dense cloud) that includes dense branches and a sub-3D point group B (sparse cloud) that includes sparse branches. Figure 92 It means from Figure 88 The diagram shows an example of a subgroup of three-dimensional points A (dense three-dimensional point group) separated from a tree structure, including dense branches. Figure 93 It means from Figure 88 The diagram shows an example of a subgroup of three-dimensional points B (a sparse three-dimensional point group) separated from the tree structure, including sparse branches.
[0805] Next, the 3D data encoding device encodes sub-3D point group A using occupancy rate encoding and sub-3D point group B using position encoding.
[0806] Furthermore, this example illustrates the application of different encoding methods (occupancy rate encoding and position encoding) as different encoding approaches. However, for example, a three-dimensional data encoding device can also use the same encoding method for sub-three-dimensional point group A and sub-three-dimensional point group B, and make the parameters used in encoding different between sub-three-dimensional point group A and sub-three-dimensional point group B.
[0807] The following describes the process of three-dimensional data encoding by a three-dimensional data encoding device. Figure 94 This is a flowchart of the three-dimensional data encoding process performed by the three-dimensional data encoding apparatus of this embodiment.
[0808] First, the three-dimensional data encoding apparatus separates the inputted three-dimensional point group into sub-three-dimensional point groups (S1701). The three-dimensional data encoding apparatus can either automatically perform this separation or perform it based on information inputted by the user. For example, the user can specify the range of the sub-three-dimensional point groups or the like. Further, as an example of automatic performance, in the case where the inputted data is LiDAR data, for example, the three-dimensional data encoding apparatus separates using distance information to each point group. Specifically, the three-dimensional data encoding apparatus separates point groups within a certain range from the measurement site from point groups outside the range. Further, the three-dimensional data encoding apparatus can also separate using information of important areas and unimportant areas.
[0809] Next, the three-dimensional data encoding apparatus encodes the sub-three-dimensional point group A by method A, thereby generating encoded data (an encoded bit stream) (S1702). Further, the three-dimensional data encoding apparatus encodes the sub-three-dimensional point group B by method B, thereby generating encoded data (S1703). In addition, the three-dimensional data encoding apparatus can also encode the sub-three-dimensional point group B by method A. In this case, the three-dimensional data encoding apparatus encodes the sub-three-dimensional point group B using parameters different from the encoding parameters used in the encoding of the sub-three-dimensional point group A. For example, the parameters can be quantization parameters. For example, the three-dimensional data encoding apparatus encodes the sub-three-dimensional point group B using a quantization parameter larger than the quantization parameter used in the encoding of the sub-three-dimensional point group A. In this case, the three-dimensional data encoding apparatus can also attach information indicating the quantization parameter used in the encoding of the sub-three-dimensional point group to the header of the encoded data of each sub-three-dimensional point group.
[0810] Next, the three-dimensional data encoding apparatus generates a bit stream by combining the encoded data obtained in step S1702 and the encoded data obtained in step S1703 (S1704).
[0811] Further, the three-dimensional data encoding apparatus can also encode information for decoding each sub-three-dimensional point group as header information of the bit stream. For example, the three-dimensional data encoding apparatus can also encode the following information.
[0812] The header information can also include information indicating the number of encoded sub-three-dimensional points. In this example, the information indicates 2.
[0813] The header information can also include information indicating the number of three-dimensional points included in each sub-three-dimensional point group and the encoding method. In this example, the information indicates the number of three-dimensional points included in the sub-three-dimensional point group A, the encoding method applied to the sub-three-dimensional point group A (method A), the number of three-dimensional points included in the sub-three-dimensional point group B, and the encoding method applied to the sub-three-dimensional point group B (method B).
[0814] The header information can also include information identifying the start position or the end position of the encoded data of each sub-three-dimensional point group.
[0815] Further, the three-dimensional data encoding apparatus can also encode the sub-three-dimensional point group A and the sub-three-dimensional point group B in parallel. Alternatively, the three-dimensional data encoding apparatus can also encode the sub-three-dimensional point group A and the sub-three-dimensional point group B sequentially.
[0816] Further, the method of separating the sub-three-dimensional point group is not limited to the above. For example, the three-dimensional data encoding apparatus changes the separation method, encodes using each of a plurality of separation methods, and calculates the encoding efficiency of the encoded data obtained using each separation method. Then, the three-dimensional data encoding apparatus selects the separation method with the highest encoding efficiency. For example, the three-dimensional data encoding apparatus can also separate the three-dimensional point group in each of a plurality of layers, calculate the encoding efficiency in each case, select the separation method (i.e., the layer in which separation is performed) with the highest encoding efficiency, and generate the sub-three-dimensional point group using the selected separation method to encode.
[0817] Further, the three-dimensional data encoding apparatus can also arrange the encoded information of the more important sub-three-dimensional point group in a position closer to the beginning of the bit stream when combining the encoded data. Thereby, the three-dimensional data decoding apparatus can obtain important information early by decoding only the bit stream at the beginning.
[0818] Next, the flow of the three-dimensional data decoding process performed by the three-dimensional data decoding apparatus will be described. Figure 95 is a flowchart of the three-dimensional data decoding process performed by the three-dimensional data decoding apparatus according to the present embodiment.
[0819] First, the three-dimensional data decoding apparatus obtains, for example, the bit stream generated by the three-dimensional data encoding apparatus described above. Next, the three-dimensional data decoding apparatus separates the encoded data of the sub-three-dimensional point group A and the encoded data of the sub-three-dimensional point group B from the obtained bit stream (S1711). Specifically, the three-dimensional data decoding apparatus decodes the information for decoding each sub-three-dimensional point group from the header information of the bit stream, and separates the encoded data of each sub-three-dimensional point group using the information.
[0820] Next, the three-dimensional data decoding apparatus obtains the sub-three-dimensional point group A by decoding the encoded data of the sub-three-dimensional point group A using the method A (S1712). Further, the three-dimensional data decoding apparatus obtains the sub-three-dimensional point group B by decoding the encoded data of the sub-three-dimensional point group B using the method B (S1713). Next, the three-dimensional data decoding apparatus combines the sub-three-dimensional point group A and the sub-three-dimensional point group B (S1714).
[0821] Further, the three-dimensional data decoding apparatus can decode the sub three-dimensional point group A and the sub three-dimensional point group B in parallel. Alternatively, the three-dimensional data decoding apparatus can decode the sub three-dimensional point group A and the sub three-dimensional point group B sequentially.
[0822] Further, the three-dimensional data decoding apparatus can decode a desired sub three-dimensional point group. For example, the three-dimensional data decoding apparatus can decode the sub three-dimensional point group A without decoding the sub three-dimensional point group B. For example, in a case where the sub three-dimensional point group A is a three-dimensional point group included in an important region of LiDAR data, the three-dimensional data decoding apparatus decodes the three-dimensional point group of the important region. The three-dimensional point group of the important region is used for self-position estimation of a vehicle or the like.
[0823] Next, a specific example of the encoding processing relating to the present embodiment will be described. Figure 96 is a flowchart of the three-dimensional data encoding processing performed by the three-dimensional data encoding apparatus relating to the present embodiment.
[0824] First, the three-dimensional data encoding apparatus separates the input three-dimensional points into a sparse three-dimensional point group and a dense three-dimensional point group (S1721). Specifically, the three-dimensional data encoding apparatus counts the number of valid leaf nodes possessed by a branch of a certain layer of an octree structure. The three-dimensional data encoding apparatus sets each branch to a dense branch or a sparse branch according to the number of valid leaf nodes of each branch. Further, the three-dimensional data encoding apparatus generates a sub three-dimensional point group (dense three-dimensional point group) that collects the dense branches and a sub three-dimensional point group (sparse three-dimensional point group) that collects the sparse branches.
[0825] Next, the three-dimensional data encoding apparatus generates encoded data by encoding the sparse three-dimensional point group (S1722). For example, the three-dimensional data encoding apparatus encodes the sparse three-dimensional point group using position encoding.
[0826] Further, the three-dimensional data encoding apparatus generates encoded data by encoding the dense three-dimensional point group (S1723). For example, the three-dimensional data encoding apparatus encodes the dense three-dimensional point group using occupancy encoding.
[0827] Next, the three-dimensional data encoding apparatus generates a bitstream by combining the encoded data of the sparse three-dimensional point group obtained in step S1722 and the encoded data of the dense three-dimensional point group obtained in step S1723 (S1724).
[0828] Further, the three-dimensional data encoding apparatus can encode information for decoding the sparse three-dimensional point group and the dense three-dimensional point group as header information of the bitstream. For example, the three-dimensional data encoding apparatus can encode the following information.
[0829] The header information can also include information indicating the number of sub three-dimensional point groups to be encoded. In this case, the information indicates 2.
[0830] The header information can also include information indicating the number of three-dimensional points included in each sub three-dimensional point group and the encoding method. In this case, the information indicates the number of three-dimensional points included in the sparse three-dimensional point group, the encoding method applied to the sparse three-dimensional point group (position encoding), the number of three-dimensional points included in the dense three-dimensional point group, and the encoding method applied to the dense three-dimensional point group (occupancy encoding).
[0831] The header information can also include information identifying the start position or the end position of the encoded data of each sub three-dimensional point group. In this case, the information indicates at least one of the start position and the end position of the encoded data of the sparse three-dimensional point group and the start position and the end position of the encoded data of the dense three-dimensional point group.
[0832] Further, the three-dimensional data encoding apparatus can encode the sparse three-dimensional point group and the dense three-dimensional point group in parallel. Alternatively, the three-dimensional data encoding apparatus can encode the sparse three-dimensional point group and the dense three-dimensional point group sequentially.
[0833] Next, a specific example of three-dimensional data decoding processing will be described. Figure 97 is a flowchart of three-dimensional data decoding processing performed by a three-dimensional data decoding apparatus relating to the present embodiment.
[0834] First, the three-dimensional data decoding apparatus obtains, for example, a bitstream generated by the three-dimensional data encoding apparatus described above. Next, the three-dimensional data decoding apparatus separates the encoded data of the sparse three-dimensional point group and the encoded data of the dense three-dimensional point group from the obtained bitstream (S1731). Specifically, the three-dimensional data decoding apparatus decodes the information used to decode each sub three-dimensional point group from the header information of the bitstream, and separates the encoded data of each sub three-dimensional point group using the information. In this case, the three-dimensional data decoding apparatus separates the encoded data of the sparse three-dimensional point group and the dense three-dimensional point group from the bitstream using the header information.
[0835] Next, the three-dimensional data decoding apparatus obtains the sparse three-dimensional point group by decoding the encoded data of the sparse three-dimensional point group (S1732). For example, the three-dimensional data decoding apparatus decodes the sparse three-dimensional point group using position decoding used to decode the encoded data that is position encoded.
[0836] Further, the three-dimensional data decoding apparatus obtains the dense three-dimensional point group by decoding the encoded data of the dense three-dimensional point group (S1733). For example, the three-dimensional data decoding apparatus decodes the dense three-dimensional point group using occupancy decoding used to decode the encoded data that is occupancy encoded.
[0837] Next, the three-dimensional data decoding apparatus combines the sparse three-dimensional point group obtained in step S1732 and the dense three-dimensional point group obtained in step S1733 (S1734).
[0838] In addition, the three-dimensional data decoding apparatus can decode the sparse three-dimensional point group and the dense three-dimensional point group in parallel. Alternatively, the three-dimensional data decoding apparatus can decode the sparse three-dimensional point group and the dense three-dimensional point group sequentially.
[0839] Further, the three-dimensional data decoding apparatus can decode a part...
Claims
1. A three-dimensional data encoding method, wherein, The information of object nodes contained in an N-ary tree structure of multiple 3D points in 3D data is encoded by referring to the information of referable nodes among multiple neighboring nodes that are spatially adjacent to each face of the object node, where N is an integer greater than 2. Generate parameters, When the parameter is the first value, the referable node includes a first node whose parent node is the same as the parent node of the object node, but does not include a second node, which is a node whose parent node is different from the parent node of the object node and whose face is spatially adjacent to the object node. When the parameter is the second value, the reference node includes the first node and the second node.
2. The three-dimensional data encoding method according to claim 1, wherein, The first node is located inside the parent node, and the second node is located outside the parent node.
3. The three-dimensional data encoding method according to claim 1 or 2, wherein, The three-dimensional data encoding method, It also generates a bitstream containing the parameters and the encoded information of the object node.
4. The three-dimensional data encoding method according to claim 1 or 2, wherein, The value of N is 8.
5. A three-dimensional data decoding method, wherein, Obtain parameters, For the information of object nodes contained in an N-ary tree structure of multiple 3D points in 3D data, decoding is performed by referring to the information of referential nodes among multiple neighboring nodes that are spatially adjacent to each face of the object node, where N is an integer greater than 2. When the parameter is the first value, the referable node includes a first node whose parent node is the same as the parent node of the object node, but does not include a second node, which is a node whose parent node is different from the parent node of the object node and whose face is spatially adjacent to the object node. When the parameter is the second value, the reference node includes the first node and the second node.
6. The three-dimensional data decoding method according to claim 5, wherein, The first node is located inside the parent node, and the second node is located outside the parent node.
7. The three-dimensional data decoding method according to claim 5 or 6, wherein, The three-dimensional data decoding method, It also obtains a bitstream containing the parameters and the encoded information of the object node.
8. The three-dimensional data decoding method according to claim 5 or 6, wherein, The value of N is 8.
9. A three-dimensional data encoding device, wherein, have: Processor; and memory, The processor uses the memory. The information of object nodes contained in an N-ary tree structure of multiple 3D points in 3D data is encoded by referring to the information of referable nodes among multiple neighboring nodes that are spatially adjacent to each face of the object node, where N is an integer greater than 2. Generate parameters, When the parameter is the first value, the referable node includes a first node whose parent node is the same as the parent node of the object node, but does not include a second node, which is a node whose parent node is different from the parent node of the object node and whose face is spatially adjacent to the object node. When the parameter is the second value, the reference node includes the first node and the second node.
10. A three-dimensional data decoding device, wherein, have: Processor; and memory, The processor uses the memory. Obtain parameters, For the information of object nodes contained in an N-ary tree structure of multiple 3D points in 3D data, decoding is performed by referring to the information of referential nodes among multiple neighboring nodes that are spatially adjacent to each face of the object node, where N is an integer greater than 2. When the parameter is the first value, the referable node includes a first node whose parent node is the same as the parent node of the object node, but does not include a second node, which is a node whose parent node is different from the parent node of the object node and whose face is spatially adjacent to the object node. When the parameter is the second value, the reference node includes the first node and the second node.
Citation Information
Patent Citations
Map display device
WO2014020663A1
Node structure for representing 3-dimensional objects using depth image
EP1321893B1
Image-encoding device, method for image encoding, image-encoding program, image-decoding device, method for image decoding, and image-decoding program
WO2013024588A1