Three-dimensional data encoding method, three-dimensional data decoding method, three-dimensional data encoding device, and three-dimensional data decoding device

The proposed three-dimensional data encoding method optimizes coding efficiency and reduces processing requirements by utilizing an N-ary tree structure that references only sibling nodes, addressing inefficiencies in existing methods.

JP7802871B2Active Publication Date: 2026-01-20PANASONIC INTELLECTUAL PROPERTY CORP OF AMERICA
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
JP2024111217
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2018-02-08
Filing Date
2024-07-10
Publication Date
2026-01-20
Estimated Expiration
2039-02-07

AI Technical Summary

Technical Problem

Existing three-dimensional data encoding methods require significant processing resources and inefficient coding efficiency, particularly when handling large volumes of point cloud data.

Method used

A three-dimensional data encoding method that utilizes an N-ary tree structure, allowing reference only to sibling nodes of a target node for encoding and decoding, thereby optimizing the use of adjacent node information and reducing processing requirements.

Benefits of technology

Improves encoding efficiency and reduces processing load by selectively referencing sibling nodes within the N-ary tree structure, enhancing the overall coding performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007802871000001
    Figure 0007802871000001
  • Figure 0007802871000002
    Figure 0007802871000002
  • Figure 0007802871000003
    Figure 0007802871000003
Patent Text Reader

Abstract

To improve encoding efficiency.SOLUTION: The three-dimensional data encoding method includes the steps of: encoding information of the target nodes included in the N (N is an integer equal to or greater than 2) tree structure of multiple 3D points included in the 3D data out of multiple adjacent nodes that are spatially adjacent to the target node by using referable node information only; and encoding the parameters. When the parameters include a first value, the referable nodes are only sibling nodes of the target node. For example, the value of parameters may suggest a referable node.SELECTED DRAWING: Figure 136
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present disclosure relates to a three-dimensional data encoding method, a three-dimensional data decoding method, a three-dimensional data encoding device, and a three-dimensional data decoding device. [Background technology]

[0002] In the future, devices and services that utilize 3D data are expected to become widespread in a wide range of fields, including computer vision for autonomous operation of automobiles or robots, map information, surveillance, infrastructure inspection, video distribution, etc. 3D data can be acquired in a variety of ways, including distance sensors such as range finders, stereo cameras, or a combination of multiple monocular cameras.

[0003] One method of representing three-dimensional data is a point cloud, which represents the shape of a three-dimensional structure using a group of points in three-dimensional space. A point cloud stores the position and color of the points. Point clouds are expected to become the mainstream method of representing three-dimensional data, but point clouds require a very large amount of data. Therefore, when storing or transmitting three-dimensional data, data compression through encoding is essential, just as with two-dimensional video images (examples include MPEG-4 AVC or HEVC standardized by MPEG).

[0004] In addition, compression of point clouds is partially supported by public libraries that perform point cloud-related processing (Point Cloud Library).

[0005] Furthermore, a technique is known in which three-dimensional map data is used to search for and display facilities located around a vehicle (see, for example, Patent Document 1). [Prior art documents] [Patent documents]

[0006] [Patent Document 1] International Publication No. 2014 / 020663 Summary of the Invention [Problem to be solved by the invention]

[0007] It is desirable to be able to improve the coding efficiency and reduce the amount of processing required when coding three-dimensional data.

[0008] An object of the present disclosure is to provide a three-dimensional data encoding method, a three-dimensional data decoding method, a three-dimensional data encoding device, or a three-dimensional data decoding device that can improve encoding efficiency and reduce the amount of processing. [Means for solving the problem]

[0009] A three-dimensional data encoding method according to one aspect of the present disclosure includes encoding information of a target node included in an N-ary tree structure of a plurality of three-dimensional points included in three-dimensional data (N is an integer of 2 or more), using only information of nodes that can be referenced from among a plurality of adjacent nodes that are spatially adjacent to the target node, and encoding parameters; generating a bitstream including the encoded information of the target node and the parameters; If the parameter contains a first value, the referenceable nodes are only sibling nodes of the target node.

[0010] A three-dimensional data decoding method according to one aspect of the present disclosure includes: From the bitstream A parameter is acquired, and information of a target node included in an N-ary tree structure (N is an integer equal to or greater than 2) of multiple three-dimensional points included in the three-dimensional data is decoded using only information of referenceable nodes among multiple adjacent nodes spatially adjacent to the target node, and if the parameter includes a first value, the referenceable nodes are only sibling nodes of the target node. [Effects of the Invention]

[0011] The present disclosure can provide a three-dimensional data encoding method, a three-dimensional data decoding method, a three-dimensional data encoding device, or a three-dimensional data decoding device that can improve encoding efficiency and reduce the amount of processing. [Brief explanation of the drawings]

[0012] [Figure 1] FIG. 1 is a diagram showing the structure of encoded three-dimensional data according to the first embodiment. [Figure 2] FIG. 2 is a diagram showing an example of a prediction structure between SPCs belonging to the lowest layer of a GOS according to the first embodiment. [Figure 3] FIG. 3 is a diagram showing an example of an inter-layer prediction structure according to the first embodiment. [Figure 4] FIG. 4 is a diagram showing an example of the coding order of the GOS according to the first embodiment. [Figure 5] FIG. 5 is a diagram showing an example of the coding order of GOS according to the first embodiment. [Figure 6] FIG. 6 is a block diagram of a three-dimensional data encoding device according to the first embodiment. [Figure 7] FIG. 7 is a flowchart of the encoding process according to the first embodiment. [Figure 8] FIG. 8 is a block diagram of a three-dimensional data decoding device according to the first embodiment. [Figure 9] FIG. 9 is a flowchart of the decoding process according to the first embodiment. [Figure 10] FIG. 10 is a diagram illustrating an example of meta information according to the first embodiment. [Figure 11] FIG. 11 is a diagram illustrating an example of the configuration of the SWLD according to the second embodiment. [Figure 12] FIG. 12 illustrates an example of the operation of the server and the client according to the second embodiment. [Figure 13] FIG. 13 is a diagram illustrating an example of the operation of the server and the client according to the second embodiment. [Figure 14] FIG. 14 is a diagram illustrating an example of the operation of the server and the client according to the second embodiment. [Figure 15] FIG. 15 illustrates an example of the operation of the server and the client according to the second embodiment. [Figure 16] FIG. 16 is a block diagram of a three-dimensional data encoding device according to the second embodiment. [Figure 17]FIG. 17 is a flowchart of the encoding process according to the second embodiment. [Figure 18] FIG. 18 is a block diagram of a three-dimensional data decoding device according to the second embodiment. [Figure 19] FIG. 19 is a flowchart of the decoding process according to the second embodiment. [Figure 20] FIG. 20 is a diagram illustrating an example of the configuration of a WLD according to the second embodiment. [Figure 21] FIG. 21 is a diagram illustrating an example of an octree structure of a WLD according to the second embodiment. [Figure 22] FIG. 22 is a diagram illustrating an example of the configuration of the SWLD according to the second embodiment. [Figure 23] FIG. 23 is a diagram illustrating an example of an octree structure of an SWLD according to the second embodiment. [Figure 24] FIG. 24 is a block diagram of a three-dimensional data creation device according to the third embodiment. [Figure 25] FIG. 25 is a block diagram of a three-dimensional data transmission device according to the third embodiment. [Figure 26] FIG. 26 is a block diagram of a three-dimensional information processing device according to the fourth embodiment. [Figure 27] FIG. 27 is a block diagram of a three-dimensional data creation device according to the fifth embodiment. [Figure 28] FIG. 28 is a diagram illustrating a configuration of a system according to the sixth embodiment. [Figure 29] FIG. 29 is a block diagram of a client device according to the sixth embodiment. [Figure 30] FIG. 30 is a block diagram of a server according to the sixth embodiment. [Figure 31] FIG. 31 is a flowchart of a three-dimensional data creation process by a client device according to the sixth embodiment. [Figure 32] FIG. 32 is a flowchart of a sensor information transmission process performed by a client device according to the sixth embodiment. [Figure 33] FIG. 33 is a flowchart of three-dimensional data creation processing by the server according to the sixth embodiment. [Figure 34] FIG. 34 is a flowchart of a three-dimensional map transmission process performed by the server according to the sixth embodiment. [Figure 35] FIG. 35 is a diagram showing a configuration of a modified example of the system according to the sixth embodiment. [Figure 36] FIG. 36 is a diagram illustrating the configurations of a server and a client device according to the sixth embodiment. [Figure 37] FIG. 37 is a block diagram of a three-dimensional data encoding device according to the seventh embodiment. [Figure 38] FIG. 38 is a diagram showing an example of a prediction residual according to the seventh embodiment. [Figure 39] FIG. 39 is a diagram illustrating an example of a volume according to the seventh embodiment. [Figure 40] FIG. 40 is a diagram showing an example of an octree representation of a volume according to the seventh embodiment. [Figure 41] FIG. 41 is a diagram showing an example of a bit string of a volume according to the seventh embodiment. [Figure 42] FIG. 42 is a diagram showing an example of an octree representation of a volume according to the seventh embodiment. [Figure 43] FIG. 43 is a diagram illustrating an example of a volume according to the seventh embodiment. [Figure 44] FIG. 44 is a diagram illustrating the intra prediction process according to the seventh embodiment. [Figure 45] FIG. 45 is a diagram for explaining the rotation and translation processing according to the seventh embodiment. [Figure 46] FIG. 46 is a diagram showing an example of the syntax of the RT application flag and the RT information according to the seventh embodiment. [Figure 47] FIG. 47 is a diagram illustrating the inter prediction process according to the seventh embodiment. [Figure 48] FIG. 48 is a block diagram of a three-dimensional data decoding device according to the seventh embodiment. [Figure 49] FIG. 49 is a flowchart of three-dimensional data encoding processing by the three-dimensional data encoding device according to the seventh embodiment. [Figure 50] FIG. 50 is a flowchart of three-dimensional data decoding processing by the three-dimensional data decoding device according to the seventh embodiment. [Figure 51] FIG. 51 is a diagram illustrating a configuration of a distribution system according to the eighth embodiment. [Figure 52] FIG. 52 is a diagram illustrating an example of a configuration of a bit stream of an encoded 3D map according to the eighth embodiment. [Figure 53] FIG. 53 is a diagram for explaining the effect of improving the coding efficiency according to the eighth embodiment. [Figure 54] FIG. 54 is a flowchart of processing by the server according to the eighth embodiment. [Figure 55] FIG. 55 is a flowchart of processing by a client according to the eighth embodiment. [Figure 56] FIG. 56 is a diagram illustrating an example of the syntax of a submap according to the eighth embodiment. [Figure 57] FIG. 57 is a diagram schematically showing the coding type switching process according to the eighth embodiment. [Figure 58] FIG. 58 is a diagram illustrating an example of the syntax of a submap according to the eighth embodiment. [Figure 59] FIG. 59 is a flowchart of three-dimensional data encoding processing according to the eighth embodiment. [Figure 60] FIG. 60 is a flowchart of three-dimensional data decoding processing according to the eighth embodiment. [Figure 61] FIG. 61 is a diagram schematically illustrating the operation of a modified example of the coding type switching process according to the eighth embodiment. [Figure 62] FIG. 62 is a diagram schematically illustrating the operation of a modified example of the coding type switching process according to the eighth embodiment. [Figure 63] FIG. 63 is a diagram schematically illustrating the operation of a modified example of the coding type switching process according to the eighth embodiment. [Figure 64] FIG. 64 is a diagram illustrating the operation of a modification of the differential value calculation process according to the eighth embodiment. [Figure 65]FIG. 65 is a diagram illustrating the operation of a variation of the differential value calculation process according to the eighth embodiment. [Figure 66] FIG. 66 is a diagram illustrating the operation of a variation of the differential value calculation process according to the eighth embodiment. [Figure 67] FIG. 67 is a diagram illustrating the operation of a variation of the differential value calculation process according to the eighth embodiment. [Figure 68] FIG. 68 is a diagram illustrating an example of the syntax of a volume according to the eighth embodiment. [Figure 69] FIG. 69 is a diagram showing an example of an important region according to the ninth embodiment. [Figure 70] FIG. 70 is a diagram showing an example of an occupancy code according to the ninth embodiment. [Figure 71] FIG. 71 is a diagram showing an example of a quadtree structure according to the ninth embodiment. [Figure 72] FIG. 72 is a diagram showing an example of an occupancy code and a location code according to the ninth embodiment. [Figure 73] FIG. 73 is a diagram illustrating an example of three-dimensional points obtained by the LiDAR according to the ninth embodiment. [Figure 74] FIG. 74 is a diagram showing an example of an octree structure according to the ninth embodiment. [Figure 75] FIG. 75 is a diagram illustrating an example of hybrid coding according to the ninth embodiment. [Figure 76] FIG. 76 is a diagram illustrating a method of switching between location coding and occupancy coding according to the ninth embodiment. [Figure 77] FIG. 77 is a diagram showing an example of a location-encoded bitstream according to the ninth embodiment. [Figure 78] FIG. 78 is a diagram showing an example of a bit stream for hybrid coding according to the ninth embodiment. [Figure 79] FIG. 79 is a diagram showing a tree structure of an occupancy code of an important 3D point according to the ninth embodiment. [Figure 80]FIG. 80 is a diagram showing a tree structure of an occupancy code of a non-important 3D point according to the ninth embodiment. [Figure 81] FIG. 81 is a diagram showing an example of a bit stream for hybrid coding according to the ninth embodiment. [Figure 82] FIG. 82 is a diagram showing an example of a bitstream including coding mode information according to the ninth embodiment. [Figure 83] FIG. 83 is a diagram illustrating an example of syntax according to the ninth embodiment. [Figure 84] FIG. 84 is a flowchart of the encoding process according to the ninth embodiment. [Figure 85] FIG. 85 is a flowchart of the node encoding process according to the ninth embodiment. [Figure 86] FIG. 86 is a flowchart of the decoding process according to the ninth embodiment. [Figure 87] FIG. 87 is a flowchart of the node decoding process according to the ninth embodiment. [Figure 88] FIG. 88 is a diagram illustrating an example of a tree structure according to the tenth embodiment. [Figure 89] FIG. 89 is a diagram illustrating an example of the number of effective leaves that each branch has according to the tenth embodiment. [Figure 90] FIG. 90 is a diagram showing an application example of the coding method according to the tenth embodiment. [Figure 91] FIG. 91 is a diagram illustrating an example of a dense branch region according to the tenth embodiment. [Figure 92] FIG. 92 is a diagram illustrating an example of a dense 3D point cloud according to the tenth embodiment. [Figure 93] FIG. 93 is a diagram illustrating an example of a sparse three-dimensional point cloud according to the tenth embodiment. [Figure 94] FIG. 94 is a flowchart of the encoding process according to the tenth embodiment. [Figure 95] FIG. 95 is a flowchart of the decoding process according to the tenth embodiment. [Figure 96] FIG. 96 is a flowchart of the encoding process according to the tenth embodiment. [Figure 97] FIG. 97 is a flowchart of the decoding process according to the tenth embodiment. [Figure 98] FIG. 98 is a flowchart of the encoding process according to the tenth embodiment. [Figure 99] FIG. 99 is a flowchart of the decoding process according to the tenth embodiment. [Figure 100] FIG. 100 is a flowchart of the separation process of three-dimensional points according to the tenth embodiment. [Figure 101] FIG. 101 is a diagram illustrating an example of syntax according to the tenth embodiment. [Figure 102] FIG. 102 is a diagram illustrating an example of dense branches according to the tenth embodiment. [Figure 103] FIG. 103 is a diagram illustrating an example of sparse branches according to the tenth embodiment. [Figure 104] FIG. 104 is a flowchart of the encoding process according to the modification of the tenth embodiment. [Figure 105] FIG. 105 is a flowchart of a decoding process according to a variation of the tenth embodiment. [Figure 106] FIG. 106 is a flowchart of a process for separating three-dimensional points according to a modification of the tenth embodiment. [Figure 107] FIG. 107 is a diagram illustrating an example of syntax according to a modification of the tenth embodiment. [Figure 108] FIG. 108 is a flowchart of the encoding process according to the tenth embodiment. [Figure 109] FIG. 109 is a flowchart of the decoding process according to the tenth embodiment. [Figure 110] FIG. 110 is a diagram illustrating an example of a tree structure according to the eleventh embodiment. [Figure 111] FIG. 111 is a diagram showing an example of an occupancy code according to the eleventh embodiment. [Figure 112] FIG. 112 is a diagram schematically illustrating the operation of the three-dimensional data encoding device according to the eleventh embodiment. [Figure 113] FIG. 113 is a diagram illustrating an example of geometric information according to the eleventh embodiment. [Figure 114] FIG. 114 is a diagram showing an example of selection of a coding table using geometric information according to the eleventh embodiment. [Figure 115] FIG. 115 is a diagram showing an example of selection of a coding table using structure information according to the eleventh embodiment. [Figure 116] FIG. 116 is a diagram showing an example of selection of a coding table using attribute information according to the eleventh embodiment. [Figure 117] FIG. 117 is a diagram showing an example of selection of a coding table using attribute information according to the eleventh embodiment. [Figure 118] FIG. 118 is a diagram showing an example of the structure of a bitstream according to the eleventh embodiment. [Figure 119] FIG. 119 is a diagram illustrating an example of a coding table according to the eleventh embodiment. [Figure 120] FIG. 120 is a diagram illustrating an example of a coding table according to the eleventh embodiment. [Figure 121] FIG. 121 is a diagram showing an example of the structure of a bitstream according to the eleventh embodiment. [Figure 122] FIG. 122 is a diagram illustrating an example of a coding table according to the eleventh embodiment. [Figure 123] FIG. 123 is a diagram illustrating an example of a coding table according to the eleventh embodiment. [Figure 124] FIG. 124 is a diagram illustrating an example of bit numbers of an occupancy code according to the eleventh embodiment. [Figure 125] FIG. 125 is a flowchart of the encoding process using geometric information according to the eleventh embodiment. [Figure 126] FIG. 126 is a flowchart of a decoding process using geometric information according to the eleventh embodiment. [Figure 127] FIG. 127 is a flowchart of the encoding process using the structure information according to the eleventh embodiment. [Figure 128] FIG. 128 is a flowchart of a decoding process using structure information according to the eleventh embodiment. [Figure 129]FIG. 129 is a flowchart of the encoding process using attribute information according to the eleventh embodiment. [Figure 130] FIG. 130 is a flowchart of a decoding process using attribute information according to the eleventh embodiment. [Figure 131] FIG. 131 is a flowchart of a coding table selection process using geometric information according to the eleventh embodiment. [Figure 132] FIG. 132 is a flowchart of the coding table selection process using the structure information according to the eleventh embodiment. [Figure 133] FIG. 133 is a flowchart of the coding table selection process using attribute information according to the eleventh embodiment. [Figure 134] FIG. 134 is a block diagram of a three-dimensional data encoding device according to the eleventh embodiment. [Figure 135] FIG. 135 is a block diagram of a three-dimensional data decoding device according to the eleventh embodiment. [Figure 136] FIG. 136 is a diagram showing the reference relationship in an octree structure according to the twelfth embodiment. [Figure 137] FIG. 137 is a diagram showing reference relationships in the spatial domain according to the twelfth embodiment. [Figure 138] FIG. 138 is a diagram illustrating an example of an adjacent reference node according to the twelfth embodiment. [Figure 139] FIG. 139 is a diagram illustrating the relationship between a parent node and a node according to the twelfth embodiment. [Figure 140] FIG. 140 is a diagram showing an example of an occupancy code of a parent node according to the twelfth embodiment. [Figure 141] FIG. 141 is a block diagram of a three-dimensional data encoding device according to the twelfth embodiment. [Figure 142] FIG. 142 is a block diagram of a three-dimensional data decoding device according to the twelfth embodiment. [Figure 143] FIG. 143 is a flowchart of three-dimensional data encoding processing according to the twelfth embodiment. [Figure 144]FIG. 144 is a flowchart of three-dimensional data decoding processing according to the twelfth embodiment. [Figure 145] FIG. 145 is a diagram illustrating an example of switching of the coding table according to the twelfth embodiment. [Figure 146] FIG. 146 is a diagram showing reference relationships in the spatial domain according to the first modification of the twelfth embodiment. [Figure 147] FIG. 147 is a diagram illustrating an example of the syntax of header information according to the first modification of the twelfth embodiment. [Figure 148] FIG. 148 is a diagram illustrating an example of the syntax of header information according to the first modification of the twelfth embodiment. [Figure 149] FIG. 149 is a diagram illustrating an example of an adjacent reference node according to the second modification of the twelfth embodiment. [Figure 150] FIG. 150 illustrates an example of a target node and adjacent nodes according to the second modification of the twelfth embodiment. [Figure 151] FIG. 151 shows the reference relationship in an octree structure according to the third modification of the twelfth embodiment. [Figure 152] FIG. 152 is a diagram showing reference relationships in the spatial domain according to the third modification of the twelfth embodiment. DETAILED DESCRIPTION OF THE INVENTION

[0013] A three-dimensional data encoding method according to one embodiment of the present disclosure encodes information about a target node included in an N-ary tree structure (N is an integer equal to or greater than 2) of multiple three-dimensional points included in the three-dimensional data, and in the encoding, allows reference to information about a first node among multiple adjacent nodes spatially adjacent to the target node, whose parent node is the same as that of the target node, and prohibits reference to information about a second node whose parent node is different from that of the target node.

[0014] According to this, the three-dimensional data encoding method can improve encoding efficiency by referencing information on a first node, among multiple adjacent nodes spatially adjacent to the target node, whose parent node is the same as the target node. Furthermore, the three-dimensional data encoding method can reduce the amount of processing by not referencing information on a second node, among multiple adjacent nodes, whose parent node is different from the target node. In this way, the three-dimensional data encoding method can improve encoding efficiency and reduce the amount of processing.

[0015] For example, the three-dimensional data encoding method may further determine whether to prohibit reference to the information of the second node, and in the encoding, switch between prohibiting and allowing reference to the information of the second node based on the result of the decision, and the three-dimensional data encoding method may further generate a bitstream including prohibition switching information that is the result of the decision and indicates whether to prohibit reference to the information of the second node.

[0016] According to this, the three-dimensional data encoding method can switch whether or not to prohibit reference to information of the second node, and the three-dimensional data decoding device can perform decoding processing appropriately using the prohibition switching information.

[0017] For example, the information of the target node may be information indicating whether or not a three-dimensional point exists in each of the child nodes belonging to the target node, the information of the first node may be information indicating whether or not a three-dimensional point exists in the first node, and the information of the second node may be information indicating whether or not a three-dimensional point exists in the second node.

[0018] For example, the encoding may involve selecting an encoding table based on whether or not a three-dimensional point exists at the first node, and entropy encoding the information of the target node using the selected encoding table.

[0019] For example, the encoding may allow reference to information of a child node of the first node among the plurality of adjacent nodes.

[0020] According to this, the three-dimensional data encoding method can refer to more detailed information of adjacent nodes, thereby improving encoding efficiency.

[0021] For example, in the encoding, the adjacent node to be referenced may be switched among the plurality of adjacent nodes depending on the spatial position of the target node within the parent node.

[0022] According to this, the three-dimensional data encoding method can refer to an appropriate adjacent node depending on the spatial position of the target node within the parent node.

[0023] A three-dimensional data decoding method according to one embodiment of the present disclosure decodes information about a target node included in an N-ary tree structure (N is an integer equal to or greater than 2) of multiple three-dimensional points included in the three-dimensional data, and in the decoding, allows reference to information about a first node among multiple adjacent nodes spatially adjacent to the target node, whose parent node is the same as the target node, and prohibits reference to information about a second node whose parent node is different from the target node.

[0024] According to this, the three-dimensional data decoding method can improve coding efficiency by referencing information on a first node, among multiple adjacent nodes spatially adjacent to the target node, whose parent node is the same as the target node. Also, the three-dimensional data decoding method can reduce the amount of processing by not referencing information on a second node, among multiple adjacent nodes, whose parent node is different from the target node. In this way, the three-dimensional data decoding method can improve coding efficiency and reduce the amount of processing.

[0025] For example, the three-dimensional data decoding method may further obtain prohibition switching information from the bitstream indicating whether to prohibit reference to the information of the second node, and in the decoding, switch between prohibiting and allowing reference to the information of the second node based on the prohibition switching information.

[0026] According to this, the three-dimensional data decoding method can appropriately perform the decoding process using the inhibition switching information.

[0027] For example, the information of the target node may be information indicating whether or not a three-dimensional point exists in each of the child nodes belonging to the target node, the information of the first node may be information indicating whether or not a three-dimensional point exists in the first node, and the information of the second node may be information indicating whether or not a three-dimensional point exists in the second node.

[0028] For example, the decoding may involve selecting an encoding table based on whether or not a three-dimensional point exists at the first node, and entropy decoding the information of the target node using the selected encoding table.

[0029] For example, in the decryption, reference to information of a child node of the first node among the plurality of adjacent nodes may be permitted.

[0030] According to this, the three-dimensional data decoding method can refer to more detailed information about adjacent nodes, thereby improving coding efficiency.

[0031] For example, in the decoding, the adjacent node to be referenced may be switched among the plurality of adjacent nodes depending on the spatial position of the target node within the parent node.

[0032] According to this, the three-dimensional data decoding method can refer to an appropriate adjacent node depending on the spatial position of the target node within the parent node.

[0033] In addition, a three-dimensional data encoding device according to one embodiment of the present disclosure includes a processor and a memory, and the processor uses the memory to encode information about a target node included in an N-ary tree structure (N is an integer greater than or equal to 2) of multiple three-dimensional points included in the three-dimensional data, and in the encoding, allows reference to information about a first node among multiple adjacent nodes spatially adjacent to the target node, whose parent node is the same as that of the target node, and prohibits reference to information about a second node whose parent node is different from that of the target node.

[0034] According to this, the three-dimensional data encoding device can improve encoding efficiency by referencing information on a first node, among multiple adjacent nodes spatially adjacent to the target node, whose parent node is the same as the target node. Furthermore, the three-dimensional data encoding device can reduce the amount of processing by not referencing information on a second node, among multiple adjacent nodes, whose parent node is different from the target node. In this way, the three-dimensional data encoding device can improve encoding efficiency and reduce the amount of processing.

[0035] In addition, a three-dimensional data decoding device according to one embodiment of the present disclosure includes a processor and a memory, and the processor uses the memory to decode information about a target node included in an N-ary tree structure (N is an integer greater than or equal to 2) of multiple three-dimensional points included in the three-dimensional data, and in the decoding, allows reference to information about a first node among multiple adjacent nodes spatially adjacent to the target node, whose parent node is the same as the target node, and prohibits reference to information about a second node whose parent node is different from the target node.

[0036] According to this, the three-dimensional data decoding device can improve coding efficiency by referencing information on a first node, among multiple adjacent nodes spatially adjacent to the target node, whose parent node is the same as the target node. Furthermore, the three-dimensional data decoding device can reduce the amount of processing by not referencing information on a second node, among multiple adjacent nodes, whose parent node is different from the target node. In this way, the three-dimensional data decoding device can improve coding efficiency and reduce the amount of processing.

[0037] These comprehensive or specific aspects may be realized as a system, a method, an integrated circuit, a computer program, or a computer-readable recording medium such as a CD-ROM, or may be realized as any combination of a system, a method, an integrated circuit, a computer program, and a recording medium.

[0038] Hereinafter, the embodiments will be described in detail with reference to the drawings. Note that each of the embodiments described below represents a specific example of the present disclosure. The numerical values, shapes, materials, components, component placement and connection configurations, steps, and step order shown in the following embodiments are merely examples and are not intended to limit the present disclosure. Furthermore, among the components in the following embodiments, components that are not described in an independent claim that represents a superordinate concept will be described as optional components.

[0039] (Embodiment 1) First, the data structure of encoded three-dimensional data (hereinafter also referred to as encoded data) according to this embodiment will be described. Fig. 1 is a diagram showing the structure of encoded three-dimensional data according to this embodiment.

[0040] In this embodiment, a three-dimensional space is divided into spaces (SPCs) corresponding to pictures in video encoding, and three-dimensional data is encoded using the spaces as units. The spaces are further divided into volumes (VLMs) corresponding to macroblocks or the like in video encoding, and prediction and conversion are performed using the VLMs as units. A volume includes a plurality of voxels (VXLs), which are the smallest units to which position coordinates can be associated. Note that prediction, like prediction performed for two-dimensional images, refers to generating predicted three-dimensional data similar to the processing unit to be processed by referring to other processing units, and encoding the difference between the predicted three-dimensional data and the processing unit to be processed. Furthermore, this prediction includes not only spatial prediction that refers to other prediction units at the same time, but also temporal prediction that refers to a prediction unit at a different time.

[0041] For example, when a three-dimensional data encoding device (hereinafter also referred to as an encoding device) encodes a three-dimensional space represented by point cloud data such as a point cloud, it encodes each point of the point cloud or multiple points contained in a voxel collectively according to the size of the voxel. By subdividing the voxels, the three-dimensional shape of the point cloud can be expressed with high precision, and by increasing the voxel size, the three-dimensional shape of the point cloud can be expressed roughly.

[0042] In the following, an example will be described in which the three-dimensional data is a point cloud, but the three-dimensional data is not limited to a point cloud and may be three-dimensional data in any format.

[0043] Alternatively, voxels with a hierarchical structure may be used. In this case, the nth layer may indicate in order whether a sample point exists in the n-1th layer or lower (a layer below the nth layer). For example, when decoding only the nth layer, if a sample point exists in the n-1th layer or lower, the sample point can be decoded by assuming that the sample point exists at the center of the voxel in the nth layer.

[0044] The encoding device also acquires point cloud data using a distance sensor, a stereo camera, a monocular camera, a gyro, an inertial sensor, or the like.

[0045] Similar to video coding, spaces are classified into at least three prediction structures, including independently decodable intra-space (I-SPC), predictive space (P-SPC), and bidirectional space (B-SPC). Spaces also have two types of time information: decoding time and display time.

[0046] As shown in Figure 1, there is a random access unit called a Group Of Space (GOS), which is a processing unit that includes multiple spaces. There is also a World (WLD), which is a processing unit that includes multiple GOS.

[0047] The spatial region occupied by the world is associated with an absolute position on Earth using GPS or latitude and longitude information. This position information is stored as meta information. Note that the meta information may be included in the encoded data or may be transmitted separately from the encoded data.

[0048] Furthermore, within a GOS, all SPCs may be three-dimensionally adjacent, or there may be SPCs that are not three-dimensionally adjacent to other SPCs.

[0049] In the following, the process of encoding, decoding, referencing, etc. of three-dimensional data included in a processing unit such as a GOS, SPC, or VLM will also be simply referred to as encoding, decoding, or referencing the processing unit, etc. The three-dimensional data included in the processing unit includes, for example, at least one pair of a spatial position such as three-dimensional coordinates and a characteristic value such as color information.

[0050] Next, the prediction structure of SPCs in a GOS will be explained. Multiple SPCs in the same GOS or multiple VLMs in the same SPC occupy different spaces, but have the same time information (decoding time and display time).

[0051] Furthermore, the first SPC in a GOS in decoding order is the I-SPC. There are two types of GOS: closed GOS and open GOS. A closed GOS is a GOS that can decode all SPCs in the GOS when decoding starts from the first I-SPC. In an open GOS, some SPCs that appear earlier in the GOS than the first I-SPC refer to a different GOS, and cannot be decoded using only that GOS.

[0052] In addition, in coded data such as map information, WLDs are sometimes decoded in the reverse order of coding, and if there is dependency between GOSs, reverse playback is difficult. Therefore, in such cases, closed GOSs are generally used.

[0053] Furthermore, the GOS has a layer structure in the height direction, and encoding or decoding is performed in order from the SPC in the lower layer.

[0054] Fig. 2 is a diagram showing an example of a prediction structure between SPCs belonging to the lowest layer of a GOS, and Fig. 3 is a diagram showing an example of a prediction structure between layers.

[0055] A GOS contains one or more I-SPCs. Objects such as people, animals, cars, bicycles, traffic lights, and landmark buildings exist in three-dimensional space, and it is particularly effective to encode small objects as I-SPCs. For example, a three-dimensional data decoding device (hereinafter also referred to as a decoding device) decodes only the I-SPCs in the GOS when decoding the GOS with low processing load or at high speed.

[0056] The encoding device may also switch the encoding interval or frequency of occurrence of I-SPC depending on the density of objects in the WLD.

[0057] 3, the encoding device or decoding device encodes or decodes multiple layers in order from the lowest layer (layer 1). This allows, for example, an autonomous vehicle to prioritize data near the ground, which contains more information.

[0058] In addition, encoded data used by drones, etc. may be encoded or decoded in order from the SPC of the highest layer in the height direction within the GOS.

[0059] Alternatively, the encoding or decoding device may encode or decode multiple layers so that the decoding device can roughly grasp the GOS and gradually increase the resolution. For example, the encoding or decoding device may encode or decode layers 3, 8, 1, 9, etc. in that order.

[0060] Next, we will explain how to handle static and dynamic objects.

[0061] In a three-dimensional space, there exist static objects or scenes such as buildings or roads (hereinafter collectively referred to as static objects), and dynamic objects such as cars or people (hereinafter referred to as dynamic objects). Object detection is performed separately by extracting feature points from point cloud data or camera images such as a stereo camera. Here, an example of a method for encoding dynamic objects will be described.

[0062] The first method is to encode static objects without distinguishing between static and dynamic objects, and the second method is to distinguish between static and dynamic objects using identification information.

[0063] For example, GOS is used as the identification unit. In this case, GOS including SPCs that constitute static objects and GOS including SPCs that constitute dynamic objects are distinguished by identification information stored within the coded data or separately from the coded data.

[0064] Alternatively, the SPC may be used as the identification unit, in which case the SPC including the VLM that constitutes a static object and the SPC including the VLM that constitutes a dynamic object are distinguished by the above-mentioned identification information.

[0065] Alternatively, the VLM or VXL may be used as the identification unit, in which case the VLM or VXL containing static objects and the VLM or VXL containing dynamic objects are distinguished by the above-mentioned identification information.

[0066] The encoding device may also encode a dynamic object as one or more VLMs or SPCs, and encode a VLM or SPC containing a static object and an SPC containing a dynamic object as different GOSs. If the size of the GOS varies depending on the size of the dynamic object, the encoding device stores the size of the GOS separately as meta information.

[0067] The encoding device may also encode static objects and dynamic objects independently of each other, and overlay the dynamic objects on a world made up of static objects. In this case, the dynamic object is made up of one or more SPCs, and each SPC corresponds to one or more SPCs that make up the static object on which the SPC is overlaid. Note that the dynamic object may be represented by one or more VLMs or VXLs instead of SPCs.

[0068] The encoding device may also encode static objects and dynamic objects as different streams.

[0069] The encoding device may also generate a GOS that includes one or more SPCs that make up a dynamic object. Furthermore, the encoding device may set the GOS (GOS_M) that includes the dynamic object and the GOS of the static object that corresponds to the spatial region of GOS_M to the same size (occupy the same spatial region). This allows superimposition processing to be performed on a GOS-by-GOS basis.

[0070] A P-SPC or B-SPC that configures a dynamic object may refer to an SPC included in a different GOS that has already been coded. In cases where the position of a dynamic object changes over time and the same dynamic object is coded as a GOS at different times, referencing across GOSs is effective from the viewpoint of compression ratio.

[0071] The encoding device may switch between the first and second methods depending on the intended use of the encoded data. For example, when the encoded three-dimensional data is used as a map, it is desirable to be able to separate dynamic objects, so the encoding device uses the second method. On the other hand, when encoding three-dimensional data of an event such as a concert or sporting event, the encoding device uses the first method if there is no need to separate dynamic objects.

[0072] The decode time and display time of a GOS or SPC can be stored in the coded data or as meta information. The time information of all static objects may be the same. In this case, the actual decode time and display time may be determined by the decoding device. Alternatively, a different value may be assigned as the decode time for each GOS or SPC, and the same value may be assigned as the display time for all. Furthermore, a model may be introduced that ensures that the decoder has a buffer of a predetermined size and can decode without failure if it reads a bitstream at a predetermined bit rate according to the decode time, as in a decoder model used in video coding, such as the HEVC HRD (Hypothetical Reference Decoder).

[0073] Next, we will explain the arrangement of GOS within a world. The coordinates of the three-dimensional space in a world are expressed by three mutually orthogonal coordinate axes (x-axis, y-axis, and z-axis). By establishing a predetermined rule for the encoding order of GOS, encoding can be performed so that spatially adjacent GOS are continuous within the encoded data. For example, in the example shown in Figure 4, GOS within the xz plane are encoded continuously. After encoding of all GOS within a certain xz plane is completed, the value of the y-axis is updated. In other words, as encoding progresses, the world extends in the y-axis direction. Furthermore, the index numbers of GOS are set in the encoding order.

[0074] Here, the three-dimensional space of the world is associated one-to-one with absolute geographical coordinates such as GPS or latitude and longitude. Alternatively, the three-dimensional space may be expressed by relative positions from a preset reference position. The directions of the x-, y-, and z-axes of the three-dimensional space are expressed as direction vectors determined based on the latitude and longitude, and the direction vectors are stored as meta information together with the encoded data.

[0075] The size of the GOS is fixed, and the encoding device stores the size as meta information. The size of the GOS may be changed depending on, for example, whether the location is an urban area or whether the location is indoors or outdoors. That is, the size of the GOS may be changed depending on the quantity or nature of objects that have information value. Alternatively, the encoding device may adaptively change the size of the GOS or the spacing between I-SPCs within the GOS depending on, for example, the density of objects within the same world. For example, the higher the object density, the smaller the GOS size and the shorter the spacing between I-SPCs within the GOS.

[0076] In the example shown in Figure 5, the third to tenth GOS regions have a high density of objects, so the GOS are subdivided to allow finer granularity for random access. Note that the seventh to tenth GOS regions are located behind the third to sixth GOS regions, respectively.

[0077] Next, the configuration and operation flow of the three-dimensional data encoding device according to this embodiment will be described. Fig. 6 is a block diagram of the three-dimensional data encoding device 100 according to this embodiment. Fig. 7 is a flowchart showing an example of the operation of the three-dimensional data encoding device 100.

[0078] 6 generates encoded three-dimensional data 112 by encoding three-dimensional data 111. This three-dimensional data encoding device 100 includes an acquisition unit 101, an encoding region determination unit 102, a division unit 103, and an encoding unit 104.

[0079] As shown in FIG. 7, first, the acquisition unit 101 acquires three-dimensional data 111, which is point cloud data (S101).

[0080] Next, the coding area determination unit 102 determines an area to be coded from the spatial area corresponding to the acquired point cloud data (S102). For example, depending on the position of the user or vehicle, the coding area determination unit 102 determines a spatial area around the position as the area to be coded.

[0081] Next, the dividing unit 103 divides the point cloud data included in the region to be coded into processing units. Here, the processing units are the above-mentioned GOS and SPC, etc. Furthermore, this region to be coded corresponds to, for example, the above-mentioned world. Specifically, the dividing unit 103 divides the point cloud data into processing units based on the size of a preset GOS or the presence or size of a dynamic object (S103). Furthermore, the dividing unit 103 determines the start position of the SPC that is the first in coding order in each GOS.

[0082] Next, the encoding unit 104 generates encoded three-dimensional data 112 by sequentially encoding the plurality of SPCs in each GOS (S104).

[0083] Although an example has been shown in which the area to be coded is divided into GOSs and SPCs and then each GOS is coded, the processing procedure is not limited to the above. For example, a procedure may be used in which the configuration of one GOS is determined, the GOS is coded, and then the configuration of the next GOS is determined.

[0084] In this way, the three-dimensional data encoding device 100 generates encoded three-dimensional data 112 by encoding three-dimensional data 111. Specifically, the three-dimensional data encoding device 100 divides the three-dimensional data into first processing units (GOS), which are random access units and each correspond to a three-dimensional coordinate, divides the first processing units (GOS) into a plurality of second processing units (SPC), and divides the second processing units (SPC) into a plurality of third processing units (VLM). Furthermore, the third processing units (VLM) include one or more voxels (VXL), which are the smallest units to which position information can be associated.

[0085] Next, the three-dimensional data encoding device 100 generates encoded three-dimensional data 112 by encoding each of the plurality of first processing units (GOS). Specifically, the three-dimensional data encoding device 100 encodes each of the plurality of second processing units (SPC) in each first processing unit (GOS). Furthermore, the three-dimensional data encoding device 100 encodes each of the plurality of third processing units (VLM) in each second processing unit (SPC).

[0086] For example, when the first processing unit (GOS) to be processed is a closed GOS, the three-dimensional data encoding device 100 encodes the second processing unit (SPC) to be processed included in the first processing unit (GOS) to be processed by referring to other second processing units (SPC) included in the first processing unit (GOS) to be processed. In other words, the three-dimensional data encoding device 100 does not refer to second processing units (SPC) included in first processing units (GOS) different from the first processing unit (GOS) to be processed.

[0087] On the other hand, if the first processing unit (GOS) to be processed is an open GOS, the second processing unit (SPC) to be processed included in the first processing unit (GOS) to be processed is encoded by referring to other second processing units (SPC) included in the first processing unit (GOS) to be processed, or second processing units (SPC) included in a first processing unit (GOS) different from the first processing unit (GOS) to be processed.

[0088] In addition, the three-dimensional data encoding device 100 selects, as the type of the second processing unit (SPC) to be processed, one of the following: a first type (I-SPC) that does not reference other second processing units (SPCs), a second type (P-SPC) that references one other second processing unit (SPC), or a third type that references two other second processing units (SPCs), and encodes the second processing unit (SPC) to be processed according to the selected type.

[0089] Next, the configuration and operation flow of the three-dimensional data decoding device according to this embodiment will be described. Fig. 8 is a block diagram of the blocks of the three-dimensional data decoding device 200 according to this embodiment. Fig. 9 is a flowchart showing an example of the operation of the three-dimensional data decoding device 200.

[0090] 8 generates decoded three-dimensional data 212 by decoding encoded three-dimensional data 211. Here, the encoded three-dimensional data 211 is, for example, the encoded three-dimensional data 112 generated by the three-dimensional data encoding device 100. This three-dimensional data decoding device 200 includes an acquisition unit 201, a decoding start GOS determination unit 202, a decoding SPC determination unit 203, and a decoding unit 204.

[0091] First, the acquisition unit 201 acquires the encoded 3D data 211 (S201). Next, the decoding start GOS determination unit 202 determines a GOS to be decoded (S202). Specifically, the decoding start GOS determination unit 202 refers to meta information stored in the encoded 3D data 211 or separately from the encoded 3D data, and determines a GOS including an SPC corresponding to a spatial position, object, or time at which decoding starts as the GOS to be decoded.

[0092] Next, the decoding SPC determination unit 203 determines the type (I, P, B) of SPC to be decoded in the GOS (S203). For example, the decoding SPC determination unit 203 determines whether to (1) decode only I-SPC, (2) decode I-SPC and P-SPC, or (3) decode all types. Note that if the type of SPC to be decoded has been determined in advance, such as when all SPCs are to be decoded, this step may not be performed.

[0093] Next, the decoding unit 204 acquires the address position in the encoded 3D data 211 where the first SPC in the GOS in decoding order (the same as the encoding order) starts, acquires the encoded data of the first SPC from the address position, and sequentially decodes each SPC in order starting from the first SPC (S204). Note that the address position is stored in meta information or the like.

[0094] In this way, the three-dimensional data decoding device 200 decodes the decoded three-dimensional data 212. Specifically, the three-dimensional data decoding device 200 generates the decoded three-dimensional data 212 of the first processing units (GOS) by decoding each of the encoded three-dimensional data 211 of the first processing units (GOS), which are random access units and each of which is associated with a three-dimensional coordinate. More specifically, the three-dimensional data decoding device 200 decodes each of the plurality of second processing units (SPC) in each first processing unit (GOS). Furthermore, the three-dimensional data decoding device 200 decodes each of the plurality of third processing units (VLM) in each second processing unit (SPC).

[0095] The following describes the meta information for random access. This meta information is generated by the three-dimensional data encoding device 100 and is included in the encoded three-dimensional data 112 (211).

[0096] In conventional random access for 2D video, decoding starts from the first frame of the random access unit that is close to the specified time. On the other hand, in the world, random access to space (coordinates, objects, etc.) is assumed in addition to time.

[0097] Therefore, to realize random access to at least three elements, coordinates, objects, and time, a table is prepared that associates each element with a GOS index number. Furthermore, the GOS index number is associated with the address of the I-SPC at the beginning of the GOS. Figure 10 shows an example of a table included in the meta information. Note that it is not necessary to use all of the tables shown in Figure 10; it is sufficient to use at least one table.

[0098] Hereinafter, as an example, random access starting from a coordinate will be described. When accessing coordinates (x2, y2, z2), first, the coordinate-GOS table is referenced and it is found that the point with coordinates (x2, y2, z2) is included in the second GOS. Next, the GOS address table is referenced and it is found that the address of the first I-SPC in the second GOS is addr(2). Therefore, the decoding unit 204 obtains data from this address and starts decoding.

[0099] The address may be an address in a logical format or a physical address on a hard disk drive or in memory. Information identifying a file segment may be used instead of the address. For example, a file segment is a unit obtained by segmenting one or more GOSs.

[0100] Furthermore, if an object spans multiple GOSs, the object-GOS table may indicate multiple GOSs to which the object belongs. If the multiple GOSs are closed GOSs, the encoding device and decoding device can encode or decode in parallel. On the other hand, if the multiple GOSs are open GOSs, the multiple GOSs can reference each other, thereby improving compression efficiency.

[0101] Examples of objects include people, animals, cars, bicycles, traffic lights, landmark buildings, etc. For example, when encoding a world, the three-dimensional data encoding device 100 can extract feature points specific to objects from a three-dimensional point cloud or the like, detect objects based on the feature points, and set the detected objects as random access points.

[0102] In this way, the three-dimensional data encoding device 100 generates first information indicating a plurality of first processing units (GOS) and three-dimensional coordinates associated with each of the plurality of first processing units (GOS). The encoded three-dimensional data 112 (211) includes this first information. The first information further indicates at least one of an object, a time, and a data storage destination associated with each of the plurality of first processing units (GOS).

[0103] The three-dimensional data decoding device 200 acquires first information from the encoded three-dimensional data 211, and uses the first information to identify the encoded three-dimensional data 211 of the first processing unit corresponding to the specified three-dimensional coordinates, object, or time, and decodes the encoded three-dimensional data 211.

[0104] Other examples of meta information will be described below. In addition to the meta information for random access, the three-dimensional data encoding device 100 may generate and store the following meta information. Furthermore, the three-dimensional data decoding device 200 may use this meta information during decoding.

[0105] When using three-dimensional data as map information, a profile may be defined depending on the application, and information indicating the profile may be included in the meta information. For example, profiles for urban areas, suburban areas, or flying objects may be defined, and the maximum or minimum size of the world, SPC, or VLM may be defined for each. For example, for urban areas, more detailed information is required than for suburban areas, so the minimum size of the VLM is set smaller.

[0106] The meta information may include a tag value indicating the type of object. This tag value is associated with the VLM, SPC, or GOS that constitutes the object. For example, a tag value may be set for each type of object, such as a tag value of "0" indicating a "person," a tag value of "1" indicating a "car," and a tag value of "2" indicating a "traffic light." Alternatively, if the type of object is difficult to determine or does not need to be determined, a tag value indicating a property such as size or whether the object is dynamic or static may be used.

[0107] The meta information may also include information indicating the range of the spatial region occupied by the world.

[0108] The meta information may also store the size of the SPC or VXL as header information common to a plurality of SPCs, such as the entire stream of coded data or an SPC in a GOS.

[0109] The meta information may also include identification information for the range sensor or camera used to generate the point cloud, or information indicating the positional accuracy of the points in the point cloud.

[0110] The meta information may also include information indicating whether the world is made up of only static objects or whether it also includes dynamic objects.

[0111] A modification of this embodiment will now be described.

[0112] The encoding device or decoding device may encode or decode two or more different SPCs or GOSs in parallel. The GOSs to be encoded or decoded in parallel can be determined based on meta-information indicating the spatial positions of the GOSs.

[0113] In cases where three-dimensional data is used as a spatial map for vehicles or flying objects moving around, or where such a spatial map is to be generated, the encoding device or decoding device may encode or decode a GOS or SPC contained in a space identified based on GPS, route information, zoom magnification, etc.

[0114] Furthermore, the decoding device may perform decoding in order from the space closest to the current location or the travel route. The encoding device or decoding device may encode or decode a space farther from the current location or the travel route by lowering the priority compared to a closer space. Here, lowering the priority means lowering the processing order, lowering the resolution (thinning out the data before processing), or lowering the image quality (increasing the encoding efficiency, for example, by increasing the quantization step), etc.

[0115] Furthermore, when decoding coded data that has been coded hierarchically in space, the decoding device may decode only the lower layers.

[0116] The decoding device may also decode data preferentially from the lowest layer depending on the zoom factor or purpose of the map.

[0117] In addition, for applications such as self-position estimation or object recognition performed when a car or robot is driving autonomously, the encoding device or decoding device may encode or decode with reduced resolution except for areas within a specific height from the road surface (area where recognition is performed).

[0118] The encoding device may also encode point clouds representing indoor and outdoor spatial shapes separately. For example, by separating the GOS representing the indoor space (indoor GOS) from the GOS representing the outdoor space (outdoor GOS), the decoding device can select the GOS to decode depending on the viewpoint position when using the encoded data.

[0119] The encoding device may also encode indoor and outdoor GOS with nearby coordinates so that they are adjacent in the encoded stream. For example, the encoding device may associate identifiers for the two and store information indicating the associated identifiers in the encoded stream or in separately stored meta information. This allows the decoding device to identify indoor and outdoor GOS with nearby coordinates by referring to the information in the meta information.

[0120] The encoding device may also switch the size of the GOS or SPC between indoor and outdoor GOS. For example, the encoding device may set a smaller GOS size indoors than outdoors. The encoding device may also change the accuracy of extracting feature points from the point cloud or the accuracy of object detection between indoor and outdoor GOS.

[0121] The encoding device may also add information to the encoded data that enables the decoding device to distinguish dynamic objects from static objects. This allows the decoding device to display dynamic objects together with red frames or explanatory text. The decoding device may also display only red frames or explanatory text instead of dynamic objects. The decoding device may also display more specific object types. For example, a red frame may be used for cars and a yellow frame for people.

[0122] Furthermore, the encoding device or decoding device may determine whether to encode or decode dynamic objects and static objects as different SPCs or GOSs depending on the frequency of appearance of dynamic objects, the ratio of static objects to dynamic objects, etc. For example, if the frequency of appearance or ratio of dynamic objects exceeds a threshold, an SPC or GOS in which dynamic objects and static objects are mixed is permitted, and if the frequency of appearance or ratio of dynamic objects does not exceed the threshold, an SPC or GOS in which dynamic objects and static objects are mixed is not permitted.

[0123] When detecting dynamic objects from two-dimensional camera image information rather than from a point cloud, the encoding device may separately acquire information for identifying the detection result (such as a frame or text) and the object position, and encode this information as part of the three-dimensional encoded data. In this case, the decoding device displays auxiliary information (such as a frame or text) indicating the dynamic object by superimposing it on the decoding result of the static object.

[0124] The encoding device may also change the density of the VXL or VLM in the SPC depending on factors such as the complexity of the shape of the static object. For example, the encoding device may set the VXL or VLM to a higher density as the shape of the static object becomes more complex. Furthermore, the encoding device may determine the quantization step, etc., used when quantizing spatial position or color information depending on the density of the VXL or VLM. For example, the encoding device may set a smaller quantization step as the VXL or VLM becomes denser.

[0125] As described above, the encoding device or decoding device according to this embodiment encodes or decodes space in units of spaces each having coordinate information.

[0126] Furthermore, the encoding device and the decoding device perform encoding and decoding in units of volumes within a space. A volume includes voxels, which are the smallest units to which position information can be associated.

[0127] The encoding device and decoding device perform encoding or decoding by associating any elements using a table that associates each element of spatial information, including coordinates, objects, and time, with a GOP, or a table that associates each element with another element. The decoding device determines coordinates using the value of a selected element, identifies a volume, voxel, or space from the coordinates, and decodes the space including the volume or voxel, or the identified space.

[0128] The encoding device also determines a volume, voxel, or space that can be selected by an element through feature point extraction or object recognition, and encodes it as a randomly accessible volume, voxel, or space.

[0129] Spaces are classified into three types: I-SPC, which can be encoded or decoded by itself; P-SPC, which is encoded or decoded by referring to any one processed space; and B-SPC, which is encoded or decoded by referring to any two processed spaces.

[0130] One or more volumes correspond to static or dynamic objects. The space containing the static objects and the space containing the dynamic objects are coded or decoded as different GOSs. That is, the SPC containing the static objects and the SPC containing the dynamic objects are assigned to different GOSs.

[0131] Dynamic objects are encoded or decoded on an object-by-object basis and associated with one or more spaces containing static objects, i.e., multiple dynamic objects are encoded individually, and the resulting encoded data for the multiple dynamic objects is associated with the SPC containing the static objects.

[0132] The encoding device and the decoding device perform encoding or decoding by increasing the priority of the I-SPC in the GOS. For example, the encoding device performs encoding so as to minimize degradation of the I-SPC (so that the original 3D data is reproduced more faithfully after decoding). Also, the decoding device decodes only the I-SPC, for example.

[0133] The encoding device may perform encoding by changing the frequency of using I-SPCs depending on the density or number (quantity) of objects in the world. In other words, the encoding device changes the frequency of selecting I-SPCs depending on the number or density of objects included in the three-dimensional data. For example, the encoding device may use I-spaces more frequently as the density of objects in the world increases.

[0134] Furthermore, the encoding device sets random access points in units of GOS, and stores information indicating the spatial region corresponding to the GOS in the header information.

[0135] The encoding device uses, for example, a default value as the spatial size of the GOS. Note that the encoding device may change the size of the GOS depending on the number (quantity) or density of objects or dynamic objects. For example, the encoding device reduces the spatial size of the GOS as the density or number of objects or dynamic objects increases.

[0136] The space or volume also includes a set of feature points derived using information obtained by sensors such as a depth sensor, a gyroscope, or a camera. The coordinates of the feature points are set at the center positions of the voxels. Furthermore, by subdividing the voxels, it is possible to achieve high accuracy of the position information.

[0137] The feature point group is derived using multiple pictures, each of which has at least two types of time information: actual time information and time information that is the same for multiple pictures associated with the space (e.g., encoding time used for rate control, etc.).

[0138] Also, encoding or decoding is performed in units of GOS, each GOS including one or more spaces.

[0139] The encoding device and the decoding device refer to the spaces in the processed GOS to predict the P space or the B space in the GOS to be processed.

[0140] Alternatively, the encoding device and the decoding device do not refer to a different GOS, but predict the P space or the B space in the GOS to be processed using the processed space in the GOS to be processed.

[0141] Furthermore, the encoding device and the decoding device transmit or receive the encoded stream in units of worlds each including one or more GOSs.

[0142] Furthermore, the GOS has a layer structure in at least one direction within a world, and the encoding device and decoding device encode or decode from the lower layer. For example, a randomly accessible GOS belongs to the lowest layer. A GOS belonging to a higher layer references a GOS belonging to the same layer or lower. In other words, the GOS is spatially divided in a predetermined direction and includes multiple layers, each containing one or more SPCs. The encoding device and decoding device encode or decode each SPC by referring to an SPC included in the same layer as the SPC or in a layer lower than the SPC.

[0143] Furthermore, the encoding device and the decoding device encode or decode consecutive GOSs within a world unit including multiple GOSs. The encoding device and the decoding device write or read information indicating the encoding or decoding order (direction) as metadata. In other words, the encoded data includes information indicating the encoding order of multiple GOSs.

[0144] Furthermore, the encoding device and the decoding device encode or decode two or more different spaces or GOSs in parallel.

[0145] The encoding device and decoding device also encode and decode spatial information (coordinates, size, etc.) of the space or GOS.

[0146] Furthermore, the encoding device and decoding device encode or decode a space or GOS included in a specific space that is specified based on external information related to its own position and / or area size, such as GPS, route information, or magnification.

[0147] The encoding device or decoding device encodes or decodes spaces farther from its own position with lower priority than spaces closer to its own position.

[0148] The encoding device sets one direction of the world according to the magnification or use, and encodes the GOS having a layer structure in that direction. The decoding device decodes the GOS having a layer structure in one direction of the world set according to the magnification or use, preferentially from the lower layer.

[0149] The encoding device varies the feature point extraction, object recognition accuracy, spatial region size, etc., included in the indoor and outdoor spaces. However, the encoding device and decoding device encode or decode the indoor GOS and outdoor GOS that are close in coordinates as adjacent in the world, and also associate their identifiers and encode or decode them.

[0150] (Embodiment 2) When using encoded point cloud data in an actual device or service, it is desirable to transmit and receive the information required for the application in order to reduce network bandwidth. However, until now, such a function has not existed in the encoding structure of 3D data, and no encoding method for this purpose has existed.

[0151] In this embodiment, we will describe a three-dimensional data encoding method and a three-dimensional data encoding device that provide the function of transmitting and receiving only the information necessary for the purpose in encoded data of a three-dimensional point cloud, as well as a three-dimensional data decoding method and a three-dimensional data decoding device that decodes the encoded data.

[0152] A voxel (VXL) having a certain amount of features or more is defined as a feature voxel (FVXL), and a world (WLD) composed of FVXL is defined as a sparse world (SWLD). Figure 11 shows an example of the configuration of a sparse world and a world. The SWLD includes FGOS, which is a GOS composed of FVXL, FSPC, which is an SPC composed of FVXL, and FVLM, which is a VLM composed of FVXL. The data structures and prediction structures of FGOS, FSPC, and FVLM may be the same as those of GOS, SPC, and VLM.

[0153] The feature is a feature that expresses three-dimensional position information of the VXL or visible light information of the VXL position, and is a feature that is often detected especially at corners and edges of three-dimensional objects. Specifically, this feature is a three-dimensional feature or visible light feature as described below, but any feature that expresses the position, brightness, color information, etc. of the VXL may be used.

[0154] As the three-dimensional feature, SHOT feature (Signature of Histograms of OrienTations), PFH feature (Point Feature Histograms), or PPF feature (Point Pair Feature) is used.

[0155] The SHOT feature is obtained by dividing the VXL area, calculating the dot product between the reference point and the normal vector of each divided area, and creating a histogram. This SHOT feature has the advantage of being highly dimensional and expressive.

[0156] The PFH feature is obtained by selecting many pairs of points near the VXL, calculating normal vectors from those two points, and creating a histogram. Because the PFH feature is a histogram feature, it is robust against some disturbances and has high expressive power.

[0157] The PPF feature is calculated using normal vectors etc. for each of two VXL points. Since all VXLs are used for this PPF feature, it is robust against occlusion.

[0158] Furthermore, as the feature amount of visible light, SIFT (Scale-Invariant Feature Transform), SURF (Speeded Up Robust Features), HOG (Histogram of Oriented Gradients), or the like, which use information such as brightness gradient information of an image, can be used.

[0159] The SWLD is generated by calculating the above feature values ​​from each VXL of the WLD and extracting the FVXL. Here, the SWLD may be updated every time the WLD is updated, or may be updated periodically after a certain period of time has elapsed, regardless of the timing of updating the WLD.

[0160] A SWLD may be generated for each feature. For example, a separate SWLD may be generated for each feature, such as SWLD1 based on SHOT features and SWLD2 based on SIFT features, and different SWLDs may be used depending on the application. Furthermore, the calculated features of each FVXL may be stored in each FVXL as feature information.

[0161] Next, we explain how to use sparse world learning (SWLD). SWLDs contain only feature voxels (FVXL), so their data size is generally smaller than WLDs, which contain all VXL.

[0162] In applications that use features to achieve certain purposes, using SWLD information instead of WLD information can reduce the time required to read from the hard disk, as well as the bandwidth and transfer time required for network transfer. For example, by storing WLD and SWLD as map information on the server and switching between WLD and SWLD as the map information to be sent in response to a client request, the network bandwidth and transfer time can be reduced. A specific example is shown below.

[0163] 12 and 13 are diagrams illustrating examples of using SWLDs and WLDs. As shown in FIG. 12, when a client 1, which is an in-vehicle device, needs map information for determining its own location, the client 1 sends a request to the server to acquire map data for self-location estimation (S301). The server transmits an SWLD to the client 1 in response to the acquisition request (S302). The client 1 determines its own location using the received SWLD (S303). In this case, the client 1 acquires VXL information around the client 1 using various methods, such as a distance sensor such as a range finder, a stereo camera, or a combination of multiple monocular cameras, and estimates its own location information from the acquired VXL information and SWLD. Here, the self-location information includes the three-dimensional location information and orientation of the client 1.

[0164] 13, when a client 2, which is an in-vehicle device, needs map information for map drawing purposes such as a three-dimensional map, the client 2 sends a request to the server to acquire map data for map drawing (S311). The server transmits a WLD to the client 2 in response to the acquisition request (S312). The client 2 uses the received WLD to draw the map (S313). In this case, the client 2 creates a rendering image using, for example, an image captured by the client 2 using a visible light camera or the like and the WLD acquired from the server, and draws the created image on the screen of a car navigation system or the like.

[0165] As described above, the server sends SWLD to the client for applications that mainly require the feature values ​​of each VXL, such as self-location estimation, and sends WLD to the client for applications that require detailed VXL information, such as map drawing. This enables efficient transmission and reception of map data.

[0166] The client may decide for itself whether it needs a SWLD or a WLD and request the server to send either a SWLD or a WLD. The server may also decide whether to send a SWLD or a WLD depending on the client or network conditions.

[0167] Next, we will explain how to switch between sending and receiving data in the sparse world (SWLD) and the world (WLD).

[0168] Whether to receive a WLD or an SWLD may be switched depending on the network bandwidth. FIG. 14 shows an example of operation in this case. For example, when a low-speed network with limited available network bandwidth, such as an LTE (Long Term Evolution) environment, is used, the client accesses the server via the low-speed network (S321) and acquires an SWLD as map information from the server (S322). On the other hand, when a high-speed network with ample network bandwidth, such as a Wi-Fi (registered trademark) environment, is used, the client accesses the server via the high-speed network (S323) and acquires a WLD from the server (S324). This allows the client to acquire appropriate map information depending on the client's network bandwidth.

[0169] Specifically, the client receives SWLD via LTE when outdoors, and acquires WLD via Wi-Fi (registered trademark) when inside a facility, etc. This allows the client to obtain more detailed map information for the indoor area.

[0170] In this way, the client may request a WLD or SWLD from the server depending on the bandwidth of the network it uses. Alternatively, the client may send information indicating the bandwidth of the network it uses to the server, and the server may send data (WLD or SWLD) suitable for the client depending on the information. Alternatively, the server may determine the network bandwidth of the client and send data (WLD or SWLD) suitable for the client.

[0171] Moreover, whether to receive a WLD or a SWLD may be switched depending on the moving speed. FIG. 15 shows an example of operation in this case. For example, when the client is moving at high speed (S331), the client receives a SWLD from the server (S332). On the other hand, when the client is moving at low speed (S333), the client receives a WLD from the server (S334). This allows the client to acquire map information suited to the speed while suppressing network bandwidth. Specifically, by receiving a SWLD with a small amount of data while traveling on a highway, the client can update rough map information at an appropriate speed. On the other hand, by receiving a WLD while traveling on an ordinary road, the client can acquire more detailed map information.

[0172] In this way, the client may request a WLD or SWLD from the server according to its own moving speed. Alternatively, the client may send information indicating its own moving speed to the server, and the server may send data (WLD or SWLD) suitable for the client according to the information. Alternatively, the server may determine the moving speed of the client and send data (WLD or SWLD) suitable for the client.

[0173] Alternatively, the client may first obtain SWLD from the server and then obtain WLD for important areas within that. For example, when obtaining map data, the client may first obtain rough map information in SWLD, then narrow down the area to areas where features such as buildings, signs, or people frequently appear, and later obtain WLD for the narrowed down area. This allows the client to obtain detailed information for the required area while reducing the amount of data received from the server.

[0174] Alternatively, the server may create a separate SWLD for each object from the WLD, and the client may receive each one depending on the application. This reduces network bandwidth. For example, the server may recognize people or cars from the WLD in advance and create a SWLD for people and a SWLD for cars. The client may receive a SWLD for people if it wants to obtain information about people around it, or a SWLD for cars if it wants to obtain information about cars. The types of SWLDs may also be distinguished by information (such as a flag or type) added to the header, etc.

[0175] Next, the configuration and operation flow of a three-dimensional data encoding device (e.g., a server) according to this embodiment will be described. Fig. 16 is a block diagram of a three-dimensional data encoding device 400 according to this embodiment. Fig. 17 is a flowchart of three-dimensional data encoding processing by the three-dimensional data encoding device 400.

[0176] 16 encodes input three-dimensional data 411 to generate encoded three-dimensional data 413 and 414, which are encoded streams. Here, the encoded three-dimensional data 413 is encoded three-dimensional data corresponding to a WLD, and the encoded three-dimensional data 414 is encoded three-dimensional data corresponding to a SWLD. This three-dimensional data encoding device 400 includes an acquisition unit 401, a coding region determination unit 402, a SWLD extraction unit 403, a WLD encoding unit 404, and a SWLD encoding unit 405.

[0177] As shown in FIG. 17, first, the acquisition unit 401 acquires input three-dimensional data 411, which is point cloud data in a three-dimensional space (S401).

[0178] Next, the coding region determination unit 402 determines a spatial region to be coded based on the spatial region in which the point cloud data exists (S402).

[0179] Next, the SWLD extraction unit 403 defines the spatial region to be coded as a WLD and calculates a feature amount from each VXL included in the WLD.The SWLD extraction unit 403 then extracts VXLs whose feature amounts are equal to or greater than a predetermined threshold, defines the extracted VXLs as FVXLs, and adds the FVXLs to the SWLD to generate extracted three-dimensional data 412 (S403).In other words, extracted three-dimensional data 412 whose feature amounts are equal to or greater than the threshold are extracted from the input three-dimensional data 411.

[0180] Next, the WLD encoding unit 404 generates encoded three-dimensional data 413 corresponding to the WLD by encoding the input three-dimensional data 411 corresponding to the WLD (S404). At this time, the WLD encoding unit 404 adds information to the header of the encoded three-dimensional data 413 to distinguish that the encoded three-dimensional data 413 is a stream including a WLD.

[0181] Furthermore, the SWLD encoding unit 405 generates encoded three-dimensional data 414 corresponding to the SWLD by encoding the extracted three-dimensional data 412 corresponding to the SWLD (S405). At this time, the SWLD encoding unit 405 adds information to the header of the encoded three-dimensional data 414 to distinguish that the encoded three-dimensional data 414 is a stream including an SWLD.

[0182] The order of the process for generating the encoded three-dimensional data 413 and the process for generating the encoded three-dimensional data 414 may be reversed. Also, some or all of these processes may be performed in parallel.

[0183] For example, a parameter called "world_type" is defined as information added to the headers of the encoded 3D data 413 and 414. world_type=0 indicates that the stream includes a WLD, and world_type=1 indicates that the stream includes a SWLD. If many other types are defined, the assigned numerical value may be increased, such as world_type=2. Furthermore, a specific flag may be included in one of the encoded 3D data 413 and 414. For example, a flag indicating that the stream includes a SWLD may be added to the encoded 3D data 414. In this case, the decoding device can determine whether the stream includes a WLD or a SWLD based on the presence or absence of the flag.

[0184] Furthermore, the encoding method used by the WLD encoding unit 404 when encoding the WLD may be different from the encoding method used by the SWLD encoding unit 405 when encoding the SWLD.

[0185] For example, in SWLD, data is thinned out, so that correlation with surrounding data may be lower than in WLD. Therefore, in the encoding method used in SWLD, inter prediction may be prioritized over intra prediction and inter prediction in comparison with the encoding method used in WLD.

[0186] Furthermore, the encoding method used for SWLD and the encoding method used for WLD may differ in the way three-dimensional positions are expressed. For example, in SWLD, the three-dimensional position of FVXL may be expressed by three-dimensional coordinates, and in WLD, the three-dimensional position may be expressed by an octree, which will be described later, or vice versa.

[0187] Furthermore, the SWLD encoding unit 405 performs encoding so that the data size of the encoded three-dimensional data 414 of SWLD is smaller than the data size of the encoded three-dimensional data 413 of WLD. For example, as described above, there is a possibility that correlation between data in SWLD is lower than that in WLD. This may result in a decrease in encoding efficiency, and the data size of the encoded three-dimensional data 414 may be larger than the data size of the encoded three-dimensional data 413 of WLD. Therefore, if the data size of the obtained encoded three-dimensional data 414 is larger than the data size of the encoded three-dimensional data 413 of WLD, the SWLD encoding unit 405 re-encodes the data to regenerate encoded three-dimensional data 414 with a reduced data size.

[0188] For example, the SWLD extraction unit 403 regenerates extracted three-dimensional data 412 with a reduced number of extracted feature points, and the SWLD encoding unit 405 encodes the extracted three-dimensional data 412. Alternatively, the degree of quantization in the SWLD encoding unit 405 may be made coarser. For example, in an octree structure described below, the degree of quantization can be made coarser by rounding the data in the lowest layer.

[0189] Furthermore, if the data size of the encoded three-dimensional data 414 of SWLD cannot be made smaller than the data size of the encoded three-dimensional data 413 of WLD, the SWLD encoding unit 405 may not generate the encoded three-dimensional data 414 of SWLD. Alternatively, the encoded three-dimensional data 413 of WLD may be copied to the encoded three-dimensional data 414 of SWLD. In other words, the encoded three-dimensional data 413 of WLD may be used as is as the encoded three-dimensional data 414 of SWLD.

[0190] Next, the configuration and operation flow of a three-dimensional data decoding device (e.g., a client) according to this embodiment will be described. Fig. 18 is a block diagram of a three-dimensional data decoding device 500 according to this embodiment. Fig. 19 is a flowchart of three-dimensional data decoding processing by the three-dimensional data decoding device 500.

[0191] 18 generates decoded three-dimensional data 512 or 513 by decoding encoded three-dimensional data 511. Here, the encoded three-dimensional data 511 is, for example, the encoded three-dimensional data 413 or 414 generated by the three-dimensional data encoding device 400.

[0192] This three-dimensional data decoding device 500 includes an acquisition unit 501 , a header analysis unit 502 , a WLD decoding unit 503 , and a SWLD decoding unit 504 .

[0193] 19, first, the acquisition unit 501 acquires encoded three-dimensional data 511 (S501). Next, the header analysis unit 502 analyzes the header of the encoded three-dimensional data 511 and determines whether the encoded three-dimensional data 511 is a stream including a WLD or a stream including a SWLD (S502). For example, the determination is made by referring to the world_type parameter described above.

[0194] If the encoded three-dimensional data 511 is a stream including a WLD (Yes in S503), the WLD decoding unit 503 decodes the encoded three-dimensional data 511 to generate decoded three-dimensional data 512 of the WLD (S504). On the other hand, if the encoded three-dimensional data 511 is a stream including an SWLD (No in S503), the SWLD decoding unit 504 decodes the encoded three-dimensional data 511 to generate decoded three-dimensional data 513 of the SWLD (S505).

[0195] Also, similarly to the encoding device, the decoding method used by the WLD decoding unit 503 when decoding the WLD may be different from the decoding method used by the SWLD decoding unit 504 when decoding the SWLD. For example, in the decoding method used for the SWLD, inter prediction, out of intra prediction and inter prediction, may be given priority over the decoding method used for the WLD.

[0196] Furthermore, the decoding method used for SWLD and the decoding method used for WLD may use different methods to represent three-dimensional positions. For example, in SWLD, the three-dimensional position of FVXL may be represented by three-dimensional coordinates, and in WLD, the three-dimensional position may be represented by an octree, which will be described later, or vice versa.

[0197] Next, we will explain the octree representation, which is a method of representing three-dimensional positions. VXL data included in three-dimensional data is converted into an octree structure and then encoded. Fig. 20 is a diagram showing an example of a VXL in a WLD. Fig. 21 is a diagram showing the octree structure of the WLD shown in Fig. 20. In the example shown in Fig. 20, there are three VXLs (hereinafter referred to as valid VXLs) VXL1 to VXL3 that contain point clouds. As shown in Fig. 21, the octree structure is composed of nodes and leaves. Each node has a maximum of eight nodes or leaves. Each leaf has VXL information. Here, among the leaves shown in Fig. 21, leaves 1, 2, and 3 represent VXL1, VXL2, and VXL3 shown in Fig. 20, respectively.

[0198] Specifically, each node and leaf corresponds to a three-dimensional position. Node 1 corresponds to the entire block shown in FIG. 20. The block corresponding to node 1 is divided into eight blocks, and of the eight blocks, the block containing a valid VXL is set as a node, and the other blocks are set as leaves. The block corresponding to the node is further divided into eight nodes or leaves, and this process is repeated for each level of the tree structure. In addition, all blocks in the lowest level are set as leaves.

[0199] FIG. 22 is a diagram showing an example of an SWLD generated from the WLD shown in FIG. 20. VXL1 and VXL2 shown in FIG. 20 are determined to be FVXL1 and FVXL2 as a result of feature extraction and are added to the SWLD. On the other hand, VXL3 is not determined to be FVXL and is not included in the SWLD. FIG. 23 is a diagram showing the octree structure of the SWLD shown in FIG. 22. In the octree structure shown in FIG. 23, leaf 3 corresponding to VXL3 shown in FIG. 21 is deleted. As a result, node 3 shown in FIG. 21 no longer has a valid VXL and is changed to a leaf. As such, the number of leaves in an SWLD is generally smaller than the number of leaves in a WLD, and the encoded 3D data of the SWLD is also smaller than the encoded 3D data of the WLD.

[0200] A modification of this embodiment will now be described.

[0201] For example, when a client such as an in-vehicle device estimates its own position, it receives a SWLD from a server and estimates its own position using the SWLD, and when detecting an obstacle, it may perform obstacle detection based on three-dimensional information of the surrounding area that it has acquired using various methods such as a distance sensor such as a range finder, a stereo camera, or a combination of multiple monocular cameras.

[0202] In addition, SWLDs generally do not contain VXL data for flat areas. Therefore, the server may store a subsampled world (subWLD) that is a subsample of the WLD for static obstacle detection, and transmit the SWLD and subWLD to the client. This allows the client to perform localization and obstacle detection while reducing network bandwidth.

[0203] Furthermore, when a client wants to quickly draw 3D map data, it may be more convenient for the map information to have a mesh structure. Therefore, the server may generate a mesh from the WLD and store it in advance as a mesh world (MWLD). For example, a client may receive an MWLD when it needs a coarse 3D drawing, and a WLD when it needs a detailed 3D drawing. This reduces network bandwidth.

[0204] Furthermore, although the server sets the VXLs among the VXLs whose feature quantities are equal to or greater than a threshold as FVXLs, FVXLs may be calculated using a different method. For example, the server may determine that the VXLs, VLMs, SPCs, or GOSs constituting a traffic light or intersection are necessary for self-localization, driving assistance, or autonomous driving, and include them in the SWLD as FVXLs, FVLMs, FSPCs, and FGOSs. This determination may also be performed manually. The FVXLs obtained by the above method may be added to the FVXLs set based on the feature quantities. That is, the SWLD extraction unit 403 may further extract data corresponding to objects having predetermined attributes from the input three-dimensional data 411 as extracted three-dimensional data 412.

[0205] Furthermore, the fact that it is necessary for such purposes may be labeled separately from the features. Furthermore, the server may separately store FVXL necessary for self-localization at traffic lights or intersections, driving assistance, autonomous driving, etc. as a higher layer (e.g., lane world) of SWLD.

[0206] The server may also add attributes to the VXL in the WLD for each random access unit or for each predetermined unit. The attributes include, for example, information indicating whether the VXL is necessary or unnecessary for self-location estimation, or information indicating whether the VXL is important as traffic information such as a traffic light or intersection. The attributes may also include a correspondence relationship with a feature (such as an intersection or road) in lane information (such as GDF: Geographic Data Files).

[0207] Furthermore, the following method may be used as a method for updating the WLD or SWLD.

[0208] Updates indicating changes in people, construction, or tree-lined streets (for trucks) are uploaded to the server as point clouds or metadata. The server updates the WLD based on the upload, and then updates the SWLD using the updated WLD.

[0209] In addition, if the client detects an inconsistency between the 3D information it generated during self-location estimation and the 3D information it received from the server, it may send the 3D information it generated to the server along with an update notification. In this case, the server updates the SWLD using the WLD. If the SWLD is not updated, the server determines that the WLD itself is out of date.

[0210] Although information for distinguishing between WLD and SWLD is added to the header information of the coded stream, if there are multiple types of worlds, such as mesh worlds or lane worlds, information for distinguishing between them may be added to the header information. Also, if there are multiple SWLDs with different features, information for distinguishing between them may be added to the header information.

[0211] Furthermore, although the SWLD is described as being composed of FVXL, it may also include VXL that has not been determined to be FVXL. For example, the SWLD may include adjacent VXL that are used when calculating the feature of FVXL. This allows the client to calculate the feature of FVXL when receiving the SWLD, even if feature information is not added to each FVXL in the SWLD. In this case, the SWLD may include information for distinguishing whether each VXL is FVXL or VXL.

[0212] As described above, the three-dimensional data encoding device 400 extracts extracted three-dimensional data 412 (second three-dimensional data) whose feature amount is greater than or equal to a threshold value from input three-dimensional data 411 (first three-dimensional data), and generates encoded three-dimensional data 414 (first encoded three-dimensional data) by encoding the extracted three-dimensional data 412.

[0213] According to this, the three-dimensional data encoding device 400 generates encoded three-dimensional data 414 by encoding data whose feature amount is equal to or greater than a threshold. This allows the amount of data to be reduced compared to when the input three-dimensional data 411 is encoded as is. Therefore, the three-dimensional data encoding device 400 can reduce the amount of data to be transmitted.

[0214] Moreover, the three-dimensional data encoding device 400 further encodes the input three-dimensional data 411 to generate encoded three-dimensional data 413 (second encoded three-dimensional data).

[0215] This allows the three-dimensional data encoding device 400 to selectively transmit the encoded three-dimensional data 413 and the encoded three-dimensional data 414 depending on, for example, the intended use.

[0216] Furthermore, the extracted three-dimensional data 412 is coded by a first coding method, and the input three-dimensional data 411 is coded by a second coding method that is different from the first coding method.

[0217] This allows the three-dimensional data encoding device 400 to use encoding methods suited to the input three-dimensional data 411 and the extracted three-dimensional data 412, respectively.

[0218] Furthermore, in the first encoding method, of intra prediction and inter prediction, inter prediction is given priority over the second encoding method.

[0219] This allows the 3D data encoding device 400 to increase the priority of inter prediction for the extracted 3D data 412, which tends to have low correlation between adjacent data.

[0220] Furthermore, the first and second encoding methods differ in the way they represent three-dimensional positions: for example, the second encoding method represents three-dimensional positions using an octree, while the first encoding method represents three-dimensional positions using three-dimensional coordinates.

[0221] This allows the three-dimensional data encoding device 400 to use a more suitable three-dimensional position representation method for three-dimensional data with different numbers of data (number of VXLs or FVXLs).

[0222] Furthermore, at least one of the encoded three-dimensional data 413 and 414 includes an identifier indicating whether the encoded three-dimensional data is encoded three-dimensional data obtained by encoding the input three-dimensional data 411, or encoded three-dimensional data obtained by encoding a portion of the input three-dimensional data 411. In other words, the identifier indicates whether the encoded three-dimensional data is encoded three-dimensional data 413 of WLD or encoded three-dimensional data 414 of SWLD.

[0223] This allows the decoding device to easily determine whether the acquired encoded three-dimensional data is encoded three-dimensional data 413 or encoded three-dimensional data 414.

[0224] Furthermore, the three-dimensional data encoding device 400 encodes the extracted three-dimensional data 412 so that the amount of data of the encoded three-dimensional data 414 is smaller than the amount of data of the encoded three-dimensional data 413 .

[0225] According to this, the three-dimensional data encoding device 400 can make the data amount of the encoded three-dimensional data 414 smaller than the data amount of the encoded three-dimensional data 413 .

[0226] Furthermore, the three-dimensional data encoding device 400 further extracts data corresponding to an object having a predetermined attribute from the input three-dimensional data 411 as extracted three-dimensional data 412. For example, the object having the predetermined attribute is an object necessary for self-position estimation, driving assistance, automatic driving, or the like, such as a traffic light or an intersection.

[0227] This allows the three-dimensional data encoding device 400 to generate encoded three-dimensional data 414 that includes data required by the decoding device.

[0228] Furthermore, the three-dimensional data encoding device 400 (server) further transmits one of the encoded three-dimensional data 413 and 414 to the client depending on the state of the client.

[0229] This allows the three-dimensional data encoding device 400 to transmit appropriate data depending on the state of the client.

[0230] The state of the client also includes the communication status of the client (for example, the network bandwidth) or the movement speed of the client.

[0231] Furthermore, the three-dimensional data encoding device 400 further transmits one of the encoded three-dimensional data 413 and 414 to the client in response to a request from the client.

[0232] This allows the three-dimensional data encoding device 400 to transmit appropriate data in response to a client request.

[0233] Furthermore, the three-dimensional data decoding device 500 according to this embodiment decodes the encoded three-dimensional data 413 or 414 generated by the three-dimensional data encoding device 400 described above.

[0234] That is, the three-dimensional data decoding device 500 decodes, by a first decoding method, encoded three-dimensional data 414 obtained by encoding extracted three-dimensional data 412 whose feature amount extracted from input three-dimensional data 411 is equal to or greater than a threshold value. Also, the three-dimensional data decoding device 500 decodes, by a second decoding method different from the first decoding method, encoded three-dimensional data 413 obtained by encoding the input three-dimensional data 411.

[0235] According to this, the three-dimensional data decoding device 500 can selectively receive the encoded three-dimensional data 414, which is generated by encoding data whose feature amount is equal to or greater than a threshold, and the encoded three-dimensional data 413, depending on, for example, the intended use. This allows the three-dimensional data decoding device 500 to reduce the amount of data to be transmitted. Furthermore, the three-dimensional data decoding device 500 can use decoding methods suitable for the input three-dimensional data 411 and the extracted three-dimensional data 412, respectively.

[0236] Furthermore, in the first decoding method, of intra prediction and inter prediction, inter prediction is given priority over the second decoding method.

[0237] This allows the 3D data decoding device 500 to increase the priority of inter prediction for extracted 3D data that is likely to have low correlation between adjacent data.

[0238] Furthermore, the first and second decoding methods differ in the way they represent three-dimensional positions. For example, the second decoding method represents three-dimensional positions using an octree, while the first decoding method represents three-dimensional positions using three-dimensional coordinates.

[0239] This allows the three-dimensional data decoding device 500 to use a more suitable three-dimensional position representation method for three-dimensional data with different numbers of data (number of VXLs or FVXLs).

[0240] Furthermore, at least one of the encoded three-dimensional data 413 and 414 includes an identifier indicating whether the encoded three-dimensional data is encoded three-dimensional data obtained by encoding the input three-dimensional data 411, or encoded three-dimensional data obtained by encoding a portion of the input three-dimensional data 411. The three-dimensional data decoding device 500 identifies the encoded three-dimensional data 413 and 414 by referring to the identifier.

[0241] This allows the three-dimensional data decoding device 500 to easily determine whether the acquired encoded three-dimensional data is the encoded three-dimensional data 413 or the encoded three-dimensional data 414.

[0242] Furthermore, the three-dimensional data decoding device 500 notifies the server of the status of the client (three-dimensional data decoding device 500). The three-dimensional data decoding device 500 receives one of the encoded three-dimensional data 413 and 414 transmitted from the server depending on the status of the client.

[0243] This allows the three-dimensional data decoding device 500 to receive appropriate data depending on the state of the client.

[0244] The state of the client also includes the communication status of the client (for example, the network bandwidth) or the movement speed of the client.

[0245] Furthermore, the three-dimensional data decoding device 500 further requests one of the encoded three-dimensional data 413 and 414 from the server, and receives one of the encoded three-dimensional data 413 and 414 transmitted from the server in response to the request.

[0246] This allows the three-dimensional data decoding device 500 to receive appropriate data according to the application.

[0247] (Embodiment 3) In this embodiment, a method for transmitting and receiving three-dimensional data between vehicles will be described. For example, three-dimensional data is transmitted and received between a vehicle and a surrounding vehicle.

[0248] 24 is a block diagram of a three-dimensional data creation device 620 according to this embodiment. This three-dimensional data creation device 620 is included in, for example, the vehicle itself, and creates denser third three-dimensional data 636 by combining received second three-dimensional data 635 with first three-dimensional data 632 created by the three-dimensional data creation device 620.

[0249] This three-dimensional data creation device 620 includes a three-dimensional data creation unit 621 , a requested range determination unit 622 , a search unit 623 , a reception unit 624 , a decoding unit 625 , and a synthesis unit 626 .

[0250] First, three-dimensional data creation unit 621 creates first three-dimensional data 632 using sensor information 631 detected by a sensor equipped in the vehicle. Next, required range determination unit 622 determines a required range, which is a three-dimensional spatial range for which data is insufficient in created first three-dimensional data 632.

[0251] Next, the search unit 623 searches for nearby vehicles that have three-dimensional data within the requested range, and transmits requested range information 633 indicating the requested range to the nearby vehicles identified by the search. Next, the receiving unit 624 receives encoded three-dimensional data 634, which is an encoded stream of the requested range, from the nearby vehicles (S624). Note that the search unit 623 may indiscriminately issue a request to all vehicles present within a specific range, and receive the encoded three-dimensional data 634 from those that respond. Furthermore, the search unit 623 may issue a request to objects other than vehicles, such as traffic lights or signs, and receive the encoded three-dimensional data 634 from the objects.

[0252] Next, the decoding unit 625 obtains second three-dimensional data 635 by decoding the received encoded three-dimensional data 634. Next, the combining unit 626 combines the first three-dimensional data 632 and the second three-dimensional data 635 to create denser third three-dimensional data 636.

[0253] Next, a description will be given of the configuration and operation of three-dimensional data transmission device 640 according to this embodiment.

[0254] The three-dimensional data transmission device 640 is included in, for example, the surrounding vehicle described above, processes the fifth three-dimensional data 652 created by the surrounding vehicle into sixth three-dimensional data 654 requested by the host vehicle, encodes the sixth three-dimensional data 654 to generate encoded three-dimensional data 634, and transmits the encoded three-dimensional data 634 to the host vehicle.

[0255] The three-dimensional data transmission device 640 includes a three-dimensional data creation unit 641 , a receiving unit 642 , an extracting unit 643 , an encoding unit 644 , and a transmitting unit 645 .

[0256] First, three-dimensional data creation unit 641 uses sensor information 651 detected by a sensor equipped in a surrounding vehicle to create fifth three-dimensional data 652. Next, receiving unit 642 receives requested range information 633 transmitted from the host vehicle.

[0257] Next, extraction unit 643 extracts three-dimensional data of the requested range indicated by requested range information 633 from fifth three-dimensional data 652, thereby processing fifth three-dimensional data 652 into sixth three-dimensional data 654. Next, encoding unit 644 encodes sixth three-dimensional data 654 to generate encoded three-dimensional data 634, which is an encoded stream. Then, transmission unit 645 transmits encoded three-dimensional data 634 to the host vehicle.

[0258] Here, we will explain an example in which the vehicle itself is equipped with a three-dimensional data creation device 620 and the surrounding vehicles are equipped with a three-dimensional data transmission device 640, but each vehicle may also have the functions of both the three-dimensional data creation device 620 and the three-dimensional data transmission device 640.

[0259] (Fourth embodiment) In this embodiment, an abnormal operation in self-location estimation based on a three-dimensional map will be described.

[0260] It is expected that applications such as self-driving cars, or autonomous movement of mobile objects such as robots, drones, and other flying objects will expand in the future. One example of a means to realize such autonomous movement is a method in which a mobile object estimates its own position within a three-dimensional map (self-location estimation) and travels according to the map.

[0261] Self-position estimation can be achieved by matching a three-dimensional map with three-dimensional information about the surroundings of the vehicle (hereinafter referred to as vehicle-detected three-dimensional data) obtained by sensors such as a rangefinder (such as LiDAR) or a stereo camera mounted on the vehicle, and estimating the vehicle's position within the three-dimensional map.

[0262] 3D maps, such as the HD maps proposed by HERE, may include not only 3D point clouds but also 2D map data such as road and intersection shape information, or real-time changing information such as traffic congestion and accidents. 3D maps are made up of multiple layers, including 3D data, 2D data, and real-time changing metadata, and devices can acquire or reference only the necessary data.

[0263] The point cloud data may be the SWLD described above, or may include point group data other than feature points. Furthermore, the transmission and reception of point cloud data is performed in units of one or more random accesses.

[0264] The following methods can be used to match the 3D map with the vehicle-detected 3D data. For example, the device compares the shapes of the point groups in each point cloud and determines that areas with high similarity between feature points are in the same location. Furthermore, if the 3D map is composed of SWLDs, the device performs matching by comparing the feature points that make up the SWLDs with the 3D feature points extracted from the vehicle-detected 3D data.

[0265] Here, to estimate the vehicle's position with high accuracy, (A) it is necessary to acquire a 3D map and 3D vehicle detection data, and (B) the accuracy of these must meet a predetermined standard. However, in the following abnormal cases, neither (A) nor (B) can be met.

[0266] (1) Three-dimensional maps cannot be obtained via communication.

[0267] (2) The 3D map does not exist, or the 3D map is obtained but is corrupted.

[0268] (3) The vehicle's sensor is out of order or the weather is bad, so the accuracy of the generated 3D data detected by the vehicle is insufficient.

[0269] The operation for dealing with these abnormal cases will be described below. Although the operation will be described below using a car as an example, the following method can be applied to any autonomously moving animal, such as a robot or a drone.

[0270] The configuration and operation of a three-dimensional information processing device according to this embodiment for dealing with abnormal cases in the three-dimensional map or the vehicle-detected three-dimensional data will be described below. Fig. 26 is a block diagram showing an example of the configuration of a three-dimensional information processing device 700 according to this embodiment.

[0271] 26 , the three-dimensional information processing device 700 is mounted on a moving object such as an automobile. The three-dimensional information processing device 700 includes a three-dimensional map acquisition unit 701, a host vehicle detection data acquisition unit 702, an abnormality case determination unit 703, a countermeasure operation determination unit 704, and an operation control unit 705.

[0272] The three-dimensional information processing device 700 may include a two-dimensional or one-dimensional sensor (not shown) for detecting structures or animals around the vehicle, such as a camera for acquiring two-dimensional images or a one-dimensional data sensor using ultrasound or a laser. The three-dimensional information processing device 700 may also include a communication unit (not shown) for acquiring the three-dimensional map via a mobile communication network such as 4G or 5G, or via vehicle-to-vehicle communication or road-to-vehicle communication.

[0273] The three-dimensional map acquisition unit 701 acquires a three-dimensional map 711 of the vicinity of the travel route. For example, the three-dimensional map acquisition unit 701 acquires the three-dimensional map 711 through a mobile communication network, vehicle-to-vehicle communication, or road-to-vehicle communication.

[0274] Next, the host vehicle detection data acquisition unit 702 acquires host vehicle detection three-dimensional data 712 based on the sensor information. For example, the host vehicle detection data acquisition unit 702 generates the host vehicle detection three-dimensional data 712 based on sensor information acquired by a sensor provided in the host vehicle.

[0275] Next, the abnormality case determination unit 703 detects an abnormality case by performing a predetermined check on at least one of the acquired three-dimensional map 711 and the host vehicle detected three-dimensional data 712. In other words, the abnormality case determination unit 703 determines whether at least one of the acquired three-dimensional map 711 and the host vehicle detected three-dimensional data 712 is abnormal.

[0276] When an abnormal case is detected, the countermeasure operation determination unit 704 determines the countermeasure operation for the abnormal case. Next, the operation control unit 705 controls the operation of each processing unit required to implement the countermeasure operation, such as the three-dimensional map acquisition unit 701.

[0277] On the other hand, if no abnormal case is detected, the three-dimensional information processing apparatus 700 ends the process.

[0278] Furthermore, the three-dimensional information processing device 700 uses the three-dimensional map 711 and the vehicle-detected three-dimensional data 712 to estimate the self-position of the vehicle having the three-dimensional information processing device 700. Next, the three-dimensional information processing device 700 automatically drives the vehicle using the result of the self-position estimation.

[0279] In this way, the three-dimensional information processing device 700 acquires map data (three-dimensional map 711) including first three-dimensional position information via a communication path. For example, the first three-dimensional position information is encoded in units of subspaces having three-dimensional coordinate information, and each is a collection of one or more subspaces, and includes multiple random access units that can be independently decoded. For example, the first three-dimensional position information is data (SWLD) in which feature points whose three-dimensional feature amounts are equal to or greater than a predetermined threshold are encoded.

[0280] Furthermore, the three-dimensional information processing device 700 generates second three-dimensional position information (subject vehicle-detected three-dimensional data 712) from the information detected by the sensor. Next, the three-dimensional information processing device 700 performs an abnormality determination process on the first three-dimensional position information or the second three-dimensional position information, thereby determining whether the first three-dimensional position information or the second three-dimensional position information is abnormal.

[0281] When the first three-dimensional position information or the second three-dimensional position information is determined to be abnormal, the three-dimensional information processing apparatus 700 determines a countermeasure action for the abnormality. Next, the three-dimensional information processing apparatus 700 performs control necessary for carrying out the countermeasure action.

[0282] This allows the three-dimensional information processing apparatus 700 to detect an abnormality in the first three-dimensional position information or the second three-dimensional position information and take appropriate action.

[0283] (Embodiment 5) In this embodiment, a method of transmitting three-dimensional data to a following vehicle will be described.

[0284] 27 is a block diagram showing an example of the configuration of a three-dimensional data creation device 810 according to this embodiment. This three-dimensional data creation device 810 is mounted on a vehicle, for example. The three-dimensional data creation device 810 transmits and receives three-dimensional data to and from an external traffic monitoring cloud, a leading vehicle, or a following vehicle, and also creates and stores three-dimensional data.

[0285] The three-dimensional data creation device 810 includes a data receiving unit 811, a communication unit 812, a reception control unit 813, a format conversion unit 814, multiple sensors 815, a three-dimensional data creation unit 816, a three-dimensional data synthesis unit 817, a three-dimensional data storage unit 818, a communication unit 819, a transmission control unit 820, a format conversion unit 821, and a data transmission unit 822.

[0286] The data receiving unit 811 receives three-dimensional data 831 from a traffic monitoring cloud or a preceding vehicle. The three-dimensional data 831 includes, for example, information such as a point cloud, visible light image, depth information, sensor position information, or speed information, including areas that cannot be detected by the sensor 815 of the vehicle itself.

[0287] The communication unit 812 communicates with the traffic monitoring cloud or the vehicle ahead, and transmits data transmission requests and the like to the traffic monitoring cloud or the vehicle ahead.

[0288] The reception control unit 813 exchanges information such as compatible formats with the communication destination via the communication unit 812, and establishes communication with the communication destination.

[0289] The format conversion unit 814 generates three-dimensional data 832 by performing format conversion or the like on the three-dimensional data 831 received by the data receiving unit 811. Furthermore, if the three-dimensional data 831 is compressed or encoded, the format conversion unit 814 performs decompression or decoding processing.

[0290] The multiple sensors 815 are a group of sensors such as LiDAR, a visible light camera, or an infrared camera that acquire information about the outside of the vehicle, and generate sensor information 833. For example, if the sensor 815 is a laser sensor such as LiDAR, the sensor information 833 is three-dimensional data such as a point cloud (point cloud data). Note that the number of sensors 815 does not need to be multiple.

[0291] The three-dimensional data creation unit 816 generates three-dimensional data 834 from the sensor information 833. The three-dimensional data 834 includes information such as a point cloud, a visible light image, depth information, sensor position information, or velocity information.

[0292] The three-dimensional data synthesis unit 817 synthesizes three-dimensional data 834 created based on the host vehicle's sensor information 833 with three-dimensional data 832 created by the traffic monitoring cloud or a preceding vehicle, etc., to construct three-dimensional data 835 that includes the space ahead of the preceding vehicle that cannot be detected by the host vehicle's sensor 815.

[0293] The three-dimensional data storage unit 818 stores the generated three-dimensional data 835 and the like.

[0294] The communication unit 819 communicates with the traffic monitoring cloud or the following vehicle, and transmits data transmission requests and the like to the traffic monitoring cloud or the following vehicle.

[0295] The transmission control unit 820 exchanges information such as supported formats with the communication destination and establishes communication with the communication destination via the communication unit 819. Furthermore, the transmission control unit 820 determines a transmission region, which is the space of the three-dimensional data to be transmitted, based on the three-dimensional data construction information of the three-dimensional data 832 generated by the three-dimensional data synthesis unit 817 and a data transmission request from the communication destination.

[0296] Specifically, in response to a data transmission request from the traffic monitoring cloud or a following vehicle, the transmission control unit 820 determines a transmission area that includes the space ahead of the vehicle that cannot be detected by the sensor of the following vehicle. The transmission control unit 820 also determines the transmission area by determining whether the transmittable space or the transmitted space has been updated based on the three-dimensional data construction information. For example, the transmission control unit 820 determines the area specified in the data transmission request and in which the corresponding three-dimensional data 835 exists as the transmission area. The transmission control unit 820 then notifies the format conversion unit 821 of the format supported by the communication destination and the transmission area.

[0297] The format conversion unit 821 converts three-dimensional data 836 of the transmission area, out of the three-dimensional data 835 stored in the three-dimensional data storage unit 818, into a format supported by the receiving side, thereby generating three-dimensional data 837. Note that the format conversion unit 821 may reduce the amount of data by compressing or encoding the three-dimensional data 837.

[0298] The data transmission unit 822 transmits three-dimensional data 837 to the traffic monitoring cloud or the following vehicle. This three-dimensional data 837 includes, for example, information such as a point cloud ahead of the vehicle, including blind spots of the following vehicle, visible light images, depth information, or sensor position information.

[0299] Although an example in which format conversion is performed by the format conversion units 814 and 821 has been described, format conversion does not necessarily have to be performed.

[0300] With this configuration, the three-dimensional data creation device 810 externally acquires three-dimensional data 831 of an area that cannot be detected by the sensor 815 of the host vehicle, and generates three-dimensional data 835 by combining the three-dimensional data 831 with three-dimensional data 834 based on sensor information 833 detected by the sensor 815 of the host vehicle. In this way, the three-dimensional data creation device 810 can generate three-dimensional data of an area that cannot be detected by the sensor 815 of the host vehicle.

[0301] In addition, in response to a data transmission request from a traffic monitoring cloud or a following vehicle, the three-dimensional data creation device 810 can transmit three-dimensional data including the space in front of the vehicle that cannot be detected by the sensors of the following vehicle to the traffic monitoring cloud or the following vehicle, etc.

[0302] (Embodiment 6) In the fifth embodiment, an example has been described in which a client device such as a vehicle transmits three-dimensional data to another vehicle or a server such as a traffic monitoring cloud. In this embodiment, the client device transmits sensor information obtained by a sensor to the server or another client device.

[0303] First, the configuration of a system according to this embodiment will be described. Fig. 28 is a diagram showing the configuration of a system for transmitting and receiving 3D maps and sensor information according to this embodiment. This system includes a server 901 and client devices 902A and 902B. When there is no need to distinguish between the client devices 902A and 902B, they will also be referred to as client device 902.

[0304] The client device 902 is, for example, an in-vehicle device mounted on a mobile object such as a vehicle. The server 901 is, for example, a traffic monitoring cloud or the like, and is capable of communicating with a plurality of client devices 902.

[0305] The server 901 transmits a three-dimensional map composed of a point cloud to the client device 902. Note that the composition of the three-dimensional map is not limited to a point cloud, and may represent other three-dimensional data such as a mesh structure.

[0306] The client device 902 transmits sensor information acquired by the client device 902 to the server 901. The sensor information includes, for example, at least one of LiDAR acquisition information, a visible light image, an infrared image, a depth image, sensor position information, and speed information.

[0307] Data transmitted between the server 901 and the client device 902 may be compressed to reduce data size, or may remain uncompressed to maintain data accuracy. When compressing data, a three-dimensional compression method based on an octree structure, for example, can be used for point clouds. Also, a two-dimensional image compression method can be used for visible light images, infrared images, and depth images. Examples of two-dimensional image compression methods include MPEG-4 AVC or HEVC standardized by MPEG.

[0308] Furthermore, the server 901 transmits the three-dimensional map managed by the server 901 to the client device 902 in response to a three-dimensional map transmission request from the client device 902. Note that the server 901 may transmit the three-dimensional map without waiting for a three-dimensional map transmission request from the client device 902. For example, the server 901 may broadcast the three-dimensional map to one or more client devices 902 that are in a predetermined space. Furthermore, the server 901 may transmit a three-dimensional map appropriate for the position of the client device 902 to the client device 902 that has received a transmission request once, at regular intervals. Furthermore, the server 901 may transmit the three-dimensional map to the client device 902 every time the three-dimensional map managed by the server 901 is updated.

[0309] The client device 902 issues a request to send a three-dimensional map to the server 901. For example, when the client device 902 wants to estimate its own position while driving, the client device 902 sends a request to send a three-dimensional map to the server 901.

[0310] In the following cases, the client device 902 may issue a request to the server 901 to transmit a three-dimensional map. If the three-dimensional map held by the client device 902 is old, the client device 902 may issue a request to the server 901 to transmit a three-dimensional map. For example, if a certain period of time has passed since the client device 902 obtained the three-dimensional map, the client device 902 may issue a request to the server 901 to transmit a three-dimensional map.

[0311] The client device 902 may issue a request to the server 901 to transmit a three-dimensional map a certain time before the client device 902 leaves the space shown in the three-dimensional map held by the client device 902. For example, when the client device 902 is located within a predetermined distance from the boundary of the space shown in the three-dimensional map held by the client device 902, the client device 902 may issue a request to the server 901 to transmit a three-dimensional map. Furthermore, when the movement route and movement speed of the client device 902 are known, the time when the client device 902 will leave the space shown in the three-dimensional map held by the client device 902 may be predicted based on these.

[0312] If the error in aligning the three-dimensional data created by the client device 902 from sensor information with the three-dimensional map is equal to or greater than a certain level, the client device 902 may issue a request to the server 901 to send the three-dimensional map.

[0313] The client device 902 transmits sensor information to the server 901 in response to a request to transmit sensor information transmitted from the server 901. Note that the client device 902 may transmit sensor information to the server 901 without waiting for a request to transmit sensor information from the server 901. For example, once the client device 902 receives a request to transmit sensor information from the server 901, the client device 902 may periodically transmit the sensor information to the server 901 for a certain period of time. Furthermore, if the error in aligning the three-dimensional data created by the client device 902 based on the sensor information with the three-dimensional map obtained from the server 901 is equal to or greater than a certain level, the client device 902 may determine that a change may have occurred in the three-dimensional map around the client device 902, and may transmit this information along with the sensor information to the server 901.

[0314] The server 901 issues a request to transmit sensor information to the client device 902. For example, the server 901 receives location information of the client device 902, such as GPS, from the client device 902. When the server 901 determines, based on the location information of the client device 902, that the client device 902 is approaching a space with little information in the three-dimensional map managed by the server 901, the server 901 issues a request to transmit sensor information to the client device 902 in order to generate a new three-dimensional map. The server 901 may also issue a request to transmit sensor information when it wants to update the three-dimensional map, when it wants to check road conditions during snowfall or a disaster, when it wants to check traffic congestion, or when it wants to check the status of incidents and accidents, etc.

[0315] Furthermore, the client device 902 may set the amount of data of the sensor information to be transmitted to the server 901 depending on the communication state or bandwidth at the time of receiving a request to transmit the sensor information from the server 901. Setting the amount of data of the sensor information to be transmitted to the server 901 means, for example, increasing or decreasing the amount of the data itself or appropriately selecting a compression method.

[0316] 29 is a block diagram showing an example configuration of a client device 902. The client device 902 receives a three-dimensional map composed of a point cloud or the like from the server 901, and estimates the self-position of the client device 902 from three-dimensional data created based on sensor information of the client device 902. The client device 902 also transmits the acquired sensor information to the server 901.

[0317] The client device 902 includes a data receiving unit 1011, a communication unit 1012, a reception control unit 1013, a format conversion unit 1014, multiple sensors 1015, a three-dimensional data creation unit 1016, a three-dimensional image processing unit 1017, a three-dimensional data storage unit 1018, a format conversion unit 1019, a communication unit 1020, a transmission control unit 1021, and a data transmission unit 1022.

[0318] The data receiving unit 1011 receives a three-dimensional map 1031 from the server 901. The three-dimensional map 1031 is data including a point cloud such as a WLD or SWLD. The three-dimensional map 1031 may include either compressed data or uncompressed data.

[0319] The communication unit 1012 communicates with the server 901 and transmits data transmission requests (for example, requests to transmit a three-dimensional map) to the server 901.

[0320] The reception control unit 1013 exchanges information such as compatible formats with the communication destination via the communication unit 1012, and establishes communication with the communication destination.

[0321] The format conversion unit 1014 generates a three-dimensional map 1032 by performing format conversion and the like on the three-dimensional map 1031 received by the data receiving unit 1011. Furthermore, if the three-dimensional map 1031 is compressed or encoded, the format conversion unit 1014 performs decompression or decoding processing. Note that if the three-dimensional map 1031 is uncompressed data, the format conversion unit 1014 does not perform decompression or decoding processing.

[0322] The multiple sensors 1015 are a group of sensors, such as LiDAR, a visible light camera, an infrared camera, or a depth sensor, that acquire information about the outside of the vehicle on which the client device 902 is mounted, and generate sensor information 1033. For example, if the sensor 1015 is a laser sensor such as LiDAR, the sensor information 1033 is three-dimensional data such as a point cloud (point cloud data). Note that the number of sensors 1015 does not need to be multiple.

[0323] The three-dimensional data creation unit 1016 creates three-dimensional data 1034 of the surroundings of the vehicle based on the sensor information 1033. For example, the three-dimensional data creation unit 1016 creates point cloud data with color information of the surroundings of the vehicle using information acquired by LiDAR and visible light images acquired by a visible light camera.

[0324] The three-dimensional image processing unit 1017 performs a process of estimating the vehicle's own position using a three-dimensional map 1032 such as a received point cloud and three-dimensional data 1034 of the surroundings of the vehicle generated from the sensor information 1033. The three-dimensional image processing unit 1017 may generate three-dimensional data 1035 of the surroundings of the vehicle by combining the three-dimensional map 1032 and the three-dimensional data 1034, and perform a process of estimating the vehicle's own position using the generated three-dimensional data 1035.

[0325] The three-dimensional data storage unit 1018 stores a three-dimensional map 1032, three-dimensional data 1034, three-dimensional data 1035, and the like.

[0326] The format conversion unit 1019 generates sensor information 1037 by converting the sensor information 1033 into a format supported by the receiving side. The format conversion unit 1019 may reduce the amount of data by compressing or encoding the sensor information 1037. The format conversion unit 1019 may omit the process if format conversion is not necessary. The format conversion unit 1019 may also control the amount of data to be transmitted in accordance with a specified transmission range.

[0327] The communication unit 1020 communicates with the server 901 and receives data transmission requests (sensor information transmission requests) and the like from the server 901.

[0328] The transmission control unit 1021 exchanges information such as compatible formats with the communication destination via the communication unit 1020, and establishes communication.

[0329] The data transmission unit 1022 transmits the sensor information 1037 to the server 901. The sensor information 1037 includes information acquired by a plurality of sensors 1015, such as information acquired by a LiDAR, a luminance image acquired by a visible light camera, an infrared image acquired by an infrared camera, a depth image acquired by a depth sensor, sensor position information, and speed information.

[0330] Next, the configuration of the server 901 will be described. Fig. 30 is a block diagram showing an example of the configuration of the server 901. The server 901 receives sensor information transmitted from the client device 902 and creates three-dimensional data based on the received sensor information. The server 901 uses the created three-dimensional data to update the three-dimensional map managed by the server 901. In addition, in response to a request from the client device 902 to transmit the three-dimensional map, the server 901 transmits the updated three-dimensional map to the client device 902.

[0331] The server 901 includes a data receiving unit 1111, a communication unit 1112, a receiving control unit 1113, a format conversion unit 1114, a three-dimensional data creation unit 1116, a three-dimensional data synthesis unit 1117, a three-dimensional data storage unit 1118, a format conversion unit 1119, a communication unit 1120, a transmission control unit 1121, and a data transmission unit 1122.

[0332] The data receiving unit 1111 receives sensor information 1037 from the client device 902. The sensor information 1037 includes, for example, information acquired by a LiDAR, a luminance image acquired by a visible light camera, an infrared image acquired by an infrared camera, a depth image acquired by a depth sensor, sensor position information, and speed information.

[0333] The communication unit 1112 communicates with the client device 902 and transmits a data transmission request (for example, a request to transmit sensor information) to the client device 902 .

[0334] The reception control unit 1113 exchanges information such as compatible formats with the communication destination via the communication unit 1112, and establishes communication.

[0335] If the received sensor information 1037 is compressed or encoded, the format conversion unit 1114 performs decompression or decoding processing to generate the sensor information 1132. Note that if the sensor information 1037 is uncompressed data, the format conversion unit 1114 does not perform decompression or decoding processing.

[0336] The three-dimensional data creation unit 1116 creates three-dimensional data 1134 of the periphery of the client device 902 based on the sensor information 1132. For example, the three-dimensional data creation unit 1116 creates point cloud data with color information of the periphery of the client device 902 using information acquired by the LiDAR and visible light images acquired by the visible light camera.

[0337] A three-dimensional data synthesis unit 1117 synthesizes three-dimensional data 1134 created based on sensor information 1132 with a three-dimensional map 1135 managed by the server 901, thereby updating the three-dimensional map 1135.

[0338] The three-dimensional data storage unit 1118 stores a three-dimensional map 1135 and the like.

[0339] The format conversion unit 1119 generates the three-dimensional map 1031 by converting the three-dimensional map 1135 into a format supported by the receiving side. The format conversion unit 1119 may reduce the amount of data by compressing or encoding the three-dimensional map 1135. The format conversion unit 1119 may also omit processing if format conversion is not necessary. The format conversion unit 1119 may also control the amount of data to be transmitted in accordance with the designation of the transmission range.

[0340] The communication unit 1120 communicates with the client device 902 and receives a data transmission request (a request to transmit a three-dimensional map) or the like from the client device 902 .

[0341] The transmission control unit 1121 exchanges information such as compatible formats with the communication destination via the communication unit 1120, and establishes communication.

[0342] The data transmission unit 1122 transmits the three-dimensional map 1031 to the client device 902. The three-dimensional map 1031 is data including a point cloud such as a WLD or SWLD. The three-dimensional map 1031 may include either compressed data or uncompressed data.

[0343] Next, a description will be given of the operational flow of the client device 902. Fig. 31 is a flowchart showing the operation of the client device 902 when acquiring a three-dimensional map.

[0344] First, the client device 902 requests the server 901 to transmit a three-dimensional map (such as a point cloud) (S1001). At this time, the client device 902 may also transmit location information of the client device 902 obtained by GPS or the like, thereby requesting the server 901 to transmit a three-dimensional map related to the location information.

[0345] Next, the client device 902 receives the three-dimensional map from the server 901 (S1002). If the received three-dimensional map is compressed data, the client device 902 decodes the received three-dimensional map to generate an uncompressed three-dimensional map (S1003).

[0346] Next, the client device 902 creates three-dimensional data 1034 of the surroundings of the client device 902 from sensor information 1033 obtained by the multiple sensors 1015 (S1004). Next, the client device 902 estimates its own position using the three-dimensional map 1032 received from the server 901 and the three-dimensional data 1034 created from the sensor information 1033 (S1005).

[0347] 32 is a flowchart showing the operation of the client device 902 when transmitting sensor information. First, the client device 902 receives a request to transmit sensor information from the server 901 (S1011). Upon receiving the transmission request, the client device 902 transmits sensor information 1037 to the server 901 (S1012). Note that, when the sensor information 1033 includes multiple pieces of information obtained by multiple sensors 1015, the client device 902 may generate the sensor information 1037 by compressing each piece of information using a compression method suitable for that piece of information.

[0348] Next, the operation flow of the server 901 will be described. Fig. 33 is a flowchart showing the operation when the server 901 acquires sensor information. First, the server 901 requests the client device 902 to transmit sensor information (S1021). Next, the server 901 receives the sensor information 1037 transmitted from the client device 902 in response to the request (S1022). Next, the server 901 creates three-dimensional data 1134 using the received sensor information 1037 (S1023). Next, the server 901 reflects the created three-dimensional data 1134 in the three-dimensional map 1135 (S1024).

[0349] 34 is a flowchart showing the operation of the server 901 when transmitting a three-dimensional map. First, the server 901 receives a request to transmit a three-dimensional map from the client device 902 (S1031). Having received the request to transmit the three-dimensional map, the server 901 transmits a three-dimensional map 1031 to the client device 902 (S1032). At this time, the server 901 may extract a three-dimensional map of the vicinity based on the location information of the client device 902 and transmit the extracted three-dimensional map. The server 901 may also compress the three-dimensional map made up of a point cloud using, for example, a compression method with an octree structure, and transmit the compressed three-dimensional map.

[0350] A modification of this embodiment will now be described.

[0351] The server 901 uses the sensor information 1037 received from the client device 902 to create three-dimensional data 1134 of the vicinity of the position of the client device 902. Next, the server 901 matches the created three-dimensional data 1134 with a three-dimensional map 1135 of the same area managed by the server 901, thereby calculating the difference between the three-dimensional data 1134 and the three-dimensional map 1135. If the difference is equal to or greater than a predetermined threshold, the server 901 determines that some abnormality has occurred in the vicinity of the client device 902. For example, when ground subsidence occurs due to a natural disaster such as an earthquake, a large difference may occur between the three-dimensional map 1135 managed by the server 901 and the three-dimensional data 1134 created based on the sensor information 1037.

[0352] The sensor information 1037 may include information indicating at least one of the sensor type, sensor performance, and sensor model number. A class ID or the like according to the sensor performance may also be added to the sensor information 1037. For example, if the sensor information 1037 is information acquired by a LiDAR, an identifier may be assigned to the sensor performance, such as Class 1 for a sensor capable of acquiring information with an accuracy of several millimeters, Class 2 for a sensor capable of acquiring information with an accuracy of several centimeters, and Class 3 for a sensor capable of acquiring information with an accuracy of several meters. The server 901 may also estimate the sensor performance information and the like from the model number of the client device 902. For example, if the client device 902 is installed in a vehicle, the server 901 may determine the sensor specification information from the vehicle model. In this case, the server 901 may acquire vehicle model information in advance, or the information may be included in the sensor information. The server 901 may also use the acquired sensor information 1037 to switch the degree of correction for the three-dimensional data 1134 created using the sensor information 1037. For example, if the sensor performance is high accuracy (class 1), the server 901 does not perform correction on the three-dimensional data 1134. If the sensor performance is low accuracy (class 3), the server 901 applies correction according to the accuracy of the sensor to the three-dimensional data 1134. For example, the server 901 increases the degree (strength) of correction as the accuracy of the sensor decreases.

[0353] The server 901 may simultaneously issue requests to send sensor information to multiple client devices 902 in a certain space. When the server 901 receives multiple pieces of sensor information from multiple client devices 902, the server 901 does not need to use all of the sensor information to create the three-dimensional data 1134, and may select the sensor information to use, for example, depending on the performance of the sensor. For example, when updating the three-dimensional map 1135, the server 901 may select high-precision sensor information (Class 1) from the multiple pieces of sensor information it has received, and create the three-dimensional data 1134 using the selected sensor information.

[0354] The server 901 is not limited to a server such as a traffic monitoring cloud, but may be another client device (mounted in a vehicle). Figure 35 is a diagram showing the system configuration in this case.

[0355] For example, client device 902C issues a request to transmit sensor information to nearby client device 902A and acquires the sensor information from client device 902A. Client device 902C then creates three-dimensional data using the acquired sensor information from client device 902A and updates the three-dimensional map of client device 902C. This allows client device 902C to generate a three-dimensional map of the space that can be acquired from client device 902A by taking advantage of the performance of client device 902C. For example, such a case is likely to occur when client device 902C has high performance.

[0356] In this case, the client device 902A that provided the sensor information is granted the right to obtain the high-precision 3D map generated by the client device 902C. The client device 902A receives the high-precision 3D map from the client device 902C in accordance with the right.

[0357] In addition, client device 902C may issue requests to send sensor information to multiple nearby client devices 902 (client device 902A and client device 902B). If the sensor of client device 902A or client device 902B is high performance, client device 902C can create three-dimensional data using the sensor information obtained by this high performance sensor.

[0358] 36 is a block diagram showing the functional configuration of the server 901 and the client device 902. The server 901 includes, for example, a 3D map compression / decoding processing unit 1201 that compresses and decodes 3D maps, and a sensor information compression / decoding processing unit 1202 that compresses and decodes sensor information.

[0359] The client device 902 includes a three-dimensional map decoding processor 1211 and a sensor information compression processor 1212. The three-dimensional map decoding processor 1211 receives encoded data of the compressed three-dimensional map and decodes the encoded data to acquire the three-dimensional map. The sensor information compression processor 1212 compresses the sensor information itself instead of three-dimensional data created from the acquired sensor information, and transmits the encoded data of the compressed sensor information to the server 901. With this configuration, the client device 902 only needs to internally include a processing unit (device or LSI) that performs processing to decode the three-dimensional map (point cloud, etc.), and does not need to internally include a processing unit that performs processing to compress the three-dimensional data of the three-dimensional map (point cloud, etc.). This allows the cost and power consumption of the client device 902 to be reduced.

[0360] As described above, the client device 902 according to this embodiment is mounted on a mobile body and generates three-dimensional data 1034 of the surroundings of the mobile body from sensor information 1033 indicating the surrounding conditions of the mobile body, which is obtained by the sensor 1015 mounted on the mobile body. The client device 902 estimates the self-position of the mobile body using the generated three-dimensional data 1034. The client device 902 transmits the acquired sensor information 1033 to the server 901 or another mobile body 902.

[0361] According to this, the client device 902 transmits the sensor information 1033 to the server 901 or the like. This may reduce the amount of data to be transmitted compared to when transmitting three-dimensional data. Furthermore, since the client device 902 does not need to perform processing such as compression or encoding of the three-dimensional data, the amount of processing by the client device 902 can be reduced. Therefore, the client device 902 can reduce the amount of data to be transmitted or simplify the device configuration.

[0362] Furthermore, the client device 902 further transmits a request to the server 901 to send a three-dimensional map, and receives a three-dimensional map 1031 from the server 901. The client device 902 estimates its own location using the three-dimensional data 1034 and the three-dimensional map 1032.

[0363] The sensor information 1033 includes at least one of information obtained by a laser sensor, a luminance image, an infrared image, a depth image, sensor position information, and sensor speed information.

[0364] The sensor information 1033 also includes information indicating the performance of the sensor.

[0365] Furthermore, the client device 902 encodes or compresses the sensor information 1033, and transmits the encoded or compressed sensor information 1037 to the server 901 or another mobile body 902. This allows the client device 902 to reduce the amount of data to be transmitted.

[0366] For example, the client device 902 includes a processor and a memory, and the processor uses the memory to perform the above-described processing.

[0367] Furthermore, server 901 according to this embodiment is capable of communicating with client device 902 mounted on the mobile object, and receives sensor information 1037 indicating the surrounding conditions of the mobile object, obtained by sensor 1015 mounted on the mobile object, from client device 902. Server 901 creates three-dimensional data 1134 of the surroundings of the mobile object from the received sensor information 1037.

[0368] According to this, the server 901 creates three-dimensional data 1134 using the sensor information 1037 transmitted from the client device 902. This may reduce the amount of data to be transmitted compared to when the client device 902 transmits the three-dimensional data. Furthermore, since the client device 902 does not need to perform processing such as compression or encoding of the three-dimensional data, the amount of processing by the client device 902 can be reduced. Therefore, the server 901 can reduce the amount of data to be transmitted or simplify the device configuration.

[0369] Furthermore, the server 901 further transmits a request to the client device 902 to transmit the sensor information.

[0370] The server 901 also updates a three-dimensional map 1135 using the created three-dimensional data 1134 and transmits the three-dimensional map 1135 to the client device 902 in response to a request from the client device 902 to transmit the three-dimensional map 1135.

[0371] The sensor information 1037 includes at least one of information obtained by a laser sensor, a luminance image, an infrared image, a depth image, sensor position information, and sensor speed information.

[0372] The sensor information 1037 also includes information indicating the performance of the sensor.

[0373] Furthermore, the server 901 further corrects the three-dimensional data in accordance with the performance of the sensor, thereby enabling the three-dimensional data creation method to improve the quality of the three-dimensional data.

[0374] Furthermore, when receiving sensor information, the server 901 receives a plurality of pieces of sensor information 1037 from a plurality of client devices 902, and selects the sensor information 1037 to be used to create the three-dimensional data 1134 based on a plurality of pieces of information indicating the performance of the sensors included in the plurality of pieces of sensor information 1037. This allows the server 901 to improve the quality of the three-dimensional data 1134.

[0375] Furthermore, the server 901 decodes or decompresses the received sensor information 1037, and creates three-dimensional data 1134 from the decoded or decompressed sensor information 1132. This allows the server 901 to reduce the amount of data to be transmitted.

[0376] For example, the server 901 includes a processor and a memory, and the processor uses the memory to perform the above-mentioned processing.

[0377] (Embodiment 7) In this embodiment, a method for encoding and decoding three-dimensional data using inter prediction processing will be described.

[0378] 37 is a block diagram of a three-dimensional data encoding device 1300 according to this embodiment. This three-dimensional data encoding device 1300 generates an encoded bitstream (hereinafter also simply referred to as a bitstream), which is an encoded signal, by encoding three-dimensional data. As shown in FIG. 37, the three-dimensional data encoding device 1300 includes a dividing unit 1301, a subtraction unit 1302, a transformation unit 1303, a quantization unit 1304, an inverse quantization unit 1305, an inverse transformation unit 1306, an addition unit 1307, a reference volume memory 1308, an intra prediction unit 1309, a reference space memory 1310, an inter prediction unit 1311, a prediction control unit 1312, and an entropy encoding unit 1313.

[0379] The dividing unit 1301 divides each space (SPC) included in the three-dimensional data into multiple volumes (VLM), which are encoding units. The dividing unit 1301 also converts the voxels in each volume into an octree representation. The dividing unit 1301 may make the spaces and volumes the same size and convert the spaces into an octree representation. The dividing unit 1301 may also add information required for the octree representation (depth information, etc.) to a bitstream header, etc.

[0380] The subtraction unit 1302 calculates the difference between the volume (volume to be coded) output from the division unit 1301 and a prediction volume generated by intra prediction or inter prediction, which will be described later, and outputs the calculated difference as a prediction residual to the conversion unit 1303. Fig. 38 is a diagram showing an example of how a prediction residual is calculated. Note that the bit strings of the volume to be coded and the prediction volume shown here are, for example, position information indicating the positions of three-dimensional points (e.g., a point cloud) included in the volume.

[0381] The octree representation and the voxel scanning order will be explained below. A volume is converted into an octree structure (octreeing) and then encoded. The octree structure consists of nodes and leaves. Each node has eight nodes or leaves, and each leaf has voxel (VXL) information. Fig. 39 is a diagram showing an example of the structure of a volume containing multiple voxels. Fig. 40 is a diagram showing an example of the volume shown in Fig. 39 converted into an octree structure. Here, among the leaves shown in Fig. 40, leaves 1, 2, and 3 represent voxels VXL1, VXL2, and VXL3 shown in Fig. 39, respectively, and represent a VXL containing a point cloud (hereinafter referred to as an effective VXL).

[0382] An octree is represented by a binary sequence of, for example, 0 and 1. For example, if a node or a valid VXL is set to value 1 and the rest to value 0, then the binary sequence shown in FIG. 40 is assigned to each node and leaf. This binary sequence is then scanned according to the scan order, either breadth-first or depth-first. For example, when scanned breadth-first, the binary sequence shown in A of FIG. 41 is obtained. When scanned depth-first, the binary sequence shown in B of FIG. 41 is obtained. The binary sequence obtained by this scan is then coded by entropy coding to reduce the amount of information.

[0383] Next, we will explain depth information in octree representation. Depth in octree representation is used to control the granularity of the point cloud information contained in the volume to be retained. Setting a larger depth allows the point cloud information to be reproduced at a finer level, but the amount of data required to represent nodes and leaves increases. Conversely, setting a smaller depth reduces the amount of data, but since multiple point cloud information with different positions and colors is considered to be in the same position and with the same color, the information contained in the original point cloud information will be lost.

[0384] For example, FIG. 42 is a diagram showing an example in which the octree with depth=2 shown in FIG. 40 is represented by an octree with depth=1. The octree shown in FIG. 42 has a smaller amount of data than the octree shown in FIG. 40. In other words, the octree shown in FIG. 42 has a smaller number of bits after binarization than the octree shown in FIG. 42. Here, leaf 1 and leaf 2 shown in FIG. 40 are represented by leaf 1 shown in FIG. 41. In other words, the information that leaf 1 and leaf 2 shown in FIG. 40 were in different positions is lost.

[0385] FIG. 43 is a diagram showing volumes corresponding to the octree shown in FIG. 42. VXL1 and VXL2 shown in FIG. 39 correspond to VXL12 shown in FIG. 43. In this case, the three-dimensional data encoding device 1300 generates color information for VXL12 shown in FIG. 43 from the color information for VXL1 and VXL2 shown in FIG. 39. For example, the three-dimensional data encoding device 1300 calculates the average value, median value, or weighted average value of the color information for VXL1 and VXL2 as the color information for VXL12. In this way, the three-dimensional data encoding device 1300 may control the reduction of data amount by changing the depth of the octree.

[0386] The three-dimensional data encoding device 1300 may set the depth information of the octree in any unit of world, space, or volume. In this case, the three-dimensional data encoding device 1300 may add the depth information to the header information of the world, the header information of the space, or the header information of the volume. Furthermore, the same value may be used as the depth information for all worlds, spaces, and volumes at different times. In this case, the three-dimensional data encoding device 1300 may add the depth information to the header information that manages the worlds of all times.

[0387] If the voxels contain color information, the transform unit 1303 applies a frequency transform, such as an orthogonal transform, to the prediction residuals of the color information of the voxels in the volume. For example, the transform unit 1303 creates a one-dimensional array by scanning the prediction residuals in a certain scan order. The transform unit 1303 then applies a one-dimensional orthogonal transform to the created one-dimensional array to convert it into the frequency domain. As a result, when the values ​​of the prediction residuals in the volume are close, the values ​​of the low-frequency components become larger and the values ​​of the high-frequency components become smaller. This allows the quantization unit 1304 to reduce the amount of code more efficiently.

[0388] Furthermore, the transform unit 1303 may use a two- or more-dimensional orthogonal transform instead of a one-dimensional one. For example, the transform unit 1303 maps prediction residuals to a two-dimensional array in a certain scan order and applies a two-dimensional orthogonal transform to the obtained two-dimensional array. The transform unit 1303 may also select an orthogonal transform method to use from a plurality of orthogonal transform methods. In this case, the 3D data encoding device 1300 adds information indicating which orthogonal transform method was used to the bitstream. The transform unit 1303 may also select an orthogonal transform method to use from a plurality of orthogonal transform methods of different dimensions. In this case, the 3D data encoding device 1300 adds information indicating which orthogonal transform method was used to the bitstream.

[0389] For example, the conversion unit 1303 matches the scan order of the prediction residuals to the scan order (breadth-first or depth-first, etc.) of the octet tree in the volume. This eliminates the need to add information indicating the scan order of the prediction residuals to the bitstream, thereby reducing overhead. The conversion unit 1303 may also apply a scan order different from the scan order of the octet tree. In this case, the 3D data encoding device 1300 adds information indicating the scan order of the prediction residuals to the bitstream. This allows the 3D data encoding device 1300 to efficiently encode the prediction residuals. The 3D data encoding device 1300 may also add information (such as a flag) indicating whether or not to apply the octet scan order to the bitstream, and add information indicating the scan order of the prediction residuals to the bitstream when the octet scan order is not applied.

[0390] The conversion unit 1303 may convert not only the prediction residual of the color information but also other attribute information of the voxels. For example, the conversion unit 1303 may convert and encode information such as reflectance obtained when the point cloud is acquired by LiDAR or the like.

[0391] If the space does not have attribute information such as color information, the conversion unit 1303 may skip the process. Furthermore, the three-dimensional data encoding device 1300 may add information (a flag) indicating whether or not the process of the conversion unit 1303 is to be skipped to the bitstream.

[0392] The quantization unit 1304 generates quantized coefficients by quantizing the frequency components of the prediction residual generated by the transform unit 1303 using the quantization control parameters. This reduces the amount of information. The generated quantized coefficients are output to the entropy coding unit 1313. The quantization unit 1304 may control the quantization control parameters on a world-by-world, space-by-space, or volume-by-volume basis. In this case, the three-dimensional data coding device 1300 adds the quantization control parameters to the respective header information, etc. The quantization unit 1304 may also control quantization by changing the weight for each frequency component of the prediction residual. For example, the quantization unit 1304 may finely quantize low-frequency components and coarsely quantize high-frequency components. In this case, the three-dimensional data coding device 1300 may add a parameter indicating the weight of each frequency component to the header.

[0393] If the space does not have attribute information such as color information, the quantization unit 1304 may skip the process. Furthermore, the three-dimensional data encoding device 1300 may add information (a flag) indicating whether or not the process of the quantization unit 1304 is to be skipped to the bitstream.

[0394] The inverse quantization unit 1305 uses the quantization control parameter to inverse quantize the quantized coefficients generated by the quantization unit 1304 to generate inverse quantized coefficients of the prediction residuals, and outputs the generated inverse quantized coefficients to the inverse transform unit 1306.

[0395] The inverse transform unit 1306 generates a post-inverse transform prediction residual by applying inverse transform to the inverse quantized coefficients generated by the inverse quantization unit 1305. This post-inverse transform prediction residual is a prediction residual generated after quantization, and therefore does not need to completely match the prediction residual output by the transform unit 1303.

[0396] The adder 1307 generates a reconstructed volume by adding the prediction residual after inverse transform applied, generated by the inverse transformer 1306, and a prediction volume generated by intra prediction or inter prediction, which will be described later, and used to generate the prediction residual before quantization. This reconstructed volume is stored in a reference volume memory 1308 or a reference space memory 1310.

[0397] The intra prediction unit 1309 generates a predicted volume of the volume to be encoded using attribute information of the adjacent volume stored in the reference volume memory 1308. The attribute information includes color information or reflectance of voxels. The intra prediction unit 1309 generates a predicted value of the color information or reflectance of the volume to be encoded.

[0398] FIG. 44 is a diagram illustrating the operation of the intra prediction unit 1309. For example, the intra prediction unit 1309 generates a prediction volume of a volume to be coded (volume idx=3) shown in FIG. 44 from an adjacent volume (volume idx=0). Here, volume idx is identifier information assigned to volumes in a space, and a different value is assigned to each volume. The order in which the volumes idx are assigned may be the same as the encoding order, or may be different from the encoding order. For example, the intra prediction unit 1309 uses the average value of color information of voxels included in volume idx=0, which is an adjacent volume, as a prediction value of color information of the volume to be coded shown in FIG. 44. In this case, a prediction residual is generated by subtracting the prediction value of color information from the color information of each voxel included in the volume to be coded. The processing of the conversion unit 1303 and subsequent processes is performed on this prediction residual. In this case, the 3D data encoding device 1300 also adds adjacent volume information and prediction mode information to the bitstream. Here, the adjacent volume information is information indicating the adjacent volume used for prediction, for example, the volume idx of the adjacent volume used for prediction. Also, the prediction mode information indicates the mode used to generate the predicted volume. The mode is, for example, an average mode that generates a predicted value from the average value of the voxels in the adjacent volume, or an intermediate value mode that generates a predicted value from the intermediate value of the voxels in the adjacent volume.

[0399] The intra prediction unit 1309 may generate a prediction volume from multiple adjacent volumes. For example, in the configuration shown in Fig. 44 , the intra prediction unit 1309 generates prediction volume 0 from the volume with volume idx=0, and generates prediction volume 1 from the volume with volume idx=1. The intra prediction unit 1309 then generates the average of prediction volume 0 and prediction volume 1 as the final prediction volume. In this case, the three-dimensional data encoding device 1300 may add multiple volume idx of the multiple volumes used to generate the prediction volume to the bitstream.

[0400] 45 is a diagram schematically illustrating inter prediction processing according to this embodiment. The inter prediction unit 1311 encodes (inter predicts) a space (SPC) at a certain time T_Cur using an encoded space at a different time T_LX. In this case, the inter prediction unit 1311 performs encoding processing by applying rotation and translation processing to the encoded space at the different time T_LX.

[0401] Furthermore, the three-dimensional data encoding device 1300 adds RT information related to the rotation and translation processing applied to the space at a different time T_LX to the bitstream. The different time T_LX is, for example, time T_L0, which is before the certain time T_Cur. In this case, the three-dimensional data encoding device 1300 may add RT information RT_L0 related to the rotation and translation processing applied to the space at time T_L0 to the bitstream.

[0402] Alternatively, the different time T_LX may be, for example, time T_L1 that is later than the certain time T_Cur. In this case, the three-dimensional data encoding device 1300 may add RT information RT_L1 related to the rotation and translation processing applied to the space of time T_L1 to the bitstream.

[0403] Alternatively, the inter prediction unit 1311 performs encoding (bi-prediction) by referencing both spaces at different times T_L0 and T_L1. In this case, the 3D data encoding device 1300 may add both pieces of RT information RT_L0 and RT_L1 related to the rotation and translation applied to the respective spaces to the bitstream.

[0404] In the above, T_L0 is a time before T_Cur and T_L1 is a time after T_Cur, but this is not necessarily limited to this. For example, T_L0 and T_L1 may both be times before T_Cur. Alternatively, T_L0 and T_L1 may both be times after T_Cur.

[0405] Furthermore, when the three-dimensional data encoding device 1300 performs encoding by referencing multiple spaces at different times, it may add RT information related to the rotation and translation applied to each space to the bitstream. For example, the three-dimensional data encoding device 1300 manages multiple referenced encoded spaces using two reference lists (an L0 list and an L1 list). If the first reference space in the L0 list is L0R0, the second reference space in the L0 list is L0R1, the first reference space in the L1 list is L1R0, and the second reference space in the L1 list is L1R1, the three-dimensional data encoding device 1300 adds RT information RT_L0R0 of L0R0, RT information RT_L0R1 of L0R1, RT information RT_L1R0 of L1R0, and RT information RT_L1R1 of L1R1 to the bitstream. For example, the three-dimensional data encoding device 1300 adds this RT information to a header or the like of the bitstream.

[0406] Furthermore, when the three-dimensional data encoding device 1300 performs encoding by referring to reference spaces at multiple different times, it determines whether rotation and translation are applied for each reference space. In this case, the three-dimensional data encoding device 1300 may add information (such as an RT application flag) indicating whether rotation and translation are applied for each reference space to header information or the like of the bitstream. For example, the three-dimensional data encoding device 1300 calculates RT information and an ICP error value using an ICP (Interactive Closest Point) algorithm for each reference space referenced from the encoding target space. If the ICP error value is equal to or less than a predetermined value, the three-dimensional data encoding device 1300 determines that rotation and translation are not necessary and sets the RT application flag to OFF. On the other hand, if the ICP error value is greater than the predetermined value, the three-dimensional data encoding device 1300 sets the RT application flag to ON and adds RT information to the bitstream.

[0407] 46 is a diagram showing an example of syntax for adding RT information and an RT application flag to a header. The number of bits allocated to each syntax element may be determined within the range that the syntax element can take. For example, if the number of reference spaces included in the reference list L0 is eight, three bits may be allocated to MaxRefSpc_l0. The number of allocated bits may be variable depending on the value that each syntax element can take, or may be fixed regardless of the value that each syntax element can take. When the number of allocated bits is fixed, the three-dimensional data encoding device 1300 may add the fixed number of bits to other header information.

[0408] Here, MaxRefSpc_l0 shown in Figure 46 indicates the number of reference spaces included in the reference list L0. RT_flag_l0[i] is the RT application flag for reference space i in reference list L0. When RT_flag_l0[i] is 1, rotation and translation are applied to reference space i. When RT_flag_l0[i] is 0, rotation and translation are not applied to reference space i.

[0409] R_l0[i] and T_l0[i] are RT information of reference space i in reference list L0. R_l0[i] is rotation information of reference space i in reference list L0. The rotation information indicates the content of the applied rotation process, such as a rotation matrix or a quaternion. T_l0[i] is translation information of reference space i in reference list L0. The translation information indicates the content of the applied translation process, such as a translation vector.

[0410] MaxRefSpc_l1 indicates the number of reference spaces included in the reference list L1. RT_flag_l1[i] is the RT application flag for reference space i in reference list L1. If RT_flag_l1[i] is 1, rotation and translation are applied to reference space i. If RT_flag_l1[i] is 0, rotation and translation are not applied to reference space i.

[0411] R_l1[i] and T_l1[i] are the RT information of the reference space i in the reference list L1. R_l1[i] is the rotation information of the reference space i in the reference list L1. The rotation information indicates the content of the applied rotation process, such as a rotation matrix or a quaternion. T_l1[i] is the translation information of the reference space i in the reference list L1. The translation information indicates the content of the applied translation process, such as a translation vector.

[0412] The inter prediction unit 1311 generates a predicted volume of the volume to be coded using information about the coded reference space stored in the reference space memory 1310. As described above, before generating a predicted volume of the volume to be coded, the inter prediction unit 1311 obtains RT information for the space to be coded and the reference space using an ICP (Interactive Closest Point) algorithm to approximate the overall positional relationship between the space to be coded and the reference space. The inter prediction unit 1311 then obtains reference space B by applying rotation and translation processing to the reference space using the obtained RT information. The inter prediction unit 1311 then generates a predicted volume of the volume to be coded in the space to be coded using information in reference space B. Here, the three-dimensional data coding device 1300 adds the RT information used to obtain reference space B to header information, etc., of the space to be coded.

[0413] In this way, the inter prediction unit 1311 can improve the accuracy of the predicted volume by applying rotation and translation processing to the reference space to bring the overall positional relationship between the encoding target space and the reference space closer together, and then generating a predicted volume using information about the reference space. Furthermore, since prediction residuals can be suppressed, the amount of coding can be reduced. Note that, while an example of performing ICP using the encoding target space and the reference space has been shown here, this is not necessarily limited to this. For example, in order to reduce the amount of processing, the inter prediction unit 1311 may obtain RT information by performing ICP using at least one of the encoding target space in which the number of voxels or point clouds has been thinned and the reference space in which the number of voxels or point clouds has been thinned.

[0414] Furthermore, if the ICP error value obtained as a result of the ICP is smaller than a predetermined first threshold, that is, for example, if the positional relationship between the encoding target space and the reference space is close, the inter prediction unit 1311 may determine that rotation and translation processing is unnecessary and may not perform rotation and translation. In this case, the 3D data encoding device 1300 may reduce overhead by not adding RT information to the bitstream.

[0415] Furthermore, if the ICP error value is greater than a predetermined second threshold, the inter prediction unit 1311 may determine that there is a large change in shape between spaces and apply intra prediction to all volumes in the encoding target space. Hereinafter, a space to which intra prediction is applied is referred to as an intra space. The second threshold is a value greater than the first threshold. Furthermore, the method is not limited to ICP, and any method for obtaining RT information from two voxel sets or two point cloud sets may be applied.

[0416] Furthermore, when the three-dimensional data includes attribute information such as shape or color, the inter prediction unit 1311 searches, for example, in the reference space for a volume having attribute information such as shape or color closest to that of the volume to be coded within the coding space, as a prediction volume for the volume to be coded within the coding space. This reference space is, for example, the reference space after the above-described rotation and translation processes have been performed. The inter prediction unit 1311 generates a prediction volume from the volume (reference volume) obtained by the search. FIG. 47 is a diagram for explaining the operation of generating a prediction volume. When encoding the volume to be coded (volume idx=0) shown in FIG. 47 using inter prediction, the inter prediction unit 1311 sequentially scans the reference volumes within the reference space and searches for the volume with the smallest prediction residual, which is the difference between the volume to be coded and the reference volume. The inter prediction unit 1311 selects the volume with the smallest prediction residual as the prediction volume. The prediction residual between the volume to be coded and the prediction volume is encoded by processing performed by the conversion unit 1303 and subsequent processes. Here, the prediction residual is the difference between the attribute information of the volume to be coded and the attribute information of the prediction volume. Furthermore, the three-dimensional data encoding device 1300 adds the volume idx of the reference volume in the reference space referenced as the prediction volume to the header of the bitstream or the like.

[0417] 47, the reference volume of volume idx=4 in the reference space L0R0 is selected as the prediction volume of the volume to be encoded. Then, the prediction residual between the volume to be encoded and the reference volume and the reference volume idx=4 are encoded and added to the bitstream.

[0418] Although an example of generating a predicted volume for attribute information has been described here, the same processing may be performed for a predicted volume for position information.

[0419] The prediction control unit 1312 controls whether to use intra prediction or inter prediction to encode the volume to be encoded. Here, a mode including intra prediction and inter prediction is referred to as a prediction mode. For example, the prediction control unit 1312 calculates, as evaluation values, the prediction residual when the volume to be encoded is predicted by intra prediction and the prediction residual when the volume to be encoded is predicted by inter prediction, and selects the prediction mode with the smaller evaluation value. The prediction control unit 1312 may calculate the actual code amount by applying orthogonal transform, quantization, and entropy coding to the prediction residual of intra prediction and the prediction residual of inter prediction, respectively, and select the prediction mode using the calculated code amount as the evaluation value. Additionally, overhead information other than the prediction residual (such as reference volume idx information) may be added to the evaluation value. Furthermore, the prediction control unit 1312 may always select intra prediction when it is predetermined that the space to be encoded is to be encoded in intra space.

[0420] The entropy coding unit 1313 generates a coded signal (coded bit stream) by variable-length coding the quantized coefficients that are input from the quantization unit 1304. Specifically, the entropy coding unit 1313, for example, binarizes the quantized coefficients and arithmetically codes the resulting binary signal.

[0421] Next, we will explain a three-dimensional data decoding device that decodes the coded signal generated by the three-dimensional data coding device 1300. Fig. 48 is a block diagram of a three-dimensional data decoding device 1400 according to this embodiment. This three-dimensional data decoding device 1400 includes an entropy decoding unit 1401, an inverse quantization unit 1402, an inverse transform unit 1403, an addition unit 1404, a reference volume memory 1405, an intra prediction unit 1406, a reference space memory 1407, an inter prediction unit 1408, and a prediction control unit 1409.

[0422] The entropy decoding unit 1401 performs variable-length decoding on the coded signal (coded bit stream). For example, the entropy decoding unit 1401 arithmetically decodes the coded signal to generate a binary signal, and generates quantization coefficients from the generated binary signal.

[0423] The inverse quantization unit 1402 inversely quantizes the quantized coefficients input from the entropy decoding unit 1401 using a quantization parameter added to the bitstream or the like, thereby generating inverse quantized coefficients.

[0424] The inverse transform unit 1403 generates prediction residuals by inverse transforming the inverse quantized coefficients input from the inverse quantization unit 1402. For example, the inverse transform unit 1403 generates prediction residuals by performing inverse orthogonal transform on the inverse quantized coefficients based on information added to the bitstream.

[0425] The adder 1404 generates a reconstructed volume by adding the prediction residual generated by the inverse transformer 1403 and the prediction volume generated by intra prediction or inter prediction. This reconstructed volume is output as decoded 3D data and is also stored in a reference volume memory 1405 or a reference space memory 1407.

[0426] The intra prediction unit 1406 generates a prediction volume by intra prediction using a reference volume in the reference volume memory 1405 and information added to the bitstream. Specifically, the intra prediction unit 1406 acquires adjacent volume information (e.g., volume idx) and prediction mode information added to the bitstream, and generates a prediction volume in the mode indicated by the prediction mode information using adjacent volumes indicated by the adjacent volume information. Note that the details of these processes are similar to the processes performed by the intra prediction unit 1309 described above, except that information added to the bitstream is used.

[0427] The inter prediction unit 1408 generates a prediction volume by inter prediction using the reference space in the reference space memory 1407 and information added to the bitstream. Specifically, the inter prediction unit 1408 applies rotation and translation processing to the reference space using RT information for each reference space added to the bitstream, and generates a prediction volume using the reference space after application. Note that, if an RT application flag for each reference space exists in the bitstream, the inter prediction unit 1408 applies rotation and translation processing to the reference space in accordance with the RT application flag. Note that the details of these processes are similar to the processes by the inter prediction unit 1311 described above, except that information added to the bitstream is used.

[0428] The prediction control unit 1409 controls whether to decode the volume to be decoded using intra prediction or inter prediction. For example, the prediction control unit 1409 selects intra prediction or inter prediction according to information that indicates the prediction mode to be used and that is added to the bitstream. Note that the prediction control unit 1409 may always select intra prediction if it has been determined in advance that the space to be decoded will be decoded using intra space.

[0429] Modifications of this embodiment will be described below. In this embodiment, an example in which rotation and translation are applied on a space-by-space basis has been described; however, rotation and translation may be applied on a smaller unit basis. For example, the three-dimensional data encoding device 1300 may divide a space into subspaces and apply rotation and translation on a subspace-by-subspace basis. In this case, the three-dimensional data encoding device 1300 generates RT information for each subspace and adds the generated RT information to a bitstream header or the like. The three-dimensional data encoding device 1300 may also apply rotation and translation on a volume-by-volume basis, which is the encoding unit. In this case, the three-dimensional data encoding device 1300 generates RT information on an encoding volume-by-volume basis and adds the generated RT information to a bitstream header or the like. Furthermore, the above may be combined. That is, the three-dimensional data encoding device 1300 may apply rotation and translation on a larger unit basis, and then apply rotation and translation on a smaller unit basis. For example, the three-dimensional data encoding device 1300 may apply rotation and translation on a space-by-space basis, and then apply different rotations and translations to each of multiple volumes included in the resulting space.

[0430] Furthermore, although the present embodiment has been described with reference to an example in which rotation and translation are applied to the reference space, this is not necessarily limited to this. For example, the three-dimensional data encoding device 1300 may change the size of the three-dimensional data by applying a scale process. Furthermore, the three-dimensional data encoding device 1300 may apply any one or two of rotation, translation, and scale. Furthermore, when applying processes in different units in multiple stages as described above, different types of processes may be applied to each unit. For example, rotation and translation may be applied in space units, and translation may be applied in volume units.

[0431] These modifications can also be applied to the three-dimensional data decoding device 1400 in the same manner.

[0432] As described above, the three-dimensional data encoding device 1300 according to this embodiment performs the following processes. FIG.

[0433] First, the three-dimensional data encoding device 1300 generates predicted position information (e.g., predicted volume) using position information of three-dimensional points included in reference three-dimensional data (e.g., reference space) at a time different from that of the target three-dimensional data (e.g., encoding target space) (S1301). Specifically, the three-dimensional data encoding device 1300 generates predicted position information by applying rotation and translation processing to the position information of three-dimensional points included in the reference three-dimensional data.

[0434] The three-dimensional data encoding device 1300 may perform the rotation and translation processing in a first unit (e.g., space) and generate the predicted position information in a second unit (e.g., volume) that is finer than the first unit. For example, the three-dimensional data encoding device 1300 searches for a volume among multiple volumes included in the reference space after the rotation and translation processing that has the smallest difference in position information from the encoding target volume included in the encoding target space, and uses the obtained volume as the predicted volume. The three-dimensional data encoding device 1300 may perform the rotation and translation processing and the generation of the predicted position information in the same unit.

[0435] In addition, the three-dimensional data encoding device 1300 may generate predicted position information by applying a first rotation and translation process to the position information of three-dimensional points included in the reference three-dimensional data in a first unit (e.g., space), and applying a second rotation and translation process to the position information of the three-dimensional points obtained by the first rotation and translation process in a second unit (e.g., volume) that is finer than the first unit.

[0436] Here, the position information and predicted position information of the 3D points are expressed in an octree structure, for example, as shown in Fig. 41. For example, the position information and predicted position information of the 3D points are expressed in a scan order that prioritizes width over depth and width in the octree structure. Alternatively, the position information and predicted position information of the 3D points are expressed in a scan order that prioritizes depth over width in the octree structure.

[0437] 46, the three-dimensional data encoding device 1300 encodes an RT application flag indicating whether or not rotation and translation processing is applied to position information of three-dimensional points included in the reference three-dimensional data. That is, the three-dimensional data encoding device 1300 generates an encoded signal (encoded bitstream) including the RT application flag. The three-dimensional data encoding device 1300 also encodes RT information indicating the details of the rotation and translation processing. That is, the three-dimensional data encoding device 1300 generates an encoded signal (encoded bitstream) including the RT information. Note that the three-dimensional data encoding device 1300 may encode the RT information when the RT application flag indicates that rotation and translation processing is to be applied, and may not need to encode the RT information when the RT application flag indicates that rotation and translation processing is not to be applied.

[0438] The three-dimensional data includes, for example, position information of the three-dimensional points and attribute information (such as color information) of each three-dimensional point. The three-dimensional data encoding device 1300 generates predicted attribute information using the attribute information of the three-dimensional points included in the reference three-dimensional data (S1302).

[0439] Next, the three-dimensional data encoding device 1300 encodes the position information of the three-dimensional point included in the target three-dimensional data using the predicted position information. For example, the three-dimensional data encoding device 1300 calculates differential position information, which is the difference between the position information of the three-dimensional point included in the target three-dimensional data and the predicted position information, as shown in Fig. 38 (S1303).

[0440] Furthermore, the three-dimensional data encoding device 1300 encodes attribute information of three-dimensional points included in the target three-dimensional data using the predicted attribute information. For example, the three-dimensional data encoding device 1300 calculates differential attribute information, which is the difference between the attribute information of three-dimensional points included in the target three-dimensional data and the predicted attribute information (S1304). Next, the three-dimensional data encoding device 1300 converts and quantizes the calculated differential attribute information (S1305).

[0441] Finally, the three-dimensional data encoding device 1300 encodes (e.g., entropy encodes) the differential position information and the quantized differential attribute information (S1306). That is, the three-dimensional data encoding device 1300 generates an encoded signal (encoded bitstream) including the differential position information and the differential attribute information.

[0442] If the three-dimensional data does not include attribute information, the three-dimensional data encoding device 1300 may not perform steps S1302, S1304, and S1305. Furthermore, the three-dimensional data encoding device 1300 may perform only one of encoding the position information of the three-dimensional points and encoding the attribute information of the three-dimensional points.

[0443] 49 is an example and is not limited to this. For example, the processing for the location information (S1301, S1303) and the processing for the attribute information (S1302, S1304, S1305) are independent of each other, and therefore may be performed in any order, or some of them may be processed in parallel.

[0444] As described above, the three-dimensional data encoding device 1300 of this embodiment generates predicted position information using position information of three-dimensional points included in reference three-dimensional data at a different time from the target three-dimensional data, and encodes differential position information that is the difference between the position information of three-dimensional points included in the target three-dimensional data and the predicted position information. This reduces the data amount of the encoded signal, thereby improving encoding efficiency.

[0445] Furthermore, the three-dimensional data encoding device 1300 of this embodiment generates predicted attribute information using attribute information of three-dimensional points included in the reference three-dimensional data, and encodes differential attribute information that is the difference between the attribute information of three-dimensional points included in the target three-dimensional data and the predicted attribute information. This reduces the data amount of the encoded signal, thereby improving encoding efficiency.

[0446] For example, the three-dimensional data encoding device 1300 includes a processor and a memory, and the processor performs the above-described processing using the memory.

[0447] FIG. 48 is a flowchart of the inter prediction process performed by the 3D data decoding device 1400.

[0448] First, the three-dimensional data decoding device 1400 decodes (for example, entropy decodes) the differential position information and differential attribute information from the coded signal (coded bit stream) (S1401).

[0449] Furthermore, the three-dimensional data decoding device 1400 decodes, from the encoded signal, an RT application flag indicating whether or not rotation and translation processing is to be applied to position information of three-dimensional points included in the reference three-dimensional data. Furthermore, the three-dimensional data decoding device 1400 decodes RT information indicating the contents of the rotation and translation processing. Note that the three-dimensional data decoding device 1400 may decode the RT information when the RT application flag indicates that rotation and translation processing is to be applied, and may not need to decode the RT information when the RT application flag indicates that rotation and translation processing is not to be applied.

[0450] Next, the three-dimensional data decoding device 1400 performs inverse quantization and inverse transformation on the decoded differential attribute information (S1402).

[0451] Next, the three-dimensional data decoding device 1400 generates predicted position information (e.g., predicted volume) using position information of three-dimensional points included in reference three-dimensional data (e.g., reference space) at a time different from that of the target three-dimensional data (e.g., decoding target space) (S1403). Specifically, the three-dimensional data decoding device 1400 generates predicted position information by applying rotation and translation processing to the position information of three-dimensional points included in the reference three-dimensional data.

[0452] More specifically, when the RT application flag indicates that rotation and translation processing is to be applied, the three-dimensional data decoding device 1400 applies rotation and translation processing to the position information of the three-dimensional points included in the reference three-dimensional data indicated by the RT information. On the other hand, when the RT application flag indicates that rotation and translation processing is not to be applied, the three-dimensional data decoding device 1400 does not apply rotation and translation processing to the position information of the three-dimensional points included in the reference three-dimensional data.

[0453] The three-dimensional data decoding device 1400 may perform the rotation and translation processing in a first unit (e.g., space) and generate the predicted position information in a second unit (e.g., volume) that is smaller than the first unit. The three-dimensional data decoding device 1400 may perform the rotation and translation processing and the generation of the predicted position information in the same unit.

[0454] In addition, the three-dimensional data decoding device 1400 may generate predicted position information by applying a first rotation and translation process to the position information of three-dimensional points included in the reference three-dimensional data in a first unit (e.g., space), and applying a second rotation and translation process to the position information of the three-dimensional points obtained by the first rotation and translation process in a second unit (e.g., volume) that is finer than the first unit.

[0455] Here, the position information and predicted position information of the 3D points are expressed in an octree structure, for example, as shown in Fig. 41. For example, the position information and predicted position information of the 3D points are expressed in a scan order that prioritizes width over depth and width in the octree structure. Alternatively, the position information and predicted position information of the 3D points are expressed in a scan order that prioritizes depth over width in the octree structure.

[0456] The three-dimensional data decoding device 1400 generates predicted attribute information using attribute information of the three-dimensional points included in the reference three-dimensional data (S1404).

[0457] Next, the three-dimensional data decoding device 1400 restores the position information of the three-dimensional points included in the target three-dimensional data by decoding the encoded position information included in the encoded signal using the predicted position information. Here, the encoded position information is, for example, differential position information, and the three-dimensional data decoding device 1400 restores the position information of the three-dimensional points included in the target three-dimensional data by adding the differential position information and the predicted position information (S1405).

[0458] Furthermore, the three-dimensional data decoding device 1400 restores the attribute information of the three-dimensional points included in the target three-dimensional data by decoding the coded attribute information included in the coded signal using the predicted attribute information. Here, the coded attribute information is, for example, differential attribute information, and the three-dimensional data decoding device 1400 restores the attribute information of the three-dimensional points included in the target three-dimensional data by adding the differential attribute information and the predicted attribute information (S1406).

[0459] Note that if the three-dimensional data does not include attribute information, the three-dimensional data decoding device 1400 does not need to perform steps S1402, S1404, and S1406. Furthermore, the three-dimensional data decoding device 1400 may perform only one of decoding the position information of the three-dimensional points and decoding the attribute information of the three-dimensional points.

[0460] 50 is an example and is not limited to this. For example, the processing for the location information (S1403, S1405) and the processing for the attribute information (S1402, S1404, S1406) are independent of each other, and therefore may be performed in any order, or some of them may be processed in parallel.

[0461] (Embodiment 8) In this embodiment, a method for representing three-dimensional points (point cloud) in encoding three-dimensional data will be described.

[0462] 51 is a block diagram showing the configuration of a three-dimensional data distribution system according to this embodiment. The distribution system shown in FIG.

[0463] Server 1501 includes a storage unit 1511 and a control unit 1512. Storage unit 1511 stores an encoded three-dimensional map 1513, which is encoded three-dimensional data.

[0464] FIG. 52 is a diagram showing an example of the bitstream configuration of the encoded 3D map 1513. The 3D map is divided into multiple sub-maps, and each sub-map is encoded. A random access header (RA) containing sub-coordinate information is attached to each sub-map. The sub-coordinate information is used to improve the encoding efficiency of the sub-map. This sub-coordinate information indicates the sub-coordinate of the sub-map. The sub-coordinate is the coordinate of the sub-map based on a reference coordinate. A 3D map including multiple sub-maps is called an overall map. The reference coordinate (e.g., the origin) in the overall map is called the reference coordinate. In other words, the sub-coordinate is the coordinate of the sub-map in the coordinate system of the overall map. In other words, the sub-coordinate indicates the offset between the coordinate system of the overall map and the coordinate system of the sub-map. The coordinate in the coordinate system of the overall map based on the reference coordinate is called the overall coordinate. The coordinate in the coordinate system of the sub-map based on the sub-coordinate is called the differential coordinate.

[0465] Client 1502 sends a message to server 1501. This message includes location information of client 1502. Control unit 1512 included in server 1501 obtains a bitstream of a submap located closest to the location of client 1502 based on the location information included in the received message. The bitstream of the submap includes subcoordinate information and is sent to client 1502. Decoder 1521 included in client 1502 uses this subcoordinate information to obtain the overall coordinates of the submap relative to the reference coordinates. Application 1522 included in client 1502 executes an application related to its own location using the obtained overall coordinates of the submap.

[0466] Furthermore, a submap indicates a partial area of ​​the overall map. Subcoordinates are the coordinates at which the submap is located in the reference coordinate space of the overall map. For example, suppose that an overall map A has submap A of AA and submap B of AB. When a vehicle wants to refer to the map of AA, it starts decoding from submap A, and when it wants to refer to the map of AB, it starts decoding from submap B. Here, submaps are random access points. Specifically, A is Osaka Prefecture, AA is Osaka City, and AB is Takatsuki City, etc.

[0467] Each submap is transmitted to the client together with sub-coordinate information, which is included in the header information of each submap, a transmission packet, or the like.

[0468] The reference coordinates that serve as the reference coordinates for the sub-coordinate information of each sub-map may be added to header information of a space higher than the sub-map, such as the header information of the overall map.

[0469] A submap may consist of one space (SPC), or it may consist of multiple SPCs.

[0470] A submap may also include a GOS (Group of Space). A submap may also consist of a world. For example, if a submap contains multiple objects, assigning the objects to different SPCs will result in the submap consisting of multiple SPCs. Assigning the objects to a single SPC will result in the submap consisting of a single SPC.

[0471] Next, the effect of improving coding efficiency when using sub-coordinate information will be described. FIG. 53 is a diagram for explaining this effect. For example, a large number of bits is required to encode 3D point A, which is located far from the reference coordinates as shown in FIG. 53. Here, the distance between the sub-coordinates and 3D point A is shorter than the distance between the reference coordinates and 3D point A. Therefore, coding efficiency can be improved by encoding the coordinates of 3D point A based on the sub-coordinates, rather than encoding the coordinates of 3D point A based on the reference coordinates. Furthermore, the sub-map bitstream includes sub-coordinate information. By sending the sub-map bitstream and the reference coordinates to the decoding side (client), the overall coordinates of the sub-map can be restored on the decoding side.

[0472] FIG. 54 is a flowchart of the processing by the server 1501 that transmits the submap.

[0473] First, the server 1501 receives a message including the location information of the client 1502 from the client 1502 (S1501). The control unit 1512 acquires an encoded bitstream of a submap based on the client's location information from the storage unit 1511 (S1502). Then, the server 1501 transmits the encoded bitstream of the submap and the reference coordinates to the client 1502 (S1503).

[0474] FIG. 55 is a flowchart of the processing by the client 1502 that receives the submap.

[0475] First, the client 1502 receives the coded bitstream of the submap and the reference coordinates transmitted from the server 1501 (S1511). Next, the client 1502 obtains the submap and subcoordinate information by decoding the coded bitstream (S1512). Next, the client 1502 restores the differential coordinates in the submap to global coordinates using the reference coordinates and the subcoordinates (S1513).

[0476] Next, an example of the syntax of information related to submaps will be described. In encoding a submap, a three-dimensional data encoding device calculates differential coordinates by subtracting sub-coordinates from the coordinates of each point cloud (three-dimensional point). The three-dimensional data encoding device then encodes the differential coordinates into a bitstream as the value of each point cloud. The encoding device also encodes sub-coordinate information indicating the sub-coordinates as header information of the bitstream. This allows a three-dimensional data decoding device to obtain the overall coordinates of each point cloud. For example, the three-dimensional data encoding device may be included in server 1501, and the three-dimensional data decoding device may be included in client 1502.

[0477] Figure 56 is a diagram showing an example of the syntax of a submap. NumOfPoint shown in Figure 56 indicates the number of point clouds included in the submap. sub_coordinate_x, sub_coordinate_y, and sub_coordinate_z are sub-coordinate information. sub_coordinate_x indicates the x-coordinate of the sub-coordinate. sub_coordinate_y indicates the y-coordinate of the sub-coordinate. sub_coordinate_z indicates the z-coordinate of the sub-coordinate.

[0478] Also, diff_x[i], diff_y[i], and diff_z[i] are the differential coordinates of the i-th point cloud in the submap. diff_x[i] indicates the differential value between the x-coordinate of the i-th point cloud in the submap and the x-coordinate of the sub-coordinate. diff_y[i] indicates the differential value between the y-coordinate of the i-th point cloud in the submap and the y-coordinate of the sub-coordinate. diff_z[i] indicates the differential value between the z-coordinate of the i-th point cloud in the submap and the z-coordinate of the sub-coordinate.

[0479] The three-dimensional data decoding device decodes point_cloud[i]_x, point_cloud[i]_y, and point_cloud[i]_z, which are the global coordinates of the i-th point cloud, using the following equations: point_cloud[i]_x is the x coordinate of the global coordinates of the i-th point cloud; point_cloud[i]_y is the y coordinate of the global coordinates of the i-th point cloud; and point_cloud[i]_z is the z coordinate of the global coordinates of the i-th point cloud.

[0480] point_cloud[i]_x=sub_coordinate_x+diff_x[i] point_cloud[i]_y=sub_coordinate_y+diff_y[i] point_cloud[i]_z=sub_coordinate_z+diff_z[i]

[0481] Next, the process of switching the application of octree coding will be described. When encoding a submap, the three-dimensional data encoding device selects whether to encode each point cloud using an octree representation (hereinafter referred to as octree coding) or to encode differential values ​​from sub-coordinates (hereinafter referred to as non-octree coding). FIG. 57 is a diagram schematically illustrating this operation. For example, if the number of point clouds in a submap is equal to or greater than a predetermined threshold, the three-dimensional data encoding device applies octree coding to the submap. If the number of point clouds in a submap is less than the threshold, the three-dimensional data encoding device applies non-octree coding to the submap. This allows the three-dimensional data encoding device to appropriately select whether to use octree coding or non-octree coding depending on the shape and density of objects included in the submap, thereby improving encoding efficiency.

[0482] Furthermore, the three-dimensional data encoding device adds information indicating whether octree encoding or non-octree encoding has been applied to the submap (hereinafter referred to as octree encoding application information) to the header of the submap, etc. This allows the three-dimensional data decoding device to determine whether the bitstream is a bitstream obtained by octree encoding the submap or a bitstream obtained by non-octree encoding the submap.

[0483] In addition, the three-dimensional data encoding device may calculate the encoding efficiency when applying octree encoding and non-octree encoding to the same point cloud, and apply the encoding method with the best encoding efficiency to the submap.

[0484] Figure 58 is a diagram showing an example of the syntax of a submap when this switching is performed. coding_type shown in Figure 58 is information indicating the coding type, and is the above-mentioned octree coding application information. coding_type=00 indicates that octree coding has been applied. coding_type=01 indicates that non-octree coding has been applied. coding_type=10 or 11 indicates that a coding method other than those mentioned above has been applied.

[0485] If the coding type is non-octree coding (non_octree), the submap includes NumOfPoint and sub-coordinate information (sub_coordinate_x, sub_coordinate_y, and sub_coordinate_z).

[0486] If the coding type is octree coding, the submap includes octree_info, which is information necessary for octree coding, such as depth information.

[0487] If the coding type is non_octree, the submap contains difference coordinates (diff_x[i], diff_y[i], and diff_z[i]).

[0488] If the coding type is octree coding, the submap contains octree_data, which is coding data for octree coding.

[0489] Although an example in which the xyz coordinate system is used as the coordinate system of the point cloud has been shown here, a polar coordinate system may also be used.

[0490] 59 is a flowchart of three-dimensional data encoding processing by a three-dimensional data encoding device. First, the three-dimensional data encoding device calculates the number of point clouds in a target submap, which is the submap to be processed (S1521). Next, the three-dimensional data encoding device determines whether the calculated number of point clouds is equal to or greater than a predetermined threshold (S1522).

[0491] If the number of point clouds is equal to or greater than the threshold (Yes in S1522), the three-dimensional data encoding device applies octree encoding to the target submap (S1523). In addition, the three-dimensional point data encoding device adds octree encoding application information indicating that octree encoding has been applied to the target submap to the header of the bitstream (S1525).

[0492] On the other hand, if the number of point clouds is less than the threshold (No in S1522), the 3D data encoding device applies non-octree encoding to the target submap (S1524).The 3D point data encoding device also adds octree encoding application information indicating that non-octree encoding has been applied to the target submap to the bitstream header (S1525).

[0493] 60 is a flowchart of three-dimensional data decoding processing by a three-dimensional data decoding device. First, the three-dimensional data decoding device decodes octree coding application information from the header of the bitstream (S1531). Next, the three-dimensional data decoding device determines whether the coding type applied to the target submap is octree coding based on the decoded octree coding application information (S1532).

[0494] If the coding type indicated by the octree coding application information is octree coding (Yes in S1532), the three-dimensional data decoding device decodes the target submap using octree decoding (S1533). On the other hand, if the coding type indicated by the octree coding application information is non-octree coding (No in S1532), the three-dimensional data decoding device decodes the target submap using non-octree decoding (S1534).

[0495] A modification of this embodiment will be described below. Figures 61 to 63 are diagrams schematically showing the operation of a modification of the coding type switching process.

[0496] As shown in Figure 61, the three-dimensional data encoding device may select whether to apply octree encoding or non-octree encoding for each space. In this case, the three-dimensional data encoding device adds octree encoding application information to the header of the space. This allows the three-dimensional data decoding device to determine for each space whether octree encoding has been applied. In this case, the three-dimensional data encoding device also sets sub-coordinates for each space and encodes the difference values ​​obtained by subtracting the sub-coordinate values ​​from the coordinates of each point cloud in the space.

[0497] This allows the three-dimensional data encoding device to appropriately switch whether or not to apply octree encoding depending on the shape of the object in the space or the number of point clouds, thereby improving encoding efficiency.

[0498] Furthermore, as shown in Figure 62, the three-dimensional data encoding device may select whether to apply octree encoding or non-octree encoding for each volume. In this case, the three-dimensional data encoding device adds octree encoding application information to the header of the volume. This allows the three-dimensional data decoding device to determine for each volume whether octree encoding has been applied. In this case, the three-dimensional data encoding device sets sub-coordinates for each volume and encodes the difference values ​​obtained by subtracting the values ​​of the sub-coordinates from the coordinates of each point cloud within the volume.

[0499] This allows the three-dimensional data encoding device to appropriately switch whether or not to apply octree encoding depending on the shape of the object in the volume or the number of point clouds, thereby improving encoding efficiency.

[0500] In the above explanation, an example of encoding the difference obtained by subtracting the sub-coordinates from the coordinates of each point cloud has been shown as non-octree encoding, but this is not necessarily limited to this, and any encoding method other than octree encoding may be used. For example, as shown in Fig. 63, the three-dimensional data encoding device may use a method (hereinafter referred to as original coordinate encoding) for encoding the value of the point cloud itself within a submap, space, or volume, rather than the difference from the sub-coordinates, as non-octree encoding.

[0501] In this case, the three-dimensional data encoding device stores information in the header indicating that original coordinate encoding has been applied to the target space (submap, space, or volume), which allows the three-dimensional data decoding device to determine whether original coordinate encoding has been applied to the target space.

[0502] Furthermore, when applying original coordinate coding, the three-dimensional data coding device may perform coding without applying quantization and arithmetic coding to the original coordinates. Furthermore, the three-dimensional data coding device may code the original coordinates with a predetermined fixed bit length. This allows the three-dimensional data coding device to generate a stream with a constant bit length at a certain timing.

[0503] In the above description, an example has been given in which the difference obtained by subtracting the sub-coordinates from the coordinates of each point cloud is coded as non-octree coding, but the present invention is not necessarily limited to this.

[0504] For example, the three-dimensional data encoding device may sequentially encode the difference values ​​between the coordinates of each point cloud. FIG. 64 is a diagram for explaining the operation in this case. For example, in the example shown in FIG. 64, when encoding point cloud PA, the three-dimensional data encoding device uses sub-coordinates as predicted coordinates and encodes the difference values ​​between the coordinates of point cloud PA and the predicted coordinates. Furthermore, when encoding point cloud PB, the three-dimensional data encoding device uses the coordinates of point cloud PA as predicted coordinates and encodes the difference values ​​between point cloud PB and the predicted coordinates. Furthermore, when encoding point cloud PC, the three-dimensional data encoding device uses point cloud PB as predicted coordinates and encodes the difference values ​​between point cloud PB and the predicted coordinates. In this way, the three-dimensional data encoding device may set a scan order for multiple point clouds and encode the difference values ​​between the coordinates of a target point cloud to be processed and the coordinates of the point cloud immediately preceding the target point cloud in the scan order.

[0505] Furthermore, in the above description, the sub-coordinates are coordinates of the lower left front corner of the sub-map, but the positions of the sub-coordinates are not limited to this. FIGS. 65 to 67 are diagrams showing other examples of the positions of the sub-coordinates. The sub-coordinates may be set to any coordinates within the target space (sub-map, space, or volume). That is, as described above, the sub-coordinates may be coordinates of the lower left front corner of the target space. As shown in FIG. 65, the sub-coordinates may be coordinates of the center of the target space. As shown in FIG. 66, the sub-coordinates may be coordinates of the upper right back corner of the target space. Furthermore, the sub-coordinates are not limited to coordinates of the lower left front or upper right back corner of the target space, but may be coordinates of any corner of the target space.

[0506] In addition, the setting position of the sub-coordinates may be the same as the coordinates of a certain point cloud in the target space (submap, space, or volume). For example, in the example shown in Figure 67, the coordinates of the sub-coordinates match the coordinates of the point cloud PD.

[0507] Furthermore, in this embodiment, an example has been shown in which the application of octree coding and the application of non-octree coding are switched, but this is not necessarily limited to this. For example, the three-dimensional data coding device may switch between applying a tree structure other than an octree and applying a non-tree structure other than the tree structure. For example, the other tree structure may be a kd tree in which division is performed using a plane perpendicular to one of the coordinate axes. Note that any method may be used as the other tree structure.

[0508] Furthermore, although the present embodiment has shown an example in which coordinate information of a point cloud is encoded, this is not necessarily limited to this. The three-dimensional data encoding device may also encode, for example, color information, three-dimensional feature quantities, or visible light feature quantities in the same manner as coordinate information. For example, the three-dimensional data encoding device may set the average value of the color information of each point cloud in a submap as sub-color information (sub-color), and encode the difference between the color information of each point cloud and the sub-color information.

[0509] Furthermore, in this embodiment, an example has been shown in which a coding method (octree coding or non-octree coding) with good coding efficiency is selected depending on the number of point clouds, etc., but this is not necessarily limited to this. For example, a three-dimensional data coding device on the server side may store bit streams of point clouds coded using octree coding, bit streams of point clouds coded using non-octree coding, and bit streams of point clouds coded using both of these coding methods, and switch the bit stream to be sent to the three-dimensional data decoding device depending on the communication environment or the processing capacity of the three-dimensional data decoding device.

[0510] Fig. 68 is a diagram showing an example of volume syntax when switching the application of octree coding. The syntax shown in Fig. 68 is basically the same as the syntax shown in Fig. 58, except that each piece of information is information on a volume basis. Specifically, NumOfPoint indicates the number of point clouds included in the volume. sub_coordinate_x, sub_coordinate_y, and sub_coordinate_z are sub-coordinate information of the volume.

[0511] Also, diff_x[i], diff_y[i], and diff_z[i] are the differential coordinates of the i-th point cloud in the volume. diff_x[i] indicates the differential value between the x-coordinate of the i-th point cloud in the volume and the x-coordinate of the sub-coordinate. diff_y[i] indicates the differential value between the y-coordinate of the i-th point cloud in the volume and the y-coordinate of the sub-coordinate. diff_z[i] indicates the differential value between the z-coordinate of the i-th point cloud in the volume and the z-coordinate of the sub-coordinate.

[0512] If the relative positions of volumes in space can be calculated, the three-dimensional data encoding device does not need to include sub-coordinate information in the volume header. In other words, the three-dimensional data encoding device may calculate the relative positions of volumes in space without including the sub-coordinate information in the header, and use the calculated positions as the sub-coordinates of each volume.

[0513] As described above, the three-dimensional data encoding device according to this embodiment determines whether or not to encode a target spatial unit among multiple spatial units (e.g., submaps, spaces, or volumes) included in three-dimensional data using an octree structure (e.g., S1522 in FIG. 59). For example, if the number of three-dimensional points included in the target spatial unit is greater than a predetermined threshold, the three-dimensional data encoding device determines to encode the target spatial unit using an octree structure. Furthermore, if the number of three-dimensional points included in the target spatial unit is equal to or less than the threshold, the three-dimensional data encoding device determines not to encode the target spatial unit using an octree structure.

[0514] If it is determined that the target space unit is to be coded using an octree structure (Yes in S1522), the three-dimensional data coding device codes the target space unit using the octree structure (S1523). On the other hand, if it is determined that the target space unit is not to be coded using an octree structure (No in S1522), the three-dimensional data coding device codes the target space unit using a method other than the octree structure (S1524). For example, in the different method, the three-dimensional data coding device codes the coordinates of three-dimensional points included in the target space unit. Specifically, in the different method, the three-dimensional data coding device codes the difference between the reference coordinates of the target space unit and the coordinates of three-dimensional points included in the target space unit.

[0515] Next, the three-dimensional data encoding device adds information indicating whether the target spatial unit has been encoded using an octree structure to the bitstream (S1525).

[0516] This allows the three-dimensional data encoding device to reduce the amount of data in the encoded signal, thereby improving encoding efficiency.

[0517] For example, the three-dimensional data encoding device includes a processor and a memory, and the processor performs the above-mentioned processing using the memory.

[0518] Furthermore, the three-dimensional data decoding device according to this embodiment decodes, from the bitstream, information indicating whether or not a target spatial unit among multiple target spatial units (e.g., submaps, spaces, or volumes) included in the three-dimensional data is to be decoded using an octree structure (e.g., S1531 in FIG. 60). If the information indicates that the target spatial unit is to be decoded using an octree structure (Yes in S1532), the three-dimensional data decoding device decodes the target spatial unit using the octree structure (S1533).

[0519] If the information indicates that the target space unit is not to be decoded using an octree structure (No in S1532), the three-dimensional data decoding device decodes the target space unit using a method other than the octree structure (S1534). For example, in the different method, the three-dimensional data decoding device decodes the coordinates of three-dimensional points included in the target space unit. Specifically, in the different method, the three-dimensional data decoding device decodes the difference between the reference coordinates of the target space unit and the coordinates of three-dimensional points included in the target space unit.

[0520] This allows the three-dimensional data decoding device to reduce the amount of data in the coded signal, thereby improving coding efficiency.

[0521] For example, the three-dimensional data decoding device includes a processor and a memory, and the processor performs the above processing using the memory.

[0522] (Embodiment 9) In this embodiment, a method for encoding a tree structure such as an octree structure will be described.

[0523] Efficiency can be improved by identifying important areas and prioritizing decoding of the three-dimensional data of the important areas.

[0524] FIG. 69 is a diagram showing an example of an important area in a three-dimensional map. An important area is, for example, an area that includes a certain number or more of three-dimensional points with large feature values ​​among the three-dimensional points in the three-dimensional map. Alternatively, an important area may be, for example, an area that includes a certain number or more of three-dimensional points required when a client, such as an in-vehicle client, performs self-location estimation. Alternatively, an important area may be a facial area in a three-dimensional model of a person. In this way, important areas can be defined for each application, and important areas may be switched depending on the application.

[0525] In this embodiment, occupancy coding and location coding are used as methods for expressing an octree structure, etc. Also, a bit string obtained by occupancy coding is called an occupancy code. A bit string obtained by location coding is called a location code.

[0526] FIG. 70 is a diagram illustrating an example of an occupancy code. FIG. 70 shows an example of an occupancy code having a quadtree structure. In FIG. 70, an occupancy code is assigned to each node. Each occupancy code indicates whether or not a child node or leaf of each node contains a 3D point. For example, in the case of a quadtree, information indicating whether each of the four child nodes or leaves of each node contains a 3D point is represented by a 4-bit occupancy code. In the case of an occupancy tree, information indicating whether or not each of the eight child nodes or leaves of each node contains a 3D point is represented by an 8-bit occupancy code. Note that, for simplicity, a quadtree structure will be used as an example, but the same applies to an occupancy tree structure. For example, as shown in FIG. 70, the occupancy code is an example of bits obtained by scanning nodes and leaves in a breadth-first manner, as described in FIG. 40 and other figures. In an occupancy code, information on multiple 3D points is decoded in a fixed order, so it is not possible to prioritize the information on any particular 3D point. The occupancy code may be a bit string obtained by scanning nodes and leaves in a depth-first manner as described with reference to FIG.

[0527] The location coding is explained below. By using the location code, important parts of the octree structure can be directly decoded. Also, important 3D points in the deep layer can be efficiently coded.

[0528] Fig. 71 is a diagram for explaining location coding, and shows an example of a quadtree structure. In the example shown in Fig. 71, 3D points A to I are represented by the quadtree structure. Furthermore, 3D points A and C are important 3D points included in an important region.

[0529] FIG. 72 is a diagram showing the occupancy code and location code representing the important 3D points A and C in the quadtree structure shown in FIG.

[0530] In location coding, in a tree structure, the index of the nodes on the path leading to the leaf to which the target 3D point belongs, which is the 3D point to be coded, and the index of the leaf are coded. Here, the index is a numerical value assigned to each node and leaf. In other words, the index is an identifier for identifying multiple child nodes of the target node. In the case of a quadtree as shown in Figure 71, the index indicates any one of 0 to 3.

[0531] For example, in the quadtree structure shown in Figure 71, if leaf A is the target three-dimensional point, leaf A is expressed as 0 → 2 → 1 → 0 → 1 → 2 → 1. Here, since the maximum value of each index is 4 (can be expressed in 2 bits) in the case of the right diagram, the number of bits required to code the location of leaf A is 7 × 2 bits = 14 bits. Similarly, when leaf C is the target for encoding, the number of bits required is 14 bits. Note that in the case of an octree, since the maximum value of each index is 8 (can be expressed in 3 bits), the number of bits required can be calculated by 3 bits × leaf depth. Note that the three-dimensional data encoding device may reduce the amount of data by binarizing each index and then entropy-encoding it.

[0532] Also, as shown in Figure 72, with occupancy codes, in order to decode leaves A and C, it is necessary to decode all nodes in the upper layer. On the other hand, with location codes, it is possible to decode only the data of leaves A and C. As a result, as shown in Figure 72, by using location codes, the number of bits can be reduced compared to occupancy codes.

[0533] Furthermore, as shown in FIG. 72, by performing dictionary compression such as LZ77 on some or all of the location codes, the amount of code can be further reduced.

[0534] Next, an example of applying location coding to 3D points (point cloud) obtained by LiDAR will be described. FIG. 73 is a diagram showing an example of 3D points obtained by LiDAR. 3D points obtained by LiDAR are sparse. In other words, when these 3D points are represented by occupancy codes, the number of zero values ​​increases. Furthermore, high 3D accuracy is required for these 3D points. In other words, the hierarchy of the octree structure becomes deeper.

[0535] Figure 74 is a diagram showing an example of such a sparse deep octree structure. The occupancy code of the octree structure shown in Figure 74 is 136 bits (= 8 bits × 17 nodes). Furthermore, since the depth is 6 and there are six three-dimensional points, the location code is 3 bits × 6 × 6 = 108 bits. In other words, the location code can reduce the amount of code by 20% compared to the occupancy code. In this way, the amount of code can be reduced by applying location coding to a sparse deep octree structure.

[0536] The code amounts of the occupancy code and the location code will be explained below. When the depth of the octree structure is 10, the maximum number of three-dimensional points is 8. 10 = 1073741824. Also, the number of bits of the occupancy code of the octet tree structure is L o is expressed as follows:

[0537] L o =8+8 2 +···+8 10 =127133512 bits

[0538] Therefore, the number of bits per 3D point is 1.143 bits. Note that in the occupancy code, this number of bits does not change even if the number of 3D points included in the octree structure changes.

[0539] On the other hand, the number of bits per 3D point in a location code is directly affected by the depth of the octree structure: specifically, the number of bits per 3D point in a location code is 3 bits x depth 10 = 30 bits.

[0540] Therefore, the number of bits of the location code of the octree structure is L l is expressed as follows:

[0541] L l =30×N

[0542] Here, N is the number of 3D points contained in the octree structure.

[0543] Therefore, N <L o When / 30=40904450.4, that is, when the number of three-dimensional points is less than 40904450, the amount of code for the location code is less than the amount of code for the occupancy code (L l <L o ).

[0544] In this way, when there are few three-dimensional points, the amount of location code is less than the amount of occupancy code, and when there are many three-dimensional points, the amount of location code is greater than the amount of occupancy code.

[0545] Therefore, the three-dimensional data encoding device may switch between using location encoding and occupancy encoding depending on the number of input three-dimensional points. In this case, the three-dimensional data encoding device may add information indicating whether location encoding or occupancy encoding was used to header information of the bitstream, etc.

[0546] Hybrid coding, which combines location coding and occupancy coding, will be described below. When encoding a dense important region, hybrid coding that combines location coding and occupancy coding is effective. Figure 75 shows an example of this. In the example shown in Figure 75, important 3D points are densely arranged. In this case, the 3D data coding device performs location coding on the upper layers with shallow depth and uses occupancy coding on the lower layers. Specifically, location coding is used up to the deepest common node, and occupancy coding is used in layers deeper than the deepest common node. Here, the deepest common node is the deepest node among the nodes that are common ancestors of multiple important 3D points.

[0547] Next, hybrid coding that prioritizes compression efficiency will be described. The three-dimensional data coding device may switch between location coding and occupancy coding according to a predetermined rule in octree coding.

[0548] Figure 76 shows an example of this rule. First, the three-dimensional data encoding device checks the proportion of nodes at each level (depth) that contain three-dimensional points. If this proportion is higher than a predetermined threshold, the three-dimensional data encoding device performs occupancy encoding on some nodes above the target level. For example, the three-dimensional data encoding device applies occupancy encoding to the levels from the target level to the deepest common node.

[0549] 76, the ratio of nodes at the third level that include three-dimensional points is higher than the threshold value. Therefore, the three-dimensional data encoding device applies occupancy encoding to the second and third levels from the third level to the deepest common node, and applies location encoding to the remaining first and fourth levels.

[0550] The calculation method for the threshold value is explained below. Each layer of the octree structure has one root node and eight child nodes. Therefore, occupancy coding requires eight bits to encode one layer of the octree structure. On the other hand, location coding requires three bits for each child node that contains a three-dimensional point. Therefore, when the number of nodes containing three-dimensional points is greater than two, occupancy coding is more effective than location coding. In other words, in this case, the threshold value is two.

[0551] An example of the structure of a bitstream generated by the above-mentioned location coding, occupancy coding, or hybrid coding will be described below.

[0552] Figure 77 is a diagram showing an example of a bitstream generated by location encoding. As shown in Figure 77, the bitstream generated by location encoding includes a header and multiple location codes. Each location encoding corresponds to one 3D point.

[0553] With this configuration, the 3D data decoding device can decode multiple 3D points individually with high accuracy. Note that Figure 77 shows an example of a bitstream in the case of a quadtree structure. In the case of an octtree structure, each index can take a value from 0 to 7.

[0554] Furthermore, the three-dimensional data encoding device may binarize an index string representing one three-dimensional point and then perform entropy encoding. For example, if the index string is 0121, the three-dimensional data encoding device may binarize 0121 to 00011001 and perform arithmetic encoding on this bit string.

[0555] Figure 78 is a diagram showing an example of a bitstream generated by hybrid coding when important 3D points are included. As shown in Figure 78, location codes in an upper layer, occupancy codes for important 3D points in a lower layer, and occupancy codes for non-important 3D points other than the important 3D points in the lower layer are arranged in this order. Note that the location code length shown in Figure 78 represents the code amount of the subsequent location code. Also, the occupancy code amount represents the code amount of the subsequent occupancy code.

[0556] This configuration allows the three-dimensional data decoding device to select different decoding schemes depending on the application.

[0557] Also, the coded data of the important 3D points is stored near the beginning of the bit stream, and the coded data of the non-important 3D points not included in the important region is stored after the coded data of the important 3D points.

[0558] Fig. 79 is a diagram showing a tree structure represented by the occupancy code of the important 3D points shown in Fig. 78. Fig. 80 is a diagram showing a tree structure represented by the occupancy code of the unimportant 3D points shown in Fig. 78. As shown in Fig. 79, the occupancy code of the important 3D points excludes information about unimportant 3D points. Specifically, since nodes 0 and 3 at depth 5 do not include any important 3D points, a value of 0 is assigned to nodes 0 and 3, indicating that no 3D points are included.

[0559] On the other hand, information about important 3D points is excluded from the occupancy code of unimportant 3D points, as shown in Fig. 80. Specifically, since node 1 at depth 5 does not include any unimportant 3D points, a value of 0 is assigned to node 1, indicating that it does not include any 3D points.

[0560] In this way, the 3D data encoding device divides the original tree structure into a first tree structure containing important 3D points and a second tree structure containing unimportant 3D points, and performs occupancy encoding on the first tree structure and the second tree structure independently, thereby enabling the 3D data decoding device to prioritize decoding of important 3D points.

[0561] Next, an example of the structure of a bitstream generated by hybrid coding that prioritizes efficiency will be described. Fig. 81 is a diagram showing an example of the structure of a bitstream generated by hybrid coding that prioritizes efficiency. As shown in Fig. 81, for each subtree, a subtree root location, an occupancy code amount, and an occupancy code are arranged in this order. The subtree location shown in Fig. 81 is the location code of the root of the subtree.

[0562] In the above configuration, when only one of location coding and occupancy coding is applied to the 8-tree structure, the following holds true.

[0563] If the length of the location encoding of the root of a subtree is equal to the depth of the octree, then the subtree has no child nodes, i.e. the location encoding has been applied to the entire tree.

[0564] If the root of the subtree is equal to the root of the octree, then occupancy coding has been applied to the entire tree.

[0565] For example, based on the above rules, a three-dimensional data decoding device can determine whether a bitstream contains location code or occupancy code.

[0566] The bitstream may also include coding mode information indicating whether location coding, occupancy coding, or hybrid coding is used. Figure 82 is a diagram showing an example of a bitstream in this case. For example, as shown in Figure 82, 2-bit coding mode information indicating the coding mode is added to the bitstream.

[0567] Note that (1) the "number of 3D points" in location coding indicates the number of subsequent 3D points. Also, (2) the "occupancy code size" in occupancy coding indicates the code size of the subsequent occupancy code. Also, (3) the "number of important subtrees" in mixed coding (important 3D points) indicates the number of subtrees containing important 3D points. Also, (4) the "number of occupancy subtrees" in mixed coding (efficiency-oriented) indicates the number of occupancy-coded subtrees.

[0568] Next, an example of syntax used to switch between applying occupancy coding and location coding will be described. Figure 83 shows this example of syntax.

[0569] 83 is a flag indicating whether the target node is a leaf or not. isleaf=1 indicates that the target node is a leaf, and isleaf=0 indicates that the target node is a node and not a leaf.

[0570] If the target node is a leaf, point_flag is added to the bitstream. point_flag is a flag indicating whether the target node (leaf) contains a 3D point. point_flag=1 indicates that the target node contains a 3D point, and point_flag=0 indicates that the target node does not contain a 3D point.

[0571] If the target node is not a leaf, coding_type is added to the bitstream. coding_type is coding type information that indicates the applied coding type. coding_type=00 indicates that location coding is applied, coding_type=01 indicates that occupancy coding is applied, and coding_type=10 or 11 indicates that other coding methods are applied.

[0572] When the encoding type is location encoding, numPoint, num_idx[i], and idx[i][j] are added to the bitstream.

[0573] numPoint indicates the number of three-dimensional points for which location encoding is performed. num_idx[i] indicates the number (depth) of indexes from the target node to the three-dimensional point i. If all the three-dimensional points for which location encoding is performed are at the same depth, then all num_idx[i] will have the same value. Therefore, before the for loop (for (i = 0; i < numPoint; i++) {) shown in FIG. 83, num_idx may be defined as a common value.

[0574] idx[i][j] indicates the value of the j-th index among the indexes from the target node to the three-dimensional point i. In the case of an octree, the number of bits of idx[i][j] is 3 bits.

[0575] As described above, an index is an identifier for identifying a plurality of child nodes of a target node. In the case of an octree, idx[i][j] indicates any one of 0 to 7. Also, in the case of an octree, there are 8 child nodes, and each child node corresponds to each of the 8 sub-blocks obtained by spatially dividing the target block corresponding to the target node into 8 parts. Therefore, idx[i][j] may be information indicating the three-dimensional position of the sub-block corresponding to the child node. For example, idx[i][j] may be a total of 3-bit information including 1-bit information indicating the position of each of x, y, and z of the sub-block.

[0576] When the encoding type is occupancy encoding, occupancy_code is added to the bitstream. occupancy_code is the occupancy code of the target node. In the case of an octree, occupancy_code is an 8-bit bit string such as the bit string "00101000".

[0577] If the value of the (i+1)th bit of occupancy_code is 1, the process moves to the child node. In other words, the child node is set as the next target node, and a bit string is generated recursively.

[0578] In this embodiment, an example has been shown in which the end of an octtree is represented by adding leaf information (isleaf, point_flag) to a bitstream, but this is not necessarily limited to this. For example, the three-dimensional data encoding device may add the maximum depth from the start node (root) of the occupancy code to the end (leaf) at which a three-dimensional point exists to the header section of the start node. The three-dimensional data encoding device may then recursively convert information of child nodes into bit strings while increasing the depth from the start node, and determine that a leaf has been reached when the depth reaches the maximum depth. Furthermore, the three-dimensional data encoding device may add information indicating the maximum depth to the first node whose coding type becomes occupancy coding, or to the start node (root) of the octtree.

[0579] As described above, the three-dimensional data encoding device may add information for switching between occupancy encoding and location encoding to the bitstream as header information for each node.

[0580] Furthermore, the three-dimensional data encoding device may entropy-encode the coding_type, numPoint, num_idx, idx, and occupancy_code of each node generated by the above method. For example, the three-dimensional data encoding device binarizes each value and then performs arithmetic encoding.

[0581] Furthermore, in the above syntax, an example is given in which a depth-first bit string in an occupancy tree structure is used as the occupancy code, but this is not necessarily limited to this. The three-dimensional data encoding device may also use a breadth-first bit string in an occupancy tree structure as the occupancy code. Even when a breadth-first bit string is used, the three-dimensional data encoding device may add information for switching between occupancy encoding and location encoding to the bit stream as header information for each node.

[0582] In this embodiment, an octal tree structure is used as an example, but this is not necessarily limited to this, and the above method may also be applied to N-ary trees (N is an integer greater than or equal to 2) such as quad trees and hexadecimal trees, or other tree structures.

[0583] An example of the flow of the encoding process for switching between the application of occupancy encoding and location encoding will be described below. Figure 84 is a flowchart of the encoding process according to this embodiment.

[0584] First, the three-dimensional data encoding device represents multiple three-dimensional points included in the three-dimensional data in an octree structure (S1601). Next, the three-dimensional data encoding device sets the root of the octree structure to the target node (S1602). Next, the three-dimensional data encoding device generates a bit string in the octree structure by performing node encoding processing on the target node (S1603). Next, the three-dimensional data encoding device generates a bit stream by entropy encoding the generated bit string (S1604).

[0585] 85 is a flowchart of the node encoding process (S1603). First, the three-dimensional data encoding device determines whether the target node is a leaf (S1611). If the target node is not a leaf (No in S1611), the three-dimensional data encoding device sets the leaf flag (isleaf) to 0 and adds the leaf flag to the bit string (S1612).

[0586] Next, the three-dimensional data encoding device determines whether the number of child nodes including the three-dimensional point is greater than a predetermined threshold (S1613). Note that the three-dimensional data encoding device may add this threshold to the bit string.

[0587] If the number of child nodes containing three-dimensional points is greater than a predetermined threshold (Yes in S1613), the three-dimensional data encoding device sets the encoding type (coding_type) to occupancy encoding and adds the encoding type to the bit string (S1614).

[0588] Next, the three-dimensional data encoding device sets occupancy encoding information and adds the occupancy encoding information to the bit string. Specifically, the three-dimensional data encoding device generates an occupancy code for the target node and adds the occupancy code to the bit string (S1615).

[0589] Next, the three-dimensional data encoding device sets the next target node according to the occupancy code (S1616). Specifically, the three-dimensional data encoding device sets the next target node from the unprocessed child node whose occupancy code is "1".

[0590] Next, the three-dimensional data encoding device performs node encoding processing on the newly set target node (S1617). That is, the processing shown in Fig. 85 is performed on the newly set target node.

[0591] If the processing of all child nodes has not been completed (No in S1618), the processing from step S1616 onwards is performed again. On the other hand, if the processing of all child nodes has been completed (Yes in S1618), the three-dimensional data encoding device ends the node encoding processing.

[0592] Also, in step S1613, if the number of child nodes containing three-dimensional points is less than or equal to a predetermined threshold (No in S1613), the three-dimensional data encoding device sets the encoding type to location encoding and adds the encoding type to the bit string (S1619).

[0593] Next, the three-dimensional data encoding device sets location encoding information and adds the location encoding information to the bit string. Specifically, the three-dimensional data encoding device generates a location code and adds the location encoding to the bit string (S1620). The location code includes numPoint, num_idx, and idx.

[0594] Also, in step S1611, if the target node is a leaf (Yes in S1611), the three-dimensional data encoding device sets the leaf flag to 1 and adds the leaf flag to the bit string (S1621). Also, the three-dimensional data encoding device sets a point flag (point_flag), which is information indicating whether the leaf includes a three-dimensional point, and adds the point flag to the bit string (S1622).

[0595] Next, an example of a flow of a decoding process in which the application of occupancy coding and location coding is switched will be described. Figure 85 is a flowchart of the decoding process according to this embodiment.

[0596] The three-dimensional data decoding device generates a bit string by entropy decoding the bit stream (S1631). Next, the three-dimensional data decoding device restores an octree structure by performing node decoding processing on the obtained bit string (S1632). Next, the three-dimensional data decoding device generates three-dimensional points from the restored octree structure (S1633).

[0597] 87 is a flowchart of the node decoding process (S1632). First, the three-dimensional data decoding device acquires (decodes) a leaf flag (isleaf) from the bit string (S1641). Next, the three-dimensional data decoding device determines whether the target node is a leaf based on the leaf flag (S1642).

[0598] If the target node is not a leaf (No in S1642), the three-dimensional data decoding device acquires the coding type (coding_type) from the bit string (S1643). The three-dimensional data decoding device determines whether the coding type is occupancy coding (S1644).

[0599] If the coding type is occupancy coding (Yes in S1644), the three-dimensional data decoding device acquires occupancy coding information from the bit string. Specifically, the three-dimensional data decoding device acquires an occupancy code from the bit string (S1645).

[0600] Next, the three-dimensional data decoding device sets the next target node according to the occupancy code (S1646). Specifically, the three-dimensional data decoding device sets the next target node from the unprocessed child node whose occupancy code is "1".

[0601] Next, the three-dimensional data decoding device performs node decoding processing on the newly set target node (S1647). That is, the processing shown in Fig. 87 is performed on the newly set target node.

[0602] If the processing of all child nodes has not been completed (No in S1648), the processing from step S1646 onwards is performed again. On the other hand, if the processing of all child nodes has been completed (Yes in S1648), the three-dimensional data decoding device ends the node decoding processing.

[0603] Furthermore, if the coding type is location coding in step S1644 (No in S1644), the three-dimensional data decoding device acquires location coding information from the bit string. Specifically, the three-dimensional data decoding device acquires a location code from the bit string (S1649). The location code includes numPoint, num_idx, and idx.

[0604] Also, if the target node is a leaf in step S1642 (Yes in S1642), the three-dimensional data decoding device obtains a point flag (point_flag), which is information indicating whether the leaf includes a three-dimensional point, from the bit string (S1650).

[0605] Although the present embodiment shows an example in which the encoding type is switched for each node, this is not necessarily limited to this. The encoding type may be fixed for each volume, space, or world. In this case, the three-dimensional data encoding device may add encoding type information to the header information of the volume, space, or world.

[0606] As described above, the three-dimensional data encoding device according to this embodiment generates first information that represents an N-ary tree structure (N is an integer equal to or greater than 2) of multiple three-dimensional points included in three-dimensional data using a first method (location encoding), and generates a bitstream that includes the first information. The first information includes three-dimensional point information (location code) corresponding to each of the multiple three-dimensional points. Each piece of three-dimensional point information includes an index (idx) corresponding to each of multiple layers in the N-ary tree structure. Each index indicates the sub-block to which the corresponding three-dimensional point belongs, out of the N sub-blocks belonging to the corresponding layer.

[0607] In other words, each piece of 3D point information indicates a path to the corresponding 3D point in the N-ary tree structure. Each index indicates a child node included in the path among the N child nodes belonging to the corresponding layer (node).

[0608] According to this, the three-dimensional data encoding method can generate a bitstream that allows three-dimensional points to be selectively decoded.

[0609] For example, the 3D point information (location code) includes information (num_idx) indicating the number of indexes included in the 3D point information. In other words, the information indicates the depth (number of layers) to the corresponding 3D point in the N-ary tree structure.

[0610] For example, the first information includes information (numPoint) indicating the number of pieces of 3D point information included in the first information. In other words, the information indicates the number of 3D points included in the N-ary tree structure.

[0611] For example, N is 8 and the index is 3 bits.

[0612] For example, a three-dimensional data encoding device has a first encoding mode for generating first information, and a second encoding mode for generating second information (occupancy code) that represents an N-ary tree structure in a second method (occupancy encoding) and generating a bit stream including the second information. The second information corresponds to each of a plurality of sub-blocks belonging to a plurality of layers in the N-ary tree structure, and includes a plurality of 1-bit pieces of information indicating whether a three-dimensional point exists in the corresponding sub-block.

[0613] For example, the three-dimensional data encoding device uses a first encoding mode when the number of three-dimensional points is equal to or less than a predetermined threshold, and uses a second encoding mode when the number of three-dimensional points is greater than the threshold, thereby reducing the amount of code in the bitstream.

[0614] For example, the first information and the second information include information (encoding mode information) indicating whether the information represents the N-ary tree structure in the first format or the second format.

[0615] For example, as shown in FIG. 75, the three-dimensional data encoding device uses the first encoding mode for part of the N-ary tree structure, and the second encoding mode for the other part of the N-ary tree structure.

[0616] For example, the three-dimensional data encoding device includes a processor and a memory, and the processor performs the above-mentioned processing using the memory.

[0617] Furthermore, the 3D data decoding device according to this embodiment acquires, from the bitstream, first information (location code) that represents an N-ary tree structure (N is an integer equal to or greater than 2) of multiple 3D points included in the 3D data in a first method (location coding). The first information includes 3D point information (location code) corresponding to each of the multiple 3D points. Each piece of 3D point information includes an index (idx) corresponding to each of multiple layers in the N-ary tree structure. Each index indicates the sub-block to which the corresponding 3D point belongs, out of the N sub-blocks belonging to the corresponding layer.

[0618] In other words, each piece of 3D point information indicates a path to the corresponding 3D point in the N-ary tree structure. Each index indicates a child node included in the path among the N child nodes belonging to the corresponding layer (node).

[0619] The three-dimensional data decoding device further uses the three-dimensional point information to restore three-dimensional points corresponding to the three-dimensional point information.

[0620] This allows the three-dimensional data decoding device to selectively decode three-dimensional points from a bitstream.

[0621] For example, the 3D point information (location code) includes information (num_idx) indicating the number of indexes included in the 3D point information. In other words, the information indicates the depth (number of layers) to the corresponding 3D point in the N-ary tree structure.

[0622] For example, the first information includes information (numPoint) indicating the number of pieces of 3D point information included in the first information. In other words, the information indicates the number of 3D points included in the N-ary tree structure.

[0623] For example, N is 8 and the index is 3 bits.

[0624] For example, the three-dimensional data decoding device further acquires second information (occupancy code) representing the N-ary tree structure in a second method (occupancy coding) from the bitstream. The three-dimensional data decoding device uses the second information to restore a plurality of three-dimensional points. The second information corresponds to each of a plurality of sub-blocks belonging to a plurality of layers in the N-ary tree structure, and includes a plurality of 1-bit pieces of information indicating whether a three-dimensional point exists in the corresponding sub-block.

[0625] For example, the first information and the second information include information (encoding mode information) indicating whether the information represents the N-ary tree structure in the first format or the second format.

[0626] For example, as shown in FIG. 75, a part of the N-ary tree structure is represented by the first method, and another part of the N-ary tree structure is represented by the second method.

[0627] For example, the three-dimensional data decoding device includes a processor and a memory, and the processor performs the above processing using the memory.

[0628] (Embodiment 10) In this embodiment, another example of a method for encoding a tree structure such as an octree structure will be described. Figure 88 is a diagram showing an example of a tree structure according to this embodiment. Note that Figure 88 shows an example of a quadtree structure.

[0629] A leaf that contains a 3D point is called a valid leaf, and a leaf that does not contain a 3D point is called an invalid leaf. A branch with the number of valid leaves greater than or equal to a threshold is called a dense branch. A branch with the number of valid leaves less than the threshold is called a sparse branch.

[0630] The three-dimensional data encoding device calculates the number of three-dimensional points (i.e., the number of valid leaves) contained in each branch in a certain layer of the tree structure. Figure 88 shows an example where the threshold is 5. In this example, there are two branches in layer 1. The left branch contains seven three-dimensional points, so it is determined to be a dense branch. The right branch contains two three-dimensional points, so it is determined to be a sparse branch.

[0631] FIG. 89 is a diagram showing an example of the number of valid leaves (3D points) in each branch of layer 5. The horizontal axis of FIG. 89 indicates the index, which is the identification number of the branch of layer 5. As shown in FIG. 89, a specific branch contains clearly more 3D points than other branches. Occupancy coding is more effective for such dense branches than for sparse branches.

[0632] The application methods of occupancy coding and location coding will be explained below. Figure 90 is a diagram showing the relationship between the number of 3D points (number of valid leaves) included in each branch in layer 5 and the coding method to be applied. As shown in Figure 90, the 3D data coding device applies occupancy coding to dense branches and location coding to sparse branches. This can improve coding efficiency.

[0633] Fig. 91 is a diagram showing an example of a dense branch region in LiDAR data. As shown in Fig. 91, the density of 3D points calculated from the number of 3D points included in each branch differs depending on the region.

[0634] In addition, separating dense 3D points (branches) from sparse 3D points (branches) has the following advantages. The closer to the LiDAR sensor, the higher the density of 3D points. Therefore, separating branches according to density enables division in the distance direction. This division is effective in certain applications. Furthermore, it is effective to use a method other than occupancy coding for sparse branches.

[0635] In this embodiment, the three-dimensional data encoding device separates an input three-dimensional point cloud into two or more sub-three-dimensional point clouds, and applies a different encoding method to each sub-three-dimensional point cloud.

[0636] For example, the three-dimensional data encoding device separates an input three-dimensional point cloud into a sub three-dimensional point cloud A (dense three-dimensional point cloud) including dense branches and a sub three-dimensional point cloud B (sparse three-dimensional point cloud) including sparse branches. Fig. 92 is a diagram showing an example of the sub three-dimensional point cloud A (dense three-dimensional point cloud) including dense branches separated from the tree structure shown in Fig. 88. Fig. 93 is a diagram showing an example of the sub three-dimensional point cloud B (sparse three-dimensional point cloud) including sparse branches separated from the tree structure shown in Fig. 88.

[0637] Next, the three-dimensional data encoding device encodes the sub-three-dimensional point cloud A using occupancy encoding, and encodes the sub-three-dimensional point cloud B using location encoding.

[0638] Here, an example has been shown in which different encoding methods (occupancy encoding and location encoding) are applied, but for example, a three-dimensional data encoding device may use the same encoding method for sub-three-dimensional point group A and sub-three-dimensional point group B, and may use different parameters for encoding sub-three-dimensional point group A and sub-three-dimensional point group B.

[0639] The flow of three-dimensional data encoding processing by the three-dimensional data encoding device will be described below. Fig. 94 is a flowchart of three-dimensional data encoding processing by the three-dimensional data encoding device according to this embodiment.

[0640] First, the three-dimensional data encoding device separates the input three-dimensional point cloud into sub three-dimensional point clouds (S1701). The three-dimensional data encoding device may perform this separation automatically or based on information input by the user. For example, the range of the sub three-dimensional point cloud may be specified by the user. As an example of automatic separation, for example, if the input data is LiDAR data, the three-dimensional data encoding device performs separation using distance information to each point cloud. Specifically, the three-dimensional data encoding device separates point clouds that are within a certain range from the measurement point and point clouds that are outside that range. The three-dimensional data encoding device may also perform separation using information such as important areas and unimportant areas.

[0641] Next, the three-dimensional data encoding device generates encoded data (encoded bitstream) by encoding sub three-dimensional point group A using technique A (S1702). The three-dimensional data encoding device also generates encoded data by encoding sub three-dimensional point group B using technique B (S1703). The three-dimensional data encoding device may encode sub three-dimensional point group B using technique A. In this case, the three-dimensional data encoding device encodes sub three-dimensional point group B using a parameter different from the encoding parameter used for encoding sub three-dimensional point group A. For example, this parameter may be a quantization parameter. For example, the three-dimensional data encoding device encodes sub three-dimensional point group B using a quantization parameter larger than the quantization parameter used for encoding sub three-dimensional point group A. In this case, the three-dimensional data encoding device may add information indicating the quantization parameter used for encoding each sub three-dimensional point group to the header of the encoded data of that sub three-dimensional point group.

[0642] Next, the three-dimensional data encoding device generates a bitstream by combining the encoded data obtained in step S1702 and the encoded data obtained in step S1703 (S1704).

[0643] Furthermore, the three-dimensional data encoding device may encode information for decoding each sub-three-dimensional point cloud as header information of the bit stream. For example, the three-dimensional data encoding device may encode the following information:

[0644] The header information may include information indicating the number of encoded sub-3D points, which in this example indicates 2.

[0645] The header information may include information indicating the number of 3D points included in each sub-3D point cloud and the encoding method. In this example, this information indicates the number of 3D points included in sub-3D point cloud A and the encoding method (method A) applied to sub-3D point cloud A, the number of 3D points included in sub-3D point cloud B and the encoding method (method B) applied to sub-3D point cloud B.

[0646] The header information may include information for identifying the start or end position of the encoded data for each sub-3D point cloud.

[0647] Furthermore, the three-dimensional data encoding device may encode the sub three-dimensional point cloud A and the sub three-dimensional point cloud B in parallel, or may encode the sub three-dimensional point cloud A and the sub three-dimensional point cloud B in sequence.

[0648] Furthermore, the method of separation into sub-3D point clouds is not limited to the above. For example, the three-dimensional data encoding device may change the separation method, perform encoding using each of multiple separation methods, and calculate the encoding efficiency of the encoded data obtained using each separation method. The three-dimensional data encoding device may then select the separation method with the highest encoding efficiency. For example, the three-dimensional data encoding device may separate the three-dimensional point cloud into each of multiple layers, calculate the encoding efficiency for each case, select the separation method (i.e., the layer to be separated) with the highest encoding efficiency, and generate and encode sub-3D point clouds using the selected separation method.

[0649] Furthermore, when combining encoded data, the three-dimensional data encoding device may position the encoded information of the more important sub-three-dimensional point clouds closer to the beginning of the bit stream, which allows the three-dimensional data decoding device to obtain important information simply by decoding the beginning bit stream, thereby enabling important information to be obtained quickly.

[0650] Next, the flow of three-dimensional data decoding processing by the three-dimensional data decoding device will be explained. Figure 95 is a flowchart of three-dimensional data decoding processing by the three-dimensional data decoding device according to this embodiment.

[0651] First, the three-dimensional data decoding device acquires, for example, a bitstream generated by the three-dimensional data encoding device. Next, the three-dimensional data decoding device separates the encoded data of sub three-dimensional point group A and the encoded data of sub three-dimensional point group B from the acquired bitstream (S1711). Specifically, the three-dimensional data decoding device decodes information for decoding each sub three-dimensional point group from header information of the bitstream, and uses the information to separate the encoded data of each sub three-dimensional point group.

[0652] Next, the three-dimensional data decoding device obtains sub three-dimensional point group A by decoding the encoded data of sub three-dimensional point group A using method A (S1712). Also, the three-dimensional data decoding device obtains sub three-dimensional point group B by decoding the encoded data of sub three-dimensional point group B using method B (S1713). Next, the three-dimensional data decoding device combines sub three-dimensional point group A and sub three-dimensional point group B (S1714).

[0653] The three-dimensional data decoding device may decode the sub three-dimensional point group A and the sub three-dimensional point group B in parallel, or may decode the sub three-dimensional point group A and the sub three-dimensional point group B in sequence.

[0654] Furthermore, the three-dimensional data decoding device may decode necessary sub-three-dimensional point clouds. For example, the three-dimensional data decoding device may decode sub-three-dimensional point cloud A but not sub-three-dimensional point cloud B. For example, if sub-three-dimensional point cloud A is a three-dimensional point cloud included in an important area of ​​LiDAR data, the three-dimensional data decoding device decodes the three-dimensional point cloud of the important area. Self-localization of a vehicle or the like is performed using the three-dimensional point cloud of the important area.

[0655] Next, a specific example of the encoding process according to this embodiment will be described. Fig. 96 is a flowchart of the three-dimensional data encoding process performed by the three-dimensional data encoding device according to this embodiment.

[0656] First, the three-dimensional data encoding device separates the input three-dimensional points into a sparse three-dimensional point cloud and a dense three-dimensional point cloud (S1721). Specifically, the three-dimensional data encoding device counts the number of valid leaves in a branch of a certain layer of the octree structure. The three-dimensional data encoding device sets each branch as a dense branch or a sparse branch depending on the number of valid leaves in each branch. Then, the three-dimensional data encoding device generates a sub-three-dimensional point cloud (dense three-dimensional point cloud) made up of dense branches and a sub-three-dimensional point cloud (sparse three-dimensional point cloud) made up of sparse branches.

[0657] Next, the three-dimensional data encoding device generates encoded data by encoding the sparse three-dimensional point cloud (S1722). For example, the three-dimensional data encoding device encodes the sparse three-dimensional point cloud using location encoding.

[0658] Furthermore, the three-dimensional data encoding device generates encoded data by encoding the dense three-dimensional point cloud (S1723). For example, the three-dimensional data encoding device encodes the dense three-dimensional point cloud using occupancy encoding.

[0659] Next, the three-dimensional data encoding device generates a bitstream by combining the encoded data of the sparse three-dimensional point cloud obtained in step S1722 and the encoded data of the dense three-dimensional point cloud obtained in step S1723 (S1724).

[0660] Furthermore, the three-dimensional data encoding device may encode information for decoding the sparse three-dimensional point cloud and the dense three-dimensional point cloud as header information of the bit stream. For example, the three-dimensional data encoding device may encode the following information:

[0661] The header information may include information indicating the number of encoded sub-3D point clouds, which in this example indicates 2.

[0662] The header information may include information indicating the number of 3D points contained in each sub-3D point cloud and the encoding method used for the sparse 3D point cloud (location encoding), and the number of 3D points contained in the dense 3D point cloud and the encoding method used for the dense 3D point cloud (occupancy encoding).

[0663] The header information may include information for identifying the start or end position of the encoded data for each sub-3D point cloud, in this example indicating at least one of the start and end positions of the encoded data for the sparse 3D point cloud and the start and end positions of the encoded data for the dense 3D point cloud.

[0664] The three-dimensional data encoding device may also encode the sparse three-dimensional point cloud and the dense three-dimensional point cloud in parallel, or may encode the sparse three-dimensional point cloud and the dense three-dimensional point cloud in sequence.

[0665] Next, a specific example of three-dimensional data decoding processing will be described. Fig. 97 is a flowchart of three-dimensional data decoding processing performed by the three-dimensional data decoding device according to this embodiment.

[0666] First, the three-dimensional data decoding device acquires, for example, a bitstream generated by the three-dimensional data encoding device. Next, the three-dimensional data decoding device separates the acquired bitstream into coded data of a sparse three-dimensional point cloud and coded data of a dense three-dimensional point cloud (S1731). Specifically, the three-dimensional data decoding device decodes information for decoding each sub-three-dimensional point cloud from header information of the bitstream, and separates the coded data of each sub-three-dimensional point cloud using the decoded information. In this example, the three-dimensional data decoding device separates the coded data of the sparse three-dimensional point cloud and the dense three-dimensional point cloud from the bitstream using the header information.

[0667] Next, the three-dimensional data decoding device obtains a sparse three-dimensional point cloud by decoding the coded data of the sparse three-dimensional point cloud (S1732). For example, the three-dimensional data decoding device decodes the sparse three-dimensional point cloud using location decoding for decoding location-coded coded data.

[0668] Furthermore, the three-dimensional data decoding device obtains a dense three-dimensional point cloud by decoding the encoded data of the dense three-dimensional point cloud (S1733). For example, the three-dimensional data decoding device decodes the dense three-dimensional point cloud using occupancy decoding for decoding occupancy-encoded encoded data.

[0669] Next, the three-dimensional data decoding device combines the sparse three-dimensional point group obtained in step S1732 with the dense three-dimensional point group obtained in step S1733 (S1734).

[0670] The three-dimensional data decoding device may decode the sparse three-dimensional point cloud and the dense three-dimensional point cloud in parallel, or may decode the sparse three-dimensional point cloud and the dense three-dimensional point cloud in sequence.

[0671] The three-dimensional data decoding device may also decode some necessary sub-three-dimensional point clouds. For example, the three-dimensional data decoding device may decode the dense three-dimensional point cloud and not decode the sparse three-dimensional data. For example, if the dense three-dimensional point cloud is a three-dimensional point cloud included in an important area of ​​LiDAR data, the three-dimensional data decoding device decodes the three-dimensional point cloud of the important area. Self-localization of a vehicle, etc., is performed using the three-dimensional point cloud of the important area.

[0672] 98 is a flowchart of the encoding process according to this embodiment. First, the three-dimensional data encoding device separates an input three-dimensional point cloud into a sparse three-dimensional point cloud and a dense three-dimensional point cloud, thereby generating the sparse three-dimensional point cloud and the dense three-dimensional point cloud (S1741).

[0673] Next, the three-dimensional data encoding device generates encoded data by encoding the dense three-dimensional point cloud (S1742). The three-dimensional data encoding device also generates encoded data by encoding the sparse three-dimensional point cloud (S1743). Finally, the three-dimensional data encoding device generates a bitstream by combining the encoded data of the sparse three-dimensional point cloud obtained in step S1742 and the encoded data of the dense three-dimensional point cloud obtained in step S1743 (S1744).

[0674] 99 is a flowchart of the decoding process according to this embodiment. First, the three-dimensional data decoding device extracts coded data of a dense three-dimensional point cloud and coded data of a sparse three-dimensional point cloud from the bit stream (S1751). Next, the three-dimensional data decoding device obtains decoded data of the dense three-dimensional point cloud by decoding the coded data of the dense three-dimensional point cloud (S1752). Furthermore, the three-dimensional data decoding device obtains decoded data of the sparse three-dimensional point cloud by decoding the coded data of the sparse three-dimensional point cloud (S1753). Next, the three-dimensional data decoding device generates a three-dimensional point cloud by combining the decoded data of the dense three-dimensional point cloud obtained in step S1752 and the decoded data of the sparse three-dimensional point cloud obtained in step S1753 (S1754).

[0675] The three-dimensional data encoding device and the three-dimensional data decoding device may encode or decode either the dense three-dimensional point cloud or the sparse three-dimensional point cloud first. Furthermore, the encoding process or the decoding process may be performed in parallel by multiple processors or the like.

[0676] The three-dimensional data encoding device may also encode one of a dense three-dimensional point cloud and a sparse three-dimensional point cloud. For example, if important information is contained in the dense three-dimensional point cloud, the three-dimensional data encoding device extracts a dense three-dimensional point cloud and a sparse three-dimensional point cloud from the input three-dimensional point cloud, encodes the dense three-dimensional point cloud, and does not encode the sparse three-dimensional point cloud. This allows the three-dimensional data encoding device to add important information to the stream while reducing the bit amount. For example, between a server and a client, when the client requests the server to transmit three-dimensional point cloud information about the client's surroundings, the server encodes the important information about the client's surroundings as a dense three-dimensional point cloud and transmits it to the client. This allows the server to transmit the information requested by the client while reducing network bandwidth.

[0677] Furthermore, the three-dimensional data decoding device may decode either the dense three-dimensional point cloud or the sparse three-dimensional point cloud. For example, if important information is contained in the dense three-dimensional point cloud, the three-dimensional data decoding device decodes the dense three-dimensional point cloud but does not decode the sparse three-dimensional point cloud. This allows the three-dimensional data decoding device to acquire necessary information while reducing the processing load of the decoding process.

[0678] Fig. 100 is a flowchart of the 3D point separation process (S1741) shown in Fig. 98. First, the 3D data encoding device sets a layer L and a threshold TH (S1761). The 3D data encoding device may add information indicating the set layer L and threshold TH to the bitstream. In other words, the 3D data encoding device may generate a bitstream including information indicating the set layer L and threshold TH.

[0679] Next, the three-dimensional data encoding device moves the position of the processing target from the root of the octtree to the first branch of layer L. In other words, the three-dimensional data encoding device selects the first branch of layer L as the branch to be processed (S1762).

[0680] Next, the three-dimensional data encoding device counts the number of valid leaves of the branch to be processed in layer L (S1763). If the number of valid leaves of the branch to be processed is greater than threshold TH (Yes in S1764), the three-dimensional data encoding device registers the branch to be processed as a dense branch in the dense three-dimensional point cloud (S1765). On the other hand, if the number of valid leaves of the branch to be processed is equal to or less than threshold TH (No in S1764), the three-dimensional data encoding device registers the branch to be processed as a sparse branch in the sparse three-dimensional point cloud (S1766).

[0681] If processing of all branches of layer L has not been completed (No in S1767), the three-dimensional data encoding device moves the position of the processing target to the next branch of layer L. In other words, the three-dimensional data encoding device selects the next branch of layer L as the branch to be processed (S1768). Then, the three-dimensional data encoding device performs the processing from step S1763 onwards for the selected next branch to be processed.

[0682] The above process is repeated until all edges in layer L have been processed (Yes in S1767).

[0683] Although the layer L and the threshold TH are set in advance in the above description, this is not necessarily limited to this. For example, the three-dimensional data encoding device may set multiple patterns of pairs of layer L and threshold TH, generate dense and sparse three-dimensional point clouds using each pair, and encode them. The three-dimensional data encoding device finally encodes the dense and sparse three-dimensional point clouds using the pair of layer L and threshold TH that has the highest encoding efficiency for the generated encoded data among the multiple pairs. This improves encoding efficiency. The three-dimensional data encoding device may also calculate the layer L and threshold TH, for example. For example, the three-dimensional data encoding device may set the layer L to half the value of the maximum number of layers included in the tree structure. The three-dimensional data encoding device may also set the threshold TH to half the value of the total number of three-dimensional points included in the tree structure.

[0684] In the above description, an example was described in which an input 3D point cloud is classified into two types, a dense 3D point cloud and a sparse 3D point cloud. However, the 3D data encoding device may classify the input 3D point cloud into three or more types of 3D point clouds. For example, if the number of valid leaves on the branch to be processed is equal to or greater than a threshold value TH1, the 3D data encoding device classifies the branch to be processed into a first dense 3D point cloud. If the number of valid leaves on the branch to be processed is less than the first threshold value TH1 and equal to or greater than a second threshold value TH2, the 3D data encoding device classifies the branch to be processed into a second dense 3D point cloud. If the number of valid leaves on the branch to be processed is less than the second threshold value TH2 and equal to or greater than a third threshold value TH3, the 3D data encoding device classifies the branch to be processed into a first sparse 3D point cloud. If the number of valid leaves on the branch to be processed is less than the threshold value TH3, the 3D data encoding device classifies the branch to be processed into a second sparse 3D point cloud.

[0685] An example of the syntax of coded data of a 3D point group according to this embodiment will be described below. Fig. 101 is a diagram showing this syntax example. pc_header() is, for example, header information of multiple input 3D points.

[0686] num_sub_pc shown in Figure 101 indicates the number of sub 3D point clouds. numPoint[i] indicates the number of 3D points contained in the i-th sub 3D point cloud. coding_type[i] is coding type information that indicates the coding type (coding method) applied to the i-th sub 3D point cloud. For example, coding_type=00 indicates that location coding is applied. coding_type=01 indicates that occupancy coding is applied. coding_type=10 or 11 indicates that another coding method is applied.

[0687] data_sub_cloud() is the coded data of the i-th sub 3D point cloud. coding_type_00_data is coded data to which a coding type of 00 has been applied, for example, location coding. coding_type_01_data is coded data to which a coding type of 01 has been applied, for example, occupancy coding.

[0688] "end_of_data" is termination information indicating the end of the encoded data. For example, a fixed bit string that is not used in the encoded data is assigned to this "end_of_data." This allows the three-dimensional data decoding device to, for example, search the bit stream for the "end_of_data" bit string, thereby skipping the decoding process of encoded data that does not need to be decoded.

[0689] The three-dimensional data encoding device may entropy-encode the encoded data generated by the above method. For example, the three-dimensional data encoding device may binarize each value and then perform computational encoding.

[0690] Furthermore, in this embodiment, examples of a quadtree structure or an octaltree structure are shown, but this is not necessarily limited to these, and the above method may also be applied to an N-ary tree (N is an integer equal to or greater than 2) such as a binary tree or a hexadecimal tree, or other tree structures.

[0691] [Variation] In the above description, as shown in Figures 93 and 94, a tree structure including dense branches and their upper layers (a tree structure from the root of the entire tree structure to the root of the dense branches) is coded, and a tree structure including sparse branches and their upper layers (a tree structure from the root of the entire tree structure to the root of the sparse branches) is coded. In this modification, the three-dimensional data coding device separates the dense branches from the sparse branches and codes the dense branches and the sparse branches. In other words, the tree structure to be coded does not include the upper layer tree structure. For example, the three-dimensional data coding device applies occupancy coding to the dense branches and location coding to the sparse branches.

[0692] Fig. 102 is a diagram showing an example of dense branches separated from the tree structure shown in Fig. 88. Fig. 103 is a diagram showing an example of sparse branches separated from the tree structure shown in Fig. 88. In this modification, the tree structures shown in Fig. 102 and Fig. 103 are each encoded.

[0693] Furthermore, the three-dimensional data en...

Claims

1. Encoding information of a target node included in an N-ary tree structure of a plurality of three-dimensional points included in the three-dimensional data using only information of nodes that can be referenced from among a plurality of adjacent nodes that are spatially adjacent to the target node; Encode the parameters, generating a bitstream including the encoded information of the target node and the parameters; If the parameter includes a first value, the referable nodes are only sibling nodes of the target node. Three-dimensional data encoding method.

2. The value of the parameter indicates the referenceable node. The three-dimensional data encoding method according to claim 1 .

3. The parent node of the target node and the sibling node is the same The three-dimensional data encoding method according to claim 2 .

4. If the parameter includes a second value, the referenceable nodes include the sibling node and other nodes. The three-dimensional data encoding method according to claim 3.

5. When the parameter includes a second value, the referable nodes include sibling nodes of the target node and nodes other than the sibling nodes. The three-dimensional data encoding method according to claim 1 .

6. The parameter is included in header information of the bitstream. The three-dimensional data encoding method according to claim 1 .

7. Obtaining parameters from the bitstream, Decoding information of a target node included in an N-ary tree structure of a plurality of three-dimensional points included in the three-dimensional data (N is an integer of 2 or more), using only information of nodes that can be referenced among a plurality of adjacent nodes that are spatially adjacent to the target node; If the parameter includes a first value, the referable nodes are only sibling nodes of the target node. Three-dimensional data decoding method.

8. The value of the parameter indicates the referenceable node. The three-dimensional data decoding method according to claim 7.

9. The parent node of the target node and the sibling node is the same The three-dimensional data decoding method according to claim 8.

10. If the parameter includes a second value, the referenceable nodes include the sibling node and other nodes. The three-dimensional data decoding method according to claim 9.

11. When the parameter includes a second value, the referable nodes include sibling nodes of the target node and nodes other than the sibling nodes. The three-dimensional data decoding method according to claim 7. The three-dimensional data decoding method according to claim 10.

12. The parameter is included in the header of the bitstream. The three-dimensional data decoding method according to claim 7.

13. a processor; a memory; The processor uses the memory to: Encoding information of a target node included in an N-ary tree structure of a plurality of three-dimensional points included in the three-dimensional data using only information of nodes that can be referenced from among a plurality of adjacent nodes that are spatially adjacent to the target node; Encode the parameters, generating a bitstream including the encoded information of the target node and the parameters; If the parameter includes a first value, the referable nodes are only sibling nodes of the target node. Three-dimensional data encoding device.

14. a processor; a memory; The processor uses the memory to: Get the parameters from the bitstream, Decoding information of a target node included in an N-ary tree structure of a plurality of three-dimensional points included in the three-dimensional data (N is an integer of 2 or more), using only information of nodes that can be referenced among a plurality of adjacent nodes that are spatially adjacent to the target node; If the parameter includes a first value, the referable nodes are only sibling nodes of the target node. Three-dimensional data decoding device.

Citation Information

Patent Citations

  • Node structure for representing three-dimensional object based on depth image

    JP2004005373A

  • Generating method of adaptive notation system of base-2n tree, and device and method for encoding / decoding three-dimensional volume data utilizing same

    JP2005259139A

  • Method and apparatus for processing a bitstream representing a 3D model

    JP2015513719A

  • JPP7521088B

  • Adaptive 2n-ary tree generating method, and method and apparatus for encoding and decoding 3D volume data using it

    US20050195191A1