Three-dimensional data encoding method, three-dimensional data decoding method, three-dimensional data encoding device, and three-dimensional data decoding device
Patent Information
- Application Number
- JP2025189156
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2018-07-09
- Filing Date
- 2025-11-10
- Publication Date
- 2026-09-18
- Estimated Expiration
- 2039-06-14
AI Technical Summary
【0011】 本開示は、処理量を低減できる三次元データ符号化方法、三次元データ復号方法、三次元データ符号化装置又は三次元データ復号装置を提供できる。
Smart Images

Figure 0007920416000001 
Figure 0007920416000002 
Figure 0007920416000003
Abstract
Description
[Technical Field]
[0001] The present disclosure relates to a three-dimensional data encoding method, a three-dimensional data decoding method, a three-dimensional data encoding device, and a three-dimensional data decoding device. [Background Art]
[0002] In the future, it is expected that devices or services utilizing three-dimensional data will become widespread in a wide range of fields including computer vision for autonomous operation of automobiles or robots, map information, monitoring, infrastructure inspection, and video distribution. Three-dimensional data is acquired by various methods such as distance sensors like range finders, stereo cameras, or combinations of multiple monocular cameras.
[0003] As one method for representing three-dimensional data, there is a representation method called point cloud that represents the shape of a three-dimensional structure by a point cloud in three-dimensional space. In a point cloud, the positions and colors of the point cloud are stored. While point cloud is expected to become mainstream as a representation method for three-dimensional data, the data amount of a point cloud is extremely large. Therefore, in the storage or transmission of three-dimensional data, compression of the data amount through encoding is essential, similar to two-dimensional moving images (for example, MPEG-4 AVC or HEVC standardized by MPEG).
[0004] Additionally, compression of point clouds is partially supported by public libraries (Point Cloud Library) and the like that perform point cloud-related processing.
[0005] Furthermore, a technique for searching for and displaying facilities located around a vehicle using three-dimensional map data is known (see, for example, Patent Document 1). [Prior Art Literature] [Patent Literature]
[0006] [Patent Document 1] International Publication No. 2014 / 020663 [Overview of the Initiative] [Problems that the invention aims to solve]
[0007] It is desirable to reduce the processing load in the encoding and decoding of three-dimensional data.
[0008] The purpose of this disclosure is to provide a three-dimensional data encoding method, a three-dimensional data decoding method, a three-dimensional data encoding device, or a three-dimensional data decoding device that can reduce the amount of processing required. [Means for solving the problem]
[0009] A three-dimensional data encoding method according to one aspect of the present disclosure generates a tree structure of a plurality of three-dimensional points contained in three-dimensional data, generates a parameter indicating a referable node, determines an adjacent occupancy pattern from among a plurality of adjacent occupancy patterns based on the occupancy status of adjacent nodes of a target node, determines a group from among a plurality of groups corresponding to the determined adjacent occupancy pattern, encodes the target node using the information of the determined group, each of the plurality of groups corresponds to one or more adjacent occupancy patterns, the number of the plurality of groups that can be determined differs depending on the value indicated by the parameter, the plurality of groups includes a first group and a second group, each adjacent occupancy pattern corresponding to the first group indicates a first number of occupied nodes, each adjacent occupancy pattern corresponding to the second group indicates a second number of occupied nodes greater than the first number, and the number of adjacent occupancy patterns corresponding to the first group is less than the number of adjacent occupancy patterns corresponding to the second group.
[0010] A three-dimensional data decoding method according to one aspect of the present disclosure involves obtaining a tree structure of multiple three-dimensional points contained in three-dimensional data, obtaining a parameter indicating a referable node, determining an adjacent occupancy pattern from among multiple adjacent occupancy patterns based on the occupancy status of adjacent nodes of the target node, determining a group from among multiple groups corresponding to the determined adjacent occupancy pattern, decoding the target node using the information of the determined group, wherein each of the multiple groups corresponds to one or more adjacent occupancy patterns, the number of determinable multiple groups differs depending on the value indicated by the parameter, the multiple groups include a first group and a second group, each adjacent occupancy pattern corresponding to the first group indicates a first number of occupant nodes, each adjacent occupancy pattern corresponding to the second group indicates a second number of occupant nodes greater than the first number, and the number of adjacent occupancy patterns corresponding to the first group is less than the number of adjacent occupancy patterns corresponding to the second group. [Effects of the Invention]
[0011] This disclosure provides a three-dimensional data encoding method, a three-dimensional data decoding method, a three-dimensional data encoding device, or a three-dimensional data decoding device that can reduce the amount of processing required. [Brief explanation of the drawing]
[0012] [Figure 1] Figure 1 is a diagram showing the structure of encoded three-dimensional data according to Embodiment 1. [Figure 2] Figure 2 shows an example of a prediction structure between SPCs belonging to the lowest layer of GOS according to Embodiment 1. [Figure 3] Figure 3 shows an example of a prediction structure between layers according to Embodiment 1. [Figure 4] Figure 4 shows an example of the encoding order of GOS according to Embodiment 1. [Figure 5] Figure 5 shows an example of the encoding order of GOS according to Embodiment 1. [Figure 6] Figure 6 is a block diagram of a three-dimensional data encoding device according to Embodiment 1. [Figure 7] FIG. 7 is a flowchart of an encoding process according to the first embodiment. [Figure 8] FIG. 8 is a block diagram of a three-dimensional data decoding apparatus according to the first embodiment. [Figure 9] FIG. 9 is a flowchart of a decoding process according to the first embodiment. [Figure 10] FIG. 10 is a diagram showing an example of meta information according to the first embodiment. [Figure 11] FIG. 11 is a diagram showing a configuration example of an SWLD according to the second embodiment. [Figure 12] FIG. 12 is a diagram showing an operation example of a server and a client according to the second embodiment. [Figure 13] FIG. 13 is a diagram showing an operation example of a server and a client according to the second embodiment. [Figure 14] FIG. 14 is a diagram showing an operation example of a server and a client according to the second embodiment. [Figure 15] FIG. 15 is a diagram showing an operation example of a server and a client according to the second embodiment. [Figure 16] FIG. 16 is a block diagram of a three-dimensional data encoding apparatus according to the second embodiment. [Figure 17] FIG. 17 is a flowchart of an encoding process according to the second embodiment. [Figure 18] FIG. 18 is a block diagram of a three-dimensional data decoding apparatus according to the second embodiment. [Figure 19] FIG. 19 is a flowchart of a decoding process according to the second embodiment. [Figure 20] FIG. 20 is a diagram showing a configuration example of a WLD according to the second embodiment. [Figure 21] FIG. 21 is a diagram showing an example of an octree structure of a WLD according to the second embodiment. [Figure 22] FIG. 22 is a diagram showing a configuration example of an SWLD according to the second embodiment. [Figure 23] FIG. 23 is a diagram showing an example of an octree structure of an SWLD according to the second embodiment. [Figure 24] Figure 24 is a block diagram of a three-dimensional data creation device according to Embodiment 3. [Figure 25] Figure 25 is a block diagram of a three-dimensional data transmission device according to Embodiment 3. [Figure 26] Figure 26 is a block diagram of a three-dimensional information processing device according to Embodiment 4. [Figure 27] Figure 27 is a block diagram of a three-dimensional data creation device according to Embodiment 5. [Figure 28] Figure 28 is a diagram showing the configuration of the system according to Embodiment 6. [Figure 29] Figure 29 is a block diagram of the client device according to Embodiment 6. [Figure 30] Figure 30 is a block diagram of the server according to Embodiment 6. [Figure 31] Figure 31 is a flowchart of the three-dimensional data creation process by the client device according to Embodiment 6. [Figure 32] Figure 32 is a flowchart of the sensor information transmission process by the client device according to Embodiment 6. [Figure 33] Figure 33 is a flowchart of the three-dimensional data creation process performed by the server according to Embodiment 6. [Figure 34] Figure 34 is a flowchart of the three-dimensional map transmission process by the server according to Embodiment 6. [Figure 35] Figure 35 shows a modified configuration of the system according to Embodiment 6. [Figure 36] Figure 36 is a diagram showing the configuration of the server and client device according to Embodiment 6. [Figure 37] Figure 37 is a block diagram of a three-dimensional data encoding device according to Embodiment 7. [Figure 38] Figure 38 shows an example of the predicted residual according to Embodiment 7. [Figure 39] Figure 39 shows an example of a volume according to Embodiment 7. [Figure 40]Figure 40 shows an example of an octree representation of a volume according to Embodiment 7. [Figure 41] Figure 41 shows an example of a bit sequence of a volume according to Embodiment 7. [Figure 42] Figure 42 shows an example of an octree representation of a volume according to Embodiment 7. [Figure 43] Figure 43 shows an example of a volume according to Embodiment 7. [Figure 44] Figure 44 is a diagram illustrating the intra-prediction process according to Embodiment 7. [Figure 45] Figure 45 is a diagram illustrating the rotation and translation processing according to Embodiment 7. [Figure 46] Figure 46 shows an example of the syntax for the RT application flag and RT information according to Embodiment 7. [Figure 47] Figure 47 is a diagram illustrating the interpretation prediction process according to Embodiment 7. [Figure 48] Figure 48 is a block diagram of a three-dimensional data decoding device according to Embodiment 7. [Figure 49] Figure 49 is a flowchart of the three-dimensional data encoding process performed by the three-dimensional data encoding device according to Embodiment 7. [Figure 50] Figure 50 is a flowchart of the three-dimensional data decoding process by the three-dimensional data decoding device according to Embodiment 7. [Figure 51] Figure 51 shows an example of a wood structure according to Embodiment 8. [Figure 52] Figure 52 shows an example of an occupancy code according to Embodiment 8. [Figure 53] Figure 53 is a schematic diagram showing the operation of the three-dimensional data encoding device according to Embodiment 8. [Figure 54] Figure 54 is a diagram showing an example of geometric information according to Embodiment 8. [Figure 55] Figure 55 shows an example of selecting an encoding table using geometric information according to Embodiment 8. [Figure 56]Figure 56 shows an example of selecting an encoding table using structural information according to Embodiment 8. [Figure 57] Figure 57 shows an example of selecting an encoding table using attribute information according to Embodiment 8. [Figure 58] Figure 58 shows an example of selecting an encoding table using attribute information according to Embodiment 8. [Figure 59] Figure 59 shows an example of the bitstream configuration according to Embodiment 8. [Figure 60] Figure 60 shows an example of an encoding table according to Embodiment 8. [Figure 61] Figure 61 shows an example of an encoding table according to Embodiment 8. [Figure 62] Figure 62 shows an example of the bitstream configuration according to Embodiment 8. [Figure 63] Figure 63 shows an example of an encoding table according to Embodiment 8. [Figure 64] Figure 64 shows an example of an encoding table according to Embodiment 8. [Figure 65] Figure 65 shows an example of the bit number of an occupancy code according to Embodiment 8. [Figure 66] Figure 66 is a flowchart of the encoding process using geometric information according to Embodiment 8. [Figure 67] Figure 67 is a flowchart of the decoding process using geometric information according to Embodiment 8. [Figure 68] Figure 68 is a flowchart of the encoding process using structural information according to Embodiment 8. [Figure 69] Figure 69 is a flowchart of the decoding process using structural information according to Embodiment 8. [Figure 70] Figure 70 is a flowchart of the encoding process using attribute information according to Embodiment 8. [Figure 71] Figure 71 is a flowchart of the decoding process using attribute information according to Embodiment 8. [Figure 72] Figure 72 is a flowchart of the coding table selection process using geometric information according to Embodiment 8. [Figure 73] Figure 73 is a flowchart of the coding table selection process using structural information according to Embodiment 8. [Figure 74] Figure 74 is a flowchart of the coding table selection process using attribute information according to Embodiment 8. [Figure 75] Figure 75 is a block diagram of a three-dimensional data encoding device according to Embodiment 8. [Figure 76] Figure 76 is a block diagram of a three-dimensional data decoding device according to Embodiment 8. [Figure 77] Figure 77 shows the reference relationships in an octave tree structure according to Embodiment 9. [Figure 78] Figure 78 is a diagram showing the reference relationship in the spatial domain according to Embodiment 9. [Figure 79] Figure 79 shows an example of an adjacent reference node according to Embodiment 9. [Figure 80] Figure 80 is a diagram showing the relationship between the parent node and the node according to Embodiment 9. [Figure 81] Figure 81 shows an example of the occupancy code of a parent node according to Embodiment 9. [Figure 82] Figure 82 is a block diagram of a three-dimensional data encoding device according to Embodiment 9. [Figure 83] Figure 83 is a block diagram of a three-dimensional data decoding device according to Embodiment 9. [Figure 84] Figure 84 is a flowchart of the three-dimensional data encoding process according to Embodiment 9. [Figure 85] Figure 85 is a flowchart of the three-dimensional data decoding process according to Embodiment 9. [Figure 86] Figure 86 shows an example of switching the encoding table according to Embodiment 9. [Figure 87] Figure 87 is a diagram showing the reference relationship in the spatial region according to Modification 1 of Embodiment 9. [Figure 88] Figure 88 shows an example of the syntax of header information according to Modification 1 of Embodiment 9. [Figure 89] Figure 89 shows an example of the syntax of header information according to Modification 1 of Embodiment 9. [Figure 90] Figure 90 shows an example of an adjacent reference node according to a modified example 2 of Embodiment 9. [Figure 91] Figure 91 shows an example of a target node and adjacent nodes according to a modified example 2 of Embodiment 9. [Figure 92] Figure 92 shows the reference relationships in an octave tree structure according to a modified example 3 of Embodiment 9. [Figure 93] Figure 93 is a diagram showing the reference relationship in the spatial region according to the modified example 3 of Embodiment 9. [Figure 94] Figure 94 shows an example of translation according to Embodiment 10. [Figure 95] Figure 95 shows an example of rotation according to Embodiment 10. [Figure 96] Figure 96 shows an example of horizontal or vertical orientation according to Embodiment 10. [Figure 97] Figure 97 shows an example of an adjacent surface according to Embodiment 10. [Figure 98] Figure 98 shows an example of translation according to Embodiment 10. [Figure 99] Figure 99 shows an example of x-axis rotation according to Embodiment 10. [Figure 100] Figure 100 shows an example of y-axis rotation according to Embodiment 10. [Figure 101] Figure 101 shows an example of z-axis rotation according to Embodiment 10. [Figure 102] Figure 102 shows an example of horizontal or vertical orientation according to Embodiment 10. [Figure 103] Figure 103 shows an example of an adjacent surface according to Embodiment 10. [Figure 104]Figure 104 shows an example of grouping adjacent occupancy patterns according to Embodiment 10. [Figure 105] Figure 105 shows an example of grouping adjacent occupancy patterns according to Embodiment 10. [Figure 106] Figure 106 shows an example of a conversion table according to Embodiment 10. [Figure 107] Figure 107 shows an example of a conversion table according to Embodiment 10. [Figure 108] Figure 108 is a diagram showing an overview of the mapping process according to Embodiment 10. [Figure 109] Figure 109 is a diagram showing an overview of the mapping process according to Embodiment 10. [Figure 110] Figure 110 is a block diagram of a three-dimensional data encoding device according to Embodiment 10. [Figure 111] Figure 111 is a block diagram of a three-dimensional data decoding device according to Embodiment 10. [Figure 112] Figure 112 is a flowchart of the three-dimensional data encoding process according to Embodiment 10. [Figure 113] Figure 113 is a flowchart of the three-dimensional data decoding process according to Embodiment 10. [Figure 114] Figure 114 is a flowchart of the three-dimensional data encoding process according to Embodiment 10. [Figure 115] Figure 115 is a flowchart of the three-dimensional data decoding process according to Embodiment 10. [Figure 116] Figure 116 is a flowchart of the three-dimensional data encoding process according to Embodiment 10. [Figure 117] Figure 117 is a flowchart of the three-dimensional data decoding process according to Embodiment 10. [Figure 118] Figure 118 shows an example of grouping adjacent occupancy patterns according to Embodiment 11. [Figure 119] Figure 119 is a flowchart of the coding table switching process according to Embodiment 11. [Figure 120] Figure 120 is a flowchart of the three-dimensional data encoding process according to Embodiment 11. [Figure 121] Figure 121 is a flowchart of the three-dimensional data decoding process according to Embodiment 11. [Figure 122] Figure 122 is a diagram illustrating the redundant coding table according to Embodiment 12. [Figure 123] Figure 123 shows static and dynamic coding tables according to Embodiment 12. [Figure 124] Figure 124 shows an example of a target node according to Embodiment 12. [Figure 125] Figure 125 is a diagram showing the operation in Case 1 according to Embodiment 12. [Figure 126] Figure 126 is a diagram showing the operation in Case 2 according to Embodiment 12. [Figure 127] Figure 127 shows an example of an adjacent node according to Embodiment 12. [Figure 128] Figure 128 shows a specific example of the number of redundant tables according to Embodiment 12. [Figure 129] Figure 129 is a diagram showing a specific example of the number of redundant tables according to Embodiment 12. [Figure 130] Figure 130 shows a specific example of a redundant table according to Embodiment 12. [Figure 131] Figure 131 shows a specific example of a redundant table according to Embodiment 12. [Figure 132] Figure 132 is a diagram showing the operation of not deleting redundant tables according to Embodiment 12. [Figure 133] Figure 133 shows an example of an encoded table with redundant tables removed according to Embodiment 12. [Figure 134] Figure 134 shows an example of the number of encoding tables when redundant tables according to Embodiment 12 are deleted. [Figure 135]Figure 135 shows the process for generating a dynamically sized coding table according to Embodiment 12. [Figure 136] Figure 136 shows an example of the table size according to Embodiment 12. [Figure 137] Figure 137 is a flowchart of the process for generating a dynamically sized coding table according to Embodiment 12. [Figure 138] Figure 138 shows an example of the table size according to Embodiment 12. [Figure 139] Figure 139 is a flowchart of the coding table switching process according to Embodiment 12. [Modes for carrying out the invention]
[0013] A three-dimensional data encoding method according to one aspect of the present disclosure encodes a target node included in an N (where N is an integer of 2 or more) subtree structure of a plurality of three-dimensional points included in three-dimensional data, encoding a first flag indicating whether the target node and its parent node refer to another node different from the target node; if the first flag indicates that the target node refers to another node, an encoding table is selected from N encoding tables according to the occupation status of the adjacent nodes of the target node, and the information of the target node is arithmetically encoded using the selected encoding table; if the first flag indicates that the target node does not refer to another node, an encoding table is selected from M encoding tables different from the N, according to the occupation status of the adjacent nodes of the target node, and the information of the target node is arithmetically encoded using the selected encoding table.
[0014] This approach reduces the number of encoding tables, thereby lowering the processing load. Furthermore, by changing the number of encoding tables depending on whether the target node and its parent node refer to other nodes that are different, the encoding tables can be configured appropriately, thus reducing the processing load while suppressing a decrease in encoding efficiency.
[0015] For example, the N items may be greater than the M items.
[0016] For example, when selecting an encoding table from the M encoding tables, a correspondence table showing the correspondence between the M encoding tables and L occupation patterns (more than M) that indicate the occupation status of the adjacent nodes may be referenced, and an encoding table may be selected from the M encoding tables according to the occupation status of the adjacent nodes.
[0017] For example, when selecting an encoding table from the M encoding tables, the encoding table may be selected from the M encoding tables according to the occupation status of the adjacent node by referring to (i) a first correspondence table showing the correspondence between L occupation patterns indicating the occupation status of the adjacent node and I encoding tables (fewer than L), and (ii) a second correspondence table showing the correspondence between the I encoding tables and M encoding tables (fewer than I).
[0018] For example, when the first flag indicates that the other node is not referenced, the occupation state of the adjacent node is one of several occupation patterns, which are a combination of the position of the target node in the parent node and the occupation states of the three adjacent nodes in the parent node. In one of the several occupation patterns, one of the three adjacent nodes is in an occupation state, and the same coding table from the M coding tables may be assigned to the several occupation patterns in which the occupied adjacent node is adjacent to the target node in a direction horizontal to the xy plane.
[0019] For example, when the first flag indicates that the other node is not referenced, the occupation state of the adjacent node is a plurality of occupation patterns represented by a combination of the position of the target node in the parent node and the occupation states of the three adjacent nodes in the parent node, and the same coding table from the M coding tables may be assigned to a plurality of occupation patterns in which one of the three adjacent nodes is in an occupation state, and the adjacent node in that occupation state is adjacent to the target node in a direction perpendicular to the xy plane.
[0020] For example, when the first flag indicates that the neighboring node does not refer to the other node, the occupation state of the neighboring node is represented by a plurality of occupation patterns, which are combinations of the position of the target node in the parent node and the occupation states of the three neighboring nodes in the parent node. The same coding table from the M coding tables may be assigned to a plurality of occupation patterns in which two of the three neighboring nodes are in an occupation state, and the plane consisting of the two occupied neighboring nodes and the target node is horizontal to the xy plane.
[0021] For example, when the first flag indicates that the neighboring node does not refer to the other node, the occupation state of the neighboring node is represented by a plurality of occupation patterns, which are combinations of the position of the target node in the parent node and the occupation states of the three neighboring nodes in the parent node. The same coding table from the M coding tables may be assigned to a plurality of occupation patterns in which two of the three neighboring nodes are in an occupation state, and the plane consisting of the two occupied neighboring nodes and the target node is perpendicular to the xy plane.
[0022] A three-dimensional data decoding method according to one aspect of the present disclosure involves decoding a target node included in an N (where N is an integer of 2 or more) subtree structure of multiple three-dimensional points included in three-dimensional data, decoding a first flag indicating whether the target node and its parent node refer to another node different from the target node, and if the first flag indicates that the target node refers to another node, selecting an encoding table from N encoding tables according to the occupation status of the adjacent nodes of the target node, and arithmetically decoding the information of the target node using the selected encoding table, and if the first flag indicates that the target node does not refer to another node, selecting an encoding table from M encoding tables different from the N, according to the occupation status of the adjacent nodes of the target node, and arithmetically decoding the information of the target node using the selected encoding table.
[0023] This approach reduces the number of encoding tables, thereby lowering the processing load. Furthermore, by changing the number of encoding tables depending on whether the target node and its parent node refer to other nodes that are different, the encoding tables can be configured appropriately, thus reducing the processing load while suppressing a decrease in encoding efficiency.
[0024] For example, the N items may be greater than the M items.
[0025] For example, when selecting an encoding table from the M encoding tables, a correspondence table showing the correspondence between the M encoding tables and L occupation patterns (more than M) that indicate the occupation status of adjacent nodes, may be referenced, and an encoding table may be selected from the M encoding tables according to the occupation status of the adjacent nodes of the target node.
[0026] For example, when selecting an encoding table from the M encoding tables, the encoding table may be selected from the M encoding tables according to the occupation status of the neighboring nodes of the target node by referring to (i) a first correspondence table showing the correspondence between L occupation patterns indicating the occupation status of neighboring nodes and I encoding tables (fewer than L), and (ii) a second correspondence table showing the correspondence between the I encoding tables and the M encoding tables (fewer than I).
[0027] For example, when the first flag indicates that the other node is not referenced, the occupation state of the adjacent node is one of several occupation patterns, which are a combination of the position of the target node in the parent node and the occupation states of the three adjacent nodes in the parent node. In one of the several occupation patterns, one of the three adjacent nodes is in an occupation state, and the same coding table from the M coding tables may be assigned to the several occupation patterns in which the occupied adjacent node is adjacent to the target node in a direction horizontal to the xy plane.
[0028] For example, when the first flag indicates that the other node is not referenced, the occupation state of the adjacent node is a plurality of occupation patterns represented by a combination of the position of the target node in the parent node and the occupation states of the three adjacent nodes in the parent node, and the same coding table from the M coding tables may be assigned to a plurality of occupation patterns in which one of the three adjacent nodes is in an occupation state, and the adjacent node in that occupation state is adjacent to the target node in a direction perpendicular to the xy plane.
[0029] For example, when the first flag indicates that the neighboring node does not refer to the other node, the occupation state of the neighboring node is represented by a plurality of occupation patterns, which are combinations of the position of the target node in the parent node and the occupation states of the three neighboring nodes in the parent node. The same coding table from the M coding tables may be assigned to a plurality of occupation patterns in which two of the three neighboring nodes are in an occupation state, and the plane consisting of the two occupied neighboring nodes and the target node is horizontal to the xy plane.
[0030] For example, when the first flag indicates that the neighboring node does not refer to the other node, the occupation state of the neighboring node is represented by a plurality of occupation patterns, which are combinations of the position of the target node in the parent node and the occupation states of the three neighboring nodes in the parent node. The same coding table from the M coding tables may be assigned to a plurality of occupation patterns in which two of the three neighboring nodes are in an occupation state, and the plane consisting of the two occupied neighboring nodes and the target node is perpendicular to the xy plane.
[0031] Furthermore, a three-dimensional data encoding device according to one aspect of the present disclosure is a three-dimensional data encoding device for encoding a plurality of three-dimensional points having attribute information, comprising a processor and a memory, wherein the processor uses the memory to encode a first flag indicating whether the target node and its parent node refer to another node different from the target node when encoding a target node included in an N (N is an integer of 2 or more) subtree structure of a plurality of three-dimensional points included in three-dimensional data, and if the first flag indicates that the target node refers to another node, the processor may select an encoding table from N encoding tables according to the occupation status of the adjacent nodes of the target node, and arithmetically encode the information of the target node using the selected encoding table, and if the first flag indicates that the target node does not refer to another node, the processor may select an encoding table from M encoding tables different from N according to the occupation status of the adjacent nodes of the target node, and arithmetically encode the information of the target node using the selected encoding table.
[0032] This approach reduces the number of encoding tables, thereby lowering the processing load. Furthermore, by changing the number of encoding tables depending on whether the target node and its parent node refer to other nodes that are different, the encoding tables can be configured appropriately, thus reducing the processing load while suppressing a decrease in encoding efficiency.
[0033] Furthermore, a three-dimensional data decoding device according to one aspect of the present disclosure is a three-dimensional data decoding device for decoding a plurality of three-dimensional points having attribute information, comprising a processor and a memory, wherein the processor uses the memory to decode a first flag indicating whether the target node and its parent node refer to another node different from the target node when decoding a target node included in an N (N is an integer of 2 or more) subtree structure of a plurality of three-dimensional points included in three-dimensional data, and if the first flag indicates that the target node refers to another node, the processor selects a coding table from N coding tables according to the occupation status of the adjacent nodes of the target node, and arithmetically decodes the information of the target node using the selected coding table, and if the first flag indicates that the target node does not refer to another node, the processor selects a coding table from M coding tables different from N according to the occupation status of the adjacent nodes of the target node, and arithmetically decodes the information of the target node using the selected coding table.
[0034] This approach reduces the number of encoding tables, thereby lowering the processing load. Furthermore, by changing the number of encoding tables depending on whether the target node and its parent node refer to other nodes that are different, the encoding tables can be configured appropriately, thus reducing the processing load while suppressing a decrease in encoding efficiency.
[0035] These comprehensive or specific embodiments may be implemented as a system, method, integrated circuit, computer program, or recording medium such as a computer-readable CD-ROM, or as any combination of a system, method, integrated circuit, computer program, and recording medium.
[0036] The embodiments will be described in detail below with reference to the drawings. Note that the embodiments described below are all specific examples of this disclosure. The numerical values, shapes, materials, components, arrangement and connection configurations of components, steps, and the order of steps shown in the following embodiments are examples only and are not intended to limit this disclosure. Furthermore, among the components in the following embodiments, those not described in the independent claim representing the highest-level concept will be described as optional components.
[0037] (Embodiment 1) First, the data structure of the encoded three-dimensional data (hereinafter also referred to as encoded data) according to this embodiment will be described. Figure 1 is a diagram showing the configuration of the encoded three-dimensional data according to this embodiment.
[0038] In this embodiment, the three-dimensional space is divided into spaces (SPCs) corresponding to pictures in video encoding, and three-dimensional data is encoded using these spaces as units. The spaces are further divided into volumes (VLMs) corresponding to macroblocks in video encoding, and prediction and transformation are performed using the VLMs as units. Each volume contains multiple voxels (VXLs), which are the smallest units to which position coordinates are associated. Prediction, similar to prediction performed on two-dimensional images, involves referencing other processing units to generate predicted three-dimensional data similar to the processing unit being processed, and then encoding the difference between this predicted three-dimensional data and the processing unit being processed. Furthermore, this prediction includes not only spatial prediction that references other prediction units at the same time, but also temporal prediction that references prediction units at different times.
[0039] For example, a three-dimensional data encoding device (hereinafter also referred to as the encoding device) encodes a three-dimensional space represented by point cloud data, such as a point cloud, by encoding each point in the point cloud, or multiple points contained within a voxel, depending on the size of the voxel. Subdividing the voxel allows for a highly accurate representation of the three-dimensional shape of the point cloud, while increasing the voxel size allows for a rougher representation of the three-dimensional shape of the point cloud.
[0040] In the following explanation, we will use the example of a point cloud as the 3D data, but the 3D data is not limited to a point cloud; any format of 3D data is acceptable.
[0041] Alternatively, a hierarchical structure of voxels may be used. In this case, for the nth-order hierarchy, it may be indicated sequentially whether or not sample points exist in the (n-1)th-order hierarchy and below (the lower layers of the nth-order hierarchy). For example, when decoding only the nth-order hierarchy, if sample points exist in the (n-1)th-order hierarchy and below, the sample points can be assumed to be at the center of the voxel of the nth-order hierarchy and decoded accordingly.
[0042] Furthermore, the encoding device acquires point cloud data using distance sensors, stereo cameras, monocular cameras, gyroscopes, or inertial sensors.
[0043] Spaces, like video encodings, are classified into at least three predictive structures, including intra-spaces (I-SPCs) that can be decoded independently, predictive spaces (P-SPCs) that allow only unidirectional referencing, and bidirectional spaces (B-SPCs) that allow bidirectional referencing. Furthermore, spaces contain two types of time information: the decoding time and the display time.
[0044] Furthermore, as shown in Figure 1, there is a processing unit called GOS (Group of Space), which is a random access unit, that contains multiple spaces. In addition, there is a processing unit called WLD (World), which contains multiple GOS.
[0045] The spatial area occupied by a world is associated with an absolute location on Earth using GPS or latitude and longitude information. This location information is stored as metadata. This metadata may be included in the encoded data or transmitted separately from the encoded data.
[0046] Furthermore, within a GOS, all SPCs may be adjacent in three dimensions, or there may be SPCs that are not adjacent in three dimensions to other SPCs.
[0047] In the following, the processing of three-dimensional data contained in processing units such as GOS, SPC, or VLM, including encoding, decoding, or referencing, will also be simply referred to as encoding, decoding, or referencing the processing unit. Furthermore, the three-dimensional data contained in the processing unit includes, for example, at least one pair of spatial position such as three-dimensional coordinates and characteristic values such as color information.
[0048] Next, we will explain the prediction structure of SPCs in GOS. Multiple SPCs within the same GOS, or multiple VLMs within the same SPC, occupy different spaces from each other, but they have the same time information (decoded time and display time).
[0049] Furthermore, the SPC that is first in the decryption order within a GOS is the I-SPC. There are also two types of GOSs: closed GOS and open GOS. A closed GOS is one in which all SPCs within the GOS can be decrypted when decryption starts from the first I-SPC. In an open GOS, some SPCs whose displayed time is earlier than the first I-SPC refer to a different GOS, and decryption cannot be performed using only that GOS.
[0050] Furthermore, with encoded data such as map information, the WLD may be decoded in the reverse direction of the encoding order, and if there are dependencies between GOSs, reverse playback becomes difficult. Therefore, in such cases, a closed GOS is generally used.
[0051] Furthermore, GOS has a layered structure in the height direction, and encoding or decoding is performed sequentially from the SPC of the lower layer.
[0052] Figure 2 shows an example of the prediction structure between SPCs belonging to the lowest layer of GOS. Figure 3 shows an example of the prediction structure between layers.
[0053] One or more I-SPCs exist within a GOS. While objects such as people, animals, cars, bicycles, traffic lights, or landmark buildings exist in three-dimensional space, it is particularly effective to encode small objects as I-SPCs. For example, a three-dimensional data decoding device (hereinafter also referred to as the decoding device) decodes only the I-SPCs within the GOS when decoding a GOS with low processing load or at high speed.
[0054] Furthermore, the encoding device may switch the encoding interval or frequency of I-SPCs according to the density of objects in the WLD.
[0055] Furthermore, in the configuration shown in Figure 3, the encoding or decoding device encodes or decodes multiple layers sequentially from the bottom layer (Layer 1). This allows for prioritizing data near the ground, which contains more information, for applications such as autonomous vehicles.
[0056] In the case of encoded data used in drones and the like, encoding or decoding may be done sequentially within the GOS, starting from the SPC layer at the top in the height direction.
[0057] Furthermore, the encoding or decoding device may encode or decode multiple layers so that the decoding device can grasp the GOS roughly and gradually increase the resolution. For example, the encoding or decoding device may encode or decode layers 3, 8, 1, 9, and so on.
[0058] Next, we will explain how to handle static and dynamic objects.
[0059] In three-dimensional space, there are static objects or scenes such as buildings or roads (hereinafter collectively referred to as static objects) and dynamic objects such as cars or people (hereinafter referred to as dynamic objects). Object detection is performed separately, for example, by extracting feature points from point cloud data or camera images such as stereo cameras. Here, we will explain an example of an encoding method for dynamic objects.
[0060] The first method is to encode static and dynamic objects without distinguishing between them. The second method is to distinguish between static and dynamic objects using identification information.
[0061] For example, GOS is used as the identification unit. In this case, GOS containing SPCs that constitute static objects and GOS containing SPCs that constitute dynamic objects are distinguished by identification information stored within the encoded data or separately from the encoded data.
[0062] Alternatively, an SPC may be used as the identification unit. In this case, an SPC containing a VLM that constitutes a static object and an SPC containing a VLM that constitutes a dynamic object are distinguished by the above identification information.
[0063] Alternatively, VLM or VXL may be used as the identification unit. In this case, VLM or VXL containing static objects and VLM or VXL containing dynamic objects are distinguished by the above identification information.
[0064] Furthermore, the encoding device may encode dynamic objects as one or more VLMs or SPCs, and encode the VLM or SPC containing static objects and the SPC containing dynamic objects as different GOSs. Also, if the size of the GOS is variable depending on the size of the dynamic objects, the encoding device stores the size of the GOS separately as metadata.
[0065] Furthermore, the encoding device may encode static objects and dynamic objects independently of each other and superimpose dynamic objects onto a world composed of static objects. In this case, a dynamic object is composed of one or more SPCs, and each SPC is associated with one or more SPCs that constitute the static object on which it is superimposed. Note that dynamic objects may be represented by one or more VLMs or VXLs instead of SPCs.
[0066] Furthermore, the encoding device may encode static objects and dynamic objects as separate streams.
[0067] Furthermore, the encoding device may generate a GOS containing one or more SPCs that constitute a dynamic object. In addition, the encoding device may set the GOS containing the dynamic object (GOS_M) and the GOS of the static object corresponding to the spatial region of GOS_M to be the same size (occupy the same spatial region). This allows superposition processing to be performed on a GOS-by-GOS basis.
[0068] The P-SPC or B-SPC that constitute a dynamic object may reference SPCs contained in different encoded GOS. In cases where the position of a dynamic object changes over time and the same dynamic object is encoded as a GOS at different times, cross-GOS references are effective from a compression standpoint.
[0069] Furthermore, the first and second methods described above may be switched depending on the intended use of the encoded data. For example, when using encoded three-dimensional data as a map, it is desirable to be able to separate dynamic objects, so the encoding device uses the second method. On the other hand, when encoding three-dimensional data of an event such as a concert or sporting event, if there is no need to separate dynamic objects, the encoding device uses the first method.
[0070] Furthermore, the decoding time and display time of GOS or SPC can be stored within the encoded data or as metadata. The time information for static objects may also be identical. In this case, the actual decoding time and display time may be determined by the decoding device. Alternatively, different values may be assigned to each GOS or SPC as the decoding time, while the same value may be assigned to all as the display time. Furthermore, a decoder model may be introduced, such as the HEVC HRD (Hypothetical Reference Decoder) in video encoding, which guarantees that decoding can be performed without failure if the decoder has a buffer of a predetermined size and reads the bitstream at a predetermined bitrate according to the decoding time.
[0071] Next, we will explain the arrangement of GOS within the world. The coordinates of the three-dimensional space in the world are represented by three mutually orthogonal coordinate axes (x-axis, y-axis, and z-axis). By establishing a predetermined rule for the coding order of GOS, coding can be performed so that spatially adjacent GOS are continuous within the coded data. For example, in the example shown in Figure 4, GOS in the xz plane are coded continuously. The value of the y-axis is updated after coding all GOS in a given xz plane is completed. That is, as coding progresses, the world expands in the y-axis direction. Also, the index numbers of the GOS are set in the coding order.
[0072] Here, the world's three-dimensional space is mapped one-to-one with geographical absolute coordinates such as GPS, latitude, and longitude. Alternatively, the three-dimensional space may be represented by relative positions from a pre-defined reference position. The directions of the x, y, and z axes of the three-dimensional space are represented as direction vectors determined based on latitude and longitude, and these direction vectors are stored as metadata along with encoded data.
[0073] Furthermore, the size of the GOS is fixed, and the encoding device stores this size as metadata. Alternatively, the size of the GOS may be switched depending on, for example, whether it is an urban area or not, or whether it is indoors or outdoors. In other words, the size of the GOS may be switched depending on the quantity or nature of objects that have informational value. Or, the encoding device may adaptively switch the size of the GOS or the spacing of I-SPCs within the GOS depending on the density of objects within the same world. For example, the encoding device may reduce the size of the GOS and shorten the spacing of I-SPCs within the GOS as the density of objects increases.
[0074] In the example in Figure 5, the GOS regions from the 3rd to the 10th are subdivided to enable fine-grained random access due to the high object density. Note that GOS regions 7 through 10 are located behind GOS regions 3 through 6, respectively.
[0075] Next, the configuration and operation flow of the three-dimensional data encoding device according to this embodiment will be described. Figure 6 is a block diagram of the three-dimensional data encoding device 100 according to this embodiment. Figure 7 is a flowchart showing an example of the operation of the three-dimensional data encoding device 100.
[0076] The three-dimensional data encoding device 100 shown in Figure 6 generates encoded three-dimensional data 112 by encoding three-dimensional data 111. This three-dimensional data encoding device 100 comprises an acquisition unit 101, an encoding region determination unit 102, a division unit 103, and an encoding unit 104.
[0077] As shown in Figure 7, first, the acquisition unit 101 acquires three-dimensional data 111, which is point cloud data (S101).
[0078] Next, the encoding region determination unit 102 determines the region to be encoded from among the spatial regions corresponding to the acquired point cloud data (S102). For example, the encoding region determination unit 102 determines the spatial region around the location of the user or vehicle as the region to be encoded.
[0079] Next, the division unit 103 divides the point cloud data included in the region to be encoded into processing units. Here, the processing units are the GOS and SPC mentioned above. The region to be encoded corresponds to, for example, the world mentioned above. Specifically, the division unit 103 divides the point cloud data into processing units based on a pre-set GOS size, or the presence or size of dynamic objects (S103). The division unit 103 also determines the starting position of the SPC that will be the first in the encoding order for each GOS.
[0080] Next, the encoding unit 104 generates encoded three-dimensional data 112 by sequentially encoding multiple SPCs within each GOS (S104).
[0081] Note that while this example shows the region to be encoded being divided into GOS and SPC before encoding each GOS, the processing procedure is not limited to the above. For example, one could determine the structure of one GOS, encode that GOS, and then determine the structure of the next GOS.
[0082] In this way, the three-dimensional data encoding device 100 generates encoded three-dimensional data 112 by encoding the three-dimensional data 111. Specifically, the three-dimensional data encoding device 100 divides the three-dimensional data into first processing units (GOS), which are random access units, each of which is associated with a three-dimensional coordinate. The first processing units (GOS) are then divided into a plurality of second processing units (SPCs), and the second processing units (SPCs) are then divided into a plurality of third processing units (VLMs). The third processing unit (VLM) also contains one or more voxels (VXLs), which are the smallest units to which positional information is associated.
[0083] Next, the three-dimensional data encoding device 100 generates encoded three-dimensional data 112 by encoding each of the multiple first processing units (GOS). Specifically, the three-dimensional data encoding device 100 encodes each of the multiple second processing units (SPC) in each first processing unit (GOS). Furthermore, the three-dimensional data encoding device 100 encodes each of the multiple third processing units (VLM) in each second processing unit (SPC).
[0084] For example, if the first processing unit (GOS) to be processed is a closed GOS, the three-dimensional data encoding device 100 encodes the second processing unit (SPC) included in the first processing unit (GOS) by referring to other second processing units (SPC) included in the first processing unit (GOS). In other words, the three-dimensional data encoding device 100 does not refer to second processing units (SPC) included in a first processing unit (GOS) different from the first processing unit (GOS) to be processed.
[0085] On the other hand, if the first processing unit (GOS) to be processed is an open GOS, the second processing unit (SPC) included in the first processing unit (GOS) to be processed is encoded by referring to another second processing unit (SPC) included in the first processing unit (GOS) to be processed, or to a second processing unit (SPC) included in a first processing unit (GOS) different from the first processing unit (GOS) to be processed.
[0086] Furthermore, the three-dimensional data encoding device 100 selects one of the following types of second processing units (SPCs) to be processed: a first type (I-SPC) that does not refer to any other second processing units (SPCs), a second type (P-SPC) that refers to one other second processing unit (SPC), and a third type that refers to two other second processing units (SPCs). The device then encodes the second processing unit (SPC) to be processed according to the selected type.
[0087] Next, the configuration and operation flow of the three-dimensional data decoding device according to this embodiment will be described. Figure 8 is a block diagram of the three-dimensional data decoding device 200 according to this embodiment. Figure 9 is a flowchart showing an example of the operation of the three-dimensional data decoding device 200.
[0088] The three-dimensional data decoding device 200 shown in Figure 8 generates decoded three-dimensional data 212 by decoding encoded three-dimensional data 211. Here, encoded three-dimensional data 211 is, for example, encoded three-dimensional data 112 generated by the three-dimensional data encoding device 100. This three-dimensional data decoding device 200 includes an acquisition unit 201, a decoding start GOS determination unit 202, a decoding SPC determination unit 203, and a decoding unit 204.
[0089] First, the acquisition unit 201 acquires encoded three-dimensional data 211 (S201). Next, the decoding start GOS determination unit 202 determines the GOS to be decoded (S202). Specifically, the decoding start GOS determination unit 202 refers to metadata stored in or separately from the encoded three-dimensional data 211 to determine the GOS to be decoded, which includes an SPC corresponding to the spatial position, object, or time to start decoding.
[0090] Next, the decryption SPC determination unit 203 determines the type of SPC (I, P, B) to be decrypted within the GOS (S203). For example, the decryption SPC determination unit 203 determines whether to (1) decrypt only I-SPCs, (2) decrypt I-SPCs and P-SPCs, or (3) decrypt all types. Note that if the type of SPC to be decrypted has been determined in advance, such as decrypting all SPCs, this step may not be performed.
[0091] Next, the decoding unit 204 obtains the address position where the first SPC in the decoding order (same as the encoding order) within the GOS starts in the encoded three-dimensional data 211, obtains the encoded data of the first SPC from that address position, and decodes each SPC sequentially starting from that first SPC (S204). Note that the above address position is stored in metadata, etc.
[0092] In this way, the three-dimensional data decoding device 200 decodes the decoded three-dimensional data 212. Specifically, the three-dimensional data decoding device 200 generates decoded three-dimensional data 212 of the first processing unit (GOS) by decoding each of the encoded three-dimensional data 211 of the first processing unit (GOS), which is a random access unit, and each of which is associated with three-dimensional coordinates. More specifically, the three-dimensional data decoding device 200 decodes each of the multiple second processing units (SPC) in each first processing unit (GOS). Furthermore, the three-dimensional data decoding device 200 decodes each of the multiple third processing units (VLM) in each second processing unit (SPC).
[0093] The metadata for random access is described below. This metadata is generated by the three-dimensional data encoding device 100 and is included in the encoded three-dimensional data 112(211).
[0094] In conventional random access to two-dimensional moving images, decoding began from the first frame of a random access unit that was near the specified time. In contrast, in the world, random access is expected not only to time but also to space (coordinates or objects, etc.).
[0095] Therefore, in order to achieve random access to at least three elements—coordinates, objects, and time—a table is prepared that associates each element with the GOS index number. Furthermore, the GOS index number is associated with the address of the I-SPC that is the starting point of the GOS. Figure 10 shows an example of a table included in the metadata. Note that it is not necessary to use all the tables shown in Figure 10; it is sufficient to use at least one table.
[0096] The following describes random access starting from coordinates as an example. When accessing coordinates (x2, y2, z2), first, the coordinate-GOS table is consulted to find that the location with coordinates (x2, y2, z2) is included in the second GOS. Next, the GOS address table is consulted to find that the address of the first I-SPC in the second GOS is addr(2). Therefore, the decoding unit 204 retrieves data from this address and begins decoding.
[0097] The address may be a logical format address or a physical address of the HDD or memory. Alternatively, information identifying a file segment may be used instead of an address. For example, a file segment is a unit formed by segmenting one or more GOSs (Global Operating Systems).
[0098] Furthermore, if an object spans multiple GOSs, the object-GOS table may indicate multiple GOSs to which the object belongs. If these multiple GOSs are closed GOSs, the encoding and decoding devices can perform encoding or decoding in parallel. On the other hand, if these multiple GOSs are open GOSs, the compression efficiency can be further improved by allowing the multiple GOSs to reference each other.
[0099] Examples of objects include people, animals, cars, bicycles, traffic lights, or landmark buildings. For example, the three-dimensional data encoding device 100 can extract feature points specific to objects from a three-dimensional point cloud or the like when encoding a world, detect objects based on these feature points, and set the detected objects as random access points.
[0100] Thus, the three-dimensional data encoding device 100 generates first information indicating a plurality of first processing units (GOS) and the three-dimensional coordinates associated with each of the plurality of first processing units (GOS). The encoded three-dimensional data 112(211) also includes this first information. Furthermore, the first information indicates at least one of the following: an object, a time, and a data storage location, associated with each of the plurality of first processing units (GOS).
[0101] The three-dimensional data decoding device 200 acquires first information from the encoded three-dimensional data 211, uses the first information to identify the encoded three-dimensional data 211 of the first processing unit corresponding to the specified three-dimensional coordinates, object, or time, and decodes the encoded three-dimensional data 211.
[0102] The following describes examples of other metadata. In addition to metadata for random access, the three-dimensional data encoding device 100 may generate and store the following metadata. The three-dimensional data decoding device 200 may also use this metadata during decoding.
[0103] When using three-dimensional data as map information, profiles may be defined according to the intended use, and information indicating the profile may be included in the metadata. For example, profiles may be defined for urban areas, suburbs, or for flying objects, and the maximum or minimum size of the world, SPC, or VLM may be defined for each. For example, for urban areas, more detailed information is required than for suburbs, so the minimum size of the VLM is set to be smaller.
[0104] Metadata may include tag values indicating the object type. These tag values are associated with the VLM, SPC, or GOS that constitute the object. For example, tag value "0" may indicate "person," tag value "1" may indicate "car," tag value "2" may indicate "traffic light," and so on, with different tag values assigned to each object type. Alternatively, if it is difficult or unnecessary to determine the object type, tag values indicating properties such as size or whether it is a dynamic or static object may be used.
[0105] Furthermore, the metadata may include information indicating the extent of the spatial region occupied by the world.
[0106] Furthermore, the metadata may include the size of the SPC or VXL as header information common to multiple SPCs, such as the entire stream of encoded data or an SPC within a GOS.
[0107] Furthermore, the metadata may include identification information for distance sensors or cameras used to generate the point cloud, or information indicating the positional accuracy of the point cloud within the point cloud.
[0108] Furthermore, the metadata may include information indicating whether the world consists solely of static objects or includes dynamic objects.
[0109] Modifications of this embodiment will be described below.
[0110] The encoding or decoding device may encode or decode two or more different SPCs or GOSs in parallel. The GOSs to be encoded or decoded in parallel can be determined based on metadata indicating the spatial location of the GOS.
[0111] In cases where three-dimensional data is used as a spatial map when a vehicle or flying object moves, or when such a spatial map is generated, the encoding or decoding device may encode or decode the GOS or SPC contained in the space identified based on GPS, route information, or zoom magnification.
[0112] Furthermore, the decoding device may perform decoding starting from the space closest to its own position or travel path. The encoding or decoding device may encode or decode spaces farther from its own position or travel path with lower priority compared to spaces closer to it. Here, lowering priority means lowering the processing order, lowering the resolution (downsampling), or lowering the image quality (increasing encoding efficiency, for example, by increasing the quantization step).
[0113] Furthermore, when a decoding device decodes encoded data that is hierarchically encoded in space, it may decode only the lower layers.
[0114] Furthermore, the decoding device may prioritize decoding from lower layers depending on the map's zoom level or intended use.
[0115] Furthermore, for applications such as self-localization or object recognition during autonomous driving of vehicles or robots, the encoding or decoding device may reduce the resolution of the area outside of the area within a specific height from the road surface (the area to be recognized) when encoding or decoding.
[0116] Furthermore, the encoding device may encode the point clouds representing the spatial shapes of the indoor and outdoor areas separately. For example, by separating the GOS representing the indoor area (indoor GOS) and the GOS representing the outdoor area (outdoor GOS), the decoding device can select the GOS to decode according to the viewpoint position when using the encoded data.
[0117] Furthermore, the encoding device may encode indoor and outdoor GOS locations with similar coordinates so that they are adjacent within the encoding stream. For example, the encoding device associates the identifiers of both locations and stores information indicating the associated identifiers within the encoding stream or in separately stored metadata. This allows the decoding device to identify indoor and outdoor GOS locations with similar coordinates by referring to the information in the metadata.
[0118] Furthermore, the encoding device may switch the size of the GOS or SPC between indoor and outdoor GOS. For example, the encoding device may set the GOS size smaller indoors than outdoors. The encoding device may also change the accuracy of extracting feature points from the point cloud or the accuracy of object detection between indoor and outdoor GOS.
[0119] Furthermore, the encoding device may add information to the encoded data that allows the decoding device to distinguish and display dynamic objects from static objects. This allows the decoding device to display dynamic objects together with a red frame or explanatory text. Alternatively, the decoding device may display only the red frame or explanatory text instead of the dynamic object. The decoding device may also display more detailed object types. For example, a red frame may be used for cars and a yellow frame for people.
[0120] Furthermore, the encoding or decoding device may decide whether to encode or decode dynamic objects and static objects as different SPCs or GOSs depending on the frequency of occurrence of dynamic objects or the ratio of static objects to dynamic objects. For example, if the frequency or ratio of occurrence of dynamic objects exceeds a threshold, an SPC or GOS containing a mixture of dynamic and static objects is permitted, while if the frequency or ratio of occurrence of dynamic objects does not exceed a threshold, an SPC or GOS containing a mixture of dynamic and static objects is not permitted.
[0121] When detecting dynamic objects from two-dimensional image information from a camera rather than a point cloud, the encoding device may separately acquire information to identify the detection result (such as a frame or text) and the object's position, and encode this information as part of the three-dimensional encoded data. In this case, the decoding device overlays auxiliary information (a frame or text) indicating the dynamic object onto the decoded result of the static object.
[0122] Furthermore, the encoding device may change the density of VXL or VLM in the SPC depending on the complexity of the shape of the static object. For example, the encoding device will set the VXL or VLM density to be denser as the shape of the static object becomes more complex. In addition, the encoding device may determine the quantization step when quantizing spatial position or color information according to the density of VXL or VLM. For example, the encoding device will set the quantization step to be smaller as the VXL or VLM density increases.
[0123] As described above, the encoding or decoding device according to this embodiment performs spatial encoding or decoding on a spatial basis that has coordinate information.
[0124] Furthermore, the encoding and decoding devices perform encoding or decoding in volume units within the space. A volume includes a voxel, which is the smallest unit to which location information is associated.
[0125] Furthermore, the encoding and decoding devices encode or decode arbitrary elements by associating each element of spatial information, including coordinates, objects, and time, with the GOP, or by associating each element with another element using a table. The decoding device determines the coordinates using the values of the selected elements, identifies a volume, voxel, or space from the coordinates, and decodes the space containing the volume or voxel, or the identified space.
[0126] Furthermore, the encoding device determines selectable volumes, voxels, or spaces based on the elements through feature point extraction or object recognition, and encodes them as randomly accessible volumes, voxels, or spaces.
[0127] Spaces are classified into three types: I-SPCs, which can be encoded or decoded on their own; P-SPCs, which are encoded or decoded by referencing any one processed space; and B-SPCs, which are encoded or decoded by referencing any two processed spaces.
[0128] One or more volumes correspond to static or dynamic objects. Spaces containing static objects and spaces containing dynamic objects are encoded or decoded as different GOSs. In other words, SPCs containing static objects and SPCs containing dynamic objects are assigned to different GOSs.
[0129] Dynamic objects are encoded or decoded individually and mapped to one or more spaces containing static objects. In other words, multiple dynamic objects are encoded individually, and the resulting encoded data of multiple dynamic objects is mapped to an SPC containing static objects.
[0130] The encoding and decoding devices prioritize the I-SPCs within the GOS when encoding or decoding. For example, the encoding device encodes in a way that minimizes I-SPC degradation (so that the original 3D data is reproduced more faithfully after decoding). The decoding device, on the other hand, decodes only the I-SPCs.
[0131] The encoding device may perform encoding by changing the frequency of using I-SPC depending on the density or number (quantity) of objects in the world. In other words, the encoding device changes the frequency of selecting I-SPC depending on the number or density of objects included in the three-dimensional data. For example, the encoding device will increase the frequency of using I-space as the density of objects in the world increases.
[0132] Furthermore, the encoding device sets random access points in GOS units and stores information indicating the spatial region corresponding to each GOS in the header information.
[0133] The encoding device uses a default value as the spatial size of the GOS. However, the encoding device may change the size of the GOS depending on the number (quantity) or density of objects or dynamic objects. For example, the encoding device will reduce the spatial size of the GOS as the density or number of objects or dynamic objects increases.
[0134] Furthermore, the space or volume includes a set of feature points derived using information obtained from sensors such as depth sensors, gyroscopes, or cameras. The coordinates of the feature points are set to the center position of the voxel. In addition, the accuracy of the positional information can be improved by subdividing the voxels.
[0135] The feature point cloud is derived using multiple pictures. Each picture has at least two types of time information: actual time information and the same time information across multiple pictures mapped to space (for example, the encoded time used for rate control, etc.).
[0136] Furthermore, encoding or decoding is performed in GOS units that contain one or more spaces.
[0137] The encoding and decoding devices refer to the spaces within the processed GOS to predict the P-space or B-space within the GOS to be processed.
[0138] Alternatively, the encoding and decoding devices do not refer to different GOSs, but instead use the processed space within the GOS to be processed to predict the P-space or B-space within the GOS to be processed.
[0139] Furthermore, the encoding and decoding devices transmit or receive encoded streams in world units containing one or more GOSs.
[0140] Furthermore, the GOS has a layered structure in at least one direction within the world, and the encoding and decoding devices encode or decode from the lower layers. For example, a randomly accessible GOS belongs to the lowest layer. A GOS belonging to a higher layer refers to a GOS belonging to the same layer or lower. In other words, the GOS is spatially divided in a predetermined direction, and each contains multiple layers, each containing one or more SPCs. The encoding and decoding devices encode or decode each SPC by referring to an SPC included in the same layer as that SPC or in a lower layer than that SPC.
[0141] Furthermore, the encoding and decoding devices sequentially encode or decode GOS within a world unit containing multiple GOS. The encoding and decoding devices write or read information indicating the encoding or decoding order (direction) as metadata. In other words, the encoded data includes information indicating the encoding order of multiple GOS.
[0142] Furthermore, the encoding device and the decoding device encode or decode two or more different spaces or GOS in parallel.
[0143] Furthermore, the encoding and decoding devices encode or decode spatial information (coordinates, size, etc.) of space or GOS.
[0144] Furthermore, the encoding and decoding devices encode or decode spaces or GOS contained within a specific space identified based on external information relating to their own position and / or area size, such as GPS, route information, or magnification.
[0145] The encoding or decoding device encodes or decodes spaces farther away from its own position with lower priority compared to spaces closer to it.
[0146] The encoding device sets one direction in the world according to the magnification or application, and encodes a GOS with a layered structure in that direction. The decoding device then decodes the GOS with a layered structure in the one direction in the world set according to the magnification or application, prioritizing from the lower layers.
[0147] The encoding device changes the accuracy of feature point extraction, object recognition, or spatial domain size between indoor and outdoor spaces. However, the encoding and decoding devices encode or decode indoor and outdoor GOS (Geoscopy) points that are close in coordinates adjacent to each other within the world, and encode or decode their identifiers in association with each other.
[0148] (Embodiment 2) When using encoded point cloud data in actual devices or services, it is desirable to send and receive necessary information depending on the application in order to reduce network bandwidth. However, until now, such functionality has not existed in the encoded structure of three-dimensional data, nor has there been an encoding method for that purpose.
[0149] This embodiment describes a three-dimensional data encoding method and a three-dimensional data encoding apparatus for providing a function to transmit and receive only the necessary information in encoded data of a three-dimensional point cloud according to its application, as well as a three-dimensional data decoding method and a three-dimensional data decoding apparatus for decoding said encoded data.
[0150] A voxel (VXL) with a certain number of features is defined as a feature voxel (FVXL), and a world (WLD) composed of FVXLs is defined as a sparse world (SWLD). Figure 11 shows examples of the configuration of sparse worlds and worlds. SWLDs include FGOS, which is a GOS composed of FVXLs; FSPC, which is a SPC composed of FVXLs; and FVLM, which is a VLM composed of FVXLs. The data structure and prediction structure of FGOS, FSPC, and FVLM may be the same as those of GOS, SPC, and VLM.
[0151] A feature is a feature that represents the three-dimensional position information of a VXL, or the visible light information of the VXL's position, and is particularly frequently detected at corners and edges of three-dimensional objects. Specifically, this feature is a three-dimensional feature or a visible light feature as shown below, but any other feature that represents the position, brightness, or color information of the VXL is acceptable.
[0152] Three-dimensional features can be obtained using SHOT features (Signature of Histograms of OrienTations), PFH features (Point Feature Histograms), or PPF features (Point Pair Feature).
[0153] SHOT features are obtained by dividing the area around VXL, calculating the dot product of the reference point and the normal vector of the divided region, and then generating a histogram. These SHOT features have the characteristics of high dimensionality and high feature representation power.
[0154] PFH features are obtained by selecting a large number of pairs of points in the vicinity of the VXL, calculating normal vectors and other parameters from these two points, and then creating a histogram. Because these PFH features are histogram features, they are robust to some disturbances and have high feature representation power.
[0155] PPF features are features calculated using normal vectors and other methods for every two VXLs. Because all VXLs are used in these PPF features, they are robust to occlusion.
[0156] Furthermore, as visible light features, SIFT (Scale-Invariant Feature Transform), SURF (Speeded Up Robust Features), or HOG (Histogram of Oriented Gradients), which use information such as the brightness gradient of the image, can be used.
[0157] SWLD is generated by calculating the above features from each VXL of WLD and extracting FVXL. Here, SWLD can be updated every time WLD is updated, or it can be updated periodically after a certain period of time regardless of when WLD is updated.
[0158] SWLDs can be generated for each feature. For example, separate SWLDs can be generated for each feature, such as SWLD1 based on SHOT features and SWLD2 based on SIFT features, and the appropriate SWLD can be used depending on the application. Alternatively, the features of each calculated FVXL can be stored as feature information within each FVXL.
[0159] Next, we will explain how to use sparse worlds (SWLDs). Because SWLDs contain only feature voxels (FVXLs), they generally have a smaller data size compared to WLDs, which contain all VXLs.
[0160] In applications that utilize features to achieve a specific objective, using SWLD information instead of WLD information can reduce read time from the hard disk, as well as bandwidth and transfer time during network transmission. For example, by storing both WLD and SWLD data on a server and switching the transmitted map information between WLD and SWLD according to client requests, network bandwidth and transfer time can be reduced. A specific example is shown below.
[0161] Figures 12 and 13 illustrate examples of SWLD and WLD usage. As shown in Figure 12, when client 1, an in-vehicle device, requires map information for self-position determination, client 1 sends a request to the server to acquire map data for self-position estimation (S301). The server sends an SWLD to client 1 in response to the acquisition request (S302). Client 1 uses the received SWLD to determine its own position (S303). At this time, client 1 acquires VXL information around client 1 using various methods such as distance sensors like rangefinders, stereo cameras, or combinations of multiple monocular cameras, and estimates its own position information from the obtained VXL information and the SWLD. Here, the self-position information includes the three-dimensional position information and orientation of client 1.
[0162] As shown in Figure 13, when client 2, an in-vehicle device, needs map information for purposes such as drawing three-dimensional maps, client 2 sends a request to the server to acquire map data for map drawing (S311). The server sends a WLD to client 2 in response to the acquisition request (S312). Client 2 uses the received WLD to perform map drawing (S313). In this case, client 2 creates a rendered image using, for example, an image taken by itself with a visible light camera and the WLD acquired from the server, and then draws the created image on the screen of a car navigation system or the like.
[0163] As described above, the server sends SWLDs to the client when primarily needing individual VXL features, such as for self-localization, and sends WLDs to the client when detailed VXL information is required, such as for map plotting. This enables efficient transmission and reception of map data.
[0164] Furthermore, the client may decide for itself whether it needs an SWLD or a WLD and request the server to send either one. The server may also decide whether to send an SWLD or a WLD based on the client or network conditions.
[0165] Next, we will explain how to switch between sending and receiving data in Sparse World (SWLD) and World (WLD) modes.
[0166] The system may switch between receiving a WLD or SWLD depending on the network bandwidth. Figure 14 shows an example of this operation. For example, when a low-speed network with limited usable network bandwidth, such as in an LTE (Long Term Evolution) environment, is used, the client accesses the server via the low-speed network (S321) and obtains an SWLD from the server as map information (S322). On the other hand, when a high-speed network with ample network bandwidth, such as in a Wi-Fi (registered trademark) environment, is used, the client accesses the server via the high-speed network (S323) and obtains a WLD from the server (S324). This allows the client to obtain appropriate map information according to the client's network bandwidth.
[0167] Specifically, the client receives the SWLD via LTE outdoors and acquires the WLD via Wi-Fi (registered trademark) when it enters an indoor facility. This allows the client to obtain more detailed indoor map information.
[0168] Thus, a client may request a WLD or SWLD from the server depending on the bandwidth of the network it is using. Alternatively, the client may send information indicating the bandwidth of the network it is using to the server, and the server may send data (WLD or SWLD) appropriate for that client based on that information. Alternatively, the server may determine the client's network bandwidth and send data (WLD or SWLD) appropriate for that client.
[0169] Furthermore, the system may switch between receiving a WLD or SWLD depending on the travel speed. Figure 15 shows an example of this operation. For example, when the client is traveling at high speed (S331), the client receives an SWLD from the server (S332). On the other hand, when the client is traveling at low speed (S333), the client receives a WLD from the server (S334). This allows the client to acquire map information appropriate to its speed while suppressing network bandwidth. Specifically, when the client is traveling on a highway, it can receive a SWLD with a small amount of data, allowing it to update rough map information at an appropriate speed. On the other hand, when the client is traveling on a general road, it can receive a WLD, allowing it to acquire more detailed map information.
[0170] Thus, the client may request a WLD or SWLD from the server according to its own movement speed. Alternatively, the client may send information indicating its movement speed to the server, and the server may send data (WLD or SWLD) appropriate to the client according to that information. Alternatively, the server may determine the client's movement speed and send data (WLD or SWLD) appropriate to the client.
[0171] Alternatively, the client may first obtain the SWLD from the server and then obtain the WLD for important areas within it. For example, when acquiring map data, the client can first obtain general map information using the SWLD, then narrow down the areas where features such as buildings, signs, or people appear frequently, and then obtain the WLD for those narrowed-down areas later. This allows the client to obtain detailed information for the necessary areas while suppressing the amount of data received from the server.
[0172] Alternatively, the server may create separate SWLDs for each object from the WLD, and the client may receive them according to its purpose. This can reduce network bandwidth usage. For example, the server may recognize people or cars in advance from the WLD and create SWLDs for people and cars. The client receives the SWLD for people if it wants to obtain information about people in the vicinity, or the SWLD for cars if it wants to obtain information about cars. Furthermore, the types of SWLDs may be distinguished by information (flags or types, etc.) added to the header.
[0173] Next, the configuration and operation flow of the three-dimensional data encoding device (e.g., a server) according to this embodiment will be described. Figure 16 is a block diagram of the three-dimensional data encoding device 400 according to this embodiment. Figure 17 is a flowchart of the three-dimensional data encoding process by the three-dimensional data encoding device 400.
[0174] The three-dimensional data encoding device 400 shown in Figure 16 generates encoded streams, encoded three-dimensional data 413 and 414, by encoding the input three-dimensional data 411. Here, encoded three-dimensional data 413 is encoded three-dimensional data corresponding to WLD, and encoded three-dimensional data 414 is encoded three-dimensional data corresponding to SWLD. This three-dimensional data encoding device 400 comprises an acquisition unit 401, an encoding region determination unit 402, an SWLD extraction unit 403, a WLD encoding unit 404, and an SWLD encoding unit 405.
[0175] As shown in Figure 17, first, the acquisition unit 401 acquires input three-dimensional data 411, which is point cloud data in three-dimensional space (S401).
[0176] Next, the encoding region determination unit 402 determines the spatial region to be encoded based on the spatial region where the point cloud data exists (S402).
[0177] Next, the SWLD extraction unit 403 defines the spatial region to be encoded as a WLD and calculates features from each VXL contained in the WLD. Then, the SWLD extraction unit 403 extracts VXLs whose features are equal to or greater than a predetermined threshold, defines the extracted VXLs as FVXLs, and adds these FVXLs to the SWLD to generate extracted three-dimensional data 412 (S403). In other words, extracted three-dimensional data 412 with features equal to or greater than the threshold is extracted from the input three-dimensional data 411.
[0178] Next, the WLD encoding unit 404 generates encoded three-dimensional data 413 corresponding to the WLD by encoding the input three-dimensional data 411 corresponding to the WLD (S404). At this time, the WLD encoding unit 404 adds information to the header of the encoded three-dimensional data 413 to distinguish that the encoded three-dimensional data 413 is a stream containing a WLD.
[0179] Furthermore, the SWLD encoding unit 405 generates encoded three-dimensional data 414 corresponding to the SWLD by encoding the extracted three-dimensional data 412 corresponding to the SWLD (S405). At this time, the SWLD encoding unit 405 adds information to the header of the encoded three-dimensional data 414 to distinguish that the encoded three-dimensional data 414 is a stream containing an SWLD.
[0180] Note that the processing order for generating encoded three-dimensional data 413 and the processing order for generating encoded three-dimensional data 414 may be reversed from the above. Also, some or all of these processes may be performed in parallel.
[0181] A parameter called "world_type" is defined as information to be added to the headers of the encoded three-dimensional data 413 and 414. If world_type=0, it indicates that the stream contains a WLD, and if world_type=1, it indicates that the stream contains an SWLD. If many other types are to be defined, the assigned number can be increased, such as world_type=2. In addition, one of the encoded three-dimensional data 413 or 414 may contain a specific flag. For example, the encoded three-dimensional data 414 may have a flag indicating that the stream contains an SWLD. In this case, the decoder can determine whether the stream contains a WLD or an SWLD based on the presence or absence of the flag.
[0182] Furthermore, the encoding method used by the WLD encoding unit 404 when encoding the WLD and the encoding method used by the SWLD encoding unit 405 when encoding the SWLD may be different.
[0183] For example, because SWLD thins out the data, it may have lower correlation with surrounding data compared to WLD. Therefore, in the encoding method used for SWLD, interpretation may be preferred over intraprediction over interprediction.
[0184] Furthermore, the encoding method used for SWLD and the encoding method used for WLD may differ in their representation of three-dimensional positions. For example, SWLD may represent the three-dimensional position of FVXL using three-dimensional coordinates, while WLD may represent the three-dimensional position using an octree, as described later, or vice versa.
[0185] Furthermore, the SWLD encoding unit 405 encodes the data such that the data size of the SWLD encoded three-dimensional data 414 is smaller than the data size of the WLD encoded three-dimensional data 413. For example, as mentioned above, SWLD may have lower correlation between data compared to WLD. This can reduce encoding efficiency, potentially causing the data size of the encoded three-dimensional data 414 to be larger than the data size of the WLD encoded three-dimensional data 413. Therefore, if the obtained encoded three-dimensional data 414 is larger than the data size of the WLD encoded three-dimensional data 413, the SWLD encoding unit 405 regenerates the encoded three-dimensional data 414 with a reduced data size by re-encoding.
[0186] For example, the SWLD extraction unit 403 regenerates the extracted three-dimensional data 412 with a reduced number of feature points, and the SWLD encoding unit 405 encodes the extracted three-dimensional data 412. Alternatively, the degree of quantization in the SWLD encoding unit 405 may be made coarser. For example, in the octree structure described later, the degree of quantization can be made coarser by rounding the data at the lowest layer.
[0187] Furthermore, if the SWLD encoding unit 405 cannot make the data size of the SWLD encoded three-dimensional data 414 smaller than the data size of the WLD encoded three-dimensional data 413, it does not need to generate the SWLD encoded three-dimensional data 414. Alternatively, the WLD encoded three-dimensional data 413 may be copied to the SWLD encoded three-dimensional data 414. In other words, the WLD encoded three-dimensional data 413 may be used as the SWLD encoded three-dimensional data 414.
[0188] Next, the configuration and operation flow of the three-dimensional data decoding device (e.g., client) according to this embodiment will be described. Figure 18 is a block diagram of the three-dimensional data decoding device 500 according to this embodiment. Figure 19 is a flowchart of the three-dimensional data decoding process by the three-dimensional data decoding device 500.
[0189] The three-dimensional data decoding device 500 shown in Figure 18 generates decoded three-dimensional data 512 or 513 by decoding encoded three-dimensional data 511. Here, encoded three-dimensional data 511 is, for example, encoded three-dimensional data 413 or 414 generated by the three-dimensional data encoding device 400.
[0190] This three-dimensional data decoding device 500 comprises an acquisition unit 501, a header analysis unit 502, a WLD decoding unit 503, and an SWLD decoding unit 504.
[0191] As shown in Figure 19, first, the acquisition unit 501 acquires the encoded three-dimensional data 511 (S501). Next, the header analysis unit 502 analyzes the header of the encoded three-dimensional data 511 and determines whether the encoded three-dimensional data 511 is a stream containing a WLD or a stream containing an SWLD (S502). For example, the world_type parameter mentioned above is referenced to make this determination.
[0192] If the encoded three-dimensional data 511 is a stream containing a WLD (Yes in S503), the WLD decoding unit 503 generates decoded three-dimensional data 512 of the WLD by decoding the encoded three-dimensional data 511 (S504). On the other hand, if the encoded three-dimensional data 511 is a stream containing a SWLD (No in S503), the SWLD decoding unit 504 generates decoded three-dimensional data 513 of the SWLD by decoding the encoded three-dimensional data 511 (S505).
[0193] Furthermore, similar to the encoding device, the decoding method used by the WLD decoding unit 503 when decoding a WLD may be different from the decoding method used by the SWLD decoding unit 504 when decoding an SWLD. For example, in the decoding method used for an SWLD, the inter-prediction method may be given priority over the intra-prediction method used for an inter-prediction method.
[0194] Furthermore, the decoding method used in SWLD and the decoding method used in WLD may differ in their representation of the three-dimensional position. For example, in SWLD, the three-dimensional position of FVXL may be represented by three-dimensional coordinates, while in WLD, the three-dimensional position may be represented by an octree, as described later, or vice versa.
[0195] Next, we will explain the octree representation, a method for representing three-dimensional positions. The VXL data contained in the three-dimensional data is converted into an octree structure and then encoded. Figure 20 shows an example of a VXL in a WLD. Figure 21 shows the octree structure of the WLD shown in Figure 20. In the example shown in Figure 20, there are three VXLs (hereinafter referred to as valid VXLs) VXL1 to VXL3 that contain point clouds. As shown in Figure 21, the octree structure consists of nodes and leaves. Each node has a maximum of eight nodes or leaves. Each leaf has VXL information. Here, among the leaves shown in Figure 21, leaves 1, 2, and 3 represent VXL1, VXL2, and VXL3 shown in Figure 20, respectively.
[0196] Specifically, each node and leaf corresponds to a three-dimensional position. Node 1 corresponds to the entire block shown in Figure 20. The block corresponding to Node 1 is divided into eight blocks, and of these eight blocks, the block containing the valid VXL is set as a node, while the other blocks are set as leaves. The block corresponding to a node is further divided into eight nodes or leaves, and this process is repeated for each level of the tree structure. In addition, all blocks at the lowest level are set as leaves.
[0197] FIG. 22 is a diagram showing an example of an SWLD generated from the WLD shown in FIG. 20. VXL1 and VXL2 shown in FIG. 20 are determined as FVXL1 and FVXL2 as a result of feature amount extraction, and are added to the SWLD. On the other hand, VXL3 is not determined as an FVXL and is not included in the SWLD. FIG. 23 is a diagram showing an octree structure of the SWLD shown in FIG. 22. In the octree structure shown in FIG. 23, leaf 3 corresponding to VXL3 shown in FIG. 21 is deleted. As a result, node 3 shown in FIG. 21 no longer has a valid VXL and is changed to a leaf. In general, the number of leaves of an SWLD is thus smaller than the number of leaves of a WLD, and the encoded three-dimensional data of the SWLD is also smaller than the encoded three-dimensional data of the WLD.
[0198] Hereinafter, modified examples of the present embodiment will be described.
[0199] For example, when a client such as an in-vehicle device performs self-position estimation, it receives an SWLD from a server, performs self-position estimation using the SWLD, and when performing obstacle detection, various methods such as a distance sensor such as a range finder, a stereo camera, or a combination of a plurality of monocular cameras may be used to perform obstacle detection based on surrounding three-dimensional information acquired by the client itself.
[0200] In addition, in general, SWLD hardly contains VXL data of flat regions. Therefore, for static obstacle detection, the server may hold a subsampled world (subWLD) obtained by subsampling a WLD, and transmit the SWLD and the subWLD to the client. This enables the client side to perform self-position estimation and obstacle detection while suppressing network bandwidth usage.
[0201] Furthermore, when clients are rendering 3D map data at high speed, it can be more convenient if the map information is in a mesh structure. Therefore, the server may generate a mesh from the World Map Data (WLD) and store it in advance as a Mesh World Data (MWLD). For example, a client can receive an MWLD if it requires a coarse 3D rendering, and a WLD if it requires a detailed 3D rendering. This can reduce network bandwidth usage.
[0202] Furthermore, the server sets the VXLs whose feature quantities are above a threshold as FVXLs, but it may calculate FVXLs using a different method. For example, the server may decide that VXLs, VLMs, SPCs, or GOSs that constitute signals or intersections are necessary for self-localization, driving assistance, or autonomous driving, and include them in the SWLD as FVXLs, FVLMs, FSPCs, or FGOSs. The above decision may also be made manually. In addition, the FVXLs obtained by the above method may be added to the FVXLs set based on feature quantities. In other words, the SWLD extraction unit 403 may further extract data corresponding to objects having predetermined attributes from the input three-dimensional data 411 as extracted three-dimensional data 412.
[0203] Furthermore, the features may be labeled separately to indicate their necessity for those applications. Additionally, the server may maintain FVXLs as a higher layer of the SWLD (e.g., lane world) necessary for self-localization of signals or intersections, driving assistance, or autonomous driving.
[0204] Furthermore, the server may also add attributes to the VXLs within the WLD for each random access unit or predetermined unit. These attributes may include, for example, information indicating whether they are necessary or unnecessary for self-localization, or information indicating whether they are important as traffic information such as signals or intersections. The attributes may also include correspondences with features (such as intersections or roads) in lane information (such as GDF: Geographic Data Files).
[0205] Additionally, the following methods may be used to update the WLD or SWLD.
[0206] Update information indicating changes such as people, construction work, or tree-lined streets (for trucks) is uploaded to the server as point cloud or metadata. Based on this upload, the server updates the WLD, and then updates the SWLD using the updated WLD.
[0207] Furthermore, if the client detects an inconsistency between the 3D information it generates during self-localization and the 3D information it receives from the server, it may send the 3D information it generates to the server along with an update notification. In this case, the server updates the SWLD using the WLD. If the SWLD is not updated, the server determines that the WLD itself is outdated.
[0208] Furthermore, while it was stated that information distinguishing between WLD and SWLD is added to the header information of the encoded stream, if there are multiple types of worlds, such as mesh worlds or lane worlds, information distinguishing between them may also be added to the header information. Also, if there are many SWLDs with different feature quantities, information distinguishing between each of them may also be added to the header information.
[0209] Furthermore, although SWLD is said to consist of FVXLs, it may also include VXLs that were not determined to be FVXLs. For example, SWLD may include adjacent VXLs used when calculating the features of FVXLs. This allows the client to calculate the features of FVXLs when it receives SWLD, even if feature information is not attached to each FVXL in SWLD. In this case, SWLD may also include information to distinguish whether each VXL is an FVXL or a VXL.
[0210] As described above, the three-dimensional data encoding device 400 extracts extracted three-dimensional data 412 (second three-dimensional data) from the input three-dimensional data 411 (first three-dimensional data) in which the feature quantity is equal to or greater than a threshold, and generates encoded three-dimensional data 414 (first encoded three-dimensional data) by encoding the extracted three-dimensional data 412.
[0211] According to this, the three-dimensional data encoding device 400 generates encoded three-dimensional data 414 by encoding data whose feature quantity is greater than or equal to a threshold. This reduces the amount of data compared to encoding the input three-dimensional data 411 as is. Therefore, the three-dimensional data encoding device 400 can reduce the amount of data transmitted.
[0212] Furthermore, the three-dimensional data encoding device 400 generates encoded three-dimensional data 413 (second encoded three-dimensional data) by encoding the input three-dimensional data 411.
[0213] According to this, the three-dimensional data encoding device 400 can selectively transmit encoded three-dimensional data 413 and encoded three-dimensional data 414, for example, depending on the intended use.
[0214] Furthermore, the extracted three-dimensional data 412 is encoded using a first encoding method, and the input three-dimensional data 411 is encoded using a second encoding method different from the first encoding method.
[0215] According to this, the three-dimensional data encoding device 400 can use encoding methods suitable for the input three-dimensional data 411 and the extracted three-dimensional data 412, respectively.
[0216] Furthermore, in the first coding method, interpretation takes precedence over intraprediction over interprediction in the second coding method.
[0217] According to this, the three-dimensional data encoding device 400 can prioritize interpretation for extracted three-dimensional data 412, where the correlation between adjacent data tends to be low.
[0218] Furthermore, the first encoding method and the second encoding method employ different representation methods for three-dimensional positions. For example, in the second encoding method, a three-dimensional position is represented by an octree, and in the first encoding method, a three-dimensional position is represented by three-dimensional coordinates.
[0219] According to this, the three-dimensional data encoding apparatus 400 can use a more suitable three-dimensional position representation method for three-dimensional data having different numbers of data (the number of VXL or FVXL).
[0220] Furthermore, at least one of the encoded three-dimensional data 413 and 414 includes an identifier indicating whether the encoded three-dimensional data is encoded three-dimensional data obtained by encoding input three-dimensional data 411, or encoded three-dimensional data obtained by encoding a part of the input three-dimensional data 411. In other words, the identifier indicates whether the encoded three-dimensional data is WLD encoded three-dimensional data 413 or SWLD encoded three-dimensional data 414.
[0221] According to this, the decoding apparatus can easily determine whether the acquired encoded three-dimensional data is the encoded three-dimensional data 413 or the encoded three-dimensional data 414.
[0222] Furthermore, the three-dimensional data encoding apparatus 400 encodes the extracted three-dimensional data 412 such that the data amount of the encoded three-dimensional data 414 is smaller than the data amount of the encoded three-dimensional data 413.
[0223] According to this, the three-dimensional data encoding apparatus 400 can make the data amount of the encoded three-dimensional data 414 smaller than the data amount of the encoded three-dimensional data 413.
[0224] Furthermore, the three-dimensional data encoding apparatus 400 further extracts, as extracted three-dimensional data 412, data corresponding to an object having a predetermined attribute from the input three-dimensional data 411. For example, the object having the predetermined attribute is an object necessary for self-position estimation, driving assist, automatic driving, or the like, such as a traffic light or an intersection.
[0225] According to this, the three-dimensional data encoding device 400 can generate encoded three-dimensional data 414 that includes the data required by the decoding device.
[0226] Furthermore, the three-dimensional data encoding device 400 (server) transmits one of the encoded three-dimensional data 413 and 414 to the client, depending on the client's status.
[0227] According to this, the three-dimensional data encoding device 400 can transmit appropriate data according to the client's status.
[0228] Furthermore, the client's status includes the client's communication status (e.g., network bandwidth) or the client's speed of movement.
[0229] Furthermore, the three-dimensional data encoding device 400 transmits one of the encoded three-dimensional data 413 and 414 to the client upon the client's request.
[0230] According to this, the three-dimensional data encoding device 400 can transmit appropriate data in response to the client's request.
[0231] Furthermore, the three-dimensional data decoding device 500 according to this embodiment decodes the encoded three-dimensional data 413 or 414 generated by the three-dimensional data encoding device 400.
[0232] In other words, the three-dimensional data decoding device 500 decodes the encoded three-dimensional data 414 obtained by encoding the extracted three-dimensional data 412, in which the feature quantities extracted from the input three-dimensional data 411 are equal to or greater than a threshold, using the first decoding method. Furthermore, the three-dimensional data decoding device 500 decodes the encoded three-dimensional data 413 obtained by encoding the input three-dimensional data 411 using a second decoding method different from the first decoding method.
[0233] According to this, the three-dimensional data decoding device 500 can selectively receive encoded three-dimensional data 414, which encodes data with feature quantities above a threshold, and encoded three-dimensional data 413, for example, depending on the intended use. This allows the three-dimensional data decoding device 500 to reduce the amount of data transmitted. Furthermore, the three-dimensional data decoding device 500 can use a decoding method suitable for the input three-dimensional data 411 and the extracted three-dimensional data 412, respectively.
[0234] Furthermore, in the first decoding method, interpretation is given priority over intraprediction in the second decoding method.
[0235] According to this, the three-dimensional data decoding device 500 can prioritize interpretation for extracted three-dimensional data where the correlation between adjacent data tends to be low.
[0236] Furthermore, the first decoding method and the second decoding method use different methods for representing three-dimensional positions. For example, in the second decoding method, the three-dimensional position is represented by an octree, while in the first decoding method, the three-dimensional position is represented by three-dimensional coordinates.
[0237] According to this, the three-dimensional data decoding device 500 can use a more suitable three-dimensional position representation method for three-dimensional data with different numbers of data (number of VXLs or FVXLs).
[0238] Furthermore, at least one of the encoded three-dimensional data 413 and 414 includes an identifier indicating whether the encoded three-dimensional data is obtained by encoding the input three-dimensional data 411 or by encoding a portion of the input three-dimensional data 411. The three-dimensional data decoding device 500 identifies the encoded three-dimensional data 413 and 414 by referring to this identifier.
[0239] According to this, the three-dimensional data decoding device 500 can easily determine whether the acquired encoded three-dimensional data is encoded three-dimensional data 413 or encoded three-dimensional data 414.
[0240] Furthermore, the three-dimensional data decoding device 500 also notifies the server of the status of the client (three-dimensional data decoding device 500). Depending on the status of the client, the three-dimensional data decoding device 500 receives one of the encoded three-dimensional data 413 and 414 transmitted from the server.
[0241] According to this, the three-dimensional data decoding device 500 can receive appropriate data according to the client's status.
[0242] Furthermore, the client's status includes the client's communication status (e.g., network bandwidth) or the client's speed of movement.
[0243] Furthermore, the three-dimensional data decoding device 500 requests one of the encoded three-dimensional data 413 and 414 from the server, and in response to the request, receives one of the encoded three-dimensional data 413 and 414 transmitted from the server.
[0244] According to this, the three-dimensional data decoding device 500 can receive appropriate data according to its intended use.
[0245] (Embodiment 3) This embodiment describes a method for transmitting and receiving three-dimensional data between vehicles. For example, three-dimensional data is transmitted and received between one vehicle and surrounding vehicles.
[0246] Figure 24 is a block diagram of the three-dimensional data creation device 620 according to this embodiment. This three-dimensional data creation device 620 creates a denser third three-dimensional data 636 by combining the received second three-dimensional data 635 with the first three-dimensional data 632 created by the three-dimensional data creation device 620, which is included in the vehicle itself.
[0247] This three-dimensional data creation device 620 comprises a three-dimensional data creation unit 621, a requested range determination unit 622, a search unit 623, a receiving unit 624, a decoding unit 625, and a synthesis unit 626.
[0248] First, the three-dimensional data creation unit 621 creates first three-dimensional data 632 using sensor information 631 detected by sensors installed in the vehicle. Next, the request range determination unit 622 determines the request range, which is the three-dimensional spatial range in which data is missing from the created first three-dimensional data 632.
[0249] Next, the search unit 623 searches for surrounding vehicles that possess three-dimensional data for the requested range, and transmits requested range information 633 indicating the requested range to the surrounding vehicles identified through the search. Next, the receiving unit 624 receives encoded three-dimensional data 634, which is an encoded stream of the requested range, from the surrounding vehicles (S624). The search unit 623 may also indiscriminately send requests to all vehicles in a specific range and receive encoded three-dimensional data 634 from those that respond. Furthermore, the search unit 623 may send requests not only to vehicles but also to objects such as traffic lights or signs and receive encoded three-dimensional data 634 from those objects.
[0250] Next, the decoding unit 625 decodes the received encoded three-dimensional data 634 to obtain the second three-dimensional data 635. Then, the combining unit 626 combines the first three-dimensional data 632 and the second three-dimensional data 635 to create a denser third three-dimensional data 636.
[0251] Next, the configuration and operation of the three-dimensional data transmission device 640 according to this embodiment will be described. Figure 25 is a block diagram of the three-dimensional data transmission device 640.
[0252] The three-dimensional data transmission device 640, for example, is included in the surrounding vehicle described above, processes the fifth three-dimensional data 652 created by the surrounding vehicle into the sixth three-dimensional data 654 requested by its own vehicle, generates encoded three-dimensional data 634 by encoding the sixth three-dimensional data 654, and transmits the encoded three-dimensional data 634 to its own vehicle.
[0253] The three-dimensional data transmission device 640 comprises a three-dimensional data creation unit 641, a receiving unit 642, an extraction unit 643, an encoding unit 644, and a transmission unit 645.
[0254] First, the three-dimensional data creation unit 641 creates fifth three-dimensional data 652 using sensor information 651 detected by sensors on surrounding vehicles. Next, the receiving unit 642 receives the requested range information 633 transmitted from its own vehicle.
[0255] Next, the extraction unit 643 processes the fifth three-dimensional data 652 into the sixth three-dimensional data 654 by extracting the three-dimensional data within the requested range indicated by the requested range information 633 from the fifth three-dimensional data 652. Next, the encoding unit 644 generates encoded three-dimensional data 634, which is an encoded stream, by encoding the sixth three-dimensional data 654. Then, the transmission unit 645 transmits the encoded three-dimensional data 634 to its own vehicle.
[0256] In this example, the vehicle itself is equipped with a three-dimensional data creation device 620, and surrounding vehicles are equipped with three-dimensional data transmission devices 640. However, each vehicle may also have the functions of both a three-dimensional data creation device 620 and a three-dimensional data transmission device 640.
[0257] (Embodiment 4) This embodiment describes the behavior of anomalies in self-localization based on a three-dimensional map.
[0258] Applications such as autonomous driving of cars, or autonomous movement of mobile objects like robots or drones, are expected to expand in the future. One example of a means to achieve such autonomous movement is for a mobile object to estimate its own position within a three-dimensional map (self-localization) and then travel according to the map.
[0259] Self-localization can be achieved by matching a three-dimensional map with three-dimensional information about the vehicle's surroundings (hereinafter referred to as "self-detection three-dimensional data") acquired by sensors such as a rangefinder (LiDAR, etc.) or stereo camera mounted on the vehicle, and estimating the vehicle's position within the three-dimensional map.
[0260] Three-dimensional maps, such as the HD maps proposed by HERE, may include not only three-dimensional point clouds but also two-dimensional map data such as road and intersection shape information, or real-time changing information such as traffic congestion and accidents. A three-dimensional map is composed of multiple layers, including three-dimensional data, two-dimensional data, and real-time changing metadata, and the device can acquire or reference only the necessary data.
[0261] The point cloud data may be SWLD as described above, or it may include point cloud data that does not contain feature points. Furthermore, the transmission and reception of point cloud data is based on one or more random access units.
[0262] The following methods can be used to match a three-dimensional map with three-dimensional vehicle detection data. For example, the device compares the shape of the point clouds in each other's point clouds and determines that areas with high similarity between feature points are in the same location. Also, if the three-dimensional map is composed of SWLDs, the device performs matching by comparing the feature points that make up the SWLD with the three-dimensional feature points extracted from the three-dimensional vehicle detection data.
[0263] Here, in order to perform self-localization with high accuracy, (A) a three-dimensional map and three-dimensional self-detection data must be acquired, and (B) the accuracy of these must meet predetermined standards. However, in the following abnormal cases, (A) or (B) cannot be met.
[0264] (1) The 3D map cannot be obtained via communication.
[0265] (2) The 3D map does not exist, or the 3D map was obtained but is corrupted.
[0266] (3) The vehicle's sensors are malfunctioning, or the accuracy of the generated 3D data for vehicle detection is insufficient due to bad weather.
[0267] The following describes the actions needed to address these abnormal cases. While a car will be used as an example, the following methods can be applied to any autonomously moving animal, such as robots or drones.
[0268] The configuration and operation of the three-dimensional information processing device according to this embodiment, for handling abnormal cases in three-dimensional maps or three-dimensional data detected by the vehicle, will be described below. Figure 26 is a block diagram showing an example configuration of the three-dimensional information processing device 700 according to this embodiment.
[0269] The three-dimensional information processing device 700 is mounted on an animal body, such as an automobile. As shown in Figure 26, the three-dimensional information processing device 700 includes a three-dimensional map acquisition unit 701, a vehicle detection data acquisition unit 702, an abnormal case determination unit 703, a response action determination unit 704, and an action control unit 705.
[0270] The three-dimensional information processing device 700 may also include two-dimensional or one-dimensional sensors (not shown) for detecting structures or animals around the vehicle, such as a camera for acquiring two-dimensional images, or a sensor for acquiring one-dimensional data using ultrasound or a laser. Furthermore, the three-dimensional information processing device 700 may also include a communication unit (not shown) for acquiring a three-dimensional map via a mobile communication network such as 4G or 5G, or via vehicle-to-vehicle communication or vehicle-to-infrastructure communication.
[0271] The 3D map acquisition unit 701 acquires a 3D map 711 of the vicinity of the travel route. For example, the 3D map acquisition unit 701 acquires the 3D map 711 via a mobile communication network, vehicle-to-vehicle communication, or vehicle-to-infrastructure communication.
[0272] Next, the vehicle detection data acquisition unit 702 acquires vehicle detection three-dimensional data 712 based on the sensor information. For example, the vehicle detection data acquisition unit 702 generates vehicle detection three-dimensional data 712 based on sensor information acquired by the sensors installed in the vehicle.
[0273] Next, the abnormal case determination unit 703 detects abnormal cases by performing predetermined checks on at least one of the acquired three-dimensional map 711 and the vehicle detection three-dimensional data 712. In other words, the abnormal case determination unit 703 determines whether at least one of the acquired three-dimensional map 711 and the vehicle detection three-dimensional data 712 is abnormal.
[0274] If an abnormal case is detected, the action determination unit 704 determines the corrective action for the abnormal case. Next, the operation control unit 705 controls the operation of each processing unit necessary for carrying out the corrective action, such as the three-dimensional map acquisition unit 701.
[0275] On the other hand, if no abnormal cases are detected, the three-dimensional information processing device 700 terminates processing.
[0276] Furthermore, the three-dimensional information processing device 700 uses the three-dimensional map 711 and the vehicle detection three-dimensional data 712 to estimate the self-position of the vehicle equipped with the three-dimensional information processing device 700. Next, the three-dimensional information processing device 700 uses the results of the self-position estimation to automatically drive the vehicle.
[0277] In this way, the three-dimensional information processing device 700 acquires map data (three-dimensional map 711) containing the first three-dimensional location information via a communication channel. For example, the first three-dimensional location information is encoded using subspaces having three-dimensional coordinate information as units, each being a collection of one or more subspaces, and containing multiple random access units, each of which can be decoded independently. For example, the first three-dimensional location information is data (SWLD) in which feature points whose three-dimensional feature quantities are greater than or equal to a predetermined threshold are encoded.
[0278] Furthermore, the three-dimensional information processing device 700 generates second three-dimensional position information (self-detection three-dimensional data 712) from the information detected by the sensor. Next, the three-dimensional information processing device 700 determines whether the first three-dimensional position information or the second three-dimensional position information is abnormal by performing an abnormality determination process on the first three-dimensional position information or the second three-dimensional position information.
[0279] If the three-dimensional information processing device 700 determines that the first three-dimensional position information or the second three-dimensional position information is abnormal, it determines a corrective action for the abnormality. Next, the three-dimensional information processing device 700 performs the necessary controls to carry out the corrective action.
[0280] As a result, the three-dimensional information processing device 700 can detect an anomaly in the first three-dimensional position information or the second three-dimensional position information and take corrective action.
[0281] (Embodiment 5) This embodiment describes a method for transmitting three-dimensional data to a following vehicle, etc.
[0282] Figure 27 is a block diagram showing an example configuration of a three-dimensional data creation device 810 according to this embodiment. This three-dimensional data creation device 810 is mounted, for example, on a vehicle. The three-dimensional data creation device 810 transmits and receives three-dimensional data with an external traffic monitoring cloud, a preceding vehicle, or a following vehicle, and also creates and stores three-dimensional data.
[0283] The three-dimensional data creation device 810 includes a data receiving unit 811, a communication unit 812, a reception control unit 813, a format conversion unit 814, a plurality of sensors 815, a three-dimensional data creation unit 816, a three-dimensional data synthesis unit 817, a three-dimensional data storage unit 818, a communication unit 819, a transmission control unit 820, a format conversion unit 821, and a data transmission unit 822.
[0284] The data receiving unit 811 receives three-dimensional data 831 from a traffic monitoring cloud or a preceding vehicle. The three-dimensional data 831 includes information such as a point cloud, visible light images, depth information, sensor position information, or speed information, including areas that cannot be detected by the vehicle's sensors 815.
[0285] The communication unit 812 communicates with the traffic monitoring cloud or the preceding vehicle and sends data transmission requests, etc., to the traffic monitoring cloud or the preceding vehicle.
[0286] The receiving control unit 813 exchanges information such as the supported format with the communication destination via the communication unit 812 and establishes communication with the communication destination.
[0287] The format conversion unit 814 generates three-dimensional data 832 by performing format conversion on the three-dimensional data 831 received by the data reception unit 811. Furthermore, if the three-dimensional data 831 is compressed or encoded, the format conversion unit 814 performs decompression or decoding.
[0288] Multiple sensors 815 are a group of sensors that acquire information from outside the vehicle, such as LiDAR, visible light cameras, or infrared cameras, and generate sensor information 833. For example, if sensor 815 is a laser sensor such as LiDAR, the sensor information 833 is three-dimensional data such as a point cloud. Note that there are not necessarily multiple sensors 815.
[0289] The three-dimensional data creation unit 816 generates three-dimensional data 834 from the sensor information 833. The three-dimensional data 834 includes information such as a point cloud, visible light image, depth information, sensor position information, or velocity information.
[0290] The three-dimensional data synthesis unit 817 synthesizes three-dimensional data 835, which includes the space in front of the preceding vehicle that cannot be detected by the vehicle's sensors 815, by combining three-dimensional data 834 created based on the vehicle's sensor information 833 with three-dimensional data 832 created by the traffic monitoring cloud or the preceding vehicle.
[0291] The three-dimensional data storage unit 818 stores the generated three-dimensional data 835, etc.
[0292] The communication unit 819 communicates with the traffic monitoring cloud or following vehicles and sends data transmission requests, etc., to the traffic monitoring cloud or following vehicles.
[0293] The transmission control unit 820 exchanges information such as the supported format with the communication destination via the communication unit 819 and establishes communication with the communication destination. The transmission control unit 820 also determines the transmission area, which is the space of the three-dimensional data to be transmitted, based on the three-dimensional data construction information of the three-dimensional data 832 generated by the three-dimensional data synthesis unit 817 and the data transmission request from the communication destination.
[0294] Specifically, the transmission control unit 820 determines a transmission area that includes the space in front of its own vehicle that cannot be detected by the sensors of the following vehicle, in response to a data transmission request from the traffic monitoring cloud or a following vehicle. The transmission control unit 820 also determines the transmission area by determining whether the transmissionable space or the transmitted space has been updated based on the three-dimensional data construction information. For example, the transmission control unit 820 determines the transmission area to be the area specified in the data transmission request and in which the corresponding three-dimensional data 835 exists. The transmission control unit 820 then notifies the format conversion unit 821 of the format supported by the communication destination and the transmission area.
[0295] The format conversion unit 821 generates three-dimensional data 837 by converting the three-dimensional data 836 in the transmission area from the three-dimensional data 835 stored in the three-dimensional data storage unit 818 to a format supported by the receiving side. The format conversion unit 821 may also reduce the amount of data by compressing or encoding the three-dimensional data 837.
[0296] The data transmission unit 822 transmits three-dimensional data 837 to a traffic monitoring cloud or following vehicles. This three-dimensional data 837 includes, for example, information such as a point cloud in front of the vehicle, including areas that are blind spots for following vehicles, visible light images, depth information, or sensor position information.
[0297] Although this example describes a case where format conversion is performed by the format conversion units 814 and 821, format conversion is not required.
[0298] With this configuration, the three-dimensional data creation device 810 acquires three-dimensional data 831 from an external source for areas that cannot be detected by the vehicle's sensors 815, and generates three-dimensional data 835 by combining the three-dimensional data 831 with three-dimensional data 834 based on sensor information 833 detected by the vehicle's sensors 815. In this way, the three-dimensional data creation device 810 can generate three-dimensional data for areas that cannot be detected by the vehicle's sensors 815.
[0299] Furthermore, the three-dimensional data creation device 810 can transmit three-dimensional data, including the space in front of its own vehicle that cannot be detected by the sensors of the following vehicle, to the traffic monitoring cloud or following vehicle in response to a data transmission request from the traffic monitoring cloud or following vehicle.
[0300] (Embodiment 6) Embodiment 5 describes an example in which a client device such as a vehicle transmits three-dimensional data to another vehicle or a server such as a traffic monitoring cloud. In this embodiment, the client device transmits sensor information obtained from the sensor to the server or another client device.
[0301] First, the system configuration according to this embodiment will be described. Figure 28 is a diagram showing the configuration of the three-dimensional map and sensor information transmission and reception system according to this embodiment. This system includes a server 901 and client devices 902A and 902B. When client devices 902A and 902B are not specifically distinguished, they will also be referred to as client device 902.
[0302] The client device 902 is, for example, an in-vehicle device mounted on a moving object such as a vehicle. The server 901 is, for example, a traffic monitoring cloud and is capable of communicating with multiple client devices 902.
[0303] Server 901 transmits a three-dimensional map composed of point clouds to client device 902. Note that the composition of the three-dimensional map is not limited to point clouds; it may also represent other three-dimensional data, such as a mesh structure.
[0304] The client device 902 transmits sensor information acquired by the client device 902 to the server 901. The sensor information includes, for example, at least one of the following: LiDAR acquisition information, visible light image, infrared image, depth image, sensor position information, and velocity information.
[0305] The data transmitted and received between the server 901 and the client device 902 may be compressed to reduce data size, or it may be left uncompressed to maintain data accuracy. When data is compressed, a three-dimensional compression method based on an octave structure, for example, can be used for point clouds. In addition, a two-dimensional image compression method can be used for visible light images, infrared images, and depth images. A two-dimensional image compression method is, for example, MPEG-4 AVC or HEVC, which are standardized by MPEG.
[0306] Furthermore, in response to a request from the client device 902 to send a 3D map, the server 901 sends a 3D map managed by the server 901 to the client device 902. The server 901 may also send a 3D map without waiting for a request from the client device 902. For example, the server 901 may broadcast a 3D map to one or more client devices 902 located in a predetermined space. Alternatively, the server 901 may send a 3D map appropriate to the location of the client device 902 at regular intervals after receiving a transmission request from the client device 902. The server 901 may also send a 3D map to the client device 902 whenever the 3D map managed by the server 901 is updated.
[0307] The client device 902 sends a request to the server 901 to send a three-dimensional map. For example, if the client device 902 wants to perform self-position estimation while driving, the client device 902 sends a request to the server 901 to send a three-dimensional map.
[0308] Furthermore, the client device 902 may request the server 901 to send a 3D map in the following cases: If the 3D map held by the client device 902 is outdated, the client device 902 may request the server 901 to send a 3D map. For example, if a certain period of time has elapsed since the client device 902 acquired the 3D map, the client device 902 may request the server 901 to send a 3D map.
[0309] Client device 902 may request server 901 to send the three-dimensional map to the server 901 a certain time before client device 902 leaves the space represented by the three-dimensional map held by client device 902. For example, client device 902 may request server 901 to send the three-dimensional map to the server 901 if it is within a predetermined distance from the boundary of the space represented by the three-dimensional map held by client device 902. Furthermore, if the movement path and speed of client device 902 are known, the time when client device 902 leaves the space represented by the three-dimensional map held by client device 902 may be predicted based on these.
[0310] If the error in the alignment between the three-dimensional data created by the client device 902 from sensor information and the three-dimensional map exceeds a certain level, the client device 902 may request the server 901 to send the three-dimensional map.
[0311] The client device 902 transmits sensor information to the server 901 in response to a request for transmission of sensor information sent from the server 901. The client device 902 may also send sensor information to the server 901 without waiting for a request for transmission of sensor information from the server 901. For example, once the client device 902 receives a request for transmission of sensor information from the server 901, it may periodically transmit sensor information to the server 901 for a certain period. Furthermore, if the error in the alignment between the three-dimensional data created by the client device 902 based on the sensor information and the three-dimensional map obtained from the server 901 exceeds a certain level, the client device 902 may determine that a change has occurred in the three-dimensional map around the client device 902 and transmit this information, along with the sensor information, to the server 901.
[0312] Server 901 requests client device 902 to transmit sensor information. For example, Server 901 receives location information of client device 902, such as GPS, from client device 902. Based on the location information of client device 902, if Server 901 determines that client device 902 is approaching an area with little information on the three-dimensional map managed by Server 901, it requests client device 902 to transmit sensor information in order to generate a new three-dimensional map. Server 901 may also request sensor information transmission if it wants to update the three-dimensional map, check road conditions during snowfall or disasters, check traffic congestion, or check incidents and accidents.
[0313] Furthermore, the client device 902 may set the amount of sensor information data to send to the server 901 depending on the communication status or bandwidth at the time of receiving the sensor information transmission request from the server 901. Setting the amount of sensor information data to send to the server 901 means, for example, increasing or decreasing the data itself, or selecting an appropriate compression method.
[0314] Figure 29 is a block diagram showing an example configuration of the client device 902. The client device 902 receives a three-dimensional map composed of a point cloud, etc., from the server 901, and estimates its own position from the three-dimensional data created based on the sensor information of the client device 902. The client device 902 also transmits the acquired sensor information to the server 901.
[0315] The client device 902 includes a data receiving unit 1011, a communication unit 1012, a reception control unit 1013, a format conversion unit 1014, a plurality of sensors 1015, a three-dimensional data creation unit 1016, a three-dimensional image processing unit 1017, a three-dimensional data storage unit 1018, a format conversion unit 1019, a communication unit 1020, a transmission control unit 1021, and a data transmission unit 1022.
[0316] The data receiving unit 1011 receives the three-dimensional map 1031 from the server 901. The three-dimensional map 1031 is data that includes point clouds such as WLD or SWLD. The three-dimensional map 1031 may contain either compressed or uncompressed data.
[0317] The communication unit 1012 communicates with the server 901 and sends data transmission requests (for example, a request to transmit a 3D map) to the server 901.
[0318] The receiving control unit 1013 exchanges information such as the supported format with the communication destination via the communication unit 1012 and establishes communication with the communication destination.
[0319] The format conversion unit 1014 generates a three-dimensional map 1032 by performing format conversion on the three-dimensional map 1031 received by the data reception unit 1011. Furthermore, if the three-dimensional map 1031 is compressed or encoded, the format conversion unit 1014 performs decompression or decoding. However, if the three-dimensional map 1031 is uncompressed data, the format conversion unit 1014 does not perform decompression or decoding.
[0320] Multiple sensors 1015 are a group of sensors that acquire external information from the vehicle on which the client device 902 is installed, such as LiDAR, visible light cameras, infrared cameras, or depth sensors, and generate sensor information 1033. For example, if sensor 1015 is a laser sensor such as LiDAR, the sensor information 1033 is three-dimensional data such as a point cloud (point cloud data). Note that there are not necessarily multiple sensors 1015.
[0321] The three-dimensional data creation unit 1016 creates three-dimensional data 1034 of the vehicle's surroundings based on the sensor information 1033. For example, the three-dimensional data creation unit 1016 uses information acquired by LiDAR and visible light images obtained by a visible light camera to create point cloud data with color information of the vehicle's surroundings.
[0322] The three-dimensional image processing unit 1017 uses the received three-dimensional map 1032, such as a point cloud, and the three-dimensional data 1034 of the vehicle's surroundings generated from sensor information 1033 to perform self-position estimation processing for the vehicle. Alternatively, the three-dimensional image processing unit 1017 may create three-dimensional data 1035 of the vehicle's surroundings by combining the three-dimensional map 1032 and the three-dimensional data 1034, and then perform self-position estimation processing using the created three-dimensional data 1035.
[0323] The three-dimensional data storage unit 1018 stores the three-dimensional map 1032, three-dimensional data 1034, and three-dimensional data 1035, etc.
[0324] The format conversion unit 1019 generates sensor information 1037 by converting the sensor information 1033 to a format supported by the receiving side. The format conversion unit 1019 may also reduce the amount of data by compressing or encoding the sensor information 1037. Furthermore, the format conversion unit 1019 may omit processing if format conversion is not necessary. The format conversion unit 1019 may also control the amount of data transmitted according to the specified transmission range.
[0325] The communication unit 1020 communicates with the server 901 and receives data transmission requests (sensor information transmission requests), etc., from the server 901.
[0326] The transmission control unit 1021 exchanges information such as the supported format with the communication destination via the communication unit 1020 and establishes communication.
[0327] The data transmission unit 1022 transmits sensor information 1037 to the server 901. The sensor information 1037 includes information acquired by multiple sensors 1015, such as information acquired by LiDAR, brightness images acquired by a visible light camera, infrared images acquired by an infrared camera, depth images acquired by a depth sensor, sensor position information, and velocity information.
[0328] Next, the configuration of server 901 will be described. Figure 30 is a block diagram showing an example configuration of server 901. Server 901 receives sensor information transmitted from client device 902 and creates three-dimensional data based on the received sensor information. Server 901 updates the three-dimensional map it manages using the created three-dimensional data. In addition, in response to a request from client device 902 to transmit the three-dimensional map, server 901 transmits the updated three-dimensional map to client device 902.
[0329] Server 901 comprises a data receiving unit 1111, a communication unit 1112, a reception control unit 1113, a format conversion unit 1114, a three-dimensional data creation unit 1116, a three-dimensional data synthesis unit 1117, a three-dimensional data storage unit 1118, a format conversion unit 1119, a communication unit 1120, a transmission control unit 1121, and a data transmission unit 1122.
[0330] The data receiving unit 1111 receives sensor information 1037 from the client device 902. The sensor information 1037 includes, for example, information acquired by LiDAR, brightness images acquired by a visible light camera, infrared images acquired by an infrared camera, depth images acquired by a depth sensor, sensor position information, and velocity information.
[0331] The communication unit 1112 communicates with the client device 902 and sends data transmission requests (for example, requests to transmit sensor information) to the client device 902.
[0332] The receiving control unit 1113 exchanges information such as the supported format with the communication destination via the communication unit 1112 and establishes communication.
[0333] The format conversion unit 1114 generates sensor information 1132 by decompressing or decoding the received sensor information 1037 if it is compressed or encoded. However, the format conversion unit 1114 does not perform decompression or decoding if the sensor information 1037 is uncompressed data.
[0334] The three-dimensional data creation unit 1116 creates three-dimensional data 1134 of the area around the client device 902 based on the sensor information 1132. For example, the three-dimensional data creation unit 1116 uses information acquired by LiDAR and visible light images obtained by a visible light camera to create point cloud data with color information of the area around the client device 902.
[0335] The three-dimensional data synthesis unit 1117 updates the three-dimensional map 1135 managed by the server 901 by synthesizing the three-dimensional data 1134, which was created based on the sensor information 1132, with the three-dimensional map 1135.
[0336] The three-dimensional data storage unit 1118 stores three-dimensional maps 1135, etc.
[0337] The format conversion unit 1119 generates a three-dimensional map 1031 by converting the three-dimensional map 1135 to a format supported by the receiving side. The format conversion unit 1119 may also reduce the amount of data by compressing or encoding the three-dimensional map 1135. Furthermore, the format conversion unit 1119 may omit processing if format conversion is not necessary. The format conversion unit 1119 may also control the amount of data transmitted according to the specified transmission range.
[0338] The communication unit 1120 communicates with the client device 902 and receives data transmission requests (such as requests to transmit a three-dimensional map) from the client device 902.
[0339] The transmission control unit 1121 exchanges information such as the supported format with the communication destination via the communication unit 1120 and establishes communication.
[0340] The data transmission unit 1122 transmits the three-dimensional map 1031 to the client device 902. The three-dimensional map 1031 is data that includes point clouds such as WLD or SWLD. The three-dimensional map 1031 may contain either compressed or uncompressed data.
[0341] Next, we will describe the operation flow of the client device 902. Figure 31 is a flowchart showing the operation of the client device 902 when acquiring a three-dimensional map.
[0342] First, the client device 902 requests the server 901 to transmit a three-dimensional map (such as a point cloud) (S1001). At this time, the client device 902 may also transmit its own location information obtained by GPS or the like, and request the server 901 to transmit a three-dimensional map related to that location information.
[0343] Next, the client device 902 receives a three-dimensional map from the server 901 (S1002). If the received three-dimensional map is compressed data, the client device 902 decodes the received three-dimensional map to generate an uncompressed three-dimensional map (S1003).
[0344] Next, the client device 902 creates three-dimensional data 1034 of the area around the client device 902 from sensor information 1033 obtained from multiple sensors 1015 (S1004). Then, the client device 902 estimates its own position using the three-dimensional map 1032 received from the server 901 and the three-dimensional data 1034 created from the sensor information 1033 (S1005).
[0345] Figure 32 is a flowchart showing the operation of the client device 902 when transmitting sensor information. First, the client device 902 receives a request to transmit sensor information from the server 901 (S1011). Upon receiving the transmission request, the client device 902 transmits sensor information 1037 to the server 901 (S1012). If the sensor information 1033 includes multiple pieces of information obtained from multiple sensors 1015, the client device 902 may generate sensor information 1037 by compressing each piece of information using a compression method suitable for each piece of information.
[0346] Next, the operation flow of server 901 will be described. Figure 33 is a flowchart showing the operation of server 901 when acquiring sensor information. First, server 901 requests client device 902 to send sensor information (S1021). Next, server 901 receives sensor information 1037 sent from client device 902 in response to the request (S1022). Next, server 901 creates three-dimensional data 1134 using the received sensor information 1037 (S1023). Next, server 901 reflects the created three-dimensional data 1134 in the three-dimensional map 1135 (S1024).
[0347] Figure 34 is a flowchart illustrating the operation of server 901 when transmitting a three-dimensional map. First, server 901 receives a request to transmit a three-dimensional map from client device 902 (S1031). Upon receiving the request to transmit a three-dimensional map, server 901 transmits the three-dimensional map 1031 to client device 902 (S1032). At this time, server 901 may extract a three-dimensional map of the vicinity of client device 902 according to its location information and transmit the extracted three-dimensional map. Alternatively, server 901 may compress the three-dimensional map composed of a point cloud using, for example, an octave tree compression method, and transmit the compressed three-dimensional map.
[0348] Modifications of this embodiment will be described below.
[0349] Server 901 uses sensor information 1037 received from client device 902 to create three-dimensional data 1134 of the area around client device 902. Next, server 901 calculates the difference between the created three-dimensional data 1134 and the three-dimensional map 1135 of the same area managed by server 901 by matching them. If the difference is greater than or equal to a predetermined threshold, server 901 determines that some kind of abnormality has occurred around client device 902. For example, when ground subsidence occurs due to a natural disaster such as an earthquake, a large difference may occur between the three-dimensional map 1135 managed by server 901 and the three-dimensional data 1134 created based on sensor information 1037.
[0350] The sensor information 1037 may include information indicating at least one of the following: the type of sensor, the performance of the sensor, and the model number of the sensor. Furthermore, a class ID corresponding to the sensor's performance may be added to the sensor information 1037. For example, if the sensor information 1037 is information acquired by a LiDAR, it is conceivable to assign identifiers to the sensor's performance, such as class 1 for sensors that can acquire information with accuracy in the millimeter range, class 2 for sensors that can acquire information with accuracy in the centimeter range, and class 3 for sensors that can acquire information with accuracy in the meter range. The server 901 may also estimate the sensor's performance information from the model number of the client device 902. For example, if the client device 902 is mounted in a vehicle, the server 901 may determine the sensor's specifications from the vehicle's make and model. In this case, the server 901 may have previously acquired information about the vehicle's make and model, or this information may be included in the sensor information. The server 901 may also use the acquired sensor information 1037 to switch the degree of correction applied to the three-dimensional data 1134 created using the sensor information 1037. For example, if the sensor performance is high precision (Class 1), the server 901 does not perform any correction on the three-dimensional data 1134. If the sensor performance is low precision (Class 3), the server 901 applies a correction to the three-dimensional data 1134 according to the accuracy of the sensor. For example, the lower the accuracy of the sensor, the stronger the degree (intensity) of the correction applied by the server 901.
[0351] Server 901 may simultaneously send requests for the transmission of sensor information to multiple client devices 902 located in a given space. When Server 901 receives multiple sensor information from multiple client devices 902, it is not necessary to use all of the sensor information to create the three-dimensional data 1134. For example, it may select which sensor information to use depending on the performance of the sensors. For example, when updating the three-dimensional map 1135, Server 901 may select high-precision sensor information (Class 1) from the multiple sensor information received and use the selected sensor information to create the three-dimensional data 1134.
[0352] Server 901 is not limited to servers such as traffic monitoring clouds, but may also be other client devices (in-vehicle). Figure 35 shows the system configuration in this case.
[0353] For example, client device 902C requests sensor information from a nearby client device 902A and obtains the sensor information from client device 902A. Then, client device 902C uses the obtained sensor information from client device 902A to create three-dimensional data and updates the three-dimensional map of client device 902C. In this way, client device 902C can generate a three-dimensional map of the space obtainable from client device 902A, taking advantage of the performance of client device 902C. For example, this case is likely to occur when client device 902C has high performance.
[0354] In this case, client device 902A, which provided the sensor information, is granted the right to acquire the high-precision three-dimensional map generated by client device 902C. Client device 902A receives the high-precision three-dimensional map from client device 902C in accordance with that right.
[0355] Furthermore, client device 902C may send requests for the transmission of sensor information to multiple nearby client devices 902 (client devices 902A and 902B). If the sensor of client device 902A or client device 902B is high-performance, client device 902C can create three-dimensional data using the sensor information obtained from this high-performance sensor.
[0356] Figure 36 is a block diagram showing the functional configuration of the server 901 and the client device 902. The server 901 includes, for example, a three-dimensional map compression / decoding processing unit 1201 that compresses and decodes three-dimensional maps, and a sensor information compression / decoding processing unit 1202 that compresses and decodes sensor information.
[0357] The client device 902 comprises a three-dimensional map decoding processing unit 1211 and a sensor information compression processing unit 1212. The three-dimensional map decoding processing unit 1211 receives encoded data of the compressed three-dimensional map, decodes the encoded data, and obtains the three-dimensional map. The sensor information compression processing unit 1212 compresses the sensor information itself instead of the three-dimensional data created from the acquired sensor information, and sends the encoded data of the compressed sensor information to the server 901. With this configuration, the client device 902 only needs to internally store a processing unit (device or LSI) that performs the processing of decoding the three-dimensional map (point cloud, etc.), and does not need to internally store a processing unit that performs the processing of compressing the three-dimensional data of the three-dimensional map (point cloud, etc.). This reduces the cost and power consumption of the client device 902.
[0358] As described above, the client device 902 according to this embodiment is mounted on a mobile body and creates three-dimensional data 1034 of the surrounding area of the mobile body from sensor information 1033 indicating the surrounding conditions of the mobile body obtained by a sensor 1015 mounted on the mobile body. The client device 902 estimates the self-position of the mobile body using the created three-dimensional data 1034. The client device 902 transmits the acquired sensor information 1033 to the server 901 or another mobile body 902.
[0359] According to this, the client device 902 transmits sensor information 1033 to the server 901, etc. This may reduce the amount of data transmitted compared to transmitting three-dimensional data. In addition, since the client device 902 does not need to perform processing such as compression or encoding of three-dimensional data, the processing load on the client device 902 can be reduced. Therefore, the client device 902 can achieve a reduction in the amount of data transmitted or a simplification of the device configuration.
[0360] Furthermore, the client device 902 sends a request to the server 901 to send a three-dimensional map, and receives the three-dimensional map 1031 from the server 901. In estimating its own position, the client device 902 uses the three-dimensional data 1034 and the three-dimensional map 1032 to estimate its own position.
[0361] Furthermore, the sensor information 1033 includes at least one of the following: information obtained from the laser sensor, brightness image, infrared image, depth image, sensor position information, and sensor velocity information.
[0362] Furthermore, sensor information 1033 includes information indicating the performance of the sensor.
[0363] Furthermore, the client device 902 encodes or compresses the sensor information 1033, and when transmitting the sensor information, it transmits the encoded or compressed sensor information 1037 to the server 901 or another mobile device 902. This allows the client device 902 to reduce the amount of data transmitted.
[0364] For example, the client device 902 includes a processor and memory, and the processor uses the memory to perform the above processing.
[0365] Furthermore, the server 901 according to this embodiment is capable of communicating with a client device 902 mounted on the mobile body, and receives sensor information 1037 from the client device 902 that indicates the surrounding conditions of the mobile body, obtained by a sensor 1015 mounted on the mobile body. The server 901 creates three-dimensional data 1134 of the surroundings of the mobile body from the received sensor information 1037.
[0366] According to this, the server 901 creates three-dimensional data 1134 using sensor information 1037 transmitted from the client device 902. This may reduce the amount of data transmitted compared to when the client device 902 transmits the three-dimensional data. In addition, since the client device 902 does not need to perform processing such as compression or encoding of the three-dimensional data, the processing load on the client device 902 can be reduced. Therefore, the server 901 can reduce the amount of data transmitted or simplify the configuration of the device.
[0367] Furthermore, the server 901 also sends a request to the client device 902 to transmit sensor information.
[0368] Furthermore, the server 901 updates the three-dimensional map 1135 using the created three-dimensional data 1134 and sends the three-dimensional map 1135 to the client device 902 in response to a request from the client device 902 to send the three-dimensional map 1135.
[0369] Furthermore, the sensor information 1037 includes at least one of the following: information obtained from the laser sensor, brightness image, infrared image, depth image, sensor position information, and sensor velocity information.
[0370] Furthermore, sensor information 1037 includes information indicating the performance of the sensor.
[0371] Furthermore, the server 901 corrects the three-dimensional data according to the performance of the sensor. This allows the three-dimensional data creation method to improve the quality of the three-dimensional data.
[0372] Furthermore, when receiving sensor information, the server 901 receives multiple pieces of sensor information 1037 from multiple client devices 902, and selects the sensor information 1037 to be used to create the three-dimensional data 1134 based on the multiple pieces of information indicating the performance of the sensors contained in the multiple pieces of sensor information 1037. In this way, the server 901 can improve the quality of the three-dimensional data 1134.
[0373] Furthermore, the server 901 decodes or decodes the received sensor information 1037 and creates three-dimensional data 1134 from the decoded or decoded sensor information 1132. This allows the server 901 to reduce the amount of data transmitted.
[0374] For example, server 901 is equipped with a processor and memory, and the processor uses the memory to perform the above processing.
[0375] (Embodiment 7) This embodiment describes a method for encoding and decoding three-dimensional data using interpretation processing.
[0376] Figure 37 is a block diagram of a three-dimensional data encoding device 1300 according to this embodiment. This three-dimensional data encoding device 1300 generates an encoded bitstream (hereinafter also simply referred to as a bitstream), which is an encoded signal, by encoding three-dimensional data. As shown in Figure 37, the three-dimensional data encoding device 1300 includes a division unit 1301, a subtraction unit 1302, a conversion unit 1303, a quantization unit 1304, an inverse quantization unit 1305, an inverse conversion unit 1306, an addition unit 1307, a reference volume memory 1308, an intra prediction unit 1309, a reference space memory 1310, an inter prediction unit 1311, a prediction control unit 1312, and an entropy encoding unit 1313.
[0377] The splitting unit 1301 divides each space (SPC) contained in the three-dimensional data into multiple volumes (VLMs), which are encoding units. The splitting unit 1301 also converts the voxels within each volume into an octree representation. The splitting unit 1301 may also make the spaces and volumes the same size and convert the spaces into an octree representation. Furthermore, the splitting unit 1301 may add information necessary for octree conversion (such as depth information) to the bitstream header, etc.
[0378] The subtraction unit 1302 calculates the difference between the volume output from the division unit 1301 (the volume to be encoded) and the predicted volume generated by the intra-prediction or inter-prediction described later, and outputs the calculated difference as the predicted residual to the conversion unit 1303. Figure 38 shows an example of the calculation of the predicted residual. The bit sequences of the volume to be encoded and the predicted volume shown here are, for example, position information indicating the positions of three-dimensional points (e.g., point clouds) included in the volume.
[0379] The following describes the octree representation and the voxel scan order. A volume is converted into an octree structure (octreeized) and then encoded. An octree structure consists of nodes and leaves. Each node has eight nodes or leaves, and each leaf has voxel (VXL) information. Figure 39 shows an example of the structure of a volume containing multiple voxels. Figure 40 shows an example of the volume shown in Figure 39 converted into an octree structure. Here, among the leaves shown in Figure 40, leaves 1, 2, and 3 represent the voxels VXL1, VXL2, and VXL3 shown in Figure 39, respectively, and represent a VXL containing a point cloud (hereinafter referred to as effective VXL).
[0380] An octree is represented, for example, by a binary sequence of 0s and 1s. For example, if nodes or valid VXLs are assigned the value 1 and all others the value 0, then each node and leaf is assigned the binary sequence shown in Figure 40. This binary sequence is then scanned according to the breadth-first or depth-first scan order. For example, if scanned in breadth-first order, the binary sequence shown in Figure 41A is obtained. If scanned in depth-first order, the binary sequence shown in Figure 41B is obtained. The binary sequence obtained by this scan is encoded by entropy coding to reduce its information content.
[0381] Next, we will explain the depth information in octree representations. In octree representations, the depth is used to control the level of granularity to which the point cloud information contained within a volume is retained. Setting a high depth allows for the reproduction of point cloud information at a finer level, but increases the amount of data required to represent nodes and leaves. Conversely, setting a low depth reduces the amount of data, but multiple point clouds with different locations and colors are treated as being at the same location and with the same color, resulting in the loss of information that the original point cloud information contained in the data.
[0382] For example, Figure 42 shows an example where the octree with depth=2 shown in Figure 40 is represented by an octree with depth=1. The octree shown in Figure 42 has less data than the octree shown in Figure 40. In other words, the octree shown in Figure 42 has fewer bits after binary conversion than the octree shown in Figure 42. Here, leaf 1 and leaf 2 shown in Figure 40 are represented by leaf 1 shown in Figure 41. In other words, the information that leaf 1 and leaf 2 shown in Figure 40 were in different positions is lost.
[0383] Figure 43 shows the volume corresponding to the octree shown in Figure 42. VXL1 and VXL2 shown in Figure 39 correspond to VXL12 shown in Figure 43. In this case, the three-dimensional data encoding device 1300 generates the color information of VXL12 shown in Figure 43 from the color information of VXL1 and VXL2 shown in Figure 39. For example, the three-dimensional data encoding device 1300 calculates the average value, median value, or weighted average value of the color information of VXL1 and VXL2 as the color information of VXL12. In this way, the three-dimensional data encoding device 1300 may control the reduction of data volume by changing the depth of the octree.
[0384] The three-dimensional data encoding device 1300 may set the depth information of the octree in units of worlds, spaces, or volumes. In this case, the three-dimensional data encoding device 1300 may add the depth information to the header information of the world, the header information of the space, or the header information of the volume. Alternatively, the same value may be used for the depth information for all worlds, spaces, and volumes of different time periods. In this case, the three-dimensional data encoding device 1300 may add the depth information to the header information that manages the world for all time periods.
[0385] If the voxels contain color information, the conversion unit 1303 applies a frequency transformation, such as an orthogonal transformation, to the predicted residuals of the color information of the voxels in the volume. For example, the conversion unit 1303 creates a one-dimensional array by scanning the predicted residuals in a certain scan order. Then, the conversion unit 1303 converts the created one-dimensional array into the frequency domain by applying a one-dimensional orthogonal transformation to it. As a result, when the predicted residual values in the volume are close, the values of the low-frequency components become larger and the values of the high-frequency components become smaller. Therefore, the quantization unit 1304 can reduce the code amount more efficiently.
[0386] Furthermore, the transformation unit 1303 may use orthogonal transformations of two or more dimensions, rather than just one dimension. For example, the transformation unit 1303 maps the predicted residuals in a certain scan order to a two-dimensional array and applies a two-dimensional orthogonal transformation to the resulting two-dimensional array. Alternatively, the transformation unit 1303 may select an orthogonal transformation method from among multiple orthogonal transformation methods. In this case, the three-dimensional data encoding device 1300 adds information to the bitstream indicating which orthogonal transformation method was used. Alternatively, the transformation unit 1303 may select an orthogonal transformation method from among multiple orthogonal transformation methods of different dimensions. In this case, the three-dimensional data encoding device 1300 adds information to the bitstream indicating which dimension's orthogonal transformation method was used.
[0387] For example, the conversion unit 1303 matches the scan order of the predicted residuals to the scan order in the octree within the volume (such as breadth-first or depth-first). This eliminates the need to add information indicating the scan order of the predicted residuals to the bitstream, thus reducing overhead. Alternatively, the conversion unit 1303 may apply a scan order different from the octree scan order. In this case, the three-dimensional data encoding device 1300 adds information indicating the scan order of the predicted residuals to the bitstream. This allows the three-dimensional data encoding device 1300 to efficiently encode the predicted residuals. Furthermore, the three-dimensional data encoding device 1300 may add information (such as a flag) to the bitstream indicating whether or not to apply the octree scan order, and if the octree scan order is not applied, it may add information indicating the scan order of the predicted residuals to the bitstream.
[0388] The conversion unit 1303 may convert not only the predicted residual of color information but also other attribute information possessed by the voxels. For example, the conversion unit 1303 may convert and encode information such as reflectance obtained when a point cloud is acquired by LiDAR or the like.
[0389] The conversion unit 1303 may skip processing if the space does not contain attribute information such as color information. The three-dimensional data encoding device 1300 may also add information (flags) to the bitstream indicating whether or not to skip processing by the conversion unit 1303.
[0390] The quantization unit 1304 generates quantization coefficients by quantizing the frequency components of the predicted residual generated by the conversion unit 1303 using quantization control parameters. This reduces the amount of information. The generated quantization coefficients are output to the entropy coding unit 1313. The quantization unit 1304 may control the quantization control parameters in world units, space units, or volume units. In this case, the three-dimensional data coding device 1300 adds the quantization control parameters to the respective header information. The quantization unit 1304 may also perform quantization control by changing the weight for each frequency component of the predicted residual. For example, the quantization unit 1304 may finely quantize low-frequency components and coarsely quantize high-frequency components. In this case, the three-dimensional data coding device 1300 may add parameters representing the weight of each frequency component to the header.
[0391] The quantization unit 1304 may skip processing if the space does not contain attribute information such as color information. The three-dimensional data encoding device 1300 may also add information (flags) to the bitstream indicating whether or not to skip processing by the quantization unit 1304.
[0392] The inverse quantization unit 1305 generates inverse quantization coefficients of the prediction residual by performing inverse quantization on the quantization coefficients generated by the quantization unit 1304 using quantization control parameters, and outputs the generated inverse quantization coefficients to the inverse transform unit 1306.
[0393] The inverse transform unit 1306 generates the predicted residual after the inverse transform by applying the inverse transform to the inverse quantization coefficients generated by the inverse quantization unit 1305. Since this predicted residual after the inverse transform is the predicted residual generated after quantization, it does not need to perfectly match the predicted residual output by the transform unit 1303.
[0394] The summing unit 1307 adds the predicted residual after applying the inverse transform, generated by the inverse transform unit 1306, to the predicted volume generated by the intra-prediction or inter-prediction described later, which was used to generate the predicted residual before quantization, to generate a reconstructed volume. This reconstructed volume is stored in the reference volume memory 1308 or the reference space memory 1310.
[0395] The intra prediction unit 1309 generates a predicted volume for the volume to be encoded using attribute information of adjacent volumes stored in the reference volume memory 1308. Attribute information includes voxel color information or reflectance. The intra prediction unit 1309 generates predicted values for the color information or reflectance of the volume to be encoded.
[0396] Figure 44 is a diagram illustrating the operation of the intra prediction unit 1309. For example, the intra prediction unit 1309 generates a predicted volume for the volume to be encoded (volume idx=3), as shown in Figure 44, from the adjacent volume (volume idx=0). Here, volume idx is identifier information attached to volumes within the space, and a different value is assigned to each volume. The order in which volume idx is assigned may be the same as the encoding order, or it may be a different order. For example, the intra prediction unit 1309 uses the average value of the color information of the voxels contained in the adjacent volume, volume idx=0, as the predicted value of the color information of the volume to be encoded, as shown in Figure 44. In this case, a prediction residual is generated by subtracting the predicted value of the color information from the color information of each voxel contained in the volume to be encoded. Processing from the conversion unit 1303 onwards is performed on this prediction residual. In this case, the three-dimensional data encoding device 1300 adds the adjacent volume information and the prediction mode information to the bitstream. Here, adjacent volume information refers to information indicating the adjacent volume used for prediction, for example, the volume IDX of the adjacent volume used for prediction. Prediction mode information refers to the mode used to generate the predicted volume. A mode is, for example, an average mode that generates predicted values from the average values of voxels within the adjacent volume, or an intermediate mode that generates predicted values from the median values of voxels within the adjacent volume.
[0397] The intra-prediction unit 1309 may generate a predicted volume from multiple adjacent volumes. For example, in the configuration shown in Figure 44, the intra-prediction unit 1309 generates predicted volume 0 from the volume with volume idx=0 and predicted volume 1 from the volume with volume idx=1. The intra-prediction unit 1309 then generates the average of predicted volume 0 and predicted volume 1 as the final predicted volume. In this case, the three-dimensional data encoding device 1300 may add multiple volume idx values from the multiple volumes used to generate the predicted volume to the bitstream.
[0398] Figure 45 schematically shows the interpretation process according to this embodiment. The interpretation unit 1311 encodes (interprets) a space (SPC) at a certain time T_Cur using an encoded space at a different time T_LX. In this case, the interpretation unit 1311 performs the encoding process by applying rotation and translation processing to the encoded space at the different time T_LX.
[0399] Furthermore, the three-dimensional data encoding device 1300 adds RT information related to rotation and translation processing applied to a space at a different time T_LX to the bitstream. A different time T_LX is, for example, a time T_L0 prior to a certain time T_Cur. In this case, the three-dimensional data encoding device 1300 may also add RT information RT_L0 related to rotation and translation processing applied to the space at time T_L0 to the bitstream.
[0400] Alternatively, a different time T_LX is, for example, a time T_L1 that is later than a certain time T_Cur. In this case, the three-dimensional data encoding device 1300 may add RT information RT_L1 related to rotation and translation processing applied to the space at time T_L1 to the bitstream.
[0401] Alternatively, the interpretation unit 1311 performs encoding (dual prediction) by referring to both spaces at different times T_L0 and T_L1. In this case, the three-dimensional data encoding device 1300 may add both RT information RT_L0 and RT_L1, which relate to rotation and translation applied to each space, to the bitstream.
[0402] In the above, T_L0 is defined as a time before T_Cur and T_L1 as a time after T_Cur, but this is not necessarily the only way. For example, both T_L0 and T_L1 may be times before T_Cur. Alternatively, both T_L0 and T_L1 may be times after T_Cur.
[0403] Furthermore, when the three-dimensional data encoding device 1300 performs encoding by referencing multiple spaces at different times, it may add rotation and translation-related RT information applied to each space to the bitstream. For example, the three-dimensional data encoding device 1300 manages the multiple encoded spaces to be referenced using two reference lists (L0 list and L1 list). If the first reference space in the L0 list is L0R0, the second reference space in the L0 list is L0R1, the first reference space in the L1 list is L1R0, and the second reference space in the L1 list is L1R1, then the three-dimensional data encoding device 1300 adds the RT information RT_L0R0 for L0R0, RT information RT_L0R1 for L0R1, RT information RT_L1R0 for L1R0, and RT information RT_L1R1 for L1R1 to the bitstream. For example, the three-dimensional data encoding device 1300 adds this RT information to the bitstream header, etc.
[0404] Furthermore, when the three-dimensional data encoding device 1300 performs encoding by referencing multiple reference spaces at different times, it determines whether or not to apply rotation and translation for each reference space. In this case, the three-dimensional data encoding device 1300 may add information (such as an RT application flag) indicating whether or not rotation and translation have been applied for each reference space to the bitstream header information, etc. For example, the three-dimensional data encoding device 1300 calculates RT information and an ICP error value for each reference space referenced from the space to be encoded using the ICP (Interactive Closest Point) algorithm. If the ICP error value is less than or equal to a predetermined value, the three-dimensional data encoding device 1300 determines that rotation and translation are not necessary and sets the RT application flag to off. On the other hand, if the ICP error value is greater than the above-mentioned predetermined value, the three-dimensional data encoding device 1300 sets the RT application flag to on and adds the RT information to the bitstream.
[0405] Figure 46 shows an example of syntax for adding RT information and RT application flags to the header. The number of bits allocated to each syntax may be determined within the range of possible values for that syntax. For example, if the reference list L0 contains 8 reference spaces, 3 bits may be allocated to MaxRefSpc_l0. The number of bits allocated may be variable depending on the possible values for each syntax, or it may be fixed regardless of the possible values. If the number of bits allocated is fixed, the three-dimensional data encoding device 1300 may add that fixed number of bits to separate header information.
[0406] Here, as shown in Figure 46, MaxRefSpc_l0 indicates the number of reference spaces included in reference list L0. RT_flag_l0[i] is the RT application flag for reference space i in reference list L0. If RT_flag_l0[i] is 1, rotation and translation are applied to reference space i. If RT_flag_l0[i] is 0, rotation and translation are not applied to reference space i.
[0407] R_l0[i] and T_l0[i] are the RT information for reference space i in reference list L0. R_l0[i] is the rotation information for reference space i in reference list L0. The rotation information indicates the content of the applied rotation operation, such as a rotation matrix or quaternion. T_l0[i] is the translation information for reference space i in reference list L0. The translation information indicates the content of the applied translation operation, such as a translation vector.
[0408] MaxRefSpc_l1 indicates the number of reference spaces included in reference list L1. RT_flag_l1[i] is the RT application flag for reference space i in reference list L1. If RT_flag_l1[i] is 1, rotation and translation are applied to reference space i. If RT_flag_l1[i] is 0, rotation and translation are not applied to reference space i.
[0409] R_l1[i] and T_l1[i] are the RT information for reference space i in reference list L1. R_l1[i] is the rotation information for reference space i in reference list L1. The rotation information indicates the content of the applied rotation operation, such as a rotation matrix or quaternion. T_l1[i] is the translation information for reference space i in reference list L1. The translation information indicates the content of the applied translation operation, such as a translation vector.
[0410] The interpretation unit 1311 generates a predicted volume of the volume to be encoded using the encoded reference space information stored in the reference space memory 1310. As described above, before generating the predicted volume of the volume to be encoded, the interpretation unit 1311 uses the Interactive Closest Point (ICP) algorithm to obtain RT information in the volume to be encoded and the reference space in order to bring the overall positional relationship between the volume to be encoded and the reference space closer together. Then, the interpretation unit 1311 obtains reference space B by applying rotation and translation processing to the reference space using the obtained RT information. After that, the interpretation unit 1311 generates a predicted volume of the volume to be encoded in the volume to be encoded using the information in reference space B. Here, the three-dimensional data encoding device 1300 adds the RT information used to obtain reference space B to the header information of the volume to be encoded, etc.
[0411] Thus, the interpretation unit 1311 can improve the accuracy of the predicted volume by applying rotation and translation processing to the reference space to bring the overall positional relationship between the space to be encoded and the reference space closer together, and then generating a predicted volume using the information from the reference space. Furthermore, since the prediction residual can be suppressed, the amount of coding can be reduced. Note that here, an example of performing ICP using the space to be encoded and the reference space has been shown, but this is not necessarily the only example. For example, in order to reduce the amount of processing, the interpretation unit 1311 may obtain RT information by performing ICP using at least one of the space to be encoded with a reduced number of voxels or point clouds, and the reference space with a reduced number of voxels or point clouds.
[0412] Furthermore, the interpretation unit 1311 may determine that rotation and translation processing is unnecessary if the ICP error value obtained as a result of ICP is smaller than a predetermined first threshold, that is, if the positional relationship between the space to be encoded and the reference space is close, and may not perform rotation and translation. In this case, the three-dimensional data encoding device 1300 may suppress overhead by not adding RT information to the bitstream.
[0413] Furthermore, if the interpretation unit 1311 determines that the shape change between spaces is large when the ICP error value is greater than a predetermined second threshold, it may apply intraprediction to all volumes of the space to be encoded. Hereinafter, the space to which intraprediction is applied will be referred to as the intraspace. The second threshold is a value greater than the first threshold mentioned above. In addition, any method that can obtain RT information from two voxel sets or two point cloud sets may be applied, not limited to ICP.
[0414] Furthermore, if the three-dimensional data includes attribute information such as shape or color, the interpretation unit 1311 searches for a volume in the reference space that has the closest attribute information (shape or color, etc.) to the volume to be encoded, for example, within the reference space, as the predicted volume for the volume to be encoded within the encoding space. This reference space is, for example, the reference space after the rotation and translation processing described above has been performed. The interpretation unit 1311 generates a predicted volume from the volume (reference volume) obtained through the search. Figure 47 is a diagram illustrating the operation of generating a predicted volume. When the interpretation unit 1311 encodes the volume to be encoded (volume idx=0) shown in Figure 47 using interpretation, it scans the reference volumes in the reference space sequentially and searches for the volume with the smallest predicted residual, which is the difference between the volume to be encoded and the reference volume. The interpretation unit 1311 selects the volume with the smallest predicted residual as the predicted volume. The predicted residual between the volume to be encoded and the predicted volume is encoded by the processing from the conversion unit 1303 onward. Here, the predicted residual is the difference between the attribute information of the volume to be encoded and the attribute information of the predicted volume. Furthermore, the three-dimensional data encoding device 1300 adds the volume idx of the reference volume in the reference space referenced as the predicted volume to the bitstream header, etc.
[0415] In the example shown in Figure 47, the reference volume with volume idx=4 in reference space L0R0 is selected as the predicted volume for the volume to be encoded. Then, the predicted residual between the volume to be encoded and the reference volume, along with the reference volume idx=4, are encoded and appended to the bitstream.
[0416] While this example demonstrates the generation of predicted volume for attribute information, similar processing may be applied to the predicted volume for location information.
[0417] The prediction control unit 1312 controls whether to encode the volume to be encoded using intra-prediction or inter-prediction. Here, the mode that includes intra-prediction and inter-prediction is called the prediction mode. For example, the prediction control unit 1312 calculates the prediction residual when the volume to be encoded is predicted using intra-prediction and the prediction residual when it is predicted using inter-prediction as evaluation values, and selects the prediction mode with the smaller evaluation value. Alternatively, the prediction control unit 1312 may calculate the actual code amount by applying orthogonal transformation, quantization, and entropy coding to the prediction residuals of intra-prediction and inter-prediction, respectively, and select the prediction mode using the calculated code amount as the evaluation value. In addition, overhead information other than the prediction residual (such as reference volume idx information) may be added to the evaluation value. Furthermore, if it is predetermined that the space to be encoded will be encoded in intra-space, the prediction control unit 1312 may always select intra-prediction.
[0418] The entropy coding unit 1313 generates an encoded signal (encoded bitstream) by variable-length encoding the quantization coefficients, which are input from the quantization unit 1304. Specifically, the entropy coding unit 1313, for example, binarizes the quantization coefficients and arithmetically encodes the resulting binary signal.
[0419] Next, a three-dimensional data decoding device that decodes the encoded signal generated by the three-dimensional data encoding device 1300 will be described. Figure 48 is a block diagram of the three-dimensional data decoding device 1400 according to this embodiment. This three-dimensional data decoding device 1400 comprises an entropy decoding unit 1401, an inverse quantization unit 1402, an inverse transform unit 1403, an adder 1404, a reference volume memory 1405, an intra prediction unit 1406, a reference space memory 1407, an inter prediction unit 1408, and a prediction control unit 1409.
[0420] The entropy decoding unit 1401 decodes the encoded signal (encoded bitstream) to a variable length. For example, the entropy decoding unit 1401 arithmetically decodes the encoded signal to generate a binary signal, and then generates quantization coefficients from the generated binary signal.
[0421] The inverse quantization unit 1402 generates inverse quantization coefficients by inverse quantizing the quantization coefficients input from the entropy decoding unit 1401 using quantization parameters added to the bitstream or the like.
[0422] The inverse transform unit 1403 generates the predicted residual by inversely transforming the inverse quantization coefficients input from the inverse quantization unit 1402. For example, the inverse transform unit 1403 generates the predicted residual by inversely transforming the inverse quantization coefficients based on the information added to the bitstream.
[0423] The summing unit 1404 adds the predicted residual generated by the inverse transform unit 1403 and the predicted volume generated by intra-prediction or inter-prediction to generate a reconstructed volume. This reconstructed volume is output as decoded three-dimensional data and stored in the reference volume memory 1405 or the reference space memory 1407.
[0424] The intra-prediction unit 1406 generates a predicted volume through intra-prediction using the reference volume in the reference volume memory 1405 and the information attached to the bitstream. Specifically, the intra-prediction unit 1406 acquires adjacent volume information (e.g., volume idx) and prediction mode information attached to the bitstream, and generates a predicted volume using the adjacent volume indicated by the adjacent volume information and the mode indicated by the prediction mode information. The details of these processes are the same as those of the intra-prediction unit 1309 described above, except that the information attached to the bitstream is used.
[0425] The interpretation unit 1408 generates a predicted volume by interpretation using the reference space in the reference space memory 1407 and the information attached to the bitstream. Specifically, the interpretation unit 1408 applies rotation and translation processing to the reference space using the RT information for each reference space attached to the bitstream, and generates a predicted volume using the reference space after processing. If an RT application flag for each reference space exists in the bitstream, the interpretation unit 1408 applies rotation and translation processing to the reference space according to the RT application flag. The details of these processes are the same as those of the interpretation unit 1311 described above, except that the information attached to the bitstream is used.
[0426] The prediction control unit 1409 controls whether to decode the volume to be decoded using intra-prediction or inter-prediction. For example, the prediction control unit 1409 selects intra-prediction or inter-prediction according to information attached to the bitstream indicating the prediction mode to be used. The prediction control unit 1409 may always select intra-prediction if it has been predetermined that the space to be decoded will be decoded in intra-space.
[0427] The following describes modifications of this embodiment. In this embodiment, an example of applying rotation and translation on a space-by-space basis has been described, but rotation and translation may be applied on a finer scale. For example, the three-dimensional data encoding device 1300 may divide the space into subspaces and apply rotation and translation on a subspace-by-subspace basis. In this case, the three-dimensional data encoding device 1300 generates RT information for each subspace and adds the generated RT information to the bitstream header, etc. Alternatively, the three-dimensional data encoding device 1300 may apply rotation and translation on a volume-by-volume basis, which is the encoding unit. In this case, the three-dimensional data encoding device 1300 generates RT information on an encoding volume-by-volume basis and adds the generated RT information to the bitstream header, etc. Furthermore, the above may be combined. That is, the three-dimensional data encoding device 1300 may apply rotation and translation on a larger scale and then apply rotation and translation on a finer scale. For example, the three-dimensional data encoding device 1300 may apply rotation and translation on a space-by-space basis and then apply different rotations and translations to each of the multiple volumes contained in the resulting space.
[0428] Furthermore, although this embodiment describes an example in which rotation and translation are applied to the reference space, it is not necessarily limited to this. For example, the three-dimensional data encoding device 1300 may change the size of the three-dimensional data by applying scaling processing, for example. Also, the three-dimensional data encoding device 1300 may apply one or two of rotation, translation, and scaling. In addition, when processing is applied in multiple stages to different units as described above, the type of processing applied to each unit may differ. For example, rotation and translation may be applied to the space unit, and translation may be applied to the volume unit.
[0429] These modifications can also be applied to the three-dimensional data decoding device 1400.
[0430] As described above, the three-dimensional data encoding device 1300 according to this embodiment performs the following processing. Figure 48 is a flowchart of the interpretation processing performed by the three-dimensional data encoding device 1300.
[0431] First, the three-dimensional data encoding device 1300 generates predicted position information (e.g., predicted volume) using the position information of three-dimensional points contained in reference three-dimensional data (e.g., reference space) at a different time from the target three-dimensional data (e.g., the space to be encoded) (S1301). Specifically, the three-dimensional data encoding device 1300 generates predicted position information by applying rotation and translation processing to the position information of three-dimensional points contained in the reference three-dimensional data.
[0432] The three-dimensional data encoding device 1300 may perform rotation and translation processing in a first unit (e.g., space) and generate predicted position information in a second unit (e.g., volume) that is finer than the first unit. For example, the three-dimensional data encoding device 1300 searches for a volume among several volumes included in the reference space after rotation and translation processing that has the smallest difference in position information between it and the volume to be encoded included in the space to be encoded, and uses the obtained volume as the predicted volume. The three-dimensional data encoding device 1300 may perform rotation and translation processing and the generation of predicted position information in the same unit.
[0433] Furthermore, the three-dimensional data encoding device 1300 may generate predicted position information by applying a first rotation and translation process in a first unit (e.g., space) to the position information of three-dimensional points included in the reference three-dimensional data, and then applying a second rotation and translation process in a second unit (e.g., volume) that is finer than the first unit to the position information of three-dimensional points obtained by the first rotation and translation process.
[0434] Here, the position information and predicted position information of a three-dimensional point are represented in an octree structure, for example, as shown in Figure 41. For example, the position information and predicted position information of a three-dimensional point are represented in a scan order that prioritizes width over depth in the octree structure. Alternatively, the position information and predicted position information of a three-dimensional point are represented in a scan order that prioritizes depth over width in the octree structure.
[0435] Furthermore, as shown in Figure 46, the three-dimensional data encoding device 1300 encodes an RT application flag indicating whether or not to apply rotation and translation processing to the position information of three-dimensional points included in the reference three-dimensional data. In other words, the three-dimensional data encoding device 1300 generates an encoded signal (encoded bitstream) that includes the RT application flag. The three-dimensional data encoding device 1300 also encodes RT information indicating the content of the rotation and translation processing. In other words, the three-dimensional data encoding device 1300 generates an encoded signal (encoded bitstream) that includes RT information. Note that the three-dimensional data encoding device 1300 encodes RT information when the RT application flag indicates that rotation and translation processing should be applied, and does not encode RT information when the RT application flag indicates that rotation and translation processing should not be applied.
[0436] Furthermore, the three-dimensional data includes, for example, positional information of three-dimensional points and attribute information (color information, etc.) of each three-dimensional point. The three-dimensional data encoding device 1300 generates predicted attribute information using the attribute information of three-dimensional points included in the reference three-dimensional data (S1302).
[0437] Next, the three-dimensional data encoding device 1300 encodes the position information of the three-dimensional points included in the target three-dimensional data using the predicted position information. For example, as shown in Figure 38, the three-dimensional data encoding device 1300 calculates differential position information, which is the difference between the position information of the three-dimensional points included in the target three-dimensional data and the predicted position information (S1303).
[0438] Furthermore, the three-dimensional data encoding device 1300 encodes the attribute information of three-dimensional points included in the target three-dimensional data using predicted attribute information. For example, the three-dimensional data encoding device 1300 calculates differential attribute information, which is the difference between the attribute information of three-dimensional points included in the target three-dimensional data and the predicted attribute information (S1304). Next, the three-dimensional data encoding device 1300 performs conversion and quantization of the calculated differential attribute information (S1305).
[0439] Finally, the three-dimensional data encoding device 1300 encodes (for example, entropy encoding) the differential position information and the quantized differential attribute information (S1306). In other words, the three-dimensional data encoding device 1300 generates an encoded signal (encoded bitstream) that includes the differential position information and the differential attribute information.
[0440] Furthermore, if the three-dimensional data does not contain attribute information, the three-dimensional data encoding device 1300 does not need to perform steps S1302, S1304, and S1305. Also, the three-dimensional data encoding device 1300 may perform only one of the following: encoding the position information of the three-dimensional points or encoding the attribute information of the three-dimensional points.
[0441] Furthermore, the processing order shown in Figure 49 is just one example and is not limited thereto. For example, the processing of location information (S1301, S1303) and the processing of attribute information (S1302, S1304, S1305) are independent of each other and may be performed in any order, or some may be processed in parallel.
[0442] As described above, in this embodiment, the three-dimensional data encoding device 1300 generates predicted position information using the position information of three-dimensional points contained in reference three-dimensional data at a different time from the target three-dimensional data, and encodes the difference in position information, which is the difference between the position information of three-dimensional points contained in the target three-dimensional data and the predicted position information. This reduces the amount of data in the encoded signal, thereby improving encoding efficiency.
[0443] Furthermore, in this embodiment, the three-dimensional data encoding device 1300 generates predicted attribute information using the attribute information of three-dimensional points included in the reference three-dimensional data, and encodes differential attribute information, which is the difference between the attribute information of three-dimensional points included in the target three-dimensional data and the predicted attribute information. This reduces the amount of data in the encoded signal, thereby improving encoding efficiency.
[0444] For example, the three-dimensional data encoding device 1300 includes a processor and memory, and the processor uses the memory to perform the above processing.
[0445] Figure 48 is a flowchart of the interpretation process performed by the three-dimensional data decoding device 1400.
[0446] First, the three-dimensional data decoding device 1400 decodes (for example, entropy decoding) the differential position information and differential attribute information from the encoded signal (encoded bitstream) (S1401).
[0447] Furthermore, the three-dimensional data decoding device 1400 decodes an RT application flag from the encoded signal, which indicates whether or not rotation and translation processing should be applied to the position information of the three-dimensional points included in the reference three-dimensional data. The three-dimensional data decoding device 1400 also decodes RT information indicating the content of the rotation and translation processing. Note that the three-dimensional data decoding device 1400 decodes the RT information when the RT application flag indicates that rotation and translation processing should be applied, and does not need to decode the RT information when the RT application flag indicates that rotation and translation processing should not be applied.
[0448] Next, the three-dimensional data decoding device 1400 performs inverse quantization and inverse transformation on the decoded difference attribute information (S1402).
[0449] Next, the three-dimensional data decoding device 1400 generates predicted position information (e.g., predicted volume) using the position information of three-dimensional points contained in reference three-dimensional data (e.g., reference space) at a different time from the target three-dimensional data (e.g., the space to be decoded) (S1403). Specifically, the three-dimensional data decoding device 1400 generates predicted position information by applying rotation and translation processing to the position information of three-dimensional points contained in the reference three-dimensional data.
[0450] More specifically, the three-dimensional data decoding device 1400 applies rotation and translation processing to the position information of three-dimensional points included in the reference three-dimensional data indicated by the RT information when the RT application flag indicates that rotation and translation processing should be applied. On the other hand, when the RT application flag indicates that rotation and translation processing should not be applied, the three-dimensional data decoding device 1400 does not apply rotation and translation processing to the position information of three-dimensional points included in the reference three-dimensional data.
[0451] The three-dimensional data decoding device 1400 may perform rotation and translation processing in a first unit (e.g., space) and generate predicted position information in a second unit (e.g., volume) that is finer than the first unit. Alternatively, the three-dimensional data decoding device 1400 may perform rotation and translation processing and generate predicted position information in the same unit.
[0452] Furthermore, the three-dimensional data decoding device 1400 may generate predicted position information by applying a first rotation and translation process in a first unit (e.g., space) to the position information of three-dimensional points included in the reference three-dimensional data, and then applying a second rotation and translation process in a second unit (e.g., volume) that is finer than the first unit to the position information of three-dimensional points obtained by the first rotation and translation process.
[0453] Here, the position information and predicted position information of a three-dimensional point are represented in an octree structure, for example, as shown in Figure 41. For example, the position information and predicted position information of a three-dimensional point are represented in a scan order that prioritizes width over depth in the octree structure. Alternatively, the position information and predicted position information of a three-dimensional point are represented in a scan order that prioritizes depth over width in the octree structure.
[0454] The three-dimensional data decoding device 1400 generates predicted attribute information using the attribute information of three-dimensional points contained in the reference three-dimensional data (S1404).
[0455] Next, the three-dimensional data decoding device 1400 decodes the encoded position information contained in the encoded signal using the predicted position information to reconstruct the position information of the three-dimensional points contained in the target three-dimensional data. Here, the encoded position information is, for example, the difference position information, and the three-dimensional data decoding device 1400 reconstructs the position information of the three-dimensional points contained in the target three-dimensional data by adding the difference position information and the predicted position information (S1405).
[0456] Furthermore, the three-dimensional data decoding device 1400 reconstructs the attribute information of three-dimensional points contained in the target three-dimensional data by decoding the encoded attribute information contained in the encoded signal using the predicted attribute information. Here, the encoded attribute information is, for example, the difference attribute information, and the three-dimensional data decoding device 1400 reconstructs the attribute information of three-dimensional points contained in the target three-dimensional data by adding the difference attribute information and the predicted attribute information (S1406).
[0457] Furthermore, if the three-dimensional data does not contain attribute information, the three-dimensional data decoding device 1400 does not need to perform steps S1402, S1404, and S1406. Also, the three-dimensional data decoding device 1400 may perform only one of the following: decoding the position information of the three-dimensional points or decoding the attribute information of the three-dimensional points.
[0458] Furthermore, the processing order shown in Figure 50 is just one example and is not limited thereto. For example, the processing of location information (S1403, S1405) and the processing of attribute information (S1402, S1404, S1406) are independent of each other and may be performed in any order, or some may be processed in parallel.
[0459] (Embodiment 8) This embodiment describes adaptive entropy coding (arithmetic coding) for occupancy codes of octrees.
[0460] Figure 51 shows an example of a quadtree structure. Figure 52 shows the occupancy code for the tree structure shown in Figure 51. Figure 53 is a schematic diagram showing the operation of a three-dimensional data encoding device in the form of this embodiment.
[0461] The three-dimensional data encoding device according to this embodiment entropy encodes 8 bits of occupancy encoding in an octree. Furthermore, the three-dimensional data encoding device updates the encoding table during the entropy encoding process of the occupancy code. In addition, the three-dimensional data encoding device uses an adaptive encoding table to utilize similarity information of three-dimensional points, rather than a single encoding table. In other words, the three-dimensional data encoding device uses multiple encoding tables.
[0462] Similarity information can include, for example, geometric information of a three-dimensional point, structural information of an octree, or attribute information of a three-dimensional point.
[0463] Although Figures 51 to 53 show a quadtree as an example, the same method can be applied to N-trees such as binary trees, octiv trees, and hexagram trees. For example, a three-dimensional data encoding device performs entropy encoding using an adaptive table (also called an encoding table) for 8-bit occupancy codes in the case of an octiv tree, 4-bit occupancy codes in the case of a quadiv tree, and 16-bit occupancy codes in the case of a hexagram tree.
[0464] The following describes adaptive entropy coding using geometric information of three-dimensional points (point cloud).
[0465] In a tree structure, if the geometric arrangement of the surrounding areas of two nodes is similar, the occupation state of the child nodes (i.e., whether or not they contain a three-dimensional point) may also be similar. Therefore, a three-dimensional data encoding device uses the geometric arrangement of the surrounding areas of the parent node to perform grouping. This allows the three-dimensional data encoding device to group the occupation states of the child nodes and use a different encoding table for each group. Thus, the encoding efficiency of entropy coding can be improved.
[0466] Figure 54 shows an example of geometric information. Geometric information includes information indicating whether each of the multiple adjacent nodes of the target node is occupied or not (i.e., whether or not it contains a three-dimensional point). For example, a three-dimensional data encoding device calculates the geometric arrangement (local geometry) around the target node using information on whether or not adjacent nodes contain a three-dimensional point (occupied or non-occupied). Adjacent nodes are, for example, nodes that are spatially surrounding the target node, or nodes that are at the same location at a different time from the target node, or that are spatially surrounding that location.
[0467] In Figure 54, hatched cubes represent target nodes to be encoded. White cubes represent adjacent nodes that also contain three-dimensional points. In Figure 54, the geometric pattern shown in (2) represents the rotated form of the geometric pattern shown in (1). Therefore, the three-dimensional data encoding device determines that these geometric patterns have high geometric similarity and performs entropy encoding on them using the same encoding table. The three-dimensional data encoding device also determines that the geometric patterns in (3) and (4) have low geometric similarity and performs entropy encoding on them using different encoding tables.
[0468] Figure 55 shows the occupancy codes of target nodes in the geometric patterns (1) to (4) shown in Figure 54, and examples of coding tables used for entropy coding. The three-dimensional data coding device determines that geometric patterns (1) and (2) belong to the same geometric group, as described above, and uses the same coding table A. The three-dimensional data coding device also uses coding tables B and C for geometric patterns (3) and (4), respectively.
[0469] Furthermore, as shown in Figure 55, the occupancy codes of target nodes in geometric patterns (1) and (2) that belong to the same geometric group may be identical.
[0470] Next, we will describe adaptive entropy coding using structure information of a tree structure. For example, the structure information includes information indicating the layer to which the target node belongs.
[0471] Figure 56 shows an example of a tree structure. Generally, the shape of local objects depends on the search scale. For example, in a tree structure, lower layers tend to be sparser than upper layers. Therefore, a three-dimensional data encoding device can improve the encoding efficiency of entropy coding by using different encoding tables for the upper and lower layers, as shown in Figure 56.
[0472] In other words, a three-dimensional data encoding device may use a different encoding table for each layer when encoding the occupancy codes of each layer. For example, for the tree structure shown in Figure 56, the three-dimensional data encoding device may use an encoding table for layer N (N=0 to 6) to perform entropy encoding for the occupancy codes of layer N. This allows the three-dimensional data encoding device to switch encoding tables according to the occurrence pattern of occupancy codes for each layer, thereby improving encoding efficiency.
[0473] Furthermore, as shown in Figure 56, the three-dimensional data encoding device may use encoding table A for the occupancy codes from layer 0 to layer 2, and encoding table B for the occupancy codes from layer 3 to layer 6. This allows the three-dimensional data encoding device to switch encoding tables according to the occurrence pattern of occupancy codes for each layer group, thereby improving encoding efficiency. The three-dimensional data encoding device may also add information about the encoding table used in each layer to the bitstream header. Alternatively, the encoding tables used in each layer may be predetermined by standards or other specifications.
[0474] Next, we will describe adaptive entropy coding using property information of three-dimensional points. For example, the property information includes information about the object containing the target node, or information about the normal vector held by the target node.
[0475] By using the attribute information of three-dimensional points, three-dimensional points with similar geometric configurations can be grouped together. For example, the normal vector representing the direction of each three-dimensional point can be used as common attribute information. By using normal vectors, it is possible to find geometric configurations related to similar occupancy codes within a tree structure.
[0476] Furthermore, color or reflectance (reflectance) may be used as attribute information. For example, a three-dimensional data encoding device may use the color or reflectance of three-dimensional points to group three-dimensional points having similar geometric arrangements, and perform processing such as switching encoding tables for each group.
[0477] Figure 57 illustrates the switching of coding tables based on normal vectors. As shown in Figure 57, different coding tables are used when the normal vectors of the target node belong to different groups of normal vectors. For example, normal vectors that fall within a predetermined range are classified into one group of normal vectors.
[0478] Furthermore, if the classification of the objects differs, the occupancy codes are also likely to differ. Therefore, the three-dimensional data encoding device may select an encoding table according to the classification of the object to which the target node belongs. Figure 58 is a diagram illustrating the switching of encoding tables based on the classification of the objects. As shown in Figure 58, different encoding tables are used when the classification of the objects differs.
[0479] The following describes an example of the bitstream configuration according to this embodiment. Figure 59 is a diagram showing an example of the bitstream configuration generated by the three-dimensional data encoding device according to this embodiment. As shown in Figure 59, the bitstream includes a set of encoding tables, a table index, and an encoding occupancy. The set of encoding tables includes a plurality of encoding tables.
[0480] The table index is an index that indicates the coding table used for the entropy coding of the subsequent coding occupancy. The coding occupancy is the occupancy code after entropy coding. Furthermore, as shown in Figure 59, the bitstream contains multiple pairs of table indexes and coding occupancies.
[0481] For example, in the example shown in Figure 59, coded occupancy 0 is data that has been entropically coded using the context model (hereinafter also referred to as context) indicated by table index 0. Coded occupancy 1 is data that has been entropically coded using the context indicated by table index 1. Alternatively, a context for coding coded occupancy 0 may be defined in advance by a standard, and the three-dimensional data decoder may use that context when decoding coded occupancy 0. This eliminates the need to add a table index to the bitstream, thus reducing overhead.
[0482] Furthermore, the three-dimensional data encoding device may add information to the header for initializing each context.
[0483] The three-dimensional data encoding device determines an encoding table using the geometric, structural, or attribute information of the target node, and encodes the occupancy code using the determined encoding table. The three-dimensional data encoding device adds the encoding result and information about the encoding table used for encoding (such as a table index) to the bitstream and transmits the bitstream to the three-dimensional data decoding device. As a result, the three-dimensional data decoding device can decode the occupancy code using the encoding table information added to the header.
[0484] Alternatively, the three-dimensional data encoding device may not add information about the encoding table used for encoding to the bitstream, and the three-dimensional data decoding device may determine the encoding table using the geometric information, structural information, or attribute information of the target node after decoding in the same way as the three-dimensional data encoding device, and decode the occupancy code using the determined encoding table. This eliminates the need to add information about the encoding table to the bitstream, thus reducing overhead.
[0485] Figures 60 and 61 show examples of coding tables. As shown in Figures 60 and 61, one coding table indicates the context model and context model type corresponding to each 8-bit occupancy code value.
[0486] As shown in the coding table in Figure 60, the same context model (context) may be applied to multiple occupancy codes. Alternatively, a different context model may be assigned to each occupancy code. This allows for the assignment of a context model according to the probability of occurrence of the occupancy code, thereby improving coding efficiency.
[0487] Furthermore, the context model type indicates, for example, whether the context model is one in which the probability table is updated according to the frequency of occurrence of occupancy codes, or whether it is a context model in which the probability table is fixed.
[0488] Next, another example of a bitstream and coding table is shown. Figure 62 shows an example of a modified bitstream configuration. As shown in Figure 62, the bitstream includes a set of coding tables and a coding occupancy. The set of coding tables includes multiple coding tables.
[0489] Figures 63 and 64 show examples of coding tables. As shown in Figures 63 and 64, one coding table indicates the context model and context model type corresponding to each bit included in the occupancy code.
[0490] Figure 65 shows an example of the relationship between an occupancy code and the bit number of the occupancy code.
[0491] Thus, a three-dimensional data encoding device may treat the occupancy code as binary data and entropy encode the occupancy code by assigning a different context model to each bit. This allows for the assignment of a context model according to the probability of occurrence of each bit of the occupancy code, thereby improving encoding efficiency.
[0492] Specifically, each bit of the occupancy code corresponds to a subblock obtained by dividing the spatial block corresponding to the target node. Therefore, coding efficiency can be improved when subblocks at the same spatial location within a block exhibit similar tendencies. For example, if the surface of the ground or road traverses a block, in an octree, the bottom four blocks will contain three-dimensional points, while the top four blocks will not. Similar patterns also appear in multiple blocks arranged horizontally. Therefore, coding efficiency can be improved by switching the context bit by bit as described above.
[0493] Alternatively, a context model may be used in which the probability table is updated according to the frequency of occurrence of each bit in the occupancy code. Alternatively, a context model may be used in which the probability table is fixed.
[0494] Next, the flow of the three-dimensional data encoding process and the three-dimensional data decoding process according to this embodiment will be described.
[0495] Figure 66 is a flowchart of a three-dimensional data coding process that includes adaptive entropy coding using geometric information.
[0496] In the decomposition process, an octave tree is generated from the initial boundary box of the three-dimensional points. The boundary box is divided according to the position of the three-dimensional points within it. Specifically, non-empty subspaces are further divided. Next, information indicating whether or not a three-dimensional point is contained in a subspace is encoded in an occupancy code. The same process is performed in the processes shown in Figures 68 and 70.
[0497] First, the three-dimensional data encoding device acquires the input three-dimensional points (S1901). Next, the three-dimensional data encoding device determines whether or not the unit length decomposition process has been completed (S1902).
[0498] If the decomposition process for unit lengths is not complete (No in S1902), the three-dimensional data encoding device generates an octave tree by performing the decomposition process on the target node (S1903).
[0499] Next, the three-dimensional data encoding device acquires geometric information (S1904) and selects an encoding table based on the acquired geometric information (S1905). Here, geometric information refers to information such as the geometric arrangement of the occupied state of the surrounding blocks of the target node, as described above.
[0500] Next, the three-dimensional data coding device uses the selected coding table to entropy code the occupancy code of the target node (S1906).
[0501] The processes described in steps S1903 to S1906 above are repeated until the unit length decomposition process is completed. If the unit length decomposition process is completed (Yes in S1902), the three-dimensional data encoding device outputs a bitstream containing the generated information (S1907).
[0502] The three-dimensional data encoding device determines an encoding table using the geometric, structural, or attribute information of the target node, and encodes the bit sequence of the occupancy code using the determined encoding table. The three-dimensional data encoding device adds the encoding result and information about the encoding table used for encoding (such as the table index) to the bitstream and transmits the bitstream to the three-dimensional data decoding device. This allows the three-dimensional data decoding device to decode the occupancy code using the encoding table information added to the header.
[0503] Alternatively, the three-dimensional data encoding device may not add information about the encoding table used for encoding to the bitstream, and the three-dimensional data decoding device may determine the encoding table using the geometric information, structural information, or attribute information of the target node after decoding in the same way as the three-dimensional data encoding device, and decode the occupancy code using the determined encoding table. This eliminates the need to add information about the encoding table to the bitstream, thus reducing overhead.
[0504] Figure 67 is a flowchart of a three-dimensional data decoding process that includes adaptive entropy decoding using geometric information.
[0505] The decomposition process included in the decoding process is the same as the decomposition process included in the encoding process described above, but differs in the following respects: The three-dimensional data decoder divides the initial bounding box using the decoded occupancy code. When the three-dimensional data decoder has finished processing a unit length, it saves the positions of the bounding boxes as three-dimensional points and positions. The same process is performed in the processes shown in Figures 69 and 71.
[0506] First, the three-dimensional data decoding device acquires the input bitstream (S1911). Next, the three-dimensional data decoding device determines whether or not the decomposition process for unit lengths has been completed (S1912).
[0507] If the decomposition process for unit lengths is not complete (No in S1912), the three-dimensional data decoding device generates an octave tree by performing the decomposition process on the target node (S1913).
[0508] Next, the three-dimensional data decoding device acquires geometric information (S1914) and selects an encoding table based on the acquired geometric information (S1915). Here, geometric information refers to information such as the geometric arrangement of the occupied state of the surrounding blocks of the target node, as described above.
[0509] Next, the three-dimensional data decoding device uses the selected coding table to entropy-decode the occupancy code of the target node (S1916).
[0510] The processes described in steps S1913 to S1916 above are repeated until the decomposition of the unit length is completed. If the decomposition of the unit length is completed (Yes in S1912), the three-dimensional data decoder outputs a three-dimensional point (S1917).
[0511] Figure 68 is a flowchart of a three-dimensional data coding process that includes adaptive entropy coding using structural information.
[0512] First, the three-dimensional data encoding device acquires the input three-dimensional points (S1921). Next, the three-dimensional data encoding device determines whether or not the unit length decomposition process has been completed (S1922).
[0513] If the decomposition process for unit lengths is not complete (No in S1922), the three-dimensional data encoding device generates an octave tree by performing the decomposition process on the target node (S1923).
[0514] Next, the three-dimensional data encoding device acquires structural information (S1924) and selects an encoding table based on the acquired structural information (S1925). Here, structural information refers to information indicating, for example, the layer to which the target node belongs, as described above.
[0515] Next, the three-dimensional data coding device uses the selected coding table to entropy code the occupancy code of the target node (S1926).
[0516] The processes described in steps S1923 to S1926 are repeated until the decomposition of the unit length is completed. If the decomposition of the unit length is completed (Yes in S1922), the three-dimensional data encoding device outputs a bitstream containing the generated information (S1927).
[0517] Figure 69 is a flowchart of a three-dimensional data decoding process that includes adaptive entropy decoding using structural information.
[0518] First, the three-dimensional data decoding device acquires the input bitstream (S1931). Next, the three-dimensional data decoding device determines whether or not the decomposition process for unit lengths has been completed (S1932).
[0519] If the decomposition process for unit lengths is not complete (No in S1932), the three-dimensional data decoding device generates an octave tree by performing the decomposition process on the target node (S1933).
[0520] Next, the three-dimensional data decoding device acquires structural information (S1934) and selects an encoding table based on the acquired structural information (S1935). Here, structural information refers to information indicating, for example, the layer to which the target node belongs, as described above.
[0521] Next, the three-dimensional data decoding device uses the selected coding table to entropy-decode the occupancy code of the target node (S1936).
[0522] The processes described in steps S1933 to S1936 are repeated until the decomposition of the unit length is completed. If the decomposition of the unit length is completed (Yes in S1932), the three-dimensional data decoder outputs a three-dimensional point (S1937).
[0523] Figure 70 is a flowchart of a three-dimensional data coding process that includes adaptive entropy coding using attribute information.
[0524] First, the three-dimensional data encoding device acquires the input three-dimensional points (S1941). Next, the three-dimensional data encoding device determines whether or not the unit length decomposition process has been completed (S1942).
[0525] If the decomposition process for unit lengths is not complete (No in S1942), the three-dimensional data encoding device generates an octave tree by performing the decomposition process on the target node (S1943).
[0526] Next, the three-dimensional data encoding device acquires attribute information (S1944) and selects an encoding table based on the acquired attribute information (S1945). Here, attribute information refers to information such as the normal vector of the target node, as described above.
[0527] Next, the three-dimensional data coding device uses the selected coding table to entropy code the occupancy code of the target node (S1946).
[0528] The processes described in steps S1943 to S1946 are repeated until the decomposition of the unit length is completed. If the decomposition of the unit length is completed (Yes in S1942), the three-dimensional data encoding device outputs a bitstream containing the generated information (S1947).
[0529] Figure 71 is a flowchart of a three-dimensional data decoding process that includes adaptive entropy decoding using attribute information.
[0530] First, the three-dimensional data decoder acquires the input bitstream (S1951). Next, the three-dimensional data decoder determines whether or not the decomposition process for unit lengths has been completed (S1952).
[0531] If the decomposition process for unit lengths is not complete (No in S1952), the three-dimensional data decoding device generates an octave tree by performing the decomposition process on the target node (S1953).
[0532] Next, the three-dimensional data decoding device acquires attribute information (S1954) and selects an encoding table based on the acquired attribute information (S1955). Here, attribute information refers to information such as the normal vector of the target node, as described above.
[0533] Next, the three-dimensional data decoding device uses the selected coding table to entropy-decode the occupancy code of the target node (S1956).
[0534] The processes described in steps S1953 to S1956 above are repeated until the decomposition of the unit length is completed. If the decomposition of the unit length is completed (Yes in S1952), the three-dimensional data decoder outputs a three-dimensional point (S1957).
[0535] Figure 72 is a flowchart of the selection process for the coding table using geometric information (S1905).
[0536] The three-dimensional data encoding device may switch the encoding table used for entropy encoding of the occupancy code using geometric information, such as information on geometric groups of a tree structure. Here, geometric group information refers to information indicating the geometric group that contains the geometric pattern of the target node.
[0537] As shown in Figure 72, if the geometric group indicated by the geometric information is geometric group 0 (Yes in S1961), the three-dimensional data encoding device selects encoding table 0 (S1962). If the geometric group indicated by the geometric information is geometric group 1 (Yes in S1963), the three-dimensional data encoding device selects encoding table 1 (S1964). Otherwise (No in S1963), the three-dimensional data encoding device selects encoding table 2 (S1965).
[0538] Note that the method of selecting the coding table is not limited to the above. For example, the three-dimensional data coding device may further switch coding tables depending on the value of the geometric group, such as using coding table 2 when the geometric group indicated by the geometric information is geometric group 2.
[0539] For example, a geometric group is determined using occupancy information indicating whether or not a point cloud is included in nodes adjacent to the target node. Geometric patterns that become the same shape by applying transformations such as rotation may also be included in the same geometric group. The 3D data encoding device may also select a geometric group using occupancy information of nodes adjacent to or around the target node, belonging to the same layer as the target node. Alternatively, the 3D data encoding device may select a geometric group using occupancy information of nodes belonging to a different layer from the target node. For example, the 3D data encoding device may select a geometric group using occupancy information of a parent node, or nodes adjacent to or around a parent node.
[0540] Furthermore, the selection process of the encoding table using geometric information in the three-dimensional data decoding device (S1915) is the same as described above.
[0541] Figure 73 is a flowchart of the selection process for the coding table using structural information (S1925).
[0542] The three-dimensional data encoding device may switch the encoding table used for entropy coding of occupancy codes using structural information, such as information about the layers of a tree structure. Here, the layer information indicates, for example, the layer to which the target node belongs.
[0543] As shown in Figure 73, if the target node belongs to layer 0 (Yes in S1971), the three-dimensional data encoding device selects encoding table 0 (S1972). If the target node belongs to layer 1 (Yes in S1973), the three-dimensional data encoding device selects encoding table 1 (S1974). Otherwise (No in S1973), the three-dimensional data encoding device selects encoding table 2 (S1975).
[0544] Note that the method of selecting the coding table is not limited to the above. For example, the three-dimensional data coding device may further switch coding tables depending on the layer to which the target node belongs, such as using coding table 2 when the target node belongs to layer 2.
[0545] Furthermore, the selection process of the coding table using structural information in the three-dimensional data decoding device (S1935) is the same as described above.
[0546] Figure 74 is a flowchart of the selection process for the coding table using attribute information (S1945).
[0547] The three-dimensional data encoding device may switch the encoding table used for entropy encoding of the occupancy code using attribute information such as information about the object to which the target node belongs, or information about the normal vector of the target node.
[0548] As shown in Figure 74, if the normal vector of the target node belongs to normal vector group 0 (Yes in S1981), the three-dimensional data encoding device selects encoding table 0 (S1982). If the normal vector of the target node belongs to normal vector group 1 (Yes in S1983), the three-dimensional data encoding device selects encoding table 1 (S1984). Otherwise (No in S1983), the three-dimensional data encoding device selects encoding table 2 (S1985).
[0549] Note that the method of selecting the coding table is not limited to the above. For example, the three-dimensional data coding device may further switch coding tables depending on the normal vector group to which the normal vector of the target node belongs, such as using coding table 2 when the normal vector of the target node belongs to normal vector group 2.
[0550] For example, a three-dimensional data encoding device selects a group of normal vectors using information about the normal vectors of the target node. For instance, a three-dimensional data encoding device determines that normal vectors whose distance from each other is below a predetermined threshold belong to the same group of normal vectors.
[0551] Furthermore, the information about the object to which the target node belongs may be, for example, information about a person, a car, or a building.
[0552] The configurations of the three-dimensional data encoding device 1900 and the three-dimensional data decoding device 1910 according to this embodiment will be described below. Figure 75 is a block diagram of the three-dimensional data encoding device 1900 according to this embodiment. The three-dimensional data encoding device 1900 shown in Figure 75 comprises an octree generation unit 1901, a similarity information calculation unit 1902, an encoding table selection unit 1903, and an entropy encoding unit 1904.
[0553] The octree generation unit 1901 generates, for example, an octree from the input three-dimensional points and generates occupancy codes for each node included in the octree. The similarity information calculation unit 1902 obtains similarity information, such as geometric information, structural information, or attribute information of the target node. The coding table selection unit 1903 selects a context to be used for entropy coding of the occupancy code according to the similarity information of the target node. The entropy coding unit 1904 generates a bitstream by entropy coding the occupancy code using the selected context. The entropy coding unit 1904 may also add information indicating the selected context to the bitstream.
[0554] Figure 76 is a block diagram of the three-dimensional data decoding device 1910 according to this embodiment. The three-dimensional data decoding device 1910 shown in Figure 76 comprises an octree generation unit 1911, a similarity information calculation unit 1912, an encoding table selection unit 1913, and an entropy decoding unit 1914.
[0555] The octree generation unit 1911 generates an octree sequentially, for example, from the lower layers to the upper layers, using information obtained from the entropy decoding unit 1914. The similarity information calculation unit 1912 obtains similarity information, which is the geometric information, structural information, or attribute information of the target node. The coding table selection unit 1913 selects a context to be used for entropy decoding of the occupancy code according to the similarity information of the target node. The entropy decoding unit 1914 generates a three-dimensional point by entropy decoding the occupancy code using the selected context. Alternatively, the entropy decoding unit 1914 may decode and obtain the selected context information attached to the bitstream and use the context indicated by that information.
[0556] As shown in Figures 63 to 65 above, multiple contexts are provided for each bit of the occupancy code. In other words, the three-dimensional data encoding device entropy encodes a bit sequence representing an N (where N is an integer greater than or equal to 2) subtree structure of multiple three-dimensional points contained in the three-dimensional data, using an encoding table selected from multiple encoding tables. The bit sequence contains N bits of information for each node in the N subtree structure. The N bits of information include N bits of 1-bit information indicating whether or not a three-dimensional point exists in each of the N child nodes of the corresponding node. In each of the multiple encoding tables, a context is provided for each bit of the N bits of information. In entropy encoding, the three-dimensional data encoding device entropy encodes each bit of the N bits of information using the context provided for that bit in the selected encoding table.
[0557] According to this, a three-dimensional data encoding device can improve encoding efficiency by switching the context for each bit.
[0558] For example, in entropy coding, a three-dimensional data encoding device selects a coding table from multiple coding tables based on whether a three-dimensional point exists in each of the multiple neighboring nodes adjacent to the target node. This allows the three-dimensional data encoding device to improve coding efficiency by switching coding tables based on whether or not a three-dimensional point exists in the neighboring nodes.
[0559] For example, in entropy coding, a three-dimensional data encoding device selects an encoding table based on an arrangement pattern that indicates the placement of adjacent nodes where a three-dimensional point exists among multiple adjacent nodes. For arrangement patterns that become identical after rotation, the same encoding table is selected. This allows the three-dimensional data encoding device to suppress the increase in the number of encoding tables.
[0560] For example, in entropy coding, a three-dimensional data encoding device selects the encoding table to use from multiple encoding tables based on the layer to which the target node belongs. This allows the three-dimensional data encoding device to improve encoding efficiency by switching encoding tables based on the layer to which the target node belongs.
[0561] For example, in entropy coding, a three-dimensional data coding device selects the coding table to use from multiple coding tables based on the normal vector of the target node. This allows the three-dimensional data coding device to improve coding efficiency by switching coding tables based on the normal vector.
[0562] For example, a three-dimensional data encoding device comprises a processor and memory, and the processor uses the memory to perform the above processing.
[0563] Furthermore, the three-dimensional data decoder performs entropy decoding of a bit sequence representing an N (where N is an integer greater than or equal to 2) subtree structure of multiple three-dimensional points contained in the three-dimensional data, using an encoding table selected from multiple encoding tables. The bit sequence contains N bits of information for each node in the N subtree structure. The N bits of information include N 1-bit pieces of information indicating whether or not a three-dimensional point exists in each of the N child nodes of the corresponding node. In each of the multiple encoding tables, a context is provided for each bit of the N bits of information. In entropy decoding, the three-dimensional data decoder performs entropy decoding of each bit of the N bits of information using the context provided for that bit in the selected encoding table.
[0564] According to this, a three-dimensional data decoding device can improve encoding efficiency by switching the context for each bit.
[0565] For example, in entropy decoding, a three-dimensional data decoding device selects a coding table from multiple coding tables based on whether a three-dimensional point exists in each of the multiple neighboring nodes adjacent to the target node. This allows the three-dimensional data decoding device to improve coding efficiency by switching coding tables based on whether or not a three-dimensional point exists in the neighboring nodes.
[0566] For example, in entropy decoding, a three-dimensional data decoder selects an encoding table based on an arrangement pattern that indicates the placement of adjacent nodes where a three-dimensional point exists among multiple adjacent nodes. For arrangement patterns that become identical after rotation, the same encoding table is selected. This allows the three-dimensional data decoder to suppress the increase in the encoding table.
[0567] For example, in entropy decoding, a three-dimensional data decoding device selects the coding table to use from multiple coding tables based on the layer to which the target node belongs. This allows the three-dimensional data decoding device to improve coding efficiency by switching coding tables based on the layer to which the target node belongs.
[0568] For example, in entropy decoding, a three-dimensional data decoding device selects the coding table to use from multiple coding tables based on the normal vector of the target node. This allows the three-dimensional data decoding device to improve coding efficiency by switching coding tables based on the normal vector.
[0569] For example, a three-dimensional data decoding device comprises a processor and memory, and the processor uses the memory to perform the above processing.
[0570] (Embodiment 9) This embodiment describes a method for controlling references during occupancy coding. While the operation of a three-dimensional data encoding device is primarily described below, similar processing may be performed in a three-dimensional data decoding device.
[0571] Figures 77 and 78 illustrate the reference relationships according to this embodiment; Figure 77 shows the reference relationships on an octave tree structure, and Figure 78 shows the reference relationships on a spatial domain.
[0572] In this embodiment, when the three-dimensional data encoding device encodes the encoding information of a node to be encoded (hereinafter referred to as the target node), it references the encoding information of each node within the parent node to which the target node belongs. However, it does not reference the encoding information of other nodes on the same layer as the parent node (hereinafter referred to as parent-adjacent nodes). In other words, the three-dimensional data encoding device sets parent-adjacent nodes to be unavailable or prohibits their reference.
[0573] Furthermore, the three-dimensional data encoding device may allow referencing the encoding information within the parent node to which the parent node belongs (hereinafter referred to as the grandparent node). In other words, the three-dimensional data encoding device may encode the target node's encoding information by referencing the encoding information of the parent node and grandparent node to which the target node belongs.
[0574] Here, encoded information refers to, for example, an occupancy code. When a three-dimensional data encoding device encodes the occupancy code of a target node, it refers to information indicating whether or not the point cloud is contained in each node within the parent node to which the target node belongs (hereinafter referred to as occupancy information). In other words, when a three-dimensional data encoding device encodes the occupancy code of a target node, it refers to the occupancy code of the parent node. On the other hand, the three-dimensional data encoding device does not refer to the occupancy information of each node within the parent-adjacent node. That is, the three-dimensional data encoding device does not refer to the occupancy code of the parent-adjacent node. Furthermore, the three-dimensional data encoding device may refer to the occupancy information of each node within the grandparent node. That is, the three-dimensional data encoding device may refer to the occupancy information of the parent node and the parent-adjacent node.
[0575] For example, when a three-dimensional data encoding device encodes the occupancy code of a target node, it switches the encoding table used when entropy encoding the target node's occupancy code using the occupancy code of the parent or grandparent node to which the target node belongs. The details of this will be described later. In this case, the three-dimensional data encoding device does not need to refer to the occupancy code of the parent or neighbor node. This allows the three-dimensional data encoding device to appropriately switch the encoding table according to the information of the parent or grandparent node's occupancy code when encoding the target node's occupancy code, thereby improving encoding efficiency. In addition, by not referring to the parent or neighbor node, the three-dimensional data encoding device can reduce the processing required to verify the information of the parent or neighbor node and the memory capacity required to store it. Furthermore, it becomes easier to scan and encode the occupancy code of each node in the octree in depth-first order.
[0576] The following describes an example of coding table switching using the parent node's occupancy code. Figure 79 shows an example of a target node and adjacent reference nodes. Figure 80 shows the relationship between a parent node and a node. Figure 81 shows an example of the parent node's occupancy code. Here, an adjacent reference node is a node that is spatially adjacent to the target node and is referenced when coding the target node. In the example shown in Figure 79, the adjacent nodes are nodes belonging to the same layer as the target node. In addition, the node X adjacent in the x direction, the node Y adjacent in the y direction, and the node Z adjacent in the z direction of the target block are used as reference adjacent nodes. In other words, one adjacent block is set as the reference adjacent block in each of the x, y, and z directions.
[0577] Note that the node numbers shown in Figure 80 are just examples, and the relationship between node numbers and node positions is not limited to these. Also, in Figure 81, node 0 is assigned to the lower bits and node 7 is assigned to the upper bits, but the assignment may be done in the reverse order. Furthermore, each node may be assigned to any bit.
[0578] The three-dimensional data encoding device determines the encoding table for entropy encoding the occupancy code of the target node using, for example, the following formula.
[0579] CodingTable=(FlagX<<2)+(FlagY<<1)+(FlagZ)
[0580] Here, CodingTable represents the coding table for the occupancy code of the target node, and is a value from 0 to 7. FlagX is the occupancy information of the neighboring node X, and is 1 if the neighboring node X contains (occupies) the point cloud, and 0 otherwise. FlagY is the occupancy information of the neighboring node Y, and is 1 if the neighboring node Y contains (occupies) the point cloud, and 0 otherwise. FlagZ is the occupancy information of the neighboring node Z, and is 1 if the neighboring node Z contains (occupies) the point cloud, and 0 otherwise.
[0581] Furthermore, since information indicating whether an adjacent node is occupied is included in the parent node's occupancy code, the three-dimensional data encoding device may select the encoding table using the value indicated in the parent node's occupancy code.
[0582] As described above, the three-dimensional data encoding device can improve encoding efficiency by switching the encoding table using information indicating whether or not point clouds are included in the adjacent nodes of the target node.
[0583] Furthermore, as shown in Figure 79, the three-dimensional data encoding device may switch adjacent reference nodes according to the spatial position of the target node within the parent node. In other words, the three-dimensional data encoding device may switch the adjacent node to be referenced from among multiple adjacent nodes according to the spatial position of the target node within the parent node.
[0584] Next, an example of the configuration of a three-dimensional data encoding device and a three-dimensional data decoding device will be described. Figure 82 is a block diagram of the three-dimensional data encoding device 2100 according to this embodiment. The three-dimensional data encoding device 2100 shown in Figure 82 comprises an octree generation unit 2101, a geometric information calculation unit 2102, an encoding table selection unit 2103, and an entropy encoding unit 2104.
[0585] The octree generation unit 2101 generates, for example, an octree from the input three-dimensional points (point cloud) and generates occupancy codes for each node included in the octree. The geometric information calculation unit 2102 obtains occupancy information indicating whether the adjacent reference nodes of the target node are occupied or not. For example, the geometric information calculation unit 2102 obtains occupancy information of adjacent reference nodes from the occupancy code of the parent node to which the target node belongs. Note that the geometric information calculation unit 2102 may switch adjacent reference nodes depending on the position of the target node within the parent node, as shown in Figure 79. Also, the geometric information calculation unit 2102 does not refer to the occupancy information of each node within the parent adjacent node.
[0586] The coding table selection unit 2103 selects a coding table to be used for entropy coding of the occupancy code of the target node using the occupancy information of adjacent reference nodes calculated by the geometric information calculation unit 2102. The entropy coding unit 2104 generates a bitstream by entropy coding the occupancy code using the selected coding table. The entropy coding unit 2104 may also add information indicating the selected coding table to the bitstream.
[0587] Figure 83 is a block diagram of the three-dimensional data decoding device 2110 according to this embodiment. The three-dimensional data decoding device 2110 shown in Figure 83 comprises an octree generation unit 2111, a geometric information calculation unit 2112, an encoding table selection unit 2113, and an entropy decoding unit 2114.
[0588] The octane tree generation unit 2111 generates an octane tree of a certain space (node) using the header information of the bitstream. For example, the octane tree generation unit 2111 generates a large space (root node) using the size of the x, y, and z axes of a certain space attached to the header information, and then generates eight small spaces A (nodes A0 to A7) by dividing that space into two along the x, y, and z axes, respectively, to generate an octane tree. Also, nodes A0 to A7 are set in order as the target nodes.
[0589] The geometric information calculation unit 2112 obtains occupancy information indicating whether the adjacent reference node of the target node is occupied or not. For example, the geometric information calculation unit 2112 obtains the occupancy information of the adjacent reference node from the occupancy code of the parent node to which the target node belongs. Note that, as shown in Figure 79, the geometric information calculation unit 2112 may switch adjacent reference nodes depending on the position of the target node within the parent node. Furthermore, the geometric information calculation unit 2112 does not refer to the occupancy information of each node within the parent adjacent node.
[0590] The coding table selection unit 2113 selects a coding table (decoding table) to be used for entropy decoding of the occupancy code of the target node using the occupancy information of adjacent reference nodes calculated by the geometric information calculation unit 2112. The entropy decoding unit 2114 generates three-dimensional points by entropy decoding the occupancy code using the selected coding table. Alternatively, the coding table selection unit 2113 may decode and obtain the information of the selected coding table attached to the bitstream, and the entropy decoding unit 2114 may use the coding table indicated by the obtained information.
[0591] Each bit of the occupancy code (8 bits) contained in the bitstream indicates whether or not the point cloud is contained in each of the eight subspaces A (nodes A0 to A7). Furthermore, the three-dimensional data decoder divides subspace node A0 into eight subspaces B (nodes B0 to B7) to generate an octave tree, and decodes the occupancy code to obtain information indicating whether or not the point cloud is contained in each node of subspace B. In this way, the three-dimensional data decoder decodes the occupancy code of each node while generating an octave tree from the large space to the small spaces.
[0592] The following describes the processing flow by the three-dimensional data encoding device and the three-dimensional data decoding device. Figure 84 is a flowchart of the three-dimensional data encoding process in the three-dimensional data encoding device. First, the three-dimensional data encoding device determines (defines) the space (target node) that contains part or all of the input three-dimensional point cloud (S2101). Next, the three-dimensional data encoding device divides the target node into eight parts to generate eight small spaces (nodes) (S2102). Next, the three-dimensional data encoding device generates an occupancy code for the target node depending on whether or not the point cloud is contained in each node (S2103).
[0593] Next, the three-dimensional data encoding device calculates (obtains) the occupancy information of the neighboring reference nodes of the target node from the occupancy code of the target node's parent node (S2104). Next, the three-dimensional data encoding device selects an encoding table to be used for entropy encoding based on the determined occupancy information of the neighboring reference nodes of the target node (S2105). Next, the three-dimensional data encoding device entropy encodes the occupancy code of the target node using the selected encoding table (S2106).
[0594] Furthermore, the three-dimensional data encoding device divides each node into eight parts and encodes the occupancy code of each node, repeating this process until the nodes can no longer be divided (S2107). In other words, the processes from steps S2102 to S2106 are repeated recursively.
[0595] Figure 85 is a flowchart of the three-dimensional data decoding method in a three-dimensional data decoding device. First, the three-dimensional data decoding device determines (defines) the space (target node) to be decoded using the bitstream header information (S2111). Next, the three-dimensional data decoding device divides the target node into eight parts to generate eight small spaces (nodes) (S2112). Next, the three-dimensional data decoding device calculates (obtains) the occupancy information of the adjacent reference nodes of the target node from the occupancy code of the parent node of the target node (S2113).
[0596] Next, the three-dimensional data decoding device selects an encoding table to be used for entropy decoding based on the occupancy information of the adjacent reference node (S2114). Then, the three-dimensional data decoding device entropy decodes the occupancy code of the target node using the selected encoding table (S2115).
[0597] Furthermore, the three-dimensional data decoding device divides each node into eight parts and decodes the occupancy code of each node, repeating this process until the nodes can no longer be divided (S2116). In other words, the process from steps S2112 to S2115 is repeated recursively.
[0598] Next, we will explain an example of switching coding tables. Figure 86 shows an example of switching coding tables. For example, as shown in coding table 0 in Figure 86, the same context model may be applied to multiple occupancy codes. Alternatively, a different context model may be assigned to each occupancy code. This allows for the assignment of a context model according to the probability of occurrence of the occupancy code, thereby improving coding efficiency. Alternatively, a context model that updates the probability table according to the frequency of occurrence of the occupancy code may be used. Or, a context model with a fixed probability table may be used.
[0599] Note that while Figure 86 shows an example where the encoding tables shown in Figures 60 and 61 are used, the encoding tables shown in Figures 63 and 64 may also be used.
[0600] The following describes Modification 1 of this embodiment. Figure 87 is a diagram showing the reference relationship in this modification. In the above embodiment, the three-dimensional data encoding device does not refer to the occupancy code of the parent neighbor node, but whether or not to refer to the occupancy code of the parent neighbor node may be switched depending on specific conditions.
[0601] For example, when a three-dimensional data encoding device performs encoding while scanning an octree in a breadth-first order, it references the occupancy information of nodes within the parent-neighbor nodes to encode the occupancy code of the target node. On the other hand, when a three-dimensional data encoding device performs encoding while scanning an octree in a depth-first order, it prohibits referencing the occupancy information of nodes within the parent-neighbor nodes. By appropriately switching which nodes can be referenced according to the scanning order (encoding order) of the nodes in the octree in this way, it is possible to improve encoding efficiency and reduce processing load.
[0602] Furthermore, the three-dimensional data encoding device may add information to the bitstream header, such as whether the octree was encoded using breadth-first or depth-first encoding. Figure 88 shows an example of the syntax of the header information in this case. The octree_scan_order shown in Figure 88 is encoding order information (encoding order flag) that indicates the encoding order of the octree. For example, if octree_scan_order is 0, it indicates breadth-first encoding, and if it is 1, it indicates depth-first encoding. As a result, the three-dimensional data decoding device can determine whether the bitstream was encoded using breadth-first or depth-first encoding by referring to octree_scan_order, and thus can decode the bitstream appropriately.
[0603] Furthermore, the three-dimensional data encoding device may add information to the bitstream header information indicating whether or not to prohibit referencing parent-neighbor nodes. Figure 89 shows an example of the header information syntax in this case. limit_refer_flag is prohibition switching information (prohibition switching flag) that indicates whether or not to prohibit referencing parent-neighbor nodes. For example, if limit_refer_flag is 1, it indicates that referencing parent-neighbor nodes is prohibited, and if it is 0, it indicates no referencing restriction (referencing parent-neighbor nodes is permitted).
[0604] In other words, the three-dimensional data encoding device decides whether or not to prohibit referencing the parent neighbor node, and based on the result of the above decision, switches between prohibiting or allowing referencing the parent neighbor node. The three-dimensional data encoding device also generates a bitstream that includes prohibition switching information indicating whether or not to prohibit referencing the parent neighbor node, which is the result of the above decision.
[0605] Furthermore, the three-dimensional data decoding device obtains prohibition switching information from the bitstream, indicating whether or not to prohibit referencing the parent neighbor node, and switches whether to prohibit or allow referencing the parent neighbor node based on the prohibition switching information.
[0606] This allows the three-dimensional data encoding device to control the referencing of parent neighbor nodes and generate a bitstream. Furthermore, the three-dimensional data decoding device can obtain information from the bitstream header indicating whether or not referencing parent neighbor nodes is prohibited.
[0607] Furthermore, while this embodiment describes occupancy coding as an example of coding that prohibits referencing parent neighbor nodes, it is not necessarily limited to this. For example, a similar method can be applied when coding other information of an octree node. For instance, the method of this embodiment may be applied when coding other attribute information such as color, normal vector, or reflectance attached to a node. A similar method can also be applied when coding a coding table or predicted values.
[0608] Next, a modified example of this embodiment 2 will be described. In the above description, an example in which three reference neighbor nodes are used was shown as shown in Figure 79, but four or more reference neighbor nodes may be used. Figure 90 is a diagram showing an example of a target node and reference neighbor nodes.
[0609] For example, the three-dimensional data encoding device calculates the encoding table for entropy encoding the occupancy code of the target node shown in Figure 90 using, for example, the following formula.
[0610] CodingTable=(FlagX0<<3)+(FlagX1<<2)+(FlagY<<1)+(FlagZ)
[0611] Here, CodingTable represents the coding table for the occupancy code of the target node, and is a value between 0 and 15. FlagXN is the occupancy information of the neighboring node XN (N=0..1), and is 1 if the neighboring node XN contains (occupies) the point cloud, and 0 otherwise. FlagY is the occupancy information of the neighboring node Y, and is 1 if the neighboring node Y contains (occupies) the point cloud, and 0 otherwise. FlagZ is the occupancy information of the neighboring node Z, and is 1 if the neighboring node Z contains (occupies) the point cloud, and 0 otherwise.
[0612] In this case, if an adjacent node, for example, adjacent node X0 in Figure 90, is unavailable (forbidden to access), the three-dimensional data encoding device may use a fixed value such as 1 (occupied) or 0 (unoccupied) as an alternative value.
[0613] Figure 91 shows an example of a target node and adjacent nodes. As shown in Figure 91, if an adjacent node is not accessible (access prohibited), the occupancy information of the adjacent node may be calculated by referring to the occupancy code of the target node's grandparent node. For example, the three-dimensional data encoding device may calculate FlagX0 in the above formula using the occupancy information of adjacent node G0 instead of adjacent node X0 shown in Figure 91, and then determine the value of the encoding table using the calculated FlagX0. Note that adjacent node G0 shown in Figure 91 is an adjacent node whose occupancy status can be determined by the occupancy code of the grandparent node. Adjacent node X1 is an adjacent node whose occupancy status can be determined by the occupancy code of the parent node.
[0614] The following describes a third modification of this embodiment. Figures 92 and 93 are diagrams showing the reference relationships related to this modification; Figure 92 shows the reference relationships on an octave tree structure, and Figure 93 shows the reference relationships on a spatial domain.
[0615] In this modified example, when the three-dimensional data encoding device encodes the encoding information of the node to be encoded (hereinafter referred to as target node 2), it refers to the encoding information of each node within the parent node to which target node 2 belongs. In other words, the three-dimensional data encoding device allows referencing information (e.g., occupancy information) of the child nodes of the first node, which has the same parent node as the target node, among multiple adjacent nodes. For example, when the three-dimensional data encoding device encodes the occupancy code of target node 2 shown in Figure 92, it refers to the nodes existing within the parent node to which target node 2 belongs, for example, the occupancy code of the target node shown in Figure 92. As shown in Figure 93, the occupancy code of the target node shown in Figure 92 indicates, for example, whether each node within the target nodes adjacent to target node 2 is occupied or not. Therefore, the three-dimensional data encoding device can switch the encoding table of the occupancy code of target node 2 according to the finer shape of the target node, thereby improving encoding efficiency.
[0616] The three-dimensional data encoding device may calculate the encoding table for entropy encoding the occupancy code of target node 2 using, for example, the following formula.
[0617] CodingTable=(FlagX1<<5)+(FlagX2<<4)+(FlagX3<<3)+(FlagX4<<2)+(FlagY<<1)+(FlagZ)
[0618] Here, CodingTable represents the coding table for the occupancy code of target node 2, and displays a value between 0 and 63. FlagXN is the occupancy information of neighboring node XN (N=1..4), showing 1 if neighboring node XN contains (occupies) the point cloud, and 0 otherwise. FlagY is the occupancy information of neighboring node Y, showing 1 if neighboring node Y contains (occupies) the point cloud, and 0 otherwise. FlagZ is the occupancy information of neighboring node Y, showing 1 if neighboring node Z contains (occupies) the point cloud, and 0 otherwise.
[0619] Furthermore, the three-dimensional data encoding device may change the method of calculating the encoding table according to the node position of target node 2 within the parent node.
[0620] Furthermore, if referencing the parent neighbor node is not prohibited, the three-dimensional data encoding device may refer to the encoding information of each node within the parent neighbor node. For example, if referencing the parent neighbor node is not prohibited, it is permitted to refer to information (e.g., occupancy information) of child nodes of a third node whose parent node is different from the target node. For example, in the example shown in Figure 91, the three-dimensional data encoding device refers to the occupancy code of neighbor node X0, whose parent node is different from the target node, and obtains the occupancy information of the child nodes of neighbor node X0. Based on the obtained occupancy information of the child nodes of neighbor node X0, the three-dimensional data encoding device switches the encoding table used for entropy encoding of the occupancy code of the target node.
[0621] As described above, the three-dimensional data encoding device according to this embodiment encodes information (e.g., occupancy codes) of target nodes included in an N (where N is an integer of 2 or more) subtree structure of multiple three-dimensional points included in three-dimensional data. As shown in Figures 77 and 78, in the above encoding, the three-dimensional data encoding device allows referencing information (e.g., occupancy information) of a first node among multiple adjacent nodes spatially adjacent to the target node, where the target node and the parent node are the same, and prohibits referencing information (e.g., occupancy information) of a second node where the target node and the parent node are different. In other words, in the above encoding, the three-dimensional data encoding device allows referencing information (e.g., occupancy codes) of the parent node, and prohibits referencing information (e.g., occupancy codes) of other nodes (parent adjacent nodes) on the same layer as the parent node.
[0622] According to this, the three-dimensional data encoding device can improve encoding efficiency by referencing information from the first node, which is one of several adjacent nodes spatially adjacent to the target node and shares the same parent node as the target node. Furthermore, the three-dimensional data encoding device can reduce processing load by not referencing information from the second node, which is one of several adjacent nodes and has a different parent node than the target node. In this way, the three-dimensional data encoding device can improve encoding efficiency and reduce processing load.
[0623] For example, the three-dimensional data encoding device further decides whether or not to prohibit referencing information from the second node, and in the encoding, it switches between prohibiting or allowing referencing information from the second node based on the result of the above decision. The three-dimensional data encoding device further generates a bitstream that includes prohibition switching information (for example, limit_refer_flag shown in Figure 89) indicating whether or not to prohibit referencing information from the second node, which is the result of the above decision.
[0624] According to this, the three-dimensional data encoding device can switch whether or not to prohibit access to the information of the second node. Furthermore, the three-dimensional data decoding device can perform decoding appropriately using the prohibition switching information.
[0625] For example, the information of the target node is information indicating whether or not a three-dimensional point exists in each of the child nodes belonging to the target node (e.g., an occupancy code), the information of the first node is information indicating whether or not a three-dimensional point exists in the first node (occupancy information of the first node), and the information of the second node is information indicating whether or not a three-dimensional point exists in the second node (occupancy information of the second node).
[0626] For example, in the above encoding, the three-dimensional data encoding device selects an encoding table based on whether or not a three-dimensional point exists at the first node, and uses the selected encoding table to entropy encode the information of the target node (e.g., occupancy code).
[0627] For example, in the above encoding, the three-dimensional data encoding device allows access to information (e.g., occupancy information) of the child nodes of the first node among multiple neighboring nodes, as shown in Figures 92 and 93.
[0628] According to this, the three-dimensional data encoding device can improve encoding efficiency because it can refer to more detailed information of adjacent nodes.
[0629] For example, as shown in Figure 79, the three-dimensional data encoding device switches the referenced neighbor node from among multiple neighboring nodes depending on the spatial position of the target node within the parent node during the encoding process.
[0630] According to this, the three-dimensional data encoding device can refer to appropriate neighboring nodes depending on the spatial position of the target node within its parent node.
[0631] For example, a three-dimensional data encoding device comprises a processor and memory, and the processor uses the memory to perform the above processing.
[0632] Furthermore, the three-dimensional data decoding device according to this embodiment decodes information (e.g., occupancy code) of a target node included in an N (where N is an integer of 2 or more) subtree structure of multiple three-dimensional points included in the three-dimensional data. As shown in Figures 77 and 78, in the above decoding, the three-dimensional data decoding device allows access to information (e.g., occupancy information) of a first node among multiple adjacent nodes spatially adjacent to the target node, where the target node and the parent node are the same, and prohibits access to information (e.g., occupancy information) of a second node where the target node and the parent node are different. In other words, in the above decoding, the three-dimensional data decoding device allows access to information (e.g., occupancy code) of the parent node, and prohibits access to information (e.g., occupancy code) of other nodes (parent adjacent nodes) on the same layer as the parent node.
[0633] According to this, the three-dimensional data decoding device can improve encoding efficiency by referencing information from the first node, which is one of several adjacent nodes spatially adjacent to the target node and shares the same parent node as the target node. Furthermore, the three-dimensional data decoding device can reduce processing load by not referencing information from the second node, which is one of several adjacent nodes and has a different parent node than the target node. Thus, the three-dimensional data decoding device can improve encoding efficiency and reduce processing load.
[0634] For example, the three-dimensional data decoding device further obtains prohibition switching information (for example, limit_refer_flag shown in Figure 89) from the bitstream, indicating whether or not to prohibit access to the information of the second node. In the decoding process, based on this prohibition switching information, the device switches between prohibiting or allowing access to the information of the second node.
[0635] According to this, the three-dimensional data decoding device can perform decoding appropriately using the prohibition switching information.
[0636] For example, the information of the target node is information indicating whether or not a three-dimensional point exists in each of the child nodes belonging to the target node (e.g., an occupancy code), the information of the first node is information indicating whether or not a three-dimensional point exists in the first node (occupancy information of the first node), and the information of the second node is information indicating whether or not a three-dimensional point exists in the second node (occupancy information of the second node).
[0637] For example, in the above decoding process, a three-dimensional data decoding device selects an encoding table based on whether or not a three-dimensional point exists at the first node, and uses the selected encoding table to entropically decode the information of the target node (e.g., an occupancy code).
[0638] For example, in the above decoding process, the three-dimensional data decoding device allows access to information (e.g., occupancy information) of the child nodes of the first node among multiple neighboring nodes, as shown in Figures 92 and 93.
[0639] According to this, the three-dimensional data decoding device can improve encoding efficiency because it can refer to more detailed information of adjacent nodes.
[0640] For example, as shown in Figure 79, the three-dimensional data decoding device switches the referenced neighbor node from among multiple neighboring nodes depending on the spatial position of the target node within the parent node during the decoding process.
[0641] According to this, the three-dimensional data decoding device can refer to the appropriate adjacent node depending on the spatial position of the target node within the parent node.
[0642] For example, a three-dimensional data decoding device comprises a processor and memory, and the processor uses the memory to perform the above processing.
[0643] (Embodiment 10) This embodiment describes a method for reducing the number of encoding tables.
[0644] If we create a coding table for each combination of the position of the target node within the parent node (8 patterns) and the occupation status patterns of the three adjacent nodes of the target node (8 patterns), then 8 × 8 = 64 coding tables are required. In the following, this combination will also be referred to as the adjacent occupation pattern. Furthermore, an occupied node will be referred to as an occupied node, and an adjacent node in an occupied state will be referred to as an adjacent occupied node.
[0645] In this embodiment, the total number of coding tables is reduced by assigning the same coding table to similar adjacent-occupancy patterns. Specifically, multiple adjacent-occupancy patterns are grouped by performing a conversion process on them. More specifically, adjacent-occupancy patterns that become the same pattern after the conversion process are grouped into the same group. Furthermore, one coding table is assigned to each group.
[0646] For example, as a transformation process, translation along the x, y, or z axis is used, as shown in Figure 94. Alternatively, rotation along the x, y, or z axis (with the x, y, or z axis as the axis) is used, as shown in Figure 95.
[0647] Furthermore, grouped adjacent occupancy patterns may be classified using the following rules. For example, as shown in Figure 96, the occupying node and the target node may be located on a plane that is horizontal or perpendicular to the coordinate plane (xy plane, yz plane, or xz plane). Alternatively, as shown in Figure 97, the adjacent plane may be the direction in which the adjacent occupying node is located relative to the target node.
[0648] Figure 98 shows an example of translation along the x, y, or z axis. Figure 99 shows an example of rotation along the x axis. Figure 100 shows an example of rotation along the y axis. Figure 101 shows an example of rotation along the z axis. Figure 102 shows an example of horizontal or perpendicular to the coordinate plane. Figure 103 shows an example of patterns on adjacent planes.
[0649] Figure 104 shows an example of dividing 64 adjacent occupancy patterns into 6 groups. In other words, 6 coding tables are used in this example. In this example, the coding table reduction is achieved by rotation along the z-axis.
[0650] Specifically, as shown in Figure 104, an adjacent occupancy pattern where the number of occupied adjacent nodes is 0 out of 3 adjacent nodes is classified as Group 0. An adjacent occupancy pattern with an occupancy of 1 and that is horizontal to the xy plane is classified as Group 1. An adjacent occupancy pattern with an occupancy of 1 and that is perpendicular to the xy plane is classified as Group 2. An adjacent occupancy pattern with an occupancy of 2 and that is perpendicular to the xy plane is classified as Group 3. An adjacent occupancy pattern with an occupancy of 2 and that is horizontal to the xy plane is classified as Group 4. An adjacent occupancy pattern with an occupancy of 3 is classified as Group 5.
[0651] Furthermore, one coding table is used for each group. Each group also includes adjacent occupancy patterns that become identical when rotated along the z-axis. Note that for group 4, translation along the z-axis is also taken into consideration.
[0652] For example, in a three-dimensional map where the xy-plane is the ground, there may be multiple buildings with similar shapes on the xy-plane. In such a case, if building A is rotated along the z-axis, it may overlap with another building B. In this case, by using the coded table of the occupancy code updated by encoding building A when encoding the occupancy code of building B, it may be possible to improve the encoding efficiency when encoding building B. Therefore, by considering the coded table for shapes rotated along the z-axis as the same group, the coded table can be updated without being affected by rotation along the z-axis, thus improving encoding efficiency. Also, for example, if building C is translated parallel to the xy-plane, it may overlap with another building D. In this case, by using the coded table of the occupancy code updated by encoding building C when encoding the occupancy code of building D, it may be possible to improve the encoding efficiency when encoding building D. Therefore, by considering the coded table related to translation parallel to the xy-plane as the same group, the coded table can be updated without being affected by translation parallel to the xy-plane, thus improving encoding efficiency.
[0653] Figure 105 shows an example of dividing 64 adjacent occupancy patterns into 8 groups. In other words, in this example, 8 coding tables are used. In this example, the coding table reduction is achieved by rotation along the z-axis and adjacent faces.
[0654] Specifically, as shown in Figure 105, an adjacent occupancy pattern where the number of occupied nodes out of three adjacent nodes is 0 is classified as Group 0. An adjacent occupancy pattern with an occupancy of 1 and horizontal to the xy plane is classified as Group 1. An adjacent occupancy pattern with an occupancy of 1 and perpendicular to the xy plane is classified as Group 2. An adjacent occupancy pattern with an occupancy of 2, perpendicular to the xy plane, and having an adjacent plane in the z direction (i.e., an adjacent occupied node exists in the z direction of the target node) is classified as Group 3. An adjacent occupancy pattern with an occupancy of 2, perpendicular to the xy plane, and having an adjacent plane in the -z direction (i.e., an adjacent occupied node exists in the -z direction of the target node) is classified as Group 4.
[0655] Adjacent occupancy patterns with 2 occupants and horizontal to the xy plane are classified as group 5. Adjacent occupancy patterns with 3 occupants and adjacent faces in the z direction are classified as group 5. Adjacent occupancy patterns with 3 occupants and adjacent faces in the -z direction are classified as group 5.
[0656] Furthermore, one coding table is used for each group. Each group also includes adjacent occupancy patterns that become identical when rotated along the z-axis. Note that for group 5, translation along the z-axis is also taken into consideration.
[0657] Mapping rules can be generated using the examples shown in Figures 104 and 105. Figure 106 shows an example of a mapping rule (transformation table) when using three coding tables for 64 adjacent-occupying patterns. In the example shown in Figure 106, each of the 64 adjacent-occupying patterns (patterns 0 to 63) is assigned one of the indices (tables 0 to 2) of the three coding tables. This rule can be represented, for example, by a lookup table (LUT).
[0658] Furthermore, mapping rules may be generated by adding new rules to a given mapping rule or by deleting parts of the rule. In other words, adjacent occupation patterns grouped by the first rule may be further grouped by the second rule. To put it another way, multiple adjacent occupation patterns may be assigned to multiple first coding tables by the first transformation table, and multiple first coding tables may be assigned to multiple second coding tables by the second transformation table, and arithmetic coding or arithmetic decoding may be performed using the second coding tables. For example, after the classification shown in Figure 105 is performed, some of the classified groups may be further merged to perform the classification shown in Figure 104.
[0659] Figure 107 shows an example of a transformation table for performing this classification. In Figure 107, coding table 1 is the index of the coding table derived from the given mapping rule, and coding table 2 is the index of the coding table representing the new mapping rule.
[0660] For example, coding table 1 is the index of the coding table obtained by the classification shown in Figure 105, and represents one of the indexes of the eight coding tables (tables 0 to 7). Similarly, coding table 2 represents one of the indexes of the six coding tables (tables 0 to 5) corresponding to the classification shown in Figure 104.
[0661] Specifically, since groups 4 and 5 shown in Figure 105 correspond to group 4 shown in Figure 104, as shown in Figure 107, tables 4 and 5 of coding table 1 are assigned to table 4 of coding table 2. Similarly, tables 6 and 7 of coding table 1 are assigned to table 5 of coding table 2.
[0662] The following is an overview of the mapping process. Mapping rules are used to find unique indexes in the coding table.
[0663] Figure 108 shows an overview of the mapping process that determines the index of the coding table from 64 adjacent occupancy patterns. As shown in Figure 108, adjacent occupancy patterns, including the location of the target node, are input to the mapping rule, and a table index (the index of the coding table) is output. The number of patterns is reduced by the mapping rule. For example, the mapping rule shown in Figure 106 is used as this mapping rule. As shown in Figure 106, the same table index is assigned to different adjacent occupancy patterns.
[0664] Next, entropy coding is performed using the coding table assigned to the obtained table index.
[0665] Figure 109 shows an overview of the mapping process when a table index is provided. As shown in Figure 109, table index 1 is input into the mapping rule and table index 2 is output. The number of patterns is reduced by the mapping rule. For example, the mapping rule shown in Figure 107 is used as this mapping rule. As shown in Figure 107, the same table index 2 is assigned to different table index 1s.
[0666] Next, entropy coding is performed using the coding table assigned to the obtained table index 2.
[0667] Next, the configuration of the three-dimensional data encoding device and the three-dimensional data decoding device according to this embodiment will be described. Figure 110 is a block diagram showing the configuration of the three-dimensional data encoding device 3600 according to this embodiment. The three-dimensional data encoding device 3600 shown in Figure 110 comprises an octree generation unit 3601, a geometric information calculation unit 3602, an index generation unit 3603, an encoding table selection unit 3604, and an entropy encoding unit 3605.
[0668] The octree generation unit 3601 generates, for example, an octree from the input three-dimensional points (point cloud) and generates occupancy codes for each node included in the octree. The geometric information calculation unit 3602 obtains occupancy information indicating whether the adjacent reference nodes of the target node are occupied or not. For example, the geometric information calculation unit 3602 calculates the occupancy information of adjacent reference nodes from the occupancy code of the parent node to which the target node belongs. The geometric information calculation unit 3602 may switch adjacent reference nodes depending on the position of the target node within the parent node. Furthermore, the geometric information calculation unit 3602 does not need to refer to the occupancy information of each node within the adjacent parent node.
[0669] The index generation unit 3603 generates an index of the coding table using adjacency information (for example, adjacency occupancy patterns).
[0670] The coding table selection unit 3604 uses the index of the coding table generated by the index generation unit 3603 to select the coding table to be used for entropy coding of the occupancy code of the target node.
[0671] Occupancy codes are encoded as decimal or binary numbers. For example, when binary encoding is used, the index of the encoding table generated by the index generation unit 3603 is used to select the binary context used for entropy encoding by the entropy encoding unit 3605. When decimal or M-ary encoding is used, the M-ary context is selected by the index of the encoding table.
[0672] The entropy coding unit 3605 generates a bitstream by entropy coding the occupancy code using the selected coding table. The entropy coding unit 3605 may also add information indicating the selected coding table to the bitstream.
[0673] Figure 111 is a block diagram of the three-dimensional data decoding device 3610 according to this embodiment. The three-dimensional data decoding device 3610 shown in Figure 111 comprises an octree generation unit 3611, a geometric information calculation unit 3612, an index generation unit 3613, an encoding table selection unit 3614, and an entropy decoding unit 3615.
[0674] The octane tree generation unit 3611 generates an octane tree of a given space (node) using the header information of the bitstream. For example, the octane tree generation unit 3611 generates a large space (root node) using the size of the space in the x, y, and z axes attached to the header information, and then generates eight small spaces A (nodes A0 to A7) by dividing that space into two in the x, y, and z axes, respectively, thereby generating an octane tree. Also, nodes A0 to A7 are set in order as the target nodes.
[0675] The geometric information calculation unit 3612 obtains occupancy information indicating whether the adjacent reference node of the target node is occupied or not. For example, the geometric information calculation unit 3612 obtains the occupancy information of the adjacent reference node from the occupancy code of the parent node to which the target node belongs. The geometric information calculation unit 3612 may switch the adjacent reference node depending on the position of the target node within the parent node. Furthermore, the geometric information calculation unit 3612 does not need to refer to the occupancy information of each node within the adjacent parent node.
[0676] The index generation unit 3613 generates an index of the coding table using adjacency information (for example, adjacency occupancy patterns).
[0677] The coding table selection unit 3614 uses the index of the coding table generated by the index generation unit 3613 to select the coding table to be used for entropy decoding of the occupancy code of the target node.
[0678] Occupancy codes are decoded as either decimal or binary numbers. For example, when binary coding is used, the index of the coding table mapped in the previous block is used to select the binary context used for entropy decoding of the next block. When decimal or M-ary coding is used, the M-ary context is selected by the index of the coding table.
[0679] The entropy decoding unit 3615 generates a three-dimensional point (point cloud) by performing entropy decoding of the occupancy code using the selected coding table. Alternatively, the entropy decoding unit 3615 may decode and obtain the information of the selected coding table attached to the bitstream, and use the coding table indicated by the obtained information.
[0680] Each bit of the occupancy code (8 bits) contained in the bitstream indicates whether or not the point cloud is contained in each of the eight subspaces A (nodes A0 to A7). Furthermore, the three-dimensional data decoder divides subspace node A0 into eight subspaces B (nodes B0 to B7) to generate an octave tree, and decodes the occupancy code to obtain information indicating whether or not the point cloud is contained in each node of subspace B. In this way, the three-dimensional data decoder decodes the occupancy code of each node while generating an octave tree from the large space to the small spaces.
[0681] The following describes the processing flow by the three-dimensional data encoding device and the three-dimensional data decoding device. Figure 112 is a flowchart of the three-dimensional data encoding process in the three-dimensional data encoding device. First, the three-dimensional data encoding device defines a space (target node) that contains part or all of the input three-dimensional point cloud (S3601). Next, the three-dimensional data encoding device divides the target node into eight parts to generate eight small spaces (nodes) (S3602). Next, the three-dimensional data encoding device generates an occupancy code for the target node depending on whether or not the point cloud is contained in each node (S3603).
[0682] Next, the three-dimensional data encoding device calculates (obtains) the occupancy information (adjacent occupancy pattern) of the target node's adjacent reference nodes from the occupancy code of the target node's parent node (S3604).
[0683] Next, the three-dimensional data coding device converts the adjacent occupancy pattern into an index in the coding table (S3605). Then, the three-dimensional data coding device selects a coding table to be used for entropy coding based on the index (S3606).
[0684] Next, the three-dimensional data coding device entropy-codes the occupancy code of the target node using the selected coding table (S3607).
[0685] Furthermore, the three-dimensional data encoding device divides each node into eight parts and encodes the occupancy code of each node, repeating this process until the nodes can no longer be divided (S3608). In other words, the processes from steps S3602 to S3607 are repeated recursively.
[0686] Figure 113 is a flowchart of the three-dimensional data decoding method in a three-dimensional data decoding device. First, the three-dimensional data decoding device defines the space to be decoded (target node) using the bitstream header information (S3611). Next, the three-dimensional data decoding device divides the target node into eight parts to generate eight small spaces (nodes) (S3612). Next, the three-dimensional data decoding device calculates (obtains) the occupancy information (adjacent occupancy pattern) of the adjacent reference nodes of the target node from the occupancy code of the parent node of the target node (S3613).
[0687] Next, the three-dimensional data decoder converts the adjacent occupancy pattern into an index in the coding table (S3614). Then, the three-dimensional data decoder selects ...
Claims
1. Generate a tree structure of multiple three-dimensional points contained in three-dimensional data, Generate parameters that indicate the nodes that can be referenced, From among multiple adjacent occupancy patterns, the adjacent occupancy pattern is determined based on the occupancy status of the adjacent nodes of the target node. From among multiple groups, a group corresponding to the determined adjacent occupancy pattern is selected. The target node is encoded using the information of the determined group, Each of the aforementioned groups corresponds to one or more adjacent occupancy patterns, Depending on the value indicated by the parameter, the number of groups that can be determined will differ. The aforementioned plurality of groups include the first group and the second group, Each adjacent occupancy pattern corresponding to the first group represents a first number of occupancy nodes, Each adjacent occupancy pattern corresponding to the second group indicates a second number of occupancy nodes that is greater than the first number, The number of adjacent occupancy patterns corresponding to the first group is less than the number of adjacent occupancy patterns corresponding to the second group. Three-dimensional data encoding method.
2. Each adjacent occupancy pattern corresponding to the same group indicates the same number of occupying nodes. The three-dimensional data encoding method according to claim 1.
3. The number of adjacent occupancy patterns is 64. The three-dimensional data encoding method according to claim 1.
4. Each of the aforementioned multiple adjacent occupancy patterns indicates the occupancy status of the six adjacent nodes adjacent to the target node. The three-dimensional data encoding method according to claim 1.
5. The six neighboring nodes adjacent to the target node include three neighboring nodes whose parent node is different from that of the target node. The three-dimensional data encoding method according to claim 4.
6. If the parameter shows a first value, the number of the multiple groups is N; if the parameter shows a second value, the number of the multiple groups is M, where M is different from N. A three-dimensional data encoding method according to any one of claims 1 to 5.
7. Obtain the tree structure of multiple three-dimensional points contained in three-dimensional data. Obtain parameters that indicate the nodes that can be referenced, From among multiple adjacent occupancy patterns, the adjacent occupancy pattern is determined based on the occupancy status of the adjacent nodes of the target node. From among multiple groups, a group corresponding to the determined adjacent occupancy pattern is selected. Using the information of the determined group, the target node is decoded. Each of the aforementioned groups corresponds to one or more adjacent occupancy patterns, Depending on the value indicated by the parameter, the number of groups that can be determined will differ. The aforementioned plurality of groups include the first group and the second group, Each adjacent occupancy pattern corresponding to the first group represents a first number of occupancy nodes, Each adjacent occupancy pattern corresponding to the second group indicates a second number of occupancy nodes that is greater than the first number, The number of adjacent occupancy patterns corresponding to the first group is less than the number of adjacent occupancy patterns corresponding to the second group. Three-dimensional data decoding method.
8. Each adjacent occupancy pattern corresponding to the same group indicates the same number of occupying nodes. The method for decoding three-dimensional data according to claim 7.
9. The number of adjacent occupancy patterns is 64. The method for decoding three-dimensional data according to claim 7.
10. Each of the aforementioned multiple adjacent occupancy patterns indicates the occupancy status of the six adjacent nodes adjacent to the target node. The method for decoding three-dimensional data according to claim 7.
11. The six neighboring nodes adjacent to the target node include three neighboring nodes whose parent node is different from that of the target node. The method for decoding three-dimensional data according to claim 10.
12. If the parameter shows a first value, the number of the multiple groups is N; if the parameter shows a second value, the number of the multiple groups is M, where M is different from N. A method for decoding three-dimensional data according to any one of claims 7 to 11.
13. Processor and Equipped with memory, The processor uses the memory to: Generate a tree structure of multiple three-dimensional points contained in three-dimensional data, Generate parameters that indicate the nodes that can be referenced, From among multiple adjacent occupancy patterns, the adjacent occupancy pattern is determined based on the occupancy status of the adjacent nodes of the target node. From among multiple groups, a group corresponding to the determined adjacent occupancy pattern is selected. The target node is encoded using the information of the determined group, Each of the aforementioned groups corresponds to one or more adjacent occupancy patterns, Depending on the value indicated by the parameter, the number of groups that can be determined will differ. The aforementioned plurality of groups include the first group and the second group, Each adjacent occupancy pattern corresponding to the first group represents a first number of occupancy nodes, Each adjacent occupancy pattern corresponding to the second group indicates a second number of occupancy nodes that is greater than the first number, The number of adjacent occupancy patterns corresponding to the first group is less than the number of adjacent occupancy patterns corresponding to the second group. Three-dimensional data encoding device.
14. Processor and Equipped with memory, The processor uses the memory to: Obtain the tree structure of multiple three-dimensional points contained in three-dimensional data. Obtain parameters that indicate the nodes that can be referenced, From among multiple adjacent occupancy patterns, the adjacent occupancy pattern is determined based on the occupancy status of the adjacent nodes of the target node. From among multiple groups, a group corresponding to the determined adjacent occupancy pattern is selected. Using the information of the determined group, the target node is decoded. Each of the aforementioned groups corresponds to one or more adjacent occupancy patterns, Depending on the value indicated by the parameter, the number of groups that can be determined will differ. The aforementioned plurality of groups include the first group and the second group, Each adjacent occupancy pattern corresponding to the first group represents a first number of occupancy nodes, Each adjacent occupancy pattern corresponding to the second group indicates a second number of occupancy nodes that is greater than the first number, The number of adjacent occupancy patterns corresponding to the first group is less than the number of adjacent occupancy patterns corresponding to the second group. Three-dimensional data decoding device.
Citation Information
Patent Citations
Method and device for binary entropy coding of point clouds
JP2021521679A
JPP7774094B
Map display device
WO2014020663A1
Methods and devices for binary entropy coding of point clouds
WO2019195920A1