Three-dimensional data encoding method, three-dimensional data decoding method, three-dimensional data encoding device, and three-dimensional data decoding device
Patent Information
- Application Number
- JP2025010021
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2019-04-25
- Filing Date
- 2025-01-23
- Publication Date
- 2026-09-30
- Estimated Expiration
- 2040-04-24
AI Technical Summary
【0015】 本開示は、適切に符号化を行える符号化方法、復号方法、符号化装置、又は復号装置を提供できる。
Smart Images

Figure 0007906065000006 
Figure 0007906065000007 
Figure 0007906065000008
Abstract
Description
[Technical Field]
[0001] This disclosure relates to an encoding method, a decoding method, an encoding device, and a decoding device. [Background technology]
[0002] In the future, devices and services utilizing three-dimensional data are expected to become widespread in a wide range of fields, including computer vision for autonomous operation of automobiles or robots, map information, monitoring, infrastructure inspection, and video distribution. Three-dimensional data can be acquired in various ways, such as using distance sensors like rangefinders, stereo cameras, or combinations of multiple monocular cameras.
[0003] One method of representing three-dimensional data is called a point cloud, which represents the shape of a three-dimensional structure using a cloud of points in three-dimensional space. In a point cloud, the position and color of the points are stored. Point clouds are expected to become the mainstream method of representing three-dimensional data, but point clouds are extremely large in size. Therefore, in the storage or transmission of three-dimensional data, data compression through encoding is essential, just as with two-dimensional moving images (for example, MPEG-4 AVC or HEVC, which are standardized by MPEG).
[0004] Furthermore, point cloud compression is partially supported by publicly available libraries (such as the Point Cloud Library) that handle point cloud-related processing.
[0005] Furthermore, there is a known technique for searching for and displaying facilities located around a vehicle using three-dimensional map data (see, for example, Patent Document 1). [Prior art documents] [Patent Documents]
[0006] [Patent Document 1] International Publication No. 2014 / 020663 [Overview of the Initiative] [Problems that the invention aims to solve]
[0007] In the encoding process of three-dimensional data, it is desirable to be able to encode the data appropriately.
[0008] This disclosure aims to provide an encoding method, a decoding method, an encoding device, or a decoding device that can perform encoding appropriately. [Means for solving the problem]
[0009] An encoding method according to one aspect of the present disclosure is an encoding method performed by an encoding device, which generates a plurality of coefficient values belonging to any of a plurality of hierarchies, performs quantization on the plurality of coefficient values, generates a plurality of quantized values, generates information indicating the total number of hierarchies to which hierarchical parameters are notified, and generates a bitstream including the hierarchical parameters, and in the quantization, in the quantization of coefficient values belonging to a hierarchy to which the hierarchical parameters are notified, derives a quantization step using the hierarchical parameters of the last hierarchy to which the hierarchical parameters are notified, in the quantization of coefficient values belonging to a hierarchy to which the hierarchical parameters are not notified, derives a quantization step using the hierarchical parameters of the last hierarchy to which the hierarchical parameters are notified.
[0010] A decoding method according to one aspect of the present disclosure is a decoding method performed by a decoding device, which obtains information indicating the total number of hierarchies for which hierarchical parameters are notified, and the hierarchical parameters, performs inverse quantization on a plurality of quantization values to generate a plurality of coefficient values, and in the inverse quantization, in the inverse quantization of coefficient values belonging to a hierarchy for which the hierarchical parameters are notified, derives a quantization step using the hierarchical parameters, and in the inverse quantization of coefficient values belonging to a hierarchy for which the hierarchical parameters are not notified, derives a quantization step using the hierarchical parameters of the last hierarchy for which the hierarchical parameters are notified.
[0011] A three-dimensional data encoding method according to one aspect of the present disclosure calculates a plurality of coefficient values from a plurality of attribute information of a plurality of three-dimensional points included in point cloud data, generates a plurality of quantized values by quantizing each of the plurality of coefficient values, generates a bitstream containing the plurality of quantized values, the plurality of coefficient values belong to one of a plurality of hierarchies, each of a predetermined number of hierarchies is assigned a quantization parameter for that hierarchy, and in the quantization, each of the plurality of coefficient values is (i) quantized using the quantization parameter if it is assigned to the hierarchy to which the coefficient value belongs, and (ii) quantized using the quantization parameter assigned to one of the predetermined number of hierarchies if it is not assigned to the hierarchy to which the coefficient value belongs.
[0012] Furthermore, a three-dimensional data encoding method according to another aspect of the present disclosure calculates a plurality of coefficient values from a plurality of attribute information of a plurality of three-dimensional points included in point cloud data, generates a plurality of quantized values by quantizing each of the plurality of coefficient values, generates a bitstream containing the plurality of quantized values, and each of the plurality of coefficient values belongs to one of a plurality of groups by being classified according to the three-dimensional space to which the three-dimensional point containing the attribute information on which the coefficient value was calculated belongs among a plurality of three-dimensional spaces, and in the quantization, each of the plurality of coefficient values is quantized using a quantization parameter for the group to which the coefficient value belongs.
[0013] A three-dimensional data decoding method according to one aspect of the present disclosure generates a plurality of coefficient values by de-quantizing each of a plurality of quantization values contained in a bitstream, calculates a plurality of attribute information of a plurality of three-dimensional points contained in point cloud data from the plurality of coefficient values, the plurality of quantization values belong to one of a plurality of hierarchies, each of a predetermined number of hierarchies is assigned a quantization parameter for that hierarchy, and in the de-quantization, each of the plurality of quantization values is (i) de-quantized using the quantization parameter if it is assigned to the hierarchy to which the quantization value belongs, and (ii) de-quantized using the quantization parameter assigned to one of the predetermined number of hierarchies if it is not assigned to the hierarchy to which the quantization value belongs.
[0014] Furthermore, a three-dimensional data decoding method according to another aspect of the present disclosure generates a plurality of coefficient values by inverse quantization of each of a plurality of quantization values contained in a bitstream, calculates a plurality of attribute information of a plurality of three-dimensional points contained in point cloud data from the plurality of coefficient values, and the plurality of quantization values belong to one of a plurality of groups by classifying them according to the three-dimensional space to which the three-dimensional point containing the attribute information on which the calculation of the quantization value was based belongs among a plurality of three-dimensional spaces, and each of a predetermined number of hierarchies in the plurality of hierarchies is assigned a quantization parameter for that hierarchy, and in the inverse quantization, each of the plurality of quantization values is inverse quantized using the quantization parameter for the hierarchy to which the quantization value belongs. [Effects of the Invention]
[0015] This disclosure can provide an encoding method, a decoding method, an encoding device, or a decoding device that can perform encoding appropriately. [Brief explanation of the drawing]
[0016] [Figure 1] Figure 1 is a diagram showing the structure of encoded three-dimensional data according to Embodiment 1. [Figure 2]FIG. 2 is a diagram showing an example of a prediction structure between SPCs belonging to the lowermost layer of a GOS according to the first embodiment. [Figure 3] FIG. 3 is a diagram showing an example of a prediction structure between layers according to the first embodiment. [Figure 4] FIG. 4 is a diagram showing an example of a coding order of a GOS according to the first embodiment. [Figure 5] FIG. 5 is a diagram showing an example of a coding order of a GOS according to the first embodiment. [Figure 6] FIG. 6 is a diagram showing an example of meta information according to the first embodiment. [Figure 7] FIG. 7 is a schematic diagram showing a state of transmission and reception of three-dimensional data between vehicles according to the second embodiment. [Figure 8] FIG. 8 is a diagram showing an example of three-dimensional data transmitted between vehicles according to the second embodiment. [Figure 9] FIG. 9 is a diagram for explaining transmission processing of three-dimensional data according to the third embodiment. [Figure 10] FIG. 10 is a diagram showing a configuration of a system according to the fourth embodiment. [Figure 11] FIG. 11 is a block diagram of a client device according to the fourth embodiment. [Figure 12] FIG. 12 is a block diagram of a server according to the fourth embodiment. [Figure 13] FIG. 13 is a flowchart of three-dimensional data creation processing by a client device according to the fourth embodiment. [Figure 14] FIG. 14 is a flowchart of sensor information transmission processing by a client device according to the fourth embodiment. [Figure 15] FIG. 15 is a flowchart of three-dimensional data creation processing by a server according to the fourth embodiment. [Figure 16] FIG. 16 is a flowchart of three-dimensional map transmission processing by a server according to the fourth embodiment. [Figure 17] FIG. 17 is a diagram showing a configuration of a modified example of the system according to the fourth embodiment. [Figure 18]Figure 18 shows the configuration of the server and client device according to Embodiment 4. [Figure 19] Figure 19 is a diagram showing the configuration of the server and client device according to Embodiment 5. [Figure 20] Figure 20 is a flowchart of the processing performed by the client device according to Embodiment 5. [Figure 21] Figure 21 is a diagram showing the configuration of the sensor information collection system according to Embodiment 5. [Figure 22] Figure 22 shows an example of a volume according to Embodiment 6. [Figure 23] Figure 23 shows an example of an octree representation of a volume according to Embodiment 6. [Figure 24] Figure 24 shows an example of a bit sequence of a volume according to Embodiment 6. [Figure 25] Figure 25 shows an example of an octree representation of a volume according to Embodiment 6. [Figure 26] Figure 26 shows an example of a volume according to Embodiment 6. [Figure 27] Figure 27 shows an example of a three-dimensional point according to Embodiment 7. [Figure 28] Figure 28 shows an example of LoD settings according to Embodiment 7. [Figure 29] Figure 29 shows an example of a threshold used to set the LoD according to Embodiment 7. [Figure 30] Figure 30 shows an example of attribute information used for the predicted value according to Embodiment 7. [Figure 31] Figure 31 shows an example of an exponential Golomb code according to Embodiment 7. [Figure 32] Figure 32 is a diagram showing the processing of exponential Golomb codes according to Embodiment 7. [Figure 33] Figure 33 shows an example of the attribute header syntax according to Embodiment 7. [Figure 34] Figure 34 shows an example of attribute data syntax according to Embodiment 7. [Figure 35] Figure 35 is a flowchart of the three-dimensional data encoding process according to Embodiment 7. [Figure 36] Figure 36 is a flowchart of the attribute information encoding process according to Embodiment 7. [Figure 37] Figure 37 is a diagram showing the processing of exponential Golomb codes according to Embodiment 7. [Figure 38] Figure 38 is a diagram showing an example of a reverse lookup table that illustrates the relationship between the remaining reference numerals and their values according to Embodiment 7. [Figure 39] Figure 39 is a flowchart of the three-dimensional data decoding process according to Embodiment 7. [Figure 40] Figure 40 is a flowchart of the attribute information decoding process according to Embodiment 7. [Figure 41] Figure 41 is a block diagram of a three-dimensional data encoding device according to Embodiment 7. [Figure 42] Figure 42 is a block diagram of a three-dimensional data encoding device according to Embodiment 7. [Figure 43] Figure 43 is a diagram illustrating the encoding of attribute information using RAHT according to Embodiment 8. [Figure 44] Figure 44 shows an example of setting the quantization scale for each hierarchy according to Embodiment 8. [Figure 45] Figure 45 shows examples of the first and second code sequences according to Embodiment 8. [Figure 46] Figure 46 shows an example of a trunket unary code according to Embodiment 8. [Figure 47] Figure 47 is a diagram illustrating the inverse Haar transform according to Embodiment 8. [Figure 48] Figure 48 shows an example of attribute information syntax according to Embodiment 8. [Figure 49] Figure 49 shows an example of coding coefficients and ZeroCnt according to Embodiment 8. [Figure 50]Figure 50 is a flowchart of the three-dimensional data encoding process according to Embodiment 8. [Figure 51] Figure 51 is a flowchart of the attribute information encoding process according to Embodiment 8. [Figure 52] Figure 52 is a flowchart of the coding coefficient coding process according to Embodiment 8. [Figure 53] Figure 53 is a flowchart of the three-dimensional data decoding process according to Embodiment 8. [Figure 54] Figure 54 is a flowchart of the attribute information decoding process according to Embodiment 8. [Figure 55] Figure 55 is a flowchart of the coding coefficient decoding process according to Embodiment 8. [Figure 56] Figure 56 is a block diagram of the attribute information encoding unit according to Embodiment 8. [Figure 57] Figure 57 is a block diagram of the attribute information decoding unit according to Embodiment 8. [Figure 58] Figure 58 shows examples of the first and second code sequences according to a modified example of Embodiment 8. [Figure 59] Figure 59 shows an example of attribute information syntax related to a modified example of Embodiment 8. [Figure 60] Figure 60 shows examples of coding coefficients, ZeroCnt and TotalZeroCnt, related to a modified example of Embodiment 8. [Figure 61] Figure 61 is a flowchart of the coding coefficient coding process according to a modified example of Embodiment 8. [Figure 62] Figure 62 is a flowchart of the coding coefficient decoding process according to a modified example of Embodiment 8. [Figure 63] Figure 63 shows an example of attribute information syntax related to a modified example of Embodiment 8. [Figure 64] Figure 64 is a block diagram showing the configuration of a three-dimensional data encoding device according to Embodiment 9. [Figure 65] Figure 65 is a block diagram showing the configuration of a three-dimensional data decoding device according to Embodiment 9. [Figure 66] Figure 66 shows an example of LoD settings according to Embodiment 9. [Figure 67] Figure 67 is a diagram showing an example of the hierarchical structure of RAHT according to Embodiment 9. [Figure 68] Figure 68 is a block diagram of a three-dimensional data encoding device according to Embodiment 9. [Figure 69] Figure 69 is a block diagram of the divided section according to Embodiment 9. [Figure 70] Figure 70 is a block diagram of the attribute information encoding unit according to Embodiment 9. [Figure 71] Figure 71 is a block diagram of a three-dimensional data decoding device according to Embodiment 9. [Figure 72] Figure 72 is a block diagram of the attribute information decoding unit according to Embodiment 9. [Figure 73] Figure 73 shows an example of setting quantization parameters in tile and slice division according to Embodiment 9. [Figure 74] Figure 74 shows an example of setting quantization parameters according to Embodiment 9. [Figure 75] Figure 75 shows an example of setting quantization parameters according to Embodiment 9. [Figure 76] Figure 76 shows an example of the syntax of the attribute information header according to Embodiment 9. [Figure 77] Figure 77 shows an example of the syntax of the attribute information header according to Embodiment 9. [Figure 78] Figure 78 shows an example of setting quantization parameters according to Embodiment 9. [Figure 79] Figure 79 shows an example of the syntax of the attribute information header according to Embodiment 9. [Figure 80] Figure 80 shows an example of the syntax of the attribute information header according to Embodiment 9. [Figure 81] Figure 81 is a flowchart of the three-dimensional data encoding process according to Embodiment 9. [Figure 82]Figure 82 is a flowchart of the attribute information encoding process according to Embodiment 9. [Figure 83] Figure 83 is a flowchart of the ΔQP determination process according to Embodiment 9. [Figure 84] Figure 84 is a flowchart of the three-dimensional data decoding process according to Embodiment 9. [Figure 85] Figure 85 is a flowchart of the attribute information decoding process according to Embodiment 9. [Figure 86] Figure 86 is a block diagram of the attribute information encoding unit according to Embodiment 9. [Figure 87] Figure 87 is a block diagram of the attribute information decoding unit according to Embodiment 9. [Figure 88] Figure 88 shows an example of setting quantization parameters according to Embodiment 9. [Figure 89] Figure 89 shows an example of the syntax of the attribute information header according to Embodiment 9. [Figure 90] Figure 90 shows an example of the syntax of the attribute information header according to Embodiment 9. [Figure 91] Figure 91 is a flowchart of the three-dimensional data encoding process according to Embodiment 9. [Figure 92] Figure 92 is a flowchart of the attribute information encoding process according to Embodiment 9. [Figure 93] Figure 93 is a flowchart of the three-dimensional data decoding process according to Embodiment 9. [Figure 94] Figure 94 is a flowchart of the attribute information decoding process according to Embodiment 9. [Figure 95] Figure 95 is a block diagram of the attribute information encoding unit according to Embodiment 9. [Figure 96] Figure 96 is a block diagram of the attribute information decoding unit according to Embodiment 9. [Figure 97] Figure 97 shows an example of the syntax of the attribute information header according to Embodiment 9. [Figure 98]Figure 98 is a graph showing the relationship between the bitrate and time of bitstream encoding according to Embodiment 10. [Figure 99] Figure 99 shows the hierarchical structure of the three-dimensional point cloud according to Embodiment 10, and the number of three-dimensional points belonging to each hierarchy. [Figure 100] Figure 100 shows a first example of classifying a single-level three-dimensional point cloud according to Embodiment 10 into sub-levels based on a specified number of three-dimensional points. [Figure 101] Figure 101 shows a second example of classifying a single-level three-dimensional point cloud according to Embodiment 10 into sub-levels based on a fixed number of three-dimensional points. [Figure 102] Figure 102 shows an example of the syntax of the attribute information header in the second example according to Embodiment 10. [Figure 103] Figure 103 shows another example of the attribute information syntax in the second example according to Embodiment 10. [Figure 104] Figure 104 shows a third example in which a single-layer three-dimensional point cloud according to Embodiment 10 is classified into a number of sub-layers different from the planned number. [Figure 105] Figure 105 shows an example of the syntax of the attribute information header in the third example according to Embodiment 10. [Figure 106] Figure 106 shows another example of the attribute information header syntax in the third example according to Embodiment 10. [Figure 107] Figure 107 shows a fourth example of classifying a single-level three-dimensional point cloud according to Embodiment 10 into sub-levels based on a specified number of three-dimensional points as a percentage. [Figure 108] Figure 108 shows an example of the syntax of the attribute information header in the fourth example according to Embodiment 10. [Figure 109] Figure 109 shows a fifth example of classifying a single-level three-dimensional point cloud into sub-levels using a Morton index according to Embodiment 10. [Figure 110] Figure 110 shows an example of the syntax of the attribute information header in the fifth example according to Embodiment 10. [Figure 111] Figure 111 shows a sixth example of classifying a single-level three-dimensional point cloud according to Embodiment 10 into sub-levels using a Morton index. [Figure 112] Figure 112 shows a sixth example of classifying a single-level three-dimensional point cloud according to Embodiment 10 into sub-levels using a Morton index. [Figure 113] Figure 113 shows a seventh example of classifying a one-tiered three-dimensional point cloud according to Embodiment 10 into sub-tiers using residuals or Delta values. [Figure 114] Figure 114 shows the arrangement of three-dimensional points when they are arranged in a two-dimensional Morton order according to Embodiment 10. [Figure 115] Figure 115 shows an example of the attribute information header syntax in the seventh example according to Embodiment 10. [Figure 116] Figure 116 shows an example of the residual bitstream syntax according to Embodiment 10. [Figure 117] Figure 117 shows the formula for calculating the encoding cost according to Embodiment 10. [Figure 118] Figure 118 is a graph showing the relationship between BPP (bits per point) and time according to Embodiment 10. [Figure 119] Figure 119 shows that the QP value applied to the encoding of attribute information according to Embodiment 10 is set for each sub-level. [Figure 120] Figure 120 shows an eighth example of classifying a three-dimensional point cloud according to Embodiment 10 into subhierarchies using Morton codes. [Figure 121] Figure 121 shows an example of the attribute information header syntax in the eighth example according to Embodiment 10. [Figure 122] Figure 122 is a flowchart of the three-dimensional data encoding process according to Embodiment 10. [Figure 123] Figure 123 is a flowchart of the three-dimensional data decoding process according to Embodiment 10. [Modes for carrying out the invention]
[0017] A three-dimensional data encoding method according to one aspect of the present disclosure calculates a plurality of coefficient values from a plurality of attribute information of a plurality of three-dimensional points included in point cloud data, generates a plurality of quantized values by quantizing each of the plurality of coefficient values, generates a bitstream containing the plurality of quantized values, the plurality of coefficient values belong to one of a plurality of hierarchies, each of a predetermined number of hierarchies is assigned a quantization parameter for that hierarchy, and in the quantization, each of the plurality of coefficient values is (i) quantized using the quantization parameter if it is assigned to the hierarchy to which the coefficient value belongs, and (ii) quantized using the quantization parameter assigned to one of the predetermined number of hierarchies if it is not assigned to the hierarchy to which the coefficient value belongs.
[0018] According to this, the three-dimensional data encoding method can switch quantization parameters for each hierarchical level, thus enabling proper encoding.
[0019] For example, the first hierarchy may be the last hierarchy among the predetermined number of hierarchy levels.
[0020] For example, in the quantization described above, if the number of the multiple levels is less than the predetermined number, the quantization parameters assigned to the predetermined number of levels that do not correspond to the multiple levels do not need to be used.
[0021] For example, the bitstream may include first information indicating a reference quantization parameter and a plurality of second pieces of information for calculating a plurality of quantization parameters for the plurality of hierarchies from the reference quantization parameter.
[0022] Furthermore, this three-dimensional data encoding method can improve encoding efficiency by encoding first information indicating a reference quantization parameter and multiple second pieces of information for calculating multiple quantization parameters from the reference quantization parameter.
[0023] Furthermore, a three-dimensional data encoding method according to another aspect of the present disclosure calculates a plurality of coefficient values from a plurality of attribute information of a plurality of three-dimensional points included in point cloud data, generates a plurality of quantized values by quantizing each of the plurality of coefficient values, generates a bitstream containing the plurality of quantized values, and each of the plurality of coefficient values belongs to one of a plurality of groups by being classified according to the three-dimensional space to which the three-dimensional point containing the attribute information on which the coefficient value was calculated belongs among a plurality of three-dimensional spaces, and in the quantization, each of the plurality of coefficient values is quantized using a quantization parameter for the group to which the coefficient value belongs.
[0024] According to this, the three-dimensional data encoding method can switch quantization parameters for each hierarchical level, thus enabling proper encoding.
[0025] A three-dimensional data decoding method according to one aspect of the present disclosure generates a plurality of coefficient values by de-quantizing each of a plurality of quantization values contained in a bitstream, calculates a plurality of attribute information of a plurality of three-dimensional points contained in point cloud data from the plurality of coefficient values, the plurality of quantization values belong to one of a plurality of hierarchies, each of a predetermined number of hierarchies is assigned a quantization parameter for that hierarchy, and in the de-quantization, each of the plurality of quantization values is (i) de-quantized using the quantization parameter if it is assigned to the hierarchy to which the quantization value belongs, and (ii) de-quantized using the quantization parameter assigned to one of the predetermined number of hierarchies if it is not assigned to the hierarchy to which the quantization value belongs.
[0026] According to this, the three-dimensional data decoding method can switch quantization parameters for each hierarchical level, thus enabling proper decoding.
[0027] For example, the first hierarchy may be the last hierarchy among the predetermined number of hierarchy levels.
[0028] For example, in the inverse quantization described above, if the number of the multiple levels is less than the predetermined number, the quantization parameters assigned to the predetermined number of levels that do not correspond to the multiple levels do not need to be used.
[0029] For example, the bitstream may include first information indicating a reference quantization parameter and a plurality of second pieces of information for calculating a plurality of quantization parameters for the plurality of hierarchies from the reference quantization parameter.
[0030] Furthermore, this three-dimensional data decoding method can appropriately decode a bitstream with improved encoding efficiency by using first information indicating a reference quantization parameter and multiple second pieces of information for calculating multiple quantization parameters from the reference quantization parameter.
[0031] Furthermore, a three-dimensional data decoding method according to another aspect of the present disclosure generates a plurality of coefficient values by inverse quantization of each of a plurality of quantization values contained in a bitstream, calculates a plurality of attribute information of a plurality of three-dimensional points contained in point cloud data from the plurality of coefficient values, and the plurality of quantization values belong to one of a plurality of groups by classifying them according to the three-dimensional space to which the three-dimensional point containing the attribute information on which the calculation of the quantization value was based belongs among a plurality of three-dimensional spaces, and each of a predetermined number of hierarchies in the plurality of hierarchies is assigned a quantization parameter for that hierarchy, and in the inverse quantization, each of the plurality of quantization values is inverse quantized using the quantization parameter for the hierarchy to which the quantization value belongs.
[0032] According to this, the three-dimensional data decoding method can switch quantization parameters for each hierarchical level, thus enabling proper decoding.
[0033] These comprehensive or specific embodiments may be implemented as a system, method, integrated circuit, computer program, or recording medium such as a computer-readable CD-ROM, or as any combination of a system, method, integrated circuit, computer program, and recording medium.
[0034] The embodiments will be described in detail below with reference to the drawings. Note that the embodiments described below are all specific examples of this disclosure. The numerical values, shapes, materials, components, arrangement and connection configurations of components, steps, and the order of steps shown in the following embodiments are examples only and are not intended to limit this disclosure. Furthermore, components in the following embodiments that are not described in an independent claim will be described as optional components.
[0035] (Embodiment 1) First, the data structure of the encoded three-dimensional data (hereinafter also referred to as encoded data) according to this embodiment will be described. Figure 1 is a diagram showing the configuration of the encoded three-dimensional data according to this embodiment.
[0036] In this embodiment, the three-dimensional space is divided into spaces (SPCs) corresponding to pictures in video encoding, and three-dimensional data is encoded using these spaces as units. The spaces are further divided into volumes (VLMs) corresponding to macroblocks in video encoding, and prediction and transformation are performed using the VLMs as units. Each volume contains multiple voxels (VXLs), which are the smallest units to which position coordinates are associated. Prediction, similar to prediction performed on two-dimensional images, involves referencing other processing units to generate predicted three-dimensional data similar to the processing unit being processed, and then encoding the difference between this predicted three-dimensional data and the processing unit being processed. Furthermore, this prediction includes not only spatial prediction that references other prediction units at the same time, but also temporal prediction that references prediction units at different times.
[0037] For example, a three-dimensional data encoding device (hereinafter also referred to as the encoding device) encodes a three-dimensional space represented by point cloud data, such as a point cloud, by encoding each point in the point cloud, or multiple points contained within a voxel, depending on the size of the voxel. Subdividing the voxel allows for a highly accurate representation of the three-dimensional shape of the point cloud, while increasing the voxel size allows for a rougher representation of the three-dimensional shape of the point cloud.
[0038] In the following explanation, we will use the example of a point cloud as the 3D data, but the 3D data is not limited to a point cloud; any format of 3D data is acceptable.
[0039] Alternatively, a hierarchical structure of voxels may be used. In this case, for the nth-order hierarchy, it may be indicated sequentially whether or not sample points exist in the (n-1)th-order hierarchy and below (the lower layers of the nth-order hierarchy). For example, when decoding only the nth-order hierarchy, if sample points exist in the (n-1)th-order hierarchy and below, the sample points can be assumed to be at the center of the voxel of the nth-order hierarchy and decoded accordingly.
[0040] Furthermore, the encoding device acquires point cloud data using distance sensors, stereo cameras, monocular cameras, gyroscopes, or inertial sensors.
[0041] Spaces, like video encodings, are classified into at least three predictive structures, including intra-spaces (I-SPCs) that can be decoded independently, predictive spaces (P-SPCs) that allow only unidirectional referencing, and bidirectional spaces (B-SPCs) that allow bidirectional referencing. Furthermore, spaces contain two types of time information: the decoding time and the display time.
[0042] Furthermore, as shown in Figure 1, there is a processing unit called GOS (Group of Space), which is a random access unit, that contains multiple spaces. In addition, there is a processing unit called WLD (World), which contains multiple GOS.
[0043] The spatial area occupied by a world is associated with an absolute location on Earth using GPS or latitude and longitude information. This location information is stored as metadata. This metadata may be included in the encoded data or transmitted separately from the encoded data.
[0044] Furthermore, within a GOS, all SPCs may be adjacent in three dimensions, or there may be SPCs that are not adjacent in three dimensions to other SPCs.
[0045] In the following, the processing of three-dimensional data contained in processing units such as GOS, SPC, or VLM, including encoding, decoding, or referencing, will also be simply referred to as encoding, decoding, or referencing the processing unit. Furthermore, the three-dimensional data contained in the processing unit includes, for example, at least one pair of spatial position such as three-dimensional coordinates and characteristic values such as color information.
[0046] Next, we will explain the prediction structure of SPCs in GOS. Multiple SPCs within the same GOS, or multiple VLMs within the same SPC, occupy different spaces from each other, but they have the same time information (decoded time and display time).
[0047] Furthermore, the SPC that is first in the decryption order within a GOS is the I-SPC. There are also two types of GOSs: closed GOS and open GOS. A closed GOS is one in which all SPCs within the GOS can be decrypted when decryption starts from the first I-SPC. In an open GOS, some SPCs whose displayed time is earlier than the first I-SPC refer to a different GOS, and decryption cannot be performed using only that GOS.
[0048] Furthermore, with encoded data such as map information, the WLD may be decoded in the reverse direction of the encoding order, and if there are dependencies between GOSs, reverse playback becomes difficult. Therefore, in such cases, a closed GOS is generally used.
[0049] Furthermore, GOS has a layered structure in the height direction, and encoding or decoding is performed sequentially from the SPC of the lower layer.
[0050] Figure 2 shows an example of the prediction structure between SPCs belonging to the lowest layer of GOS. Figure 3 shows an example of the prediction structure between layers.
[0051] One or more I-SPCs exist within a GOS. While objects such as people, animals, cars, bicycles, traffic lights, or landmark buildings exist in three-dimensional space, it is particularly effective to encode small objects as I-SPCs. For example, a three-dimensional data decoding device (hereinafter also referred to as the decoding device) decodes only the I-SPCs within the GOS when decoding a GOS with low processing load or at high speed.
[0052] Furthermore, the encoding device may switch the encoding interval or frequency of I-SPCs according to the density of objects in the WLD.
[0053] Furthermore, in the configuration shown in Figure 3, the encoding or decoding device encodes or decodes multiple layers sequentially from the bottom layer (Layer 1). This allows for prioritizing data near the ground, which contains more information, for applications such as autonomous vehicles.
[0054] Furthermore, in the case of encoded data used in drones and the like, encoding or decoding may be done sequentially within the GOS, starting from the SPC layer at the top in the height direction.
[0055] Furthermore, the encoding or decoding device may encode or decode multiple layers so that the decoding device can grasp the GOS roughly and gradually increase the resolution. For example, the encoding or decoding device may encode or decode layers 3, 8, 1, 9, and so on.
[0056] Next, we will explain how to handle static and dynamic objects.
[0057] In three-dimensional space, there are static objects or scenes such as buildings or roads (hereinafter collectively referred to as static objects) and dynamic objects such as cars or people (hereinafter referred to as dynamic objects). Object detection is performed separately, for example, by extracting feature points from point cloud data or camera images such as stereo cameras. Here, we will explain an example of an encoding method for dynamic objects.
[0058] The first method is to encode static and dynamic objects without distinguishing between them. The second method is to distinguish between static and dynamic objects using identification information.
[0059] For example, GOS is used as the identification unit. In this case, GOS containing SPCs that constitute static objects and GOS containing SPCs that constitute dynamic objects are distinguished by identification information stored within the encoded data or separately from the encoded data.
[0060] Alternatively, an SPC may be used as the identification unit. In this case, an SPC containing a VLM that constitutes a static object and an SPC containing a VLM that constitutes a dynamic object are distinguished by the above identification information.
[0061] Alternatively, VLM or VXL may be used as the identification unit. In this case, VLM or VXL containing static objects and VLM or VXL containing dynamic objects are distinguished by the above identification information.
[0062] Furthermore, the encoding device may encode dynamic objects as one or more VLMs or SPCs, and encode the VLM or SPC containing static objects and the SPC containing dynamic objects as different GOSs. Also, if the size of the GOS is variable depending on the size of the dynamic objects, the encoding device stores the size of the GOS separately as metadata.
[0063] Furthermore, the encoding device may encode static objects and dynamic objects independently of each other and superimpose dynamic objects onto a world composed of static objects. In this case, a dynamic object is composed of one or more SPCs, and each SPC is associated with one or more SPCs that constitute the static object on which it is superimposed. Note that dynamic objects may be represented by one or more VLMs or VXLs instead of SPCs.
[0064] Furthermore, the encoding device may encode static objects and dynamic objects as separate streams.
[0065] Furthermore, the encoding device may generate a GOS containing one or more SPCs that constitute a dynamic object. In addition, the encoding device may set the GOS containing the dynamic object (GOS_M) and the GOS of the static object corresponding to the spatial region of GOS_M to be the same size (occupy the same spatial region). This allows superposition processing to be performed on a GOS-by-GOS basis.
[0066] The P-SPC or B-SPC that constitute a dynamic object may reference SPCs contained in different encoded GOS. In cases where the position of a dynamic object changes over time and the same dynamic object is encoded as a GOS at different times, cross-GOS references are effective from a compression standpoint.
[0067] Furthermore, the first and second methods described above may be switched depending on the intended use of the encoded data. For example, when using encoded three-dimensional data as a map, it is desirable to be able to separate dynamic objects, so the encoding device uses the second method. On the other hand, when encoding three-dimensional data of an event such as a concert or sporting event, if there is no need to separate dynamic objects, the encoding device uses the first method.
[0068] Furthermore, the decoding time and display time of GOS or SPC can be stored within the encoded data or as metadata. The time information for static objects may also be identical. In this case, the actual decoding time and display time may be determined by the decoding device. Alternatively, different values may be assigned to each GOS or SPC as the decoding time, while the same value may be assigned to all as the display time. Furthermore, a decoder model may be introduced, such as the HEVC HRD (Hypothetical Reference Decoder) in video encoding, which guarantees that decoding can be performed without failure if the decoder has a buffer of a predetermined size and reads the bitstream at a predetermined bitrate according to the decoding time.
[0069] Next, we will explain the arrangement of GOS within the world. The coordinates of the three-dimensional space in the world are represented by three mutually orthogonal coordinate axes (x-axis, y-axis, and z-axis). By establishing a predetermined rule for the coding order of GOS, coding can be performed so that spatially adjacent GOS are continuous within the coded data. For example, in the example shown in Figure 4, GOS in the xz plane are coded continuously. The value of the y-axis is updated after coding all GOS in a given xz plane is completed. That is, as coding progresses, the world expands in the y-axis direction. Also, the index numbers of the GOS are set in the coding order.
[0070] Here, the world's three-dimensional space is mapped one-to-one with geographical absolute coordinates such as GPS, latitude, and longitude. Alternatively, the three-dimensional space may be represented by relative positions from a pre-defined reference position. The directions of the x, y, and z axes of the three-dimensional space are represented as direction vectors determined based on latitude and longitude, and these direction vectors are stored as metadata along with encoded data.
[0071] Furthermore, the size of the GOS is fixed, and the encoding device stores this size as metadata. Alternatively, the size of the GOS may be switched depending on, for example, whether it is an urban area or not, or whether it is indoors or outdoors. In other words, the size of the GOS may be switched depending on the quantity or nature of objects that have informational value. Or, the encoding device may adaptively switch the size of the GOS or the spacing of I-SPCs within the GOS depending on the density of objects within the same world. For example, the encoding device may reduce the size of the GOS and shorten the spacing of I-SPCs within the GOS as the density of objects increases.
[0072] In the example in Figure 5, the GOS regions from the 3rd to the 10th are subdivided to enable fine-grained random access due to the high object density. Note that GOS regions 7 through 10 are located behind GOS regions 3 through 6, respectively.
[0073] (Embodiment 2) This embodiment describes a method for transmitting and receiving three-dimensional data between vehicles.
[0074] Figure 7 is a schematic diagram showing the transmission and reception of three-dimensional data 607 between the vehicle 600 and the surrounding vehicle 601.
[0075] When acquiring three-dimensional data using sensors mounted on the vehicle 600 (such as distance sensors like rangefinders, stereo cameras, or combinations of multiple monocular cameras), an occlusion region (hereinafter referred to as the occlusion region 604) occurs where three-dimensional data cannot be created, even though it is within the sensor detection range 602 of the vehicle 600, due to obstacles such as surrounding vehicles 601. Furthermore, while increasing the space in which three-dimensional data can be acquired improves the accuracy of autonomous operation, the sensor detection range of the vehicle 600 alone is finite.
[0076] The sensor detection range 602 of the vehicle 600 includes the area 603 where three-dimensional data can be acquired and the occlusion area 604. The area in which the vehicle 600 wants to acquire three-dimensional data includes the sensor detection range 602 of the vehicle 600 and the other areas. In addition, the sensor detection range 605 of the surrounding vehicle 601 includes the occlusion area 604 and the area 606 that is not included in the sensor detection range 602 of the vehicle 600.
[0077] The surrounding vehicle 601 transmits the information it detects to its own vehicle 600. By acquiring the information detected by the surrounding vehicle 601, such as the vehicle in front, the own vehicle 600 can acquire three-dimensional data 607 of the occlusion region 604 and the region 606 outside the sensor detection range 602 of the own vehicle 600. The own vehicle 600 uses the information acquired by the surrounding vehicle 601 to supplement the three-dimensional data of the occlusion region 604 and the region 606 outside the sensor detection range.
[0078] The uses of three-dimensional data in the autonomous operation of vehicles or robots include self-localization, detection of surrounding conditions, or both. For example, for self-localization, three-dimensional data generated by the vehicle 600 based on sensor information from the vehicle 600 is used. For detection of surrounding conditions, in addition to the three-dimensional data generated by the vehicle 600, three-dimensional data acquired from surrounding vehicles 601 is also used.
[0079] The surrounding vehicle 601 that transmits the three-dimensional data 607 to the vehicle 600 may be determined according to the state of the vehicle 600. For example, this surrounding vehicle 601 is the vehicle in front when the vehicle 600 is moving straight, the vehicle coming from behind when the vehicle 600 is turning right, and the vehicle behind when the vehicle 600 is reversing. Alternatively, the driver of the vehicle 600 may directly specify the surrounding vehicle 601 that transmits the three-dimensional data 607 to the vehicle 600.
[0080] Furthermore, the vehicle 600 may search for surrounding vehicles 601 that possess three-dimensional data for areas within the space where it wants to acquire three-dimensional data 607, but which it cannot acquire itself. Areas that cannot be acquired by the vehicle 600 include occlusion areas 604 or areas 606 outside the sensor detection range 602.
[0081] Furthermore, the vehicle 600 may identify the occlusion region 604 based on the sensor information of the vehicle 600. For example, the vehicle 600 may identify the occlusion region 604 as an area within the sensor detection range 602 of the vehicle 600 where three-dimensional data cannot be created.
[0082] The following describes an example of operation when the vehicle transmitting the three-dimensional data 607 is the vehicle in front. Figure 8 shows an example of the three-dimensional data transmitted in this case.
[0083] As shown in Figure 8, the three-dimensional data 607 transmitted from the preceding vehicle is, for example, a sparse world (SWLD) of a point cloud. In other words, the preceding vehicle creates three-dimensional data of a WLD (point cloud) from information detected by its own sensors, and then creates three-dimensional data of an SWLD (point cloud) by extracting data from the WLD that have a feature value above a threshold. The preceding vehicle then transmits the created SWLD three-dimensional data to its own vehicle 600.
[0084] Vehicle 600 receives the SWLD and merges it into the point cloud created by vehicle 600.
[0085] The transmitted SWLD contains information about its absolute coordinates (the SWLD's position in the coordinate system of the three-dimensional map). Vehicle 600 can perform the merge process by overwriting the point cloud it generates based on these absolute coordinates.
[0086] The SWLD transmitted from the surrounding vehicle 601 may be the SWLD of area 606 outside the sensor detection range 602 of the vehicle 600 and within the sensor detection range 605 of the surrounding vehicle 601, or the SWLD of the occlusion area 604 for the vehicle 600, or both. Alternatively, the transmitted SWLD may be the SWLD of the area used by the surrounding vehicle 601 for detecting the surrounding conditions.
[0087] Furthermore, surrounding vehicles 601 may change the density of the transmitted point cloud according to the communication time based on the speed difference between their own vehicle 600 and surrounding vehicles 601. For example, if the speed difference is large and the communication time is short, surrounding vehicles 601 may reduce the density (amount of data) of the point cloud by extracting three-dimensional points with large feature quantities from the SWLD.
[0088] Furthermore, detecting the surrounding environment involves determining the presence or absence of people, vehicles, and road construction equipment, identifying their type, and detecting their position, direction of movement, and speed of movement.
[0089] Furthermore, the vehicle 600 may acquire braking information of the surrounding vehicle 601 in addition to, or instead of, the three-dimensional data 607 generated by the surrounding vehicle 601. Here, braking information of the surrounding vehicle 601 refers to information indicating, for example, whether the accelerator or brake of the surrounding vehicle 601 was pressed, or to what extent.
[0090] Furthermore, in the point cloud generated by each vehicle, the three-dimensional space is subdivided into random access units to accommodate low-latency communication between vehicles. On the other hand, in the case of three-dimensional maps and other map data downloaded from a server, the three-dimensional space is divided into larger random access units compared to the case of vehicle-to-vehicle communication.
[0091] Data from areas prone to occlusion, such as the area in front of the preceding vehicle or the area behind the following vehicle, is divided into small random access units for low-latency data.
[0092] At high speeds, the importance of the front view increases, so each vehicle creates SWLDs in a narrowed field of view range using fine random access units.
[0093] If the SWLD created by the preceding vehicle for transmission includes an area where the vehicle 600 can acquire a point cloud, the preceding vehicle may reduce the amount of transmission by removing the point cloud in that area.
[0094] (Embodiment 3) This embodiment describes a method for transmitting three-dimensional data to a following vehicle. Figure 9 shows an example of the target space of the three-dimensional data to be transmitted to a following vehicle.
[0095] Vehicle 801 transmits three-dimensional data, such as point clouds, contained in a rectangular space 802 with width W, height H, and depth D located at a distance L from the vehicle 801 in front of it, to a traffic monitoring cloud that monitors road conditions or to a following vehicle at time intervals of Δt.
[0096] If a vehicle or person enters space 802 from the outside, causing a change in the three-dimensional data contained in space 802 that has been previously transmitted, vehicle 801 will also transmit the three-dimensional data of the space that has been changed.
[0097] Note that while Figure 9 shows an example where the shape of space 802 is a rectangular prism, space 802 does not necessarily have to be a rectangular prism; it only needs to include the space on the road ahead that is a blind spot for following vehicles.
[0098] It is desirable that the distance L be set to a distance at which a following vehicle can safely stop after receiving the three-dimensional data. For example, the distance L is set to the sum of the distance the following vehicle travels while it is receiving the three-dimensional data, the distance the following vehicle travels before it begins to decelerate in response to the received data, and the distance required for the following vehicle to safely stop after it begins to decelerate. Since these distances change with speed, the distance L may also change according to the vehicle's speed V, such that L = a × V + b (where a and b are constants).
[0099] The width W is set to a value greater than at least the width of the lane in which the vehicle 801 is traveling. More preferably, the width W is set to a size that includes adjacent spaces such as the left and right lanes or shoulders.
[0100] The depth D can be a fixed value, but it can also change according to the vehicle's speed V, as in D = c × V + d (where c and d are constants). Furthermore, by setting D such that D > V × Δt, the transmitted space can overlap with previously transmitted space. This allows vehicle 801 to more reliably transmit the space on the track without any omissions to following vehicles, etc.
[0101] In this way, by limiting the three-dimensional data transmitted by vehicle 801 to a space useful to following vehicles, the capacity of the transmitted three-dimensional data can be effectively reduced, thereby achieving lower communication latency and lower costs.
[0102] (Embodiment 4) Embodiment 3 describes an example in which a client device, such as a vehicle, transmits three-dimensional data to another vehicle or a server, such as a traffic monitoring cloud. In this embodiment, the client device transmits sensor information obtained from the sensor to the server or another client device.
[0103] First, the system configuration according to this embodiment will be described. Figure 10 shows the configuration of the three-dimensional map and sensor information transmission and reception system according to this embodiment. This system includes a server 901 and client devices 902A and 902B. When client devices 902A and 902B are not specifically distinguished, they will also be referred to as client device 902.
[0104] The client device 902 is, for example, an in-vehicle device mounted on a moving object such as a vehicle. The server 901 is, for example, a traffic monitoring cloud and is capable of communicating with multiple client devices 902.
[0105] Server 901 transmits a three-dimensional map composed of point clouds to client device 902. Note that the composition of the three-dimensional map is not limited to point clouds; it may also represent other three-dimensional data, such as a mesh structure.
[0106] The client device 902 transmits sensor information acquired by the client device 902 to the server 901. The sensor information includes, for example, at least one of the following: LiDAR acquisition information, visible light image, infrared image, depth image, sensor position information, and velocity information.
[0107] The data transmitted and received between the server 901 and the client device 902 may be compressed to reduce data size, or it may be left uncompressed to maintain data accuracy. When data is compressed, a three-dimensional compression method based on an octave structure, for example, can be used for point clouds. In addition, a two-dimensional image compression method can be used for visible light images, infrared images, and depth images. A two-dimensional image compression method is, for example, MPEG-4 AVC or HEVC, which are standardized by MPEG.
[0108] Furthermore, in response to a request from the client device 902 to send a 3D map, the server 901 sends a 3D map managed by the server 901 to the client device 902. The server 901 may also send a 3D map without waiting for a request from the client device 902. For example, the server 901 may broadcast a 3D map to one or more client devices 902 located in a predetermined space. Alternatively, the server 901 may send a 3D map appropriate to the location of the client device 902 at regular intervals after receiving a transmission request from the client device 902. The server 901 may also send a 3D map to the client device 902 whenever the 3D map managed by the server 901 is updated.
[0109] The client device 902 sends a request to the server 901 to send a three-dimensional map. For example, if the client device 902 wants to perform self-position estimation while driving, the client device 902 sends a request to the server 901 to send a three-dimensional map.
[0110] Furthermore, the client device 902 may request the server 901 to send a 3D map in the following cases: If the 3D map held by the client device 902 is outdated, the client device 902 may request the server 901 to send a 3D map. For example, if a certain period of time has elapsed since the client device 902 acquired the 3D map, the client device 902 may request the server 901 to send a 3D map.
[0111] Client device 902 may request server 901 to send the three-dimensional map to the server 901 a certain time before client device 902 leaves the space represented by the three-dimensional map held by client device 902. For example, client device 902 may request server 901 to send the three-dimensional map to the server 901 if it is within a predetermined distance from the boundary of the space represented by the three-dimensional map held by client device 902. Furthermore, if the movement path and speed of client device 902 are known, the time when client device 902 leaves the space represented by the three-dimensional map held by client device 902 may be predicted based on these.
[0112] If the error in the alignment between the three-dimensional data created by the client device 902 from sensor information and the three-dimensional map exceeds a certain level, the client device 902 may request the server 901 to send the three-dimensional map.
[0113] The client device 902 transmits sensor information to the server 901 in response to a request for transmission of sensor information sent from the server 901. The client device 902 may also send sensor information to the server 901 without waiting for a request for transmission of sensor information from the server 901. For example, once the client device 902 receives a request for transmission of sensor information from the server 901, it may periodically transmit sensor information to the server 901 for a certain period. Furthermore, if the error in the alignment between the three-dimensional data created by the client device 902 based on the sensor information and the three-dimensional map obtained from the server 901 exceeds a certain level, the client device 902 may determine that a change has occurred in the three-dimensional map around the client device 902 and transmit this information, along with the sensor information, to the server 901.
[0114] Server 901 requests client device 902 to transmit sensor information. For example, Server 901 receives location information of client device 902, such as GPS, from client device 902. Based on the location information of client device 902, if Server 901 determines that client device 902 is approaching an area with little information on the three-dimensional map managed by Server 901, it requests client device 902 to transmit sensor information in order to generate a new three-dimensional map. Server 901 may also request sensor information transmission if it wants to update the three-dimensional map, check road conditions during snowfall or disasters, check traffic congestion, or check incidents and accidents.
[0115] Furthermore, the client device 902 may set the amount of sensor information data to send to the server 901 depending on the communication status or bandwidth at the time of receiving the sensor information transmission request from the server 901. Setting the amount of sensor information data to send to the server 901 means, for example, increasing or decreasing the data itself, or selecting an appropriate compression method.
[0116] Figure 11 is a block diagram showing an example configuration of the client device 902. The client device 902 receives a three-dimensional map composed of a point cloud, etc., from the server 901, and estimates its own position from the three-dimensional data created based on the sensor information of the client device 902. The client device 902 also transmits the acquired sensor information to the server 901.
[0117] The client device 902 includes a data receiving unit 1011, a communication unit 1012, a reception control unit 1013, a format conversion unit 1014, a plurality of sensors 1015, a three-dimensional data creation unit 1016, a three-dimensional image processing unit 1017, a three-dimensional data storage unit 1018, a format conversion unit 1019, a communication unit 1020, a transmission control unit 1021, and a data transmission unit 1022.
[0118] The data receiving unit 1011 receives the three-dimensional map 1031 from the server 901. The three-dimensional map 1031 is data that includes point clouds such as WLD or SWLD. The three-dimensional map 1031 may contain either compressed or uncompressed data.
[0119] The communication unit 1012 communicates with the server 901 and sends data transmission requests (for example, a request to transmit a 3D map) to the server 901.
[0120] The receiving control unit 1013 exchanges information such as the supported format with the communication destination via the communication unit 1012 and establishes communication with the communication destination.
[0121] The format conversion unit 1014 generates a three-dimensional map 1032 by performing format conversion on the three-dimensional map 1031 received by the data reception unit 1011. Furthermore, if the three-dimensional map 1031 is compressed or encoded, the format conversion unit 1014 performs decompression or decoding. However, if the three-dimensional map 1031 is uncompressed data, the format conversion unit 1014 does not perform decompression or decoding.
[0122] Multiple sensors 1015 are a group of sensors that acquire external information from the vehicle on which the client device 902 is installed, such as LiDAR, visible light cameras, infrared cameras, or depth sensors, and generate sensor information 1033. For example, if sensor 1015 is a laser sensor such as LiDAR, the sensor information 1033 is three-dimensional data such as a point cloud (point cloud data). Note that there are not necessarily multiple sensors 1015.
[0123] The three-dimensional data creation unit 1016 creates three-dimensional data 1034 of the vehicle's surroundings based on the sensor information 1033. For example, the three-dimensional data creation unit 1016 uses information acquired by LiDAR and visible light images obtained by a visible light camera to create point cloud data with color information of the vehicle's surroundings.
[0124] The three-dimensional image processing unit 1017 uses the received three-dimensional map 1032, such as a point cloud, and the three-dimensional data 1034 of the vehicle's surroundings generated from sensor information 1033 to perform self-position estimation processing for the vehicle. Alternatively, the three-dimensional image processing unit 1017 may create three-dimensional data 1035 of the vehicle's surroundings by combining the three-dimensional map 1032 and the three-dimensional data 1034, and then perform self-position estimation processing using the created three-dimensional data 1035.
[0125] The three-dimensional data storage unit 1018 stores the three-dimensional map 1032, three-dimensional data 1034, and three-dimensional data 1035, etc.
[0126] The format conversion unit 1019 generates sensor information 1037 by converting the sensor information 1033 to a format supported by the receiving side. The format conversion unit 1019 may also reduce the amount of data by compressing or encoding the sensor information 1037. Furthermore, the format conversion unit 1019 may omit processing if format conversion is not necessary. The format conversion unit 1019 may also control the amount of data transmitted according to the specified transmission range.
[0127] The communication unit 1020 communicates with the server 901 and receives data transmission requests (sensor information transmission requests), etc., from the server 901.
[0128] The transmission control unit 1021 exchanges information such as the supported format with the communication destination via the communication unit 1020 and establishes communication.
[0129] The data transmission unit 1022 transmits sensor information 1037 to the server 901. The sensor information 1037 includes information acquired by multiple sensors 1015, such as information acquired by LiDAR, brightness images acquired by a visible light camera, infrared images acquired by an infrared camera, depth images acquired by a depth sensor, sensor position information, and velocity information.
[0130] Next, the configuration of server 901 will be described. Figure 12 is a block diagram showing an example configuration of server 901. Server 901 receives sensor information transmitted from client device 902 and creates three-dimensional data based on the received sensor information. Server 901 updates the three-dimensional map it manages using the created three-dimensional data. In addition, in response to a request from client device 902 to transmit the three-dimensional map, server 901 transmits the updated three-dimensional map to client device 902.
[0131] Server 901 comprises a data receiving unit 1111, a communication unit 1112, a reception control unit 1113, a format conversion unit 1114, a three-dimensional data creation unit 1116, a three-dimensional data synthesis unit 1117, a three-dimensional data storage unit 1118, a format conversion unit 1119, a communication unit 1120, a transmission control unit 1121, and a data transmission unit 1122.
[0132] The data receiving unit 1111 receives sensor information 1037 from the client device 902. The sensor information 1037 includes, for example, information acquired by LiDAR, brightness images acquired by a visible light camera, infrared images acquired by an infrared camera, depth images acquired by a depth sensor, sensor position information, and velocity information.
[0133] The communication unit 1112 communicates with the client device 902 and sends data transmission requests (for example, requests to transmit sensor information) to the client device 902.
[0134] The receiving control unit 1113 exchanges information such as the supported format with the communication destination via the communication unit 1112 and establishes communication.
[0135] The format conversion unit 1114 generates sensor information 1132 by decompressing or decoding the received sensor information 1037 if it is compressed or encoded. However, the format conversion unit 1114 does not perform decompression or decoding if the sensor information 1037 is uncompressed data.
[0136] The three-dimensional data creation unit 1116 creates three-dimensional data 1134 of the area around the client device 902 based on the sensor information 1132. For example, the three-dimensional data creation unit 1116 uses information acquired by LiDAR and visible light images obtained by a visible light camera to create point cloud data with color information of the area around the client device 902.
[0137] The three-dimensional data synthesis unit 1117 updates the three-dimensional map 1135 managed by the server 901 by synthesizing the three-dimensional data 1134, which was created based on the sensor information 1132, with the three-dimensional map 1135.
[0138] The three-dimensional data storage unit 1118 stores three-dimensional maps 1135, etc.
[0139] The format conversion unit 1119 generates a three-dimensional map 1031 by converting the three-dimensional map 1135 to a format supported by the receiving side. The format conversion unit 1119 may also reduce the amount of data by compressing or encoding the three-dimensional map 1135. Furthermore, the format conversion unit 1119 may omit processing if format conversion is not necessary. The format conversion unit 1119 may also control the amount of data transmitted according to the specified transmission range.
[0140] The communication unit 1120 communicates with the client device 902 and receives data transmission requests (such as requests to transmit a three-dimensional map) from the client device 902.
[0141] The transmission control unit 1121 exchanges information such as the supported format with the communication destination via the communication unit 1120 and establishes communication.
[0142] The data transmission unit 1122 transmits the three-dimensional map 1031 to the client device 902. The three-dimensional map 1031 is data that includes point clouds such as WLD or SWLD. The three-dimensional map 1031 may contain either compressed or uncompressed data.
[0143] Next, we will describe the operation flow of the client device 902. Figure 13 is a flowchart showing the operation of the client device 902 when acquiring a three-dimensional map.
[0144] First, the client device 902 requests the server 901 to transmit a three-dimensional map (such as a point cloud) (S1001). At this time, the client device 902 may also transmit its own location information obtained by GPS or the like, and request the server 901 to transmit a three-dimensional map related to that location information.
[0145] Next, the client device 902 receives a three-dimensional map from the server 901 (S1002). If the received three-dimensional map is compressed data, the client device 902 decodes the received three-dimensional map to generate an uncompressed three-dimensional map (S1003).
[0146] Next, the client device 902 creates three-dimensional data 1034 of the area around the client device 902 from sensor information 1033 obtained from multiple sensors 1015 (S1004). Then, the client device 902 estimates its own position using the three-dimensional map 1032 received from the server 901 and the three-dimensional data 1034 created from the sensor information 1033 (S1005).
[0147] Figure 14 is a flowchart showing the operation of the client device 902 when transmitting sensor information. First, the client device 902 receives a request to transmit sensor information from the server 901 (S1011). Upon receiving the transmission request, the client device 902 transmits the sensor information 1037 to the server 901 (S1012). If the sensor information 1033 includes multiple pieces of information obtained from multiple sensors 1015, the client device 902 may generate the sensor information 1037 by compressing each piece of information using a compression method suitable for each piece of information.
[0148] Next, the operation flow of server 901 will be described. Figure 15 is a flowchart showing the operation of server 901 when acquiring sensor information. First, server 901 requests client device 902 to send sensor information (S1021). Next, server 901 receives sensor information 1037 sent from client device 902 in response to the request (S1022). Next, server 901 creates three-dimensional data 1134 using the received sensor information 1037 (S1023). Next, server 901 reflects the created three-dimensional data 1134 in the three-dimensional map 1135 (S1024).
[0149] Figure 16 is a flowchart illustrating the operation of server 901 when transmitting a three-dimensional map. First, server 901 receives a request to transmit a three-dimensional map from client device 902 (S1031). Upon receiving the request to transmit a three-dimensional map, server 901 transmits the three-dimensional map 1031 to client device 902 (S1032). At this time, server 901 may extract a three-dimensional map of the vicinity of client device 902 according to its location information and transmit the extracted three-dimensional map. Alternatively, server 901 may compress the three-dimensional map composed of a point cloud using, for example, an octave tree compression method, and transmit the compressed three-dimensional map.
[0150] Modifications of this embodiment will be described below.
[0151] Server 901 uses sensor information 1037 received from client device 902 to create three-dimensional data 1134 of the area around client device 902. Next, server 901 calculates the difference between the created three-dimensional data 1134 and the three-dimensional map 1135 of the same area managed by server 901 by matching them. If the difference is greater than or equal to a predetermined threshold, server 901 determines that some kind of abnormality has occurred around client device 902. For example, when ground subsidence occurs due to a natural disaster such as an earthquake, a large difference may occur between the three-dimensional map 1135 managed by server 901 and the three-dimensional data 1134 created based on sensor information 1037.
[0152] The sensor information 1037 may include information indicating at least one of the following: the type of sensor, the performance of the sensor, and the model number of the sensor. Furthermore, a class ID corresponding to the sensor's performance may be added to the sensor information 1037. For example, if the sensor information 1037 is information acquired by a LiDAR, it is conceivable to assign identifiers to the sensor's performance, such as class 1 for sensors that can acquire information with accuracy in the millimeter range, class 2 for sensors that can acquire information with accuracy in the centimeter range, and class 3 for sensors that can acquire information with accuracy in the meter range. The server 901 may also estimate the sensor's performance information from the model number of the client device 902. For example, if the client device 902 is mounted in a vehicle, the server 901 may determine the sensor's specifications from the vehicle's make and model. In this case, the server 901 may have previously acquired information about the vehicle's make and model, or this information may be included in the sensor information. The server 901 may also use the acquired sensor information 1037 to switch the degree of correction applied to the three-dimensional data 1134 created using the sensor information 1037. For example, if the sensor performance is high precision (Class 1), the server 901 does not perform any correction on the three-dimensional data 1134. If the sensor performance is low precision (Class 3), the server 901 applies a correction to the three-dimensional data 1134 according to the accuracy of the sensor. For example, the lower the accuracy of the sensor, the stronger the degree (intensity) of the correction applied by the server 901.
[0153] Server 901 may simultaneously send requests for the transmission of sensor information to multiple client devices 902 located in a given space. When Server 901 receives multiple sensor information from multiple client devices 902, it is not necessary to use all of the sensor information to create the three-dimensional data 1134. For example, it may select which sensor information to use depending on the performance of the sensors. For example, when updating the three-dimensional map 1135, Server 901 may select high-precision sensor information (Class 1) from the multiple sensor information received and use the selected sensor information to create the three-dimensional data 1134.
[0154] Server 901 is not limited to servers such as traffic monitoring clouds, but may also be other client devices (in-vehicle). Figure 17 shows the system configuration in this case.
[0155] For example, client device 902C requests sensor information from a nearby client device 902A and obtains the sensor information from client device 902A. Then, client device 902C uses the obtained sensor information from client device 902A to create three-dimensional data and updates the three-dimensional map of client device 902C. In this way, client device 902C can generate a three-dimensional map of the space obtainable from client device 902A, taking advantage of the performance of client device 902C. For example, this case is likely to occur when client device 902C has high performance.
[0156] In this case, client device 902A, which provided the sensor information, is granted the right to acquire the high-precision three-dimensional map generated by client device 902C. Client device 902A receives the high-precision three-dimensional map from client device 902C in accordance with that right.
[0157] Furthermore, client device 902C may send requests for the transmission of sensor information to multiple nearby client devices 902 (client devices 902A and 902B). If the sensor of client device 902A or client device 902B is high-performance, client device 902C can create three-dimensional data using the sensor information obtained from this high-performance sensor.
[0158] Figure 18 is a block diagram showing the functional configuration of server 901 and client device 902. Server 901 includes, for example, a three-dimensional map compression / decoding processing unit 1201 that compresses and decodes three-dimensional maps, and a sensor information compression / decoding processing unit 1202 that compresses and decodes sensor information.
[0159] The client device 902 comprises a three-dimensional map decoding processing unit 1211 and a sensor information compression processing unit 1212. The three-dimensional map decoding processing unit 1211 receives encoded data of the compressed three-dimensional map, decodes the encoded data, and obtains the three-dimensional map. The sensor information compression processing unit 1212 compresses the sensor information itself instead of the three-dimensional data created from the acquired sensor information, and sends the encoded data of the compressed sensor information to the server 901. With this configuration, the client device 902 only needs to internally store a processing unit (device or LSI) that performs the processing of decoding the three-dimensional map (point cloud, etc.), and does not need to internally store a processing unit that performs the processing of compressing the three-dimensional data of the three-dimensional map (point cloud, etc.). This reduces the cost and power consumption of the client device 902.
[0160] As described above, the client device 902 according to this embodiment is mounted on a mobile body and creates three-dimensional data 1034 of the surrounding area of the mobile body from sensor information 1033 indicating the surrounding conditions of the mobile body obtained by a sensor 1015 mounted on the mobile body. The client device 902 estimates the self-position of the mobile body using the created three-dimensional data 1034. The client device 902 transmits the acquired sensor information 1033 to the server 901 or another mobile body 902.
[0161] According to this, the client device 902 transmits sensor information 1033 to the server 901, etc. This may reduce the amount of data transmitted compared to transmitting three-dimensional data. In addition, since the client device 902 does not need to perform processing such as compression or encoding of three-dimensional data, the processing load on the client device 902 can be reduced. Therefore, the client device 902 can achieve a reduction in the amount of data transmitted or a simplification of the device configuration.
[0162] Furthermore, the client device 902 sends a request to the server 901 to send a three-dimensional map, and receives the three-dimensional map 1031 from the server 901. In estimating its own position, the client device 902 uses the three-dimensional data 1034 and the three-dimensional map 1032 to estimate its own position.
[0163] Furthermore, the sensor information 1033 includes at least one of the following: information obtained from the laser sensor, brightness image, infrared image, depth image, sensor position information, and sensor velocity information.
[0164] Furthermore, sensor information 1033 includes information indicating the performance of the sensor.
[0165] Furthermore, the client device 902 encodes or compresses the sensor information 1033, and when transmitting the sensor information, it transmits the encoded or compressed sensor information 1037 to the server 901 or another mobile device 902. This allows the client device 902 to reduce the amount of data transmitted.
[0166] For example, the client device 902 includes a processor and memory, and the processor uses the memory to perform the above processing.
[0167] Furthermore, the server 901 according to this embodiment is capable of communicating with a client device 902 mounted on the mobile body, and receives sensor information 1037 from the client device 902 that indicates the surrounding conditions of the mobile body, obtained by a sensor 1015 mounted on the mobile body. The server 901 creates three-dimensional data 1134 of the surroundings of the mobile body from the received sensor information 1037.
[0168] According to this, the server 901 creates three-dimensional data 1134 using sensor information 1037 transmitted from the client device 902. This may reduce the amount of data transmitted compared to when the client device 902 transmits the three-dimensional data. In addition, since the client device 902 does not need to perform processing such as compression or encoding of the three-dimensional data, the processing load on the client device 902 can be reduced. Therefore, the server 901 can reduce the amount of data transmitted or simplify the configuration of the device.
[0169] Furthermore, the server 901 also sends a request to the client device 902 to transmit sensor information.
[0170] Furthermore, the server 901 updates the three-dimensional map 1135 using the created three-dimensional data 1134 and sends the three-dimensional map 1135 to the client device 902 in response to a request from the client device 902 to send the three-dimensional map 1135.
[0171] Furthermore, the sensor information 1037 includes at least one of the following: information obtained from the laser sensor, brightness image, infrared image, depth image, sensor position information, and sensor velocity information.
[0172] Furthermore, sensor information 1037 includes information indicating the performance of the sensor.
[0173] Furthermore, the server 901 corrects the three-dimensional data according to the performance of the sensor. This allows the three-dimensional data creation method to improve the quality of the three-dimensional data.
[0174] Furthermore, when receiving sensor information, the server 901 receives multiple pieces of sensor information 1037 from multiple client devices 902, and selects the sensor information 1037 to be used to create the three-dimensional data 1134 based on the multiple pieces of information indicating the performance of the sensors contained in the multiple pieces of sensor information 1037. In this way, the server 901 can improve the quality of the three-dimensional data 1134.
[0175] Furthermore, the server 901 decodes or decodes the received sensor information 1037 and creates three-dimensional data 1134 from the decoded or decoded sensor information 1132. This allows the server 901 to reduce the amount of data transmitted.
[0176] For example, server 901 is equipped with a processor and memory, and the processor uses the memory to perform the above processing.
[0177] (Embodiment 5) This embodiment describes a modified version of Embodiment 4 described above. Figure 19 is a diagram showing the configuration of the system according to this embodiment. The system shown in Figure 19 includes a server 2001, a client device 2002A, and a client device 2002B.
[0178] Client devices 2002A and 2002B are mounted on a moving object such as a vehicle and transmit sensor information to server 2001. Server 2001 transmits a three-dimensional map (point cloud) to client devices 2002A and 2002B.
[0179] Client device 2002A comprises a sensor information acquisition unit 2011, a storage unit 2012, and a data transmission feasibility determination unit 2013. The configuration of client device 2002B is similar. Furthermore, in the following description, unless otherwise specified, client device 2002A and client device 2002B will also be referred to as client device 2002.
[0180] Figure 20 is a flowchart showing the operation of the client device 2002 according to this embodiment.
[0181] The sensor information acquisition unit 2011 acquires various sensor information using sensors (sensor group) mounted on the mobile body. In other words, the sensor information acquisition unit 2011 acquires sensor information indicating the surrounding conditions of the mobile body obtained by sensors (sensor group) mounted on the mobile body. The sensor information acquisition unit 2011 also stores the acquired sensor information in the storage unit 2012. This sensor information includes at least one of LiDAR acquisition information, visible light images, infrared images, and depth images. The sensor information may also include at least one of sensor position information, velocity information, acquisition time information, and acquisition location information. Sensor position information indicates the position of the sensor that acquired the sensor information. Velocity information indicates the velocity of the mobile body when the sensor acquired the sensor information. Acquisition time information indicates the time when the sensor information was acquired by the sensor. Acquisition location information indicates the position of the mobile body or sensor when the sensor information was acquired by the sensor.
[0182] Next, the data transmission feasibility determination unit 2013 determines whether the mobile device (client device 2002) is in an environment where it can transmit sensor information to the server 2001 (S2002). For example, the data transmission feasibility determination unit 2013 may use information such as GPS to identify the location and time of the client device 2002 and determine whether data can be transmitted. Alternatively, the data transmission feasibility determination unit 2013 may determine whether data can be transmitted based on whether it can connect to a specific access point.
[0183] If the client device 2002 determines that the moving object is in an environment where it can transmit sensor information to the server 2001 (Yes in S2002), it transmits the sensor information to the server 2001 (S2003). In other words, as soon as the client device 2002 is in a situation where it can transmit sensor information to the server 2001, it transmits the sensor information it holds to the server 2001. For example, suppose a millimeter-wave access point capable of high-speed communication is installed at an intersection. When the client device 2002 enters the intersection, it transmits the sensor information it holds to the server 2001 at high speed using millimeter-wave communication.
[0184] Next, the client device 2002 deletes the sensor information already transmitted to the server 2001 from the storage unit 2012 (S2004). The client device 2002 may also delete sensor information that has not yet been transmitted to the server 2001 if it meets certain conditions. For example, the client device 2002 may delete the sensor information from the storage unit 2012 when the acquisition time of the stored sensor information becomes older than a certain time from the current time. In other words, the client device 2002 may delete the sensor information from the storage unit 2012 if the difference between the time the sensor information was acquired by the sensor and the current time exceeds a predetermined time. Furthermore, the client device 2002 may delete the sensor information from the storage unit 2012 when the acquisition location of the stored sensor information is more than a certain distance from the current location. In other words, the client device 2002 may delete the sensor information from the storage unit 2012 if the difference between the position of the moving object or sensor when the sensor information was acquired by the sensor and the current position of the moving object or sensor exceeds a predetermined distance. This makes it possible to reduce the capacity of the storage unit 2012 of the client device 2002.
[0185] If client device 2002 has not finished acquiring sensor information (No in S2005), client device 2002 repeats the processing from step S2001 onwards. If client device 2002 has finished acquiring sensor information (Yes in S2005), client device 2002 terminates processing.
[0186] Furthermore, the client device 2002 may select the sensor information to send to the server 2001 according to the communication status. For example, if high-speed communication is possible, the client device 2002 will prioritize sending sensor information with a large size stored in the memory unit 2012 (e.g., LiDAR acquisition information). If high-speed communication is difficult, the client device 2002 will send sensor information with a small size stored in the memory unit 2012 and a high priority (e.g., visible light images). This allows the client device 2002 to efficiently send the sensor information stored in the memory unit 2012 to the server 2001 according to the network status.
[0187] Furthermore, the client device 2002 may obtain time information indicating the current time and location information indicating the current location from the server 2001. The client device 2002 may also determine the acquisition time and location of sensor information based on the acquired time and location information. In other words, the client device 2002 may obtain time information from the server 2001 and generate acquisition time information using the acquired time information. Furthermore, the client device 2002 may obtain location information from the server 2001 and generate acquisition location information using the acquired location information.
[0188] For example, regarding time information, the server 2001 and client device 2002 synchronize their time using a mechanism such as NTP (Network Time Protocol) or PTP (Precision Time Protocol). This allows client device 2002 to obtain accurate time information. Furthermore, since the server 2001 can synchronize time between multiple client devices, the time in sensor information acquired by different client devices 2002 can be synchronized. Therefore, the server 2001 can handle sensor information that indicates the synchronized time. Note that any method other than NTP or PTP can be used for time synchronization. Also, GPS information may be used as the above time information and location information.
[0189] Server 2001 may obtain sensor information from multiple client devices 2002 by specifying the time or location. For example, in the event of an accident, to find clients that were nearby, Server 2001 broadcasts a sensor information transmission request to multiple client devices 2002, specifying the time and location of the accident. Client devices 2002 that have sensor information for the corresponding time and location then transmit the sensor information to Server 2001. In other words, Client device 2002 receives a sensor information transmission request from Server 2001 that includes specification information specifying the location and time. If Client device 2002 determines that the storage unit 2012 has stored sensor information obtained at the location and time indicated by the specification information, and that the moving object is in an environment where it can transmit sensor information to Server 2001, it transmits the sensor information obtained at the location and time indicated by the specification information to Server 2001. As a result, Server 2001 can obtain sensor information related to the occurrence of an accident from multiple client devices 2002 and use it for accident analysis, etc.
[0190] Furthermore, the client device 2002 may refuse to transmit sensor information when it receives a request from the server 2001 to transmit sensor information. Alternatively, the client device 2002 may pre-configure which of the multiple sensor information requests it can transmit. Or, the server 2001 may query the client device 2002 each time to determine whether or not to transmit sensor information.
[0191] Furthermore, client devices 2002 that transmit sensor information to server 2001 may be awarded points. These points can be used to pay for things like gasoline, electric vehicle (EV) charging fees, highway tolls, or rental car fees. Also, after acquiring sensor information, server 2001 may delete information that identifies the client device 2002 that sent the sensor information. For example, this information could be the network address of client device 2002. This anonymizes the sensor information, allowing users of client device 2002 to confidently transmit sensor information from client device 2002 to server 2001. Server 2001 may also consist of multiple servers. For example, by sharing sensor information among multiple servers, even if one server fails, other servers can communicate with client device 2002. This prevents service interruptions due to server failures.
[0192] Furthermore, the specified location in the sensor information transmission request indicates the location where the accident occurred, and may differ from the location of the client device 2002 at the specified time specified in the sensor information transmission request. Therefore, the server 2001 can request information acquisition from client devices 2002 located within a specified range, such as within XXm of the location, by specifying a range such as within XXm of the location. Similarly, for the specified time, the server 2001 may specify a range such as within N seconds before or after a certain time. This allows the server 2001 to acquire sensor information from client devices 2002 that were located within XXm of absolute position S at "time: tN to t+N". When the client device 2002 transmits three-dimensional data such as LiDAR, it may transmit data generated immediately after time t.
[0193] Furthermore, the server 2001 may separately specify information indicating the location of the client device 2002 from which sensor information is to be acquired, and the location where the sensor information is desired. For example, the server 2001 specifies that sensor information including at least the range YYm from absolute position S should be acquired from a client device 2002 located within XXm of absolute position S. When the client device 2002 selects the three-dimensional data to transmit, it selects one or more randomly accessible units of three-dimensional data so as to include at least the sensor information within the specified range. Also, when the client device 2002 transmits a visible light image, it may transmit multiple temporally consecutive image data, including at least the frame immediately before or after time t.
[0194] If the client device 2002 can utilize multiple physical networks, such as 5G, WiFi, or multiple modes in 5G, for transmitting sensor information, the client device 2002 may select the network to use according to the priority notified by the server 2001. Alternatively, the client device 2002 may select a network that can secure appropriate bandwidth based on the size of the data to be transmitted. Alternatively, the client device 2002 may select a network to use based on the cost of data transmission, etc. Furthermore, the transmission request from the server 2001 may include information indicating a transmission deadline, such as transmitting if the client device 2002 can start transmitting by time T. If sufficient sensor information is not obtained within the deadline, the server 2001 may issue another transmission request.
[0195] The sensor information may include compressed or uncompressed sensor data, along with header information indicating the characteristics of the sensor data. The client device 2002 may transmit the header information to the server 2001 via a different physical network or communication protocol than the sensor data. For example, the client device 2002 transmits the header information to the server 2001 prior to transmitting the sensor data. The server 2001 determines whether to acquire the sensor data from the client device 2002 based on the analysis results of the header information. For example, the header information may include information indicating the LiDAR point cloud acquisition density, elevation angle, or frame rate, or the resolution, signal-to-noise ratio, or frame rate of the visible light image. This allows the server 2001 to acquire sensor information from the client device 2002 that has sensor data of the determined quality.
[0196] As described above, the client device 2002 is mounted on the mobile body and acquires sensor information indicating the surrounding conditions of the mobile body obtained by sensors mounted on the mobile body, and stores the sensor information in the storage unit 2012. The client device 2002 determines whether the mobile body is in an environment where it can transmit sensor information to the server 2001, and if it determines that the mobile body is in an environment where it can transmit sensor information to the server, it transmits the sensor information to the server 2001.
[0197] Furthermore, the client device 2002 creates three-dimensional data of the surroundings of the moving object from the sensor information and uses the created three-dimensional data to estimate the self-position of the moving object.
[0198] Furthermore, the client device 2002 sends a request to the server 2001 to send a 3D map, and receives the 3D map from the server 2001. In estimating its own position, the client device 2002 uses the 3D data and the 3D map to estimate its own position.
[0199] Furthermore, the processing performed by the client device 2002 may be implemented as an information transmission method in the client device 2002.
[0200] Furthermore, the client device 2002 may include a processor and a memory, and the processor may use the memory to execute the above processing.
[0201] Next, the sensor information collection system according to the present embodiment will be described. Figure 21 is a diagram showing the configuration of the sensor information collection system according to the present embodiment. As shown in Figure 21, the sensor information collection system according to the present embodiment includes a terminal 2021A, a terminal 2021B, a communication device 2022A, a communication device 2022B, a network 2023, a data collection server 2024, a map server 2025, and a client device 2026. Note that when the terminal 2021A and the terminal 2021B are not particularly distinguished, they are also referred to as the terminal 2021. When the communication device 2022A and the communication device 2022B are not particularly distinguished, they are also referred to as the communication device 2022.
[0202] The data collection server 2024 collects data such as sensor data obtained by sensors included in the terminal 2021 as position-related data associated with positions in a three-dimensional space.
[0203] Sensor data is, for example, data acquired by using a sensor included in the terminal 2021, such as the state surrounding the terminal 2021 or the internal state of the terminal 2021. The terminal 2021 transmits sensor data collected from one or more sensor devices located at positions where direct communication with the terminal 2021 is possible, or communication is possible via one or more relay devices using the same communication scheme, to the data collection server 2024.
[0204] Data included in the position-related data may include, for example, information indicating the operating state, operation logs, service usage status of the terminal itself or devices included in the terminal. Furthermore, data included in the position-related data may include information that associates the identifier of the terminal 2021 with the position or movement route of the terminal 2021.
[0205] The position indicating information included in the position-related data is associated with position indicating information in three-dimensional data such as three-dimensional map data, for example. Details of the position indicating information will be described later.
[0206] In addition to position information which is information indicating a position, the position-related data may include at least one of the aforementioned time information and information indicating an attribute of data included in the position-related data or a type (for example, a model number) of a sensor that generated the data. The position information and the time information may be stored in a header area of the position-related data or a header area of a frame storing the position-related data. Further, the position information and the time information may be transmitted and / or stored separately from the position-related data as metadata associated with the position-related data.
[0207] The map server 2025 is, for example, connected to a network 2023, and transmits three-dimensional data such as three-dimensional map data in response to a request from another device such as a terminal 2021. Further, as described in each of the foregoing embodiments, the map server 2025 may have a function of updating three-dimensional data using sensor information transmitted from the terminal 2021, or the like.
[0208] The data collection server 2024 is, for example, connected to a network 2023, collects position-related data from another device such as a terminal 2021, and stores the collected position-related data in a storage device inside itself or in another server. Further, the data collection server 2024 transmits the collected position-related data or metadata of three-dimensional map data generated based on the position-related data, or the like, to the terminal 2021 in response to a request from the terminal 2021.
[0209] Network 2023 is a communication network, such as the Internet. Terminal 2021 is connected to Network 2023 via communication device 2022. Communication device 2022 communicates with terminal 2021 by switching between one or more communication methods. Communication device 2022 is, for example, (1) a base station such as LTE (Long Term Evolution), (2) an access point (AP) such as WiFi or millimeter wave communication, (3) a gateway for an LPWA (Low Power Wide Area) Network such as SIGFOX, LoRaWAN or Wi-SUN, or (4) a communication satellite that communicates using a satellite communication method such as DVB-S2.
[0210] The base station may communicate with terminal 2021 using a method classified as NB-IoT (Narrow Band-IoT) or LPWA such as LTE-M, or it may communicate with terminal 2021 while switching between these methods.
[0211] Here, we take an example where terminal 2021 has the function to communicate with communication device 2022 that uses two types of communication methods, and communicates with map server 2025 or data collection server 2024 using one of these communication methods, or by switching between multiple communication methods and the communication device 2022 that is the direct communication partner. However, the configuration of the sensor information collection system and terminal 2021 is not limited to this. For example, terminal 2021 may not have the function to communicate using multiple communication methods, but may have the function to communicate using one of the communication methods. Also, terminal 2021 may support three or more communication methods. Furthermore, each terminal 2021 may support different communication methods.
[0212] Terminal 2021 has the configuration of, for example, the client device 902 shown in Figure 11. Terminal 2021 performs position estimation, such as its own position, using the received three-dimensional data. Terminal 2021 also generates position-related data by associating sensor data acquired from sensors with position information obtained through position estimation processing.
[0213] The location information added to location-related data indicates, for example, the position in the coordinate system used in the three-dimensional data. For example, the location information is a coordinate value expressed as latitude and longitude. In this case, terminal 2021 may include information indicating the coordinate system on which the coordinate value is based, and the three-dimensional data used for position estimation, along with the coordinate value itself. The coordinate value may also include altitude information.
[0214] Furthermore, location information may be associated with data units or spatial units that can be used to encode the three-dimensional data described above. These units include, for example, WLD, GOS, SPC, VLM, or VXL. In this case, location information is represented by an identifier that identifies a data unit such as an SPC that corresponds to location-related data. In addition to the identifier that identifies a data unit such as an SPC, location information may also include information indicating three-dimensional data encoded from the three-dimensional space containing the data unit such as the SPC, or information indicating a detailed location within the SPC. Information indicating three-dimensional data is, for example, the file name of the three-dimensional data.
[0215] Thus, by generating location-related data associated with location information based on position estimation using three-dimensional data, this system can add more accurate location information to sensor information than when adding location information based on the self-position of a client device (terminal 2021) acquired using GPS. As a result, even when other devices use the location-related data in other services, it may be possible to more accurately identify the location corresponding to the location-related data in real space by performing position estimation based on the same three-dimensional data.
[0216] In this embodiment, the example of data transmitted from terminal 2021 being location-related data was used for explanation. However, data transmitted from terminal 2021 may not be associated with location information. In other words, the transmission and reception of three-dimensional data or sensor data described in other embodiments may be performed via the network 2023 described in this embodiment.
[0217] Next, we will describe different examples of location information that indicates a position in three-dimensional or two-dimensional real space or map space. The location information attached to location-related data may also be information that indicates the relative position to a feature point in three-dimensional data. Here, the feature point that serves as the basis for the location information is, for example, a feature point that is encoded as SWLD and notified to terminal 2021 as three-dimensional data.
[0218] Information indicating the relative position to a feature point may be represented, for example, by a vector from the feature point to the point indicated by the position information, and may also be information indicating the direction and distance from the feature point to the point indicated by the position information. Alternatively, information indicating the relative position to a feature point may be information indicating the displacement amounts of the X, Y, and Z axes from the feature point to the point indicated by the position information. Furthermore, information indicating the relative position to a feature point may be information indicating the distance from each of three or more feature points to the point indicated by the position information. Note that the relative position may not be the relative position of the point indicated by the position information expressed with respect to each feature point, but rather the relative position of each feature point expressed with respect to the point indicated by the position information. An example of position information based on the relative position to a feature point includes information for identifying the reference feature point and information indicating the relative position of the point indicated by the position information with respect to that feature point. Furthermore, if information indicating the relative position to a feature point is provided separately from the three-dimensional data, the information indicating the relative position to a feature point may include the coordinate axes used to derive the relative position, information indicating the type of three-dimensional data, and / or information indicating the magnitude per unit quantity (scale, etc.) of the value of the information indicating the relative position.
[0219] Furthermore, the location information may include information indicating the relative position of multiple feature points to each feature point. When the location information is represented by the relative position to multiple feature points, terminal 2021, which is trying to identify the location indicated by the location information in real space, may calculate candidate points for the location indicated by the location information from the position of each feature point estimated from the sensor data, and determine that the point obtained by averaging the calculated candidate points is the point indicated by the location information. With this configuration, the influence of errors when estimating the position of feature points from sensor data can be reduced, and thus the accuracy of estimating the point indicated by the location information in real space can be improved. In addition, if the location information includes information indicating the relative position to multiple feature points, even if there are feature points that cannot be detected due to constraints such as the type or performance of the sensors that terminal 2021 has, it is possible to estimate the value of the point indicated by the location information if even one of the multiple feature points can be detected.
[0220] Points identifiable from sensor data can be used as feature points. Points identifiable from sensor data are, for example, points or points within a region that satisfy predetermined conditions for feature point detection, such as the aforementioned three-dimensional features or visible light data features being above a threshold.
[0221] Furthermore, markers placed in real space may be used as feature points. In this case, the markers only need to be detectable and their location determined from data acquired using sensors such as LiDAR or cameras. For example, markers can be represented by changes in color or brightness values (reflectance), or by three-dimensional shapes (such as bumps and depressions). Alternatively, coordinate values indicating the position of the marker, or a two-dimensional code or barcode generated from the identifier of the marker may be used.
[0222] Furthermore, a light source that transmits optical signals may be used as a marker. When an optical signal light source is used as a marker, not only information for obtaining location, such as coordinate values or identifiers, but also other data may be transmitted by the optical signal. For example, the optical signal may include information such as the content of the service corresponding to the location of the marker, an address such as a URL for obtaining the content, or an identifier of a wireless communication device for receiving the service, and a wireless communication method for connecting to the wireless communication device. By using an optical communication device (light source) as a marker, it becomes easier to transmit data other than location information, and it becomes possible to dynamically switch such data.
[0223] Terminal 2021 grasps the correspondence between feature points in different data sets, for example, by using an identifier commonly used between the data sets, or by using information or a table that shows the correspondence between feature points in the data sets. If there is no information showing the correspondence between feature points, terminal 2021 may determine that the feature point that is closest when the coordinates of a feature point in one three-dimensional data set are converted to a position in the three-dimensional data space of the other is the corresponding feature point.
[0224] As described above, when using location information based on relative position, even between terminals 2021 or services using different three-dimensional data, the location indicated by the location information can be identified or estimated based on common feature points included in or associated with each three-dimensional data. As a result, it becomes possible to identify or estimate the same location with higher accuracy between terminals 2021 or services using different three-dimensional data.
[0225] Furthermore, even when using map data or 3D data represented using different coordinate systems, the impact of errors associated with coordinate system transformations can be reduced, enabling the integration of services based on more accurate location information.
[0226] Hereinafter, an example of functions provided by the data collection server 2024 will be described. The data collection server 2024 may transfer the received position-related data to another data server. When there are a plurality of data servers, the data collection server 2024 determines to which data server the received position-related data is to be transferred, and transfers the position-related data to the data server determined as the transfer destination.
[0227] The data collection server 2024 determines the transfer destination, for example, based on transfer destination server determination rules set in advance in the data collection server 2024. The transfer destination server determination rules are set, for example, in a transfer destination table that associates an identifier associated with each terminal 2021 with a transfer destination data server.
[0228] The terminal 2021 adds an identifier associated with the terminal 2021 to the position-related data to be transmitted, and transmits the position-related data to the data collection server 2024. The data collection server 2024 specifies a transfer destination data server corresponding to the identifier added to the position-related data based on the transfer destination server determination rules using the transfer destination table or the like, and transmits the position-related data to the specified data server. Further, the transfer destination server determination rule may be specified by a determination condition using the time or place where the position-related data is acquired. Here, the identifier associated with the transmission source terminal 2021 described above is, for example, an identifier unique to each terminal 2021, or an identifier indicating a group to which the terminal 2021 belongs.
[0229] Furthermore, the destination table does not necessarily have to directly associate identifiers associated with the source terminal with destination data servers. For example, data collection server 2024 maintains a management table that stores tag information assigned to each identifier unique to terminal 2021, and a destination table that associates said tag information with destination data servers. Data collection server 2024 may use the management table and the destination table to determine the destination data server based on the tag information. Here, the tag information is, for example, management control information or service provision control information assigned to the type, model number, owner, group to which it belongs, or other identifiers corresponding to the identifier of terminal 2021. Also, instead of identifiers associated with the source terminal 2021, a unique identifier for each sensor may be used in the destination table. Furthermore, the rules for determining the destination server may be set from client device 2026.
[0230] The data collection server 2024 may determine multiple data servers as destinations and transfer the received location-related data to those multiple data servers. With this configuration, for example, when automatically backing up location-related data, or when it is necessary to send location-related data to data servers that provide each service in order to use the location-related data in common across different services, the intended data transfer can be achieved by changing the settings of the data collection server 2024. As a result, the man-hours required for system construction and modification can be reduced compared to setting destinations for location-related data on individual terminals 2021.
[0231] The data acquisition server 2024 may, in response to a transfer request signal received from a data server, register the data server specified in the transfer request signal as a new transfer destination and transfer any subsequent location-related data to that data server.
[0232] The data collection server 2024 may store the location-related data received from terminal 2021 in a recording device, and in response to a transmission request signal received from terminal 2021 or the data server, it may transmit the location-related data specified in the transmission request signal to the requesting terminal 2021 or data server.
[0233] The data collection server 2024 may determine whether it is possible to provide location-related data to the requesting data server or terminal 2021, and if it is determined that it is possible to provide the data, it may transfer or transmit the location-related data to the requesting data server or terminal 2021.
[0234] If the data collection server 2024 receives a request for current location-related data from client device 2026, it may request terminal 2021 to send location-related data even if it is not the timing for terminal 2021 to send location-related data, and terminal 2021 may send location-related data in response to that request.
[0235] In the above description, it was assumed that terminal 2021 transmits location information data to the data collection server 2024. However, the data collection server 2024 may also have functions necessary for collecting location-related data from terminal 2021, such as functions for managing terminal 2021, or functions used when collecting location-related data from terminal 2021.
[0236] The data collection server 2024 may also have the function of sending a data request signal to terminal 2021 requesting the transmission of location information data and collecting location-related data.
[0237] The data collection server 2024 has pre-registered management information, such as an address for communicating with the terminal 2021 that is the target of data collection, or a terminal 2021-specific identifier. Based on the registered management information, the data collection server 2024 collects location-related data from terminal 2021. The management information may also include information such as the type of sensor that terminal 2021 has, the number of sensors that terminal 2021 has, and the communication method that terminal 2021 supports.
[0238] The data collection server 2024 may collect information from terminal 2021, such as its operating status or current location.
[0239] The registration of management information may be performed by the client device 2026, or the registration process may be initiated when terminal 2021 sends a registration request to the data collection server 2024. The data collection server 2024 may have a function to control communication with terminal 2021.
[0240] The communication between the data collection server 2024 and the terminal 2021 may be a dedicated line provided by a service provider such as an MNO (Mobile Network Operator) or MVNO (Mobile Virtual Network Operator), or a virtual dedicated line configured with a VPN (Virtual Private Network). This configuration allows for secure communication between the terminal 2021 and the data collection server 2024.
[0241] The data collection server 2024 may have a function to authenticate terminal 2021 or a function to encrypt data transmitted to and from terminal 2021. Here, the authentication process of terminal 2021 or the data encryption process is performed using an identifier unique to terminal 2021 or an identifier unique to a group of terminals including multiple terminals 2021, which has been shared in advance between the data collection server 2024 and terminal 2021. This identifier is, for example, the IMSI (International Mobile Subscriber Identity), which is a unique number stored on a SIM (Subscriber Identity Module) card. The identifier used for authentication and the identifier used for data encryption may be the same or different.
[0242] Authentication or data encryption between the data collection server 2024 and terminal 2021 can be provided if both the data collection server 2024 and terminal 2021 have the functionality to perform such processing, and is independent of the communication method used by the relaying communication device 2022. Therefore, a common authentication or encryption process can be used without considering the communication method used by terminal 2021, improving the convenience of system construction for users. However, "independent of the communication method used by the relaying communication device 2022" means that it is not essential to change it according to the communication method. In other words, for the purpose of improving transmission efficiency or ensuring security, the authentication or data encryption process between the data collection server 2024 and terminal 2021 may be switched according to the communication method used by the relaying device.
[0243] The data collection server 2024 may provide the client device 2026 with a UI that manages data collection rules, such as the type of location-related data to be collected from terminal 2021 and the data collection schedule. This allows the user to specify the terminal 2021 from which to collect data, as well as the data collection time and frequency, using the client device 2026. The data collection server 2024 may also specify a map area from which to collect data and collect location-related data from terminal 2021 included in that area.
[0244] When data collection rules are managed on a per-terminal basis (terminal 2021), the client device 2026 displays a list of the terminals or sensors to be managed on its screen. The user then sets whether data collection is necessary or the collection schedule for each item in the list.
[0245] When specifying an area on a map from which data should be collected, the client device 2026 displays, for example, a two-dimensional or three-dimensional map of the area to be managed. The user selects the area from which data should be collected on the displayed map. The area selected on the map may be a circular or rectangular area centered on a point specified on the map, or a circular or rectangular area that can be identified by dragging. Alternatively, the client device 2026 may select an area using predefined units such as a city, an area within a city, a block, or a major road. Furthermore, instead of specifying an area using a map, an area may be set by entering latitude and longitude values, or an area may be selected from a list of candidate areas derived based on entered text information. The text information may include, for example, the names of regions, cities, or landmarks.
[0246] Furthermore, data collection may be performed while dynamically changing the specified area by having the user specify one or more terminals 2021 and set conditions such as within a 100-meter radius around those terminals 2021.
[0247] Furthermore, if the client device 2026 is equipped with sensors such as a camera, a map area may be specified based on the real-world position of the client device 2026 obtained from the sensor data. For example, the client device 2026 may estimate its own position using the sensor data and specify a data collection area within a predetermined distance from a point on the map corresponding to the estimated position, or within a distance specified by the user. Alternatively, the client device 2026 may specify the sensing area of the sensor, i.e., the area corresponding to the acquired sensor data, as the data collection area. Or, the client device 2026 may specify a data collection area based on the position corresponding to the sensor data specified by the user. The estimation of the map area or position corresponding to the sensor data may be performed by the client device 2026 or by the data collection server 2024.
[0248] When specifying an area on a map, the data collection server 2024 may collect the current location information of each terminal 2021 to identify terminals 2021 within the specified area and request the identified terminals 2021 to transmit location-related data. Alternatively, instead of the data collection server 2024 identifying terminals 2021 within the area, the data collection server 2024 may send information indicating the specified area to terminals 2021, and terminals 2021 may determine whether they are within the specified area and transmit location-related data if they are determined to be within the specified area.
[0249] The data collection server 2024 transmits data such as lists or maps to the client device 2026 to provide the aforementioned UI (User Interface) in the application run by the client device 2026. The data collection server 2024 may transmit not only data such as lists or maps, but also the application program to the client device 2026. Furthermore, the aforementioned UI may be provided as content created in HTML or similar format that can be displayed in a browser. Note that some data, such as map data, may be provided by servers other than the data collection server 2024, such as the map server 2025.
[0250] When the client device 2026 receives input that indicates completion of input, such as when the user presses a settings button, it sends the entered information as configuration information to the data collection server 2024. Based on the configuration information received from the client device 2026, the data collection server 2024 sends a signal to each terminal 2021 requesting location-related data or notifying them of location-related data collection rules, and then collects location-related data.
[0251] Next, we will describe an example of controlling the operation of terminal 2021 based on additional information added to three-dimensional or two-dimensional map data.
[0252] In this configuration, object information indicating the location of a power supply unit, such as a wireless power supply antenna or power supply coil embedded in a road or parking lot, is included in or associated with three-dimensional data and provided to terminal 2021, such as a car or drone.
[0253] A vehicle or drone that has acquired object information in order to charge will automatically move its position so that the charging part, such as the charging antenna or charging coil, is positioned opposite the area indicated by the object information, and will begin charging. In the case of a vehicle or drone that does not have an autonomous driving function, the driver or operator will be presented with the direction to move or the operation to perform using images or sounds displayed on the screen. When it is determined that the position of the charging part, calculated based on the estimated self-position, has entered the area indicated by the object information or within a predetermined distance from that area, the images or sounds presented will switch to instructions to stop driving or operation, and charging will begin.
[0254] Furthermore, the object information may not be information indicating the location of the power supply unit, but rather information indicating a region where, if the charging unit is placed within that region, a charging efficiency of a predetermined threshold or higher can be obtained. The location of the object information may be represented by the center point of the region indicated by the object information, or by a region or line in a two-dimensional plane, or by a region, line or plane in three-dimensional space.
[0255] This configuration allows for the determination of the power supply antenna's location, which cannot be determined from LiDAR sensing data or camera footage. This enables more precise alignment between the wireless charging antenna on a vehicle or other device (2021) and the wireless power supply antenna embedded in the road or other infrastructure. As a result, the charging speed during wireless charging can be shortened, and charging efficiency can be improved.
[0256] The object information may include objects other than the power supply antenna. For example, the three-dimensional data may include the location of the AP for millimeter-wave wireless communication as object information. This allows terminal 2021 to know the AP's location in advance, and then direct the beam's direction towards the object information to initiate communication. As a result, improvements in communication quality can be achieved, such as increased transmission speed, reduced time to initiate communication, and extended communication period.
[0257] The object information may include information indicating the type of object corresponding to the object information. Furthermore, the object information may include information indicating the processing that terminal 2021 should perform if terminal 2021 is located within a real-space region corresponding to the three-dimensional data position of the object information, or within a predetermined distance from that region.
[0258] Object information may be provided from a server different from the server providing the 3D data. When object information is provided separately from the 3D data, object groups containing object information used in the same service may be provided as separate data depending on the type of service or equipment.
[0259] The three-dimensional data used in combination with object information may be either point cloud data from a WLD or feature point data from a SWLD.
[0260] (Embodiment 6) The following describes the octree representation and the voxel scan order. A volume is converted into an octree structure (octreeized) and then encoded. An octree structure consists of nodes and leaves. Each node has eight nodes or leaves, and each leaf has voxel (VXL) information. Figure 22 shows an example of the structure of a volume containing multiple voxels. Figure 23 shows an example of the volume shown in Figure 22 converted into an octree structure. Here, among the leaves shown in Figure 23, leaves 1, 2, and 3 represent the voxels VXL1, VXL2, and VXL3 shown in Figure 22, respectively, and represent a VXL containing a point cloud (hereinafter referred to as effective VXL).
[0261] An octree is represented, for example, by a binary sequence of 0s and 1s. For example, if nodes or valid VXLs are assigned the value 1 and all others the value 0, then each node and leaf is assigned the binary sequence shown in Figure 23. This binary sequence is then scanned according to the breadth-first or depth-first scan order. For example, if scanned in breadth-first order, the binary sequence shown in Figure 24A is obtained. If scanned in depth-first order, the binary sequence shown in Figure 24B is obtained. The binary sequence obtained by this scan is then encoded by entropy coding to reduce its information content.
[0262] Next, we will explain the depth information in octree representations. In octree representations, the depth is used to control the level of granularity to which the point cloud information contained within a volume is retained. Setting a high depth allows for the reproduction of point cloud information at a finer level, but increases the amount of data required to represent nodes and leaves. Conversely, setting a low depth reduces the amount of data, but multiple point clouds with different locations and colors are treated as being at the same location and with the same color, resulting in the loss of information that the original point cloud information contained in the data.
[0263] For example, Figure 25 shows an example where the octree with depth=2 shown in Figure 23 is represented by an octree with depth=1. The octree shown in Figure 25 has less data than the octree shown in Figure 23. In other words, the octree shown in Figure 25 has fewer bits after binary conversion than the octree shown in Figure 25. Here, leaf 1 and leaf 2 shown in Figure 23 are represented by leaf 1 shown in Figure 24. In other words, the information that leaf 1 and leaf 2 shown in Figure 23 were in different positions is lost.
[0264] Figure 26 shows the volume corresponding to the octree shown in Figure 25. VXL1 and VXL2 shown in Figure 22 correspond to VXL12 shown in Figure 26. In this case, the three-dimensional data encoding device generates the color information of VXL12 shown in Figure 26 from the color information of VXL1 and VXL2 shown in Figure 22. For example, the three-dimensional data encoding device calculates the average value, median value, or weighted average value of the color information of VXL1 and VXL2 as the color information of VXL12. In this way, the three-dimensional data encoding device may control the reduction of data volume by changing the depth of the octree.
[0265] The three-dimensional data encoding device may set the depth information of the octree in units of worlds, spaces, or volumes. The device may also add the depth information to the world header information, space header information, or volume header information. Furthermore, the same value may be used for depth information across all worlds, spaces, and volumes at different time points. In this case, the device may add the depth information to the header information that manages all worlds at all time points.
[0266] (Embodiment 7) Three-dimensional point cloud information includes geometry and attribute information. Geometry includes coordinates (x, y, and z coordinates) relative to a given point. When encoding geometry, instead of directly encoding the coordinates of each three-dimensional point, a method is used to reduce the amount of encoding by representing the position of each three-dimensional point using an octave tree and encoding the information in the octave tree.
[0267] On the other hand, attribute information includes information such as color information (RGB, YUV, etc.), reflectance, and normal vector for each three-dimensional point. For example, a three-dimensional data encoding device can encode attribute information using a different encoding method than that used for positional information.
[0268] This embodiment describes a method for encoding attribute information. In this embodiment, integer values are used as the values of the attribute information. For example, if each color component of the RGB or YUV color information is 8-bit precision, each color component can take an integer value between 0 and 255. If the reflectance value is 10-bit precision, the reflectance value can take an integer value between 0 and 1023. If the bit precision of the attribute information is decimal precision, the three-dimensional data encoding device may multiply the attribute information value by a scale value and then round it to an integer value. The three-dimensional data encoding device may also add this scale value to the bitstream header, etc.
[0269] One possible method for encoding attribute information of a three-dimensional point is to calculate a predicted value for the attribute information of the three-dimensional point and encode the difference (prediction residual) between the original attribute information value and the predicted value. For example, if the attribute information value of a three-dimensional point p is Ap and the predicted value is Pp, the three-dimensional data encoding device encodes the absolute difference Diffp = |Ap - Pp|. In this case, if the predicted value Pp can be generated with high accuracy, the value of the absolute difference Diffp will become smaller. Therefore, for example, the amount of encoding can be reduced by entropy encoding the absolute difference Diffp using an encoding table where the number of generated bits decreases as the value becomes smaller.
[0270] One possible method for generating predicted attribute information is to use the attribute information of a reference three-dimensional point, which is another three-dimensional point located around the target three-dimensional point to be encoded. Here, a reference three-dimensional point is a three-dimensional point located within a predetermined distance range from the target three-dimensional point. For example, if there is a target three-dimensional point p=(x1,y1,z1) and a three-dimensional point q=(x2,y2,z2), the three-dimensional data encoding device calculates the Euclidean distance d(p,q) between the three-dimensional point p and the three-dimensional point q as shown in (Equation A1).
[0271]
number
[0272] The 3D data encoding device determines that the position of 3D point q is close to the position of target 3D point p if the Euclidean distance d(p, q) is smaller than a predetermined threshold THd, and decides to use the attribute information value of 3D point q to generate the predicted attribute information value of target 3D point p. Note that the distance calculation method may be other; for example, the Mahalanobis distance may be used. Furthermore, the 3D data encoding device may decide not to use 3D points outside a predetermined distance range from the target 3D point in the prediction process. For example, if a 3D point r exists and the distance d(p, r) between target 3D p and 3D point r is greater than or equal to the threshold THd, the 3D data encoding device may decide not to use 3D point r in the prediction. Note that the 3D data encoding device may add information indicating the threshold THd to the bitstream header, etc.
[0273] Figure 27 shows an example of a three-dimensional point. In this example, the distance d(p, q) between the target three-dimensional point p and the three-dimensional point q is smaller than the threshold THd. Therefore, the three-dimensional data encoding device determines that the three-dimensional point q is the reference three-dimensional point of the target three-dimensional point p, and decides to use the value of the attribute information Aq of the three-dimensional point q to generate the predicted value Pp of the attribute information Ap of the target three-dimensional point p.
[0274] On the other hand, the distance d(p,r) between the target three-dimensional point p and the three-dimensional point r is greater than or equal to the threshold THd. Therefore, the three-dimensional data encoding device determines that the three-dimensional point r is not a reference three-dimensional point of the target three-dimensional point p, and determines that it will not use the value of the attribute information Ar of the three-dimensional point r to generate the predicted value Pp of the attribute information Ap of the target three-dimensional point p.
[0275] Furthermore, when a three-dimensional data encoding device encodes the attribute information of a target three-dimensional point using predicted values, it uses a three-dimensional point whose attribute information has already been encoded and decoded as a reference three-dimensional point. Similarly, when a three-dimensional data decoding device decodes the attribute information of a target three-dimensional point to be decoded using predicted values, it uses a three-dimensional point whose attribute information has already been decoded as a reference three-dimensional point. This allows the same predicted values to be generated during encoding and decoding, so that the bitstream of three-dimensional points generated during encoding can be correctly decoded on the decoding side.
[0276] Furthermore, when encoding attribute information of three-dimensional points, it is conceivable to classify each three-dimensional point into multiple levels using its positional information before encoding. Here, each classified level is called LoD (Level of Detail). The method for generating LoDs will be explained using Figure 28.
[0277] First, the 3D data encoding device selects an initial point a0 and assigns it to LoD0. Next, the 3D data encoding device extracts point a1 whose distance from point a0 is greater than the LoD0 threshold Thres_LoD[0] and assigns it to LoD0. Next, the 3D data encoding device extracts point a2 whose distance from point a1 is greater than the LoD0 threshold Thres_LoD[0] and assigns it to LoD0. In this way, the 3D data encoding device configures LoD0 such that the distance between each point in LoD0 is greater than the threshold Thres_LoD[0].
[0278] Next, the 3D data encoding device selects point b0, which has not yet been assigned a Level of Direction (LoD), and assigns it to LoD1. Next, the 3D data encoding device extracts point b1, which is farther from point b0 than the LoD1 threshold Thres_LoD[1] and has not yet been assigned a Level of Direction (LoD), and assigns it to LoD1. Next, the 3D data encoding device extracts point b2, which is farther from point b1 than the LoD1 threshold Thres_LoD[1] and has not yet been assigned a Level of Direction (LoD), and assigns it to LoD1. In this way, the 3D data encoding device configures LoD1 such that the distance between each point in LoD1 is greater than the threshold Thres_LoD[1].
[0279] Next, the 3D data encoding device selects point c0, which has not yet been assigned a LoD, and assigns it to LoD2. Next, the 3D data encoding device extracts point c1, which is farther from point c0 than the LoD2 threshold Thres_LoD[2] and has not yet been assigned a LoD, and assigns it to LoD2. Next, the 3D data encoding device extracts point c2, which is farther from point c1 than the LoD2 threshold Thres_LoD[2] and has not yet been assigned a LoD, and assigns it to LoD2. In this way, the 3D data encoding device configures LoD2 such that the distance between each point in LoD2 is greater than the threshold Thres_LoD[2]. For example, as shown in Figure 29, the thresholds Thres_LoD[0], Thres_LoD[1], and Thres_LoD[2] for each LoD are set.
[0280] Furthermore, the three-dimensional data encoding device may add information indicating the threshold for each LoD to the bitstream header, etc. For example, in the example shown in Figure 29, the three-dimensional data encoding device may add the thresholds Thres_LoD[0], Thres_LoD[1], and Thres_LoD[2] to the header.
[0281] Furthermore, the three-dimensional data encoding device may assign all three-dimensional points that have not yet been assigned a LoD to the lowest layer of the LoD. In this case, the three-dimensional data encoding device can reduce the amount of code in the header by not adding the threshold of the lowest layer of the LoD to the header. For example, in the example shown in Figure 29, the three-dimensional data encoding device adds the thresholds Thres_LoD[0] and Thres_LoD[1] to the header, but does not add Thres_LoD[2] to the header. In this case, the three-dimensional data decoding device may estimate the value of Thres_LoD[2] to be 0. The three-dimensional data encoding device may also add the number of LoD layers to the header. This allows the three-dimensional data decoding device to determine the lowest layer of the LoD using the number of LoD layers.
[0282] Furthermore, by setting the threshold values for each layer of the LoD to be larger for higher layers, as shown in Figure 29, higher layers (layers closer to LoD0) become sparse point groups with greater distances between 3D points, while lower layers become dense point groups with closer distances between 3D points. In the example shown in Figure 29, LoD0 is the top layer.
[0283] Furthermore, the method for selecting the initial three-dimensional points when setting each LoD may depend on the coding order during positional information coding. For example, the three-dimensional data encoding device selects the first three-dimensional point coded during positional information coding as the initial point a0 of LoD0, and then uses initial point a0 as the base point to select points a1 and a2 to construct LoD0. Then, the three-dimensional data encoding device may select the three-dimensional point that was coded earliest among the three-dimensional points not belonging to LoD0 as the initial point b0 of LoD1. In other words, the three-dimensional data encoding device may select the three-dimensional point that was coded earliest among the three-dimensional points not belonging to the upper layers of LoDn (LoD0 to LoDn-1) as the initial point n0 of LoDn. As a result, the three-dimensional data decoding device can construct the same LoD as during coding by using the same initial point selection method during decoding, and thus can decode the bitstream appropriately. Specifically, the three-dimensional data decoding device selects the three-dimensional point that was coded earliest among the three-dimensional points not belonging to the upper layers of LoDn as the initial point n0 of LoDn.
[0284] The following describes a method for generating predicted attribute information of three-dimensional points using LoD information. For example, when a three-dimensional data encoding device encodes three-dimensional points sequentially starting from those contained in LoD0, it generates the target three-dimensional points contained in LoD1 using the encoded and decoded (hereinafter simply referred to as "encoded") attribute information contained in LoD0 and LoD1. In this way, the three-dimensional data encoding device generates predicted attribute information of three-dimensional points contained in LoDn using the encoded attribute information contained in LoDn' (n'<=n). In other words, the three-dimensional data encoding device does not use the attribute information of three-dimensional points contained in lower layers of LoDn to calculate predicted attribute information of three-dimensional points contained in LoDn.
[0285] For example, a three-dimensional data encoding device generates predicted attribute values for a three-dimensional point by calculating the average of the attribute values of N or fewer encoded three-dimensional points surrounding the target three-dimensional point to be encoded. Alternatively, the three-dimensional data encoding device may add the value of N to the bitstream header, etc. Furthermore, the three-dimensional data encoding device may change the value of N for each three-dimensional point and add the value of N to each three-dimensional point. This allows for the selection of an appropriate N for each three-dimensional point, thereby improving the accuracy of the predicted values and reducing the prediction residual. Alternatively, the three-dimensional data encoding device may add the value of N to the bitstream header and fix the value of N within the bitstream. This eliminates the need to encode or decode the value of N for each three-dimensional point, thus reducing processing load. Finally, the three-dimensional data encoding device may encode the value of N separately for each Level of Data (LoD). This allows for the selection of an appropriate N for each LoD, improving encoding efficiency.
[0286] Alternatively, the three-dimensional data encoding device may calculate the predicted value of the attribute information of a three-dimensional point by using the weighted average of the attribute information of N surrounding encoded three-dimensional points. For example, the three-dimensional data encoding device calculates weights using the distance information between the target three-dimensional point and the N surrounding three-dimensional points.
[0287] When a 3D data encoding device encodes the value of N separately for each Level of Data (LoD), for example, it may set a larger value of N for higher layers of the LoD and a smaller value of N for lower layers. In the higher layers of the LoD, the distance between the 3D points is greater, so setting a larger value of N and selecting multiple surrounding 3D points for averaging may improve prediction accuracy. Conversely, in the lower layers of the LoD, the distance between the 3D points is smaller, so setting a smaller value of N allows for efficient prediction while reducing the processing load of averaging.
[0288] Figure 30 shows an example of attribute information used for prediction values. As described above, the prediction value of point P included in LoDN is generated using the encoded surrounding points P' included in LoDN'(N'<=N). Here, the surrounding points P' are selected based on their distance from point P. For example, the prediction value of the attribute information of point b2 shown in Figure 30 is generated using the attribute information of points a0, a1, a2, b0, and b1.
[0289] The surrounding points selected change depending on the value of N mentioned above. For example, if N=5, a0, a1, a2, b0, and b1 are selected as surrounding points of point b2. If N=4, points a0, a1, a2, and b1 are selected based on distance information.
[0290] The predicted value is calculated using a distance-dependent weighted average. For example, in the example shown in Figure 30, the predicted value a2p for point a2 is calculated using a weighted average of the attribute information of points a0 and a1, as shown in (Equation A2) and (Equation A3). i This is the attribute information value of point ai.
[0291]
number
[0292] Furthermore, the predicted value b2p for point b2 is calculated by the weighted average of the attribute information of points a0, a1, a2, b0, and b1, as shown in (Equations A4) to (Equations A6). i This is the attribute information value of point bi.
[0293]
number
[0294] Furthermore, the three-dimensional data encoding device may calculate the difference (prediction residual) between the attribute information value of a three-dimensional point and the predicted value generated from surrounding points, and then quantize the calculated prediction residual. For example, the three-dimensional data encoding device performs quantization by dividing the prediction residual by a quantization scale (also called a quantization step). In this case, the smaller the quantization scale, the smaller the error that may occur due to quantization (quantization error). Conversely, the larger the quantization scale, the larger the quantization error.
[0295] Furthermore, the 3D data encoding device may change the quantization scale used for each Level of Data (LoD). For example, the 3D data encoding device may use a smaller quantization scale in the upper layers and a larger quantization scale in the lower layers. Since the attribute information values of 3D points belonging to the upper layers may be used as predicted values for the attribute information of 3D points belonging to the lower layers, encoding efficiency can be improved by reducing the quantization scale in the upper layers to suppress quantization errors that may occur in the upper layers and increasing the accuracy of the predicted values. The 3D data encoding device may also add the quantization scale used for each LoD to the header or other components. This allows the 3D data decoding device to correctly decode the quantization scale, thus enabling proper decoding of the bitstream.
[0296] Furthermore, the three-dimensional data encoding device may convert the signed integer value (signed quantized value), which is the prediction residual after quantization, into an unsigned integer value (unsigned quantized value). This eliminates the need to consider the occurrence of negative integers when entropy encoding the prediction residual. Note that the three-dimensional data encoding device does not necessarily need to convert the signed integer value into an unsigned integer value; for example, the sign bit may be entropy encoded separately.
[0297] The predicted residual is calculated by subtracting the predicted value from the original value. For example, the predicted residual a2r for point a2 is calculated by subtracting the predicted value a2p for point a2 from the attribute information value A2 for point a2, as shown in (Equation A7). The predicted residual b2r for point b2 is calculated by subtracting the predicted value b2p for point b2 from the attribute information value B2 for point b2, as shown in (Equation A8).
[0298] a²r = A² - a²p ... (Equation A7) b²r = B² - b²p ... (Equation A8)
[0299] Furthermore, the predicted residuals are quantized by dividing them by the QS (Quantization Step). For example, the quantized value a2q of point a2 is calculated by (Equation A9). The quantized value b2q of point b2 is calculated by (Equation A10). Here, QS_LoD0 is the QS for LoD0, and QS_LoD1 is the QS for LoD1. In other words, the QS may be changed according to the LoD.
[0300] a2q = a2r / QS_LoD0 ... (Equation A9) b2q=b2r / QS_LoD1 (Formula A10)
[0301] Furthermore, the three-dimensional data encoding device converts the quantized value, which is a signed integer, into an unsigned integer as follows: If the signed integer value a2q is less than 0, the three-dimensional data encoding device sets the unsigned integer value a2u to -1-(2×a2q). If the signed integer value a2q is 0 or greater, the three-dimensional data encoding device sets the unsigned integer value a2u to 2×a2q.
[0302] Similarly, the three-dimensional data encoding device sets the unsigned integer value b2u to -1-(2×b2q) if the signed integer value b2q is less than 0. The three-dimensional data encoding device sets the unsigned integer value b2u to 2×b2q if the signed integer value b2q is 0 or greater.
[0303] Furthermore, the three-dimensional data encoding device may encode the predicted residuals (unsigned integer values) after quantization using entropy coding. For example, the unsigned integer values may be binarized and then binary arithmetic coding may be applied.
[0304] In this case, the three-dimensional data encoding device may switch the binarization method depending on the value of the predicted residual. For example, if the predicted residual pu is smaller than the threshold R_TH, the three-dimensional data encoding device binarizes the predicted residual pu with a fixed number of bits required to represent the threshold R_TH. If the predicted residual pu is greater than or equal to the threshold R_TH, the three-dimensional data encoding device binarizes the binarized data of the threshold R_TH and the value of (pu-R_TH) using an exponential Golomb or the like.
[0305] For example, if the threshold R_TH is 63 and the predicted residual pu is less than 63, the three-dimensional data encoding device will binarize the predicted residual pu using 6 bits. If the predicted residual pu is 63 or greater, the three-dimensional data encoding device will perform arithmetic encoding by binarizing the binary data of the threshold R_TH (111111) and (pu-63) using an exponential golomb.
[0306] In a more specific example, if the predicted residual pu is 32, the three-dimensional data encoding device generates 6 bits of binary data (100000) and arithmetically encodes this bit sequence. Similarly, if the predicted residual pu is 66, the three-dimensional data encoding device generates binary data of the threshold R_TH (111111) and a bit sequence (00100) representing the value 3 (66-63) in exponential golombs, and arithmetically encodes this bit sequence (111111+00100).
[0307] In this way, the three-dimensional data encoding device can encode data while suppressing a rapid increase in the number of binarized bits when the predicted residual becomes large, by switching the binarization method according to the magnitude of the predicted residual. The three-dimensional data encoding device may also add a threshold R_TH to the bitstream header or the like.
[0308] For example, when encoding is performed at a high bit rate, i.e., when the quantization scale is small, the quantization error is small and the prediction accuracy is high, which may result in a smaller prediction residual. Therefore, in this case, the three-dimensional data encoding device sets a large threshold R_TH. This reduces the likelihood of encoding binarized data with a threshold R_TH, improving encoding efficiency. Conversely, when encoding is performed at a low bit rate, i.e., when the quantization scale is large, the quantization error is large and the prediction accuracy is poor, which may result in a larger prediction residual. Therefore, in this case, the three-dimensional data encoding device sets a small threshold R_TH. This prevents a rapid increase in the bit length of the binarized data.
[0309] Furthermore, the 3D data encoding device may switch the threshold R_TH for each Level of Data (LoD) and add the LoD threshold R_TH to the header, etc. In other words, the 3D data encoding device may switch the binarization method for each LoD. For example, in the upper layers, the distance between 3D points is large, so the prediction accuracy may be poor and as a result the prediction residual may be large. Therefore, the 3D data encoding device can prevent a rapid increase in the bit length of the binarized data by setting a small threshold R_TH for the upper layers. Also, in the lower layers, the distance between 3D points is small, so the prediction accuracy may be high and as a result the prediction residual may be small. Therefore, the 3D data encoding device can improve encoding efficiency by setting a large threshold R_TH for each layer.
[0310] Figure 31 is a diagram showing an example of an exponential Golomb code, illustrating the relationship between the value before binarization (multi-level) and the bit after binarization (code). Note that the 0s and 1s shown in Figure 31 may also be inverted.
[0311] Furthermore, the three-dimensional data encoding device applies arithmetic coding to the binarized data of the prediction residuals. This improves coding efficiency. Note that when applying arithmetic coding, the probability trends of the occurrence of 0 and 1 for each bit may differ between the n-bit code (the n-bit binarized portion of the binarized data) and the remaining code (the portion binarized using exponential golombs). Therefore, the three-dimensional data encoding device may switch the method of applying arithmetic coding between the n-bit code and the remaining code.
[0312] For example, a three-dimensional data encoding device performs arithmetic encoding of an n-bit code using a different encoding table (probability table) for each bit. In this case, the three-dimensional data encoding device may change the number of encoding tables used for each bit. For example, the three-dimensional data encoding device uses one encoding table to perform arithmetic encoding of the first bit b0 of an n-bit code. The three-dimensional data encoding device then uses two encoding tables for the next bit b1. Furthermore, the three-dimensional data encoding device switches the encoding table used for arithmetic encoding of bit b1 depending on the value of b0 (0 or 1). Similarly, the three-dimensional data encoding device uses four encoding tables for the next bit b2. Furthermore, the three-dimensional data encoding device switches the encoding table used for arithmetic encoding of bit b2 depending on the values of b0 and b1 (0 to 3).
[0313] Thus, when the three-dimensional data encoding device arithmetically encodes each bit bn-1 of an n-bit code, 2 n-1 The system uses a set of encoding tables. Furthermore, the three-dimensional data encoding device switches the encoding table used depending on the value (generation pattern) of the bits prior to bn-1. This allows the three-dimensional data encoding device to use the appropriate encoding table for each bit, thereby improving encoding efficiency.
[0314] Note that the three-dimensional data encoding device may reduce the number of encoding tables used for each bit. For example, when arithmetically encoding each bit bn-1, the three-dimensional data encoding device may switch between 2 m encoding tables according to the value (occurrence pattern) of m bits (m<n-1) preceding bn-1. This makes it possible to improve encoding efficiency while reducing the number of encoding tables used for each bit. Note that the three-dimensional data encoding device may update the occurrence probabilities of 0 and 1 in each encoding table according to the actually occurring value of binarized data. Furthermore, the three-dimensional data encoding device may fix the occurrence probabilities of 0 and 1 in encoding tables for some bits. This makes it possible to suppress the number of updates to the occurrence probabilities, thereby reducing the processing amount.
[0315] For example, when an n-bit code is b0b1b2…bn-1, one encoding table (CTb0) is provided for b0. Two encoding tables (CTb10, CTb11) are provided for b1, and the encoding table to be used is switched according to the value (0 to 1) of b0. Four encoding tables (CTb20, CTb21, CTb22, CTb23) are provided for b2, and the encoding table to be used is switched according to the values (0 to 3) of b0 and b1. The encoding table for bn-1 has 2 n-1 encoding tables (CTbn0, CTbn1, …, CTbn(2 n-1 -1)), and the encoding table to be used is switched according to the value (0 to 2 n-1 -1) of b0b1…bn-2.
[0316] Note that the three-dimensional data encoding device may apply m-ary arithmetic encoding (m=2 n -1) that sets values from 0 to 2 n n-1 for an n-bit code without binarization. Furthermore, when the three-dimensional data encoding device arithmetically encodes an n-bit code using m-ary, the three-dimensional data decoding device may also restore the n-bit code through m-ary arithmetic decoding.
[0317] Figure 32 illustrates the processing when the remaining code is an exponential Golomb code, for example. The remaining code, which is the part binarized using exponential Golomb, includes a prefix part and a suffix part, as shown in Figure 32. For example, a three-dimensional data encoding device switches the encoding table between the prefix part and the suffix part. That is, the three-dimensional data encoding device arithmetically encodes each bit in the prefix part using the encoding table for the prefix, and arithmetically encodes each bit in the suffix part using the encoding table for the suffix.
[0318] The three-dimensional data encoding device may update the probability of occurrence of 0 and 1 in each encoding table according to the actual values of the binarized data that have occurred. Alternatively, the three-dimensional data encoding device may fix the probability of occurrence of 0 and 1 in either encoding table. This reduces the number of times the probability of occurrence is updated, thereby reducing the processing load. For example, the three-dimensional data encoding device may update the probability of occurrence for the prefix part and fix the probability of occurrence for the suffix part.
[0319] Furthermore, the three-dimensional data encoding device decodes the predicted residual after quantization by inverse quantization and reconstruction, and uses the decoded value, which is the decoded predicted residual, for predictions beyond the three-dimensional point to be encoded. Specifically, the three-dimensional data encoding device calculates the inverse quantized value by multiplying the predicted residual (quantized value) after quantization by the quantization scale, and obtains the decoded value (reconstructed value) by adding the inverse quantized value and the predicted value.
[0320] For example, the inverse quantization value a2iq of point a2 is calculated using the quantization value a2q of point a2 by (Equation A11). The inverse quantization value b2iq of point b2 is calculated using the quantization value b2q of point b2 by (Equation A12). Here, QS_LoD0 is the QS for LoD0, and QS_LoD1 is the QS for LoD1. In other words, the QS may be changed according to the LoD.
[0321] a2iq=a2q×QS_LoD0 (Formula A11) b2iq=b2q×QS_LoD1 (Formula A12)
[0322] For example, the decoded value a2rec of point a2 is calculated by adding the predicted value a2p of point a2 to the inverse quantized value a2iq of point a2, as shown in (Equation A13). The decoded value b2rec of point b2 is calculated by adding the predicted value b2p of point b2 to the inverse quantized value b2iq of point b2, as shown in (Equation A14).
[0323] a2rec=a2iq+a2p (formula A13) b2rec=b2iq+b2p (formula A14)
[0324] The following describes an example of bitstream syntax according to this embodiment. Figure 33 shows an example of attribute header syntax according to this embodiment. The attribute header is header information for attribute information. As shown in Figure 33, the attribute header includes hierarchy number information (NumLoD), three-dimensional point number information (NumOfPoint[i]), hierarchy threshold (Thres_Lod[i]), surrounding point number information (NumNeighorPoint[i]), prediction threshold (THd[i]), quantization scale (QS[i]), and binarization threshold (R_TH[i]).
[0325] The number of levels (NumLoD) indicates the number of levels of the Level of Data (LoD) used.
[0326] The three-dimensional point number information (NumOfPoint[i]) indicates the number of three-dimensional points belonging to hierarchy i. The three-dimensional data encoding device may also add total three-dimensional point number information (AllNumOfPoint), indicating the total number of three-dimensional points, to a separate header. In this case, the three-dimensional data encoding device does not need to add NumOfPoint[NumLoD-1], indicating the number of three-dimensional points belonging to the lowest layer, to the header. In this case, the three-dimensional data decoding device can calculate NumOfPoint[NumLoD-1] using (Equation A15). This reduces the amount of code in the header.
[0327] [Numerical]
[0328] The hierarchical threshold (Thres_Lod[i]) is a threshold used for setting level i. The three-dimensional data encoding device and the three-dimensional data decoding device configure LoDi such that the distance between each point in LoDi is greater than the threshold Thres_LoD[i]. In addition, the three-dimensional data encoding device does not need to add the value of Thres_Lod[NumLoD-1] (the lowest layer) to the header. In this case, the three-dimensional data decoding device estimates the value of Thres_Lod[NumLoD-1] as 0. This can reduce the coding amount of the header.
[0329] The surrounding point count information (NumNeighorPoint[i]) indicates the upper limit of the number of surrounding points used for generating predicted values of three-dimensional points belonging to level i. When the number of surrounding points M is less than NumNeighorPoint[i] (M<NumNeighorPoint[i]), the three-dimensional data encoding device may calculate the predicted value using M surrounding points. In addition, when it is not necessary to set different values of NumNeighorPoint[i] for each LoD, the three-dimensional data encoding device may add one piece of surrounding point count information (NumNeighorPoint) used for all LoDs to the header.
[0330] The prediction threshold (THd[i]) indicates the upper limit of the distance between surrounding three-dimensional points used for predicting a target three-dimensional point to be encoded or decoded at level i and the target three-dimensional point. The three-dimensional data encoding device and the three-dimensional data decoding device do not use three-dimensional points whose distance from the target three-dimensional point is greater than THd[i] for prediction. Note that when it is not necessary to set different values of THd[i] for each LoD, the three-dimensional data encoding device may add one prediction threshold (THd) used for all LoDs to the header.
[0331] The quantization scale (QS[i]) indicates the quantization scale used in quantization and inverse quantization for level i.
[0332] The binarization threshold (R_TH[i]) is a threshold used to switch the binarization method for the predicted residuals of three-dimensional points belonging to hierarchy i. For example, if the predicted residual is less than the threshold R_TH, the three-dimensional data encoder binarizes the predicted residual pu with a fixed number of bits. If the predicted residual is greater than or equal to the threshold R_TH, it binarizes the binarized data of threshold R_TH and the value of (pu - R_TH) using exponential golomb. If it is not necessary to switch the value of R_TH[i] for each LoD, the three-dimensional data encoder may add a single binarization threshold (R_TH) used for all LoDs to the header.
[0333] Note that R_TH[i] may be the maximum value that can be represented by n bits. For example, R_TH is 63 for 6 bits and 255 for 8 bits. Alternatively, instead of encoding the maximum value that can be represented by n bits as the binarization threshold, the 3D data encoder may encode the number of bits. For example, the 3D data encoder may add the value 6 to the header when R_TH[i]=63, and the value 8 when R_TH[i]=255. Alternatively, the 3D data encoder may define the minimum number of bits that represent R_TH[i] (minimum number of bits) and add the relative number of bits from the minimum value to the header. For example, the 3D data encoder may add the value 0 to the header when R_TH[i]=63 and the minimum number of bits is 6, and add the value 2 to the header when R_TH[i]=255 and the minimum number of bits is 6.
[0334] Furthermore, the three-dimensional data encoding device may entropy encode at least one of NumLoD, Thres_Lod[i], NumNeighborPoint[i], THd[i], QS[i], and R_TH[i] and add it to the header. For example, the three-dimensional data encoding device may binarize each value and then arithmetic encode it. Alternatively, the three-dimensional data encoding device may encode each value with a fixed length to reduce processing load.
[0335] Furthermore, the three-dimensional data encoding device does not need to include at least one of NumLoD, Thres_Lod[i], NumNeighborPoint[i], THd[i], QS[i], and R_TH[i] in the header. For example, at least one of these values may be defined in the profile or level of a standard or similar specification. This can reduce the number of bits in the header.
[0336] Figure 34 shows an example of the syntax of attribute data according to this embodiment. This attribute data includes encoded data of attribute information for multiple three-dimensional points. As shown in Figure 34, the attribute data includes an n-bit code and a remaining code.
[0337] An n-bit code is the encoded data or a portion thereof of the predicted residual of the attribute information value. The bit length of the n-bit code depends on the value of R_TH[i]. For example, if the value of R_TH[i] is 63, the n-bit code is 6 bits, and if the value of R_TH[i] is 255, the n-bit code is 8 bits.
[0338] The remaining code is the coded data of the predicted residual of the attribute information value, coded using exponential golomb. This remaining code is coded or coded when the n-bit code is the same as R_TH[i]. The three-dimensional data decoder decodes the predicted residual by adding the value of the n-bit code and the value of the remaining code. If the n-bit code is not the same as R_TH[i], the remaining code does not need to be coded or coded.
[0339] The following describes the processing flow in the three-dimensional data encoding device. Figure 35 is a flowchart of the three-dimensional data encoding process performed by the three-dimensional data encoding device.
[0340] First, the three-dimensional data encoding device encodes the positional information (geometry) (S3001). For example, three-dimensional data encoding is performed using an octave representation.
[0341] The three-dimensional data encoding device reassigns the attribute information of the original three-dimensional point to the changed three-dimensional point if the position of the three-dimensional point changes due to quantization or the like after encoding the position information (S3002). For example, the three-dimensional data encoding device performs the reassignment by interpolating the value of the attribute information according to the amount of change in position. For example, the three-dimensional data encoding device detects N three-dimensional points that are close to the changed three-dimensional position and performs a weighted average of the attribute information values of the N three-dimensional points. For example, in the weighted average, the three-dimensional data encoding device determines the weights based on the distance from the changed three-dimensional position to each of the N three-dimensional points. Then, the three-dimensional data encoding device determines the value obtained by the weighted average as the attribute information value of the changed three-dimensional point. Furthermore, if two or more three-dimensional points change to the same three-dimensional position due to quantization or the like, the three-dimensional data encoding device may assign the average value of the attribute information of the two or more three-dimensional points before the change as the attribute information value of the changed three-dimensional point.
[0342] Next, the three-dimensional data encoding device encodes the reassigned attribute information (Attribute) (S3003). For example, if the three-dimensional data encoding device encodes multiple types of attribute information, it may encode the multiple types of attribute information sequentially. For example, if the three-dimensional data encoding device encodes color and reflectance as attribute information, it may generate a bitstream in which the encoded result of reflectance is appended after the encoded result of color. Note that the order of the multiple encoded results of attribute information appended to the bitstream is not limited to this order and may be any order.
[0343] Furthermore, the three-dimensional data encoding device may add information to the header or elsewhere indicating the starting location of the encoded data for each attribute information within the bitstream. This allows the three-dimensional data decoding device to selectively decode the attribute information that needs to be decoded, thus omitting the decoding process for attribute information that does not need to be decoded. Therefore, the processing load of the three-dimensional data decoding device can be reduced. In addition, the three-dimensional data encoding device may encode multiple types of attribute information in parallel and integrate the encoding results into a single bitstream. This allows the three-dimensional data encoding device to encode multiple types of attribute information at high speed.
[0344] Figure 36 is a flowchart of the attribute information encoding process (S3003). First, the three-dimensional data encoding device sets the Level of Data (LoD) (S3011). In other words, the three-dimensional data encoding device assigns each three-dimensional point to one of several LoDs.
[0345] Next, the three-dimensional data encoding device starts a loop for each Level of Data (LoD) (S3012). In other words, the three-dimensional data encoding device repeatedly performs the processes in steps S3013 to S3021 for each LoD.
[0346] Next, the three-dimensional data encoding device starts a loop for each three-dimensional point (S3013). In other words, the three-dimensional data encoding device repeats the process from steps S3014 to S3020 for each three-dimensional point.
[0347] First, the three-dimensional data encoding device searches for multiple surrounding points, which are three-dimensional points that exist around the target three-dimensional point to be processed, in order to calculate the predicted value of the target three-dimensional point (S3014). Next, the three-dimensional data encoding device calculates the weighted average of the attribute information values of the multiple surrounding points and sets the obtained value as the predicted value P (S3015). Next, the three-dimensional data encoding device calculates the prediction residual, which is the difference between the attribute information of the target three-dimensional point and the predicted value (S3016). Next, the three-dimensional data encoding device calculates the quantized value by quantizing the prediction residual (S3017). Next, the three-dimensional data encoding device arithmetically encodes the quantized value (S3018).
[0348] Furthermore, the three-dimensional data encoding device calculates the inverse quantized value by inverse quantizing the quantized value (S3019). Next, the three-dimensional data encoding device generates the decoded value by adding the predicted value to the inverse quantized value (S3020). Next, the three-dimensional data encoding device terminates the loop for three-dimensional points (S3021). Furthermore, the three-dimensional data encoding device terminates the loop for LoD (Line of Data) (S3022).
[0349] The following describes the three-dimensional data decoding process in a three-dimensional data decoding device that decodes the bitstream generated by the three-dimensional data encoding device described above.
[0350] The three-dimensional data decoder generates decoded binarized data by arithmetic decoding the binarized attribute information data within the bitstream generated by the three-dimensional data encoder in the same manner as the three-dimensional data encoder. If the three-dimensional data encoder switches the application method of arithmetic coding between the n-bit binarized portion (n-bit code) and the exponential golomb binarized portion (remaining code), the three-dimensional data decoder will perform decoding accordingly when applying arithmetic decoding.
[0351] For example, a three-dimensional data decoder performs arithmetic decoding of an n-bit code using a different coding table (decoding table) for each bit. In this case, the three-dimensional data decoder may change the number of coding tables used for each bit. For example, the first bit b0 of an n-bit code is decoded using one coding table. The three-dimensional data decoder then uses two coding tables for the next bit b1. Furthermore, the three-dimensional data decoder switches the coding table used for arithmetic decoding of bit b1 depending on the value of b0 (0 or 1). Similarly, the three-dimensional data decoder uses four coding tables for the next bit b2. Furthermore, the three-dimensional data decoder switches the coding table used for arithmetic decoding of bit b2 depending on the values of b0 and b1 (0 to 3).
[0352] As described above, when arithmetically decoding each bit bn-1 of an n-bit code, the three-dimensional data decoding device performs 2 n-1 encoding tables are used. Further, the three-dimensional data decoding device switches the encoding table to be used according to the values (occurrence patterns) of bits before bn-1. Accordingly, the three-dimensional data decoding device can properly decode a bit stream with improved encoding efficiency by using an appropriate encoding table for each bit.
[0353] Note that the three-dimensional data decoding device may reduce the number of encoding tables used for each bit. For example, when arithmetically decoding each bit bn-1, the three-dimensional data decoding device performs 2 in accordance with values (occurrence patterns) of m bits (m < n-1) before bn-1 m encoding tables may be switched. Accordingly, the three-dimensional data decoding device can properly decode a bit stream with improved encoding efficiency while suppressing the number of encoding tables used for each bit. Note that the three-dimensional data decoding device may update the occurrence probabilities of 0 and 1 in each encoding table according to the actually occurring values of binarized data. Further, the three-dimensional data decoding device may fix the occurrence probabilities of 0 and 1 in encoding tables for some bits. Accordingly, the number of updates of the occurrence probabilities can be suppressed, so the amount of processing can be reduced.
[0354] For example, when an n-bit code is b0b1b2...bn-1, the number of encoding tables for b0 is one (CTb0). The number of encoding tables for b1 is two (CTb10, CTb11). Further, the encoding table is switched according to the value (0 to 1) of b0. The number of encoding tables for b2 is four (CTb20, CTb21, CTb22, CTb23). Further, the encoding table is switched according to the values (0 to 3) of b0 and b1. The number of encoding tables for bn-1 is 2 n-1 (CTbn0, CTbn1, ..., CTbn(2 n-1 -1)). Further, the encoding table is switched according to the values (0 to 2 n-1 -1) of b0b1...bn-2.
[0355] Figure 37 illustrates, for example, the processing when the remaining code is an exponential Golomb code. The portion (remaining code) that the three-dimensional data encoding device has binarized and encoded using exponential Golomb includes a prefix section and a suffix section, as shown in Figure 37. For example, the three-dimensional data decoding device switches the encoding table between the prefix section and the suffix section. That is, the three-dimensional data decoding device arithmetically decodes each bit in the prefix section using the encoding table for the prefix, and arithmetically decodes each bit in the suffix section using the encoding table for the suffix.
[0356] The three-dimensional data decoder may update the probability of occurrence of 0 and 1 in each encoding table according to the value of the binarized data generated during decoding. Alternatively, the three-dimensional data decoder may fix the probability of occurrence of 0 and 1 in either encoding table. This reduces the number of updates to the occurrence probability, thereby reducing the processing load. For example, the three-dimensional data decoder may update the occurrence probability for the prefix section and fix the occurrence probability for the suffix section.
[0357] Furthermore, the three-dimensional data decoder decodes the quantized prediction residuals (unsigned integer values) by multi-leveling the binarized data of the arithmetic-decoded prediction residuals according to the encoding method used by the three-dimensional data encoding device. The three-dimensional data decoder first calculates the value of the decoded n-bit code by arithmetic decoding the binarized data of the n-bit code. Next, the three-dimensional data decoder compares the value of the n-bit code with the value of R_TH.
[0358] The three-dimensional data decoder determines that if the value of the n-bit code matches the value of R_TH, the next bit encoded with exponential golom exists, and arithmetic decoding is performed on the remaining code, which is the binarized data encoded with exponential golom. The three-dimensional data decoder then calculates the value of the remaining code from the decoded remaining code using a reverse lookup table that shows the relationship between the remaining code and its value. Figure 38 is a diagram showing an example of a reverse lookup table that shows the relationship between the remaining code and its value. Next, the three-dimensional data decoder obtains the multi-level quantized predicted residual by adding the obtained value of the remaining code to R_TH.
[0359] On the other hand, if the value of the n-bit code and the value of R_TH do not match (the value is smaller than R_TH), the three-dimensional data decoder uses the value of the n-bit code as the predicted residual after multi-level quantization. This allows the three-dimensional data decoder to appropriately decode the bitstream generated by the three-dimensional data encoding device by switching the binarization method according to the value of the predicted residual.
[0360] Furthermore, if the threshold R_TH is attached to the bitstream header, the three-dimensional data decoder may decode the value of the threshold R_TH from the header and switch the decoding method using the decoded value of the threshold R_TH. Also, if the threshold R_TH is attached to the header for each Level of Data (LoD), the three-dimensional data decoder may switch the decoding method using the decoded threshold R_TH for each LoD.
[0361] For example, if the threshold R_TH is 63 and the value of the decoded n-bit code is 63, the three-dimensional data decoder obtains the value of the remaining code by decoding the remaining code using exponential golomb. For example, in the example shown in Figure 38, the remaining code is 00100, and the value of the remaining code is obtained as 3. Next, the three-dimensional data decoder obtains the predicted residual value of 66 by adding the threshold R_TH value of 63 and the remaining code value of 3.
[0362] Furthermore, if the value of the decoded n-bit code is 32, the three-dimensional data decoder sets the value of the n-bit code (32) to the value of the predicted residual.
[0363] Furthermore, the three-dimensional data decoder converts the decoded quantized prediction residual from an unsigned integer value to a signed integer value, for example, by the reverse of the processing performed in the three-dimensional data encoding device. This allows the three-dimensional data decoder to properly decode bitstreams generated without considering the occurrence of negative integers when entropy coding the prediction residual. Note that the three-dimensional data decoder does not necessarily need to convert unsigned integer values to signed integer values; for example, when decoding a bitstream generated by separately entropy coding the sign bit, the sign bit may be decoded.
[0364] The three-dimensional data decoder generates decoded values by decoding the predicted residuals after quantization, which have been converted to signed integer values, through inverse quantization and reconstruction. The three-dimensional data decoder also uses the generated decoded values to predict the three-dimensional points and beyond that of the target of decoding. Specifically, the three-dimensional data decoder calculates the inverse quantized value by multiplying the predicted residuals after quantization by the decoded quantization scale, and then obtains the decoded value by adding the inverse quantized value and the predicted value.
[0365] The decoded unsigned integer value (unsigned quantized value) is converted to a signed integer value by the following process: If the LSB (least significant bit) of the decoded unsigned integer value a2u is 1, the three-dimensional data decoder sets the signed integer value a2q to -((a2u+1)>>1). If the LSB of the unsigned integer value a2u is not 1, the three-dimensional data decoder sets the signed integer value a2q to (a2u>>1).
[0366] Similarly, the 3D data decoder sets the signed integer b2q to -((b2u+1)>>1) if the LSB of the decoded unsigned integer b2u is 1. The 3D data decoder sets the signed integer b2q to (b2u>>1) if the LSB of the unsigned integer n2u is not 1.
[0367] Furthermore, the details of the inverse quantization and reconstruction process using the three-dimensional data decoding device are the same as those of the inverse quantization and reconstruction process using the three-dimensional data encoding device.
[0368] The following describes the processing flow in the three-dimensional data decoding device. Figure 39 is a flowchart of the three-dimensional data decoding process by the three-dimensional data decoding device. First, the three-dimensional data decoding device decodes the position information (geometry) from the bitstream (S3031). For example, the three-dimensional data decoding device performs decoding using an octave tree representation.
[0369] Next, the three-dimensional data decoder decodes attribute information from the bitstream (S3032). For example, if the three-dimensional data decoder decodes multiple types of attribute information, it may decode the multiple types of attribute information in order. For example, if the three-dimensional data decoder decodes color and reflectance as attribute information, it decodes the color encoding result and the reflectance encoding result in the order in which they are added to the bitstream. For example, if the reflectance encoding result is added after the color encoding result in the bitstream, the three-dimensional data decoder decodes the color encoding result, and then decodes the reflectance encoding result. The three-dimensional data decoder may decode the encoding results of the attribute information added to the bitstream in any order.
[0370] Furthermore, the three-dimensional data decoding device may obtain information indicating the start location of the encoded data for each attribute information within the bitstream by decoding the header, etc. This allows the three-dimensional data decoding device to selectively decode the attribute information that needs to be decoded, thus omitting the decoding process for attribute information that does not need to be decoded. Therefore, the processing load of the three-dimensional data decoding device can be reduced. In addition, the three-dimensional data decoding device may decode multiple types of attribute information in parallel and integrate the decoding results into a single three-dimensional point cloud. This allows the three-dimensional data decoding device to decode multiple types of attribute information at high speed.
[0371] Figure 40 is a flowchart of the attribute information decoding process (S3032). First, the three-dimensional data decoding device sets the Level of Direction (LoD) (S3041). That is, the three-dimensional data decoding device assigns each of the multiple three-dimensional points having decoded position information to one of the multiple LoDs. For example, this assignment method is the same as the assignment method used in the three-dimensional data encoding device.
[0372] Next, the three-dimensional data decoding device starts a loop for each Level of Data (LoD) (S3042). In other words, the three-dimensional data decoding device repeats the process from steps S3043 to S3049 for each LoD.
[0373] Next, the three-dimensional data decoding device starts a loop for each three-dimensional point (S3043). In other words, the three-dimensional data decoding device repeats the process from steps S3044 to S3048 for each three-dimensional point.
[0374] First, the three-dimensional data decoding device searches for multiple surrounding points, which are three-dimensional points that exist around the target three-dimensional point to be processed, in order to calculate the predicted value of the target three-dimensional point (S3044). Next, the three-dimensional data decoding device calculates the weighted average of the attribute information values of the multiple surrounding points and sets the obtained value as the predicted value P (S3045). These processes are the same as those performed in the three-dimensional data encoding device.
[0375] Next, the three-dimensional data decoding device arithmetically decodes the quantized values from the bitstream (S3046). The three-dimensional data decoding device also calculates the inverse quantized values by inverse quantizing the decoded quantized values (S3047). Next, the three-dimensional data decoding device generates the decoded values by adding the predicted values to the inverse quantized values (S3048). Next, the three-dimensional data decoding device terminates the loop for each three-dimensional point (S3049). The three-dimensional data decoding device also terminates the loop for each Line of Data (LoD) (S3050).
[0376] Next, the configuration of the three-dimensional data encoding device and the three-dimensional data decoding device according to this embodiment will be described. Figure 41 is a block diagram showing the configuration of the three-dimensional data encoding device 3000 according to this embodiment. This three-dimensional data encoding device 3000 includes a position information encoding unit 3001, an attribute information reassignment unit 3002, and an attribute information encoding unit 3003.
[0377] The attribute information encoding unit 3003 encodes the position information (geometry) of multiple three-dimensional points included in the input point cloud. The attribute information reassignment unit 3002 reassigns the attribute information values of multiple three-dimensional points included in the input point cloud using the encoding and decoding results of the position information. The attribute information encoding unit 3003 encodes the reassigned attribute information. The three-dimensional data encoding device 3000 generates a bitstream containing the encoded position information and encoded attribute information.
[0378] Figure 42 is a block diagram showing the configuration of a three-dimensional data decoding device 3010 according to this embodiment. This three-dimensional data decoding device 3010 includes a position information decoding unit 3011 and an attribute information decoding unit 3012.
[0379] The position information decoding unit 3011 decodes the position information (geometry) of multiple three-dimensional points from the bitstream. The attribute information decoding unit 3012 decodes the attribute information (attribute) of multiple three-dimensional points from the bitstream. The three-dimensional data decoding device 3010 generates an output point group by combining the decoded position information and the decoded attribute information.
[0380] (Embodiment 8) Below, we will describe another method for encoding attribute information of three-dimensional points, using RAHT (Region Adaptive Hierarchical Transform). Figure 43 is a diagram illustrating attribute information encoding using RAHT.
[0381] First, the three-dimensional data encoding device generates Morton codes based on the positional information of the three-dimensional points and sorts the attribute information of the three-dimensional points in order of their Morton codes. For example, the three-dimensional data encoding device may sort in ascending order of the Morton codes. Note that the sorting order is not limited to Morton code order; other orders may be used.
[0382] Next, the three-dimensional data encoding device generates high-frequency and low-frequency components of hierarchical L by applying a Haar transform to the attribute information of two adjacent three-dimensional points in Morton coding order. For example, the three-dimensional data encoding device may use a 2x2 matrix Haar transform. The generated high-frequency components are included in the coding coefficients as high-frequency components of hierarchical L, and the generated low-frequency components are used as input values for the higher hierarchical L+1 above hierarchical L.
[0383] The three-dimensional data encoding device generates high-frequency components of hierarchical L using attribute information of hierarchical L, and then proceeds to process hierarchical L+1. In the processing of hierarchical L+1, the three-dimensional data encoding device generates high-frequency and low-frequency components of hierarchical L+1 by applying the Haar transform to the two low-frequency components obtained by the Haar transform of the attribute information of hierarchical L. The generated high-frequency components are included in the encoding coefficients as high-frequency components of hierarchical L+1, and the generated low-frequency components are used as input values for the higher hierarchical L+2 above hierarchical L+1.
[0384] The three-dimensional data encoding device repeats this hierarchical processing until only one low-frequency component is input to a hierarchical level, at which point it determines that it has reached the highest level, Lmax. The three-dimensional data encoding device includes the low-frequency component of level Lmax-1, which is input to level Lmax, as an encoding coefficient. Then, it quantizes the values of the low-frequency or high-frequency components included in the encoding coefficient and encodes them using entropy encoding or the like.
[0385] Furthermore, if only one three-dimensional point exists as two adjacent three-dimensional points when applying the Haar transform, the three-dimensional data encoding device may use the attribute information value of that single existing three-dimensional point as the input value for the higher level.
[0386] In this way, the three-dimensional data encoding device applies the Haar transform hierarchically to the input attribute information to generate high-frequency and low-frequency components of the attribute information, and then performs encoding by applying quantization and other methods described later. This improves encoding efficiency.
[0387] If the attribute information is N-dimensional, the three-dimensional data encoding device may apply the Haar transform independently to each dimension and calculate the corresponding encoding coefficient. For example, if the attribute information is color information (RGB or YUV, etc.), the three-dimensional data encoding device applies the Haar transform to each component and calculates the corresponding encoding coefficient.
[0388] The three-dimensional data encoding device may apply the Haar transform in the order of hierarchy L, L+1, ..., hierarchy Lmax. The closer to hierarchy Lmax, the more encoding coefficients are generated that contain low-frequency components of the input attribute information.
[0389] The w0 and w1 shown in Figure 43 are weights assigned to each three-dimensional point. For example, the three-dimensional data encoding device may calculate the weights based on distance information between two adjacent three-dimensional points to which the Haar transform is applied. For example, the three-dimensional data encoding device may improve encoding efficiency by increasing the weight as the distance decreases. Note that the three-dimensional data encoding device may calculate these weights using other methods, or may not use weights at all.
[0390] In the example shown in Figure 43, the input attribute information is a0, a1, a2, a3, a4, and a5. Furthermore, among the coding coefficients after the Haar transformation, Ta1, Ta5, Tb1, Tb3, Tc1, and d0 are coded. Other coding coefficients (b0, b2, c0, etc.) are intermediate values and are not coded.
[0391] Specifically, in the example shown in Figure 43, a Haar transform is applied to a0 and a1 to generate a high-frequency component Ta1 and a low-frequency component b0. Here, when weights w0 and w1 are equal, the low-frequency component b0 is the average value of a0 and a1, and the high-frequency component Ta1 is the difference between a0 and a1.
[0392] Since a2 does not have a corresponding attribute, a2 is used as is for b1. Similarly, since a3 does not have a corresponding attribute, a3 is used as is for b2. Furthermore, a Haar transform is performed on a4 and a5 to generate the high-frequency component Ta5 and the low-frequency component b3.
[0393] In layer L+1, a Haar transform is applied to b0 and b1, generating a high-frequency component Tb1 and a low-frequency component c0. Similarly, a Haar transform is applied to b2 and b3, generating a high-frequency component Tb3 and a low-frequency component c1.
[0394] In the Lmax-1 hierarchical level, a Haar transform is applied to c0 and c1, generating a high-frequency component Tc1 and a low-frequency component d0.
[0395] A three-dimensional data encoding device may encode the encoded coefficients after applying the Haar transform by quantizing them. For example, a three-dimensional data encoding device performs quantization by dividing the encoded coefficients by a quantization scale (also called a quantization step (QS)). In this case, the smaller the quantization scale, the smaller the error that can occur due to quantization (quantization error). Conversely, the larger the quantization scale, the larger the quantization error.
[0396] Furthermore, the 3D data encoding device may vary the quantization scale value for each layer. Figure 44 shows an example of setting the quantization scale for each layer. For example, the 3D data encoding device may use a smaller quantization scale for higher layers and a larger quantization scale for lower layers. The coding coefficients of 3D points belonging to higher layers contain more low-frequency components than those of lower layers, and are therefore likely to be components important for human visual characteristics, etc. Therefore, by reducing the quantization scale of higher layers and suppressing quantization errors that may occur in higher layers, visual degradation can be suppressed and coding efficiency can be improved.
[0397] Furthermore, the three-dimensional data encoding device may add the quantization scale for each layer to the header, etc. This allows the three-dimensional data decoding device to correctly decode the quantization scale and appropriately decode the bitstream.
[0398] Furthermore, the three-dimensional data encoding device may adaptively switch the quantization scale value according to the importance of the three-dimensional points being encoded. For example, the three-dimensional data encoding device may use a smaller quantization scale for highly important three-dimensional points and a larger quantization scale for less important three-dimensional points. For example, the three-dimensional data encoding device may calculate importance from weights during the Haar transform. For example, the three-dimensional data encoding device may calculate the quantization scale using the sum of w0 and w1. By reducing the quantization scale for highly important three-dimensional points in this way, the quantization error can be reduced, and the encoding efficiency can be improved.
[0399] Furthermore, the QS value can be reduced for higher layers. This results in a larger QW value for higher layers, which improves prediction efficiency by suppressing the quantization error of the three-dimensional points.
[0400] Here, the quantized coding coefficient Ta1q of the coding coefficient Ta1 of attribute information a1 is expressed as Ta1 / QS_L. Note that QS may be the same value at all or some of the levels.
[0401] QW (Quantization Weight) is a value that represents the importance of the three-dimensional point to be encoded. For example, the sum of w0 and w1 mentioned above may be used as QW. As a result, the QW value increases with higher layers, and prediction efficiency can be improved by suppressing the quantization error of those three-dimensional points.
[0402] For example, a three-dimensional data encoding device may initially initialize the QW values of all three-dimensional points to 1 and then update the QW of each three-dimensional point using the w0 and w1 values from the Haar transform. Alternatively, the three-dimensional data encoding device may not initialize the QW of all three-dimensional points to 1, but instead change the initial values according to the layer. For example, setting larger initial QW values for higher layers reduces the quantization scale of the higher layers. This reduces prediction errors in the higher layers, thereby improving the prediction accuracy of the lower layers and improving encoding efficiency. Note that a three-dimensional data encoding device does not necessarily have to use QW.
[0403] When using QW, the quantized value Ta1q of Ta1 is calculated by (Equation K1) and (Equation K2).
[0404]
number
[0405] Furthermore, the three-dimensional data encoding device scans and encodes the encoded coefficients (unsigned integer values) after quantization in a certain order. For example, the three-dimensional data encoding device encodes multiple three-dimensional points sequentially from the three-dimensional points in the upper layers toward the lower layers.
[0406] For example, in the case shown in Figure 43, the three-dimensional data encoding device encodes multiple three-dimensional points in the order of d0q, Tc1q, Tb1q, Tb3q, Ta1q, and Ta5q, which are included in the upper layer Lmax. Here, the lower the layer L, the more likely the encoded coefficient after quantization is to become zero. The following factors can be cited as reasons for this.
[0407] The coding coefficients of the lower layer L tend to be zero for some three-dimensional points because they exhibit higher frequency components than the upper layers. Furthermore, due to the switching of quantization scales according to the importance mentioned above, the quantization scale becomes larger in the lower layers, making it easier for the coded coefficients after quantization to become zero.
[0408] Thus, the lower the layer, the more likely the quantized coding coefficients are to be zero, and the more likely consecutive values of zero are to occur in the first code sequence. Figure 45 shows examples of the first and second code sequences.
[0409] The three-dimensional data encoding device counts the number of times the value 0 occurs in the first code sequence and encodes the number of consecutive occurrences of 0 instead of the consecutive values of 0. In other words, the three-dimensional data encoding device generates the second code sequence by replacing the encoding coefficients of consecutive values of 0 in the first code sequence with the number of consecutive 0s (ZeroCnt). This improves encoding efficiency when the encoding coefficients after quantization are consecutive and encoding the number of consecutive 0s is done rather than encoding a large number of 0s.
[0410] Furthermore, the three-dimensional data encoding device may entropy encode the value of ZeroCnt. For example, the three-dimensional data encoding device may binarize the value of ZeroCnt using a trunket unary code with a total number of encoded three-dimensional points T, and then arithmetic encode each bit after binarization. Figure 46 shows an example of a trunket unary code when the total number of encoded three-dimensional points is T. In this case, the three-dimensional data encoding device may improve encoding efficiency by using a different encoding table for each bit. For example, the three-dimensional data encoding device may use encoding table 1 for the first bit, encoding table 2 for the second bit, and encoding table 3 for subsequent bits. In this way, the three-dimensional data encoding device can improve encoding efficiency by switching encoding tables for each bit.
[0411] Furthermore, the three-dimensional data encoding device may binarize ZeroCnt using Exponential-Golomb and then perform arithmetic encoding. This improves efficiency compared to binarization-arithmetic encoding using trunket-unary coding when ZeroCnt values tend to be large. The three-dimensional data encoding device may also add a flag to the header to switch between using trunket-unary coding and Exponential-Golomb. This allows the three-dimensional data encoding device to improve encoding efficiency by selecting the optimal binarization method. The three-dimensional data decoding device can then correctly decode the bitstream by referring to the flag included in the header and switching the binarization method.
[0412] The three-dimensional data decoder may convert the decoded quantized coding coefficients from unsigned integer values to signed integer values in the reverse manner of the method used by the three-dimensional data decoder. This allows the three-dimensional data decoder to properly decode the generated bitstream without considering the occurrence of negative integers when the coding coefficients are entropy coded. However, the three-dimensional data decoder does not necessarily need to convert the coding coefficients from unsigned integer values to signed integer values. For example, if the three-dimensional data decoder decodes a bitstream that contains separately entropy coded bits, it may decode those coded bits.
[0413] The three-dimensional data decoder decodes the quantized coding coefficients, which have been converted to signed integer values, by inverse quantization and inverse Haar transform. The decoder also uses the decoded coding coefficients to predict the points beyond the three-dimensional point being decoded. Specifically, the decoder calculates the inverse quantized value by multiplying the quantized coding coefficient by the decoded quantization scale. Next, the decoder obtains the decoded value by applying the inverse Haar transform, described later, to the inverse quantized value.
[0414] For example, a three-dimensional data decoding device converts a decoded unsigned integer value into a signed integer value in the following way: If the LSB (least significant bit) of the decoded unsigned integer value a2u is 1, the signed integer value Ta1q is set to -((a2u+1)>>1). If the LSB of the decoded unsigned integer value a2u is not 1 (i.e., it is 0), the signed integer value Ta1q is set to (a2u>>1).
[0415] Furthermore, the inverse quantization value of Ta1 is expressed as Ta1q × QS_L, where Ta1q is the quantization value of Ta1, and QS_L is the quantization step of hierarchy L.
[0416] Furthermore, the QS may be the same value at all or some of the layers. The three-dimensional data encoding device may also add information indicating the QS to the header, etc. This allows the three-dimensional data decoding device to correctly perform inverse quantization using the same QS used by the three-dimensional data encoding device.
[0417] Next, we will explain the inverse Haar transform. Figure 47 is a diagram illustrating the inverse Haar transform. The three-dimensional data decoding device decodes the attribute values of three-dimensional points by applying the inverse Haar transform to the coding coefficients after inverse quantization.
[0418] First, the three-dimensional data decoding device generates Morton codes based on the positional information of the three-dimensional points and sorts the three-dimensional points in order of their Morton codes. For example, the three-dimensional data decoding device may sort the points in ascending order of their Morton codes. Note that the sorting order is not limited to Morton code order; other orders may be used.
[0419] Next, the three-dimensional data decoder reconstructs the attribute information of adjacent three-dimensional points in Morton coding order in layer L by applying an inverse Haar transform to the coding coefficients containing the low-frequency components of layer L+1 and the coding coefficients containing the high-frequency components of layer L. For example, the three-dimensional data decoder may use an inverse Haar transform of a 2x2 matrix. The reconstructed attribute information of layer L is used as the input value for the lower layer L-1.
[0420] The three-dimensional data decoder repeats this hierarchical processing until all attribute information at the lowest layer has been decoded, at which point it terminates. If, when applying the inverse Haar transform, only one three-dimensional point exists as two adjacent three-dimensional points at layer L-1, the three-dimensional data decoder may substitute the value of the encoded component at layer L for the attribute value of the single existing three-dimensional point. This allows the three-dimensional data decoder to apply the Haar transform to all values of the input attribute information, enabling it to correctly decode a bitstream with improved encoding efficiency.
[0421] If the attribute information is N-dimensional, the three-dimensional data decoder may apply the inverse Haar transform independently to each dimension and decode the respective coding coefficients. For example, if the attribute information is color information (RGB or YUV, etc.), the three-dimensional data decoder applies the inverse Haar transform to the coding coefficients for each component and decodes the respective attribute values.
[0422] The three-dimensional data decoder may apply the inverse Haar transform in the order of hierarchy Lmax, L+1, ..., hierarchy L. Also, w0 and w1 shown in Figure 47 are weights assigned to each three-dimensional point. For example, the three-dimensional data decoder may calculate the weights based on distance information between two adjacent three-dimensional points to which the inverse Haar transform is applied. For example, the three-dimensional data encoder may decode a bitstream with improved encoding efficiency by increasing the weight as the distance decreases.
[0423] In the example shown in Figure 47, the coding coefficients after inverse quantization are Ta1, Ta5, Tb1, Tb3, Tc1, and d0, and the decoded values are a0, a1, a2, a3, a4, and a5.
[0424] Figure 48 shows an example of the syntax for attribute information (attribute_data). Attribute information (attribute_data) includes the number of zeros (ZeroCnt), the attribute dimension (attribute_dimension), and the coding coefficient (value[j][i]).
[0425] The number of consecutive zeros (ZeroCnt) indicates the number of consecutive values of 0 in the coded coefficients after quantization. Note that the three-dimensional data encoding device may also perform arithmetic coding after binarizing the ZeroCnt.
[0426] Furthermore, as shown in Figure 48, the three-dimensional data encoding device may determine whether the layer L (layerL) to which the encoding coefficient belongs is equal to or greater than a predetermined threshold TH_layer, and switch the information to be added to the bitstream depending on the determination result. For example, if the determination result is true, the three-dimensional data encoding device adds all the encoding coefficients of the attribute information to the bitstream. Alternatively, if the determination result is false, the three-dimensional data encoding device may add some of the encoding coefficients to the bitstream.
[0427] Specifically, if the judgment result is true, the three-dimensional data encoding device adds the encoded result of the RGB or YUV three-dimensional color information to the bitstream. If the judgment result is false, the three-dimensional data encoding device may add some of the color information, such as G or Y, to the bitstream, and leave the other components omitted. In this way, the three-dimensional data encoding device can improve encoding efficiency by not adding some of the encoding coefficients of layers containing encoding coefficients that represent high-frequency components that are less visually noticeable to the bitstream (layers smaller than TH_layer).
[0428] The attribute dimension (attribute_dimension) indicates the number of dimensions in the attribute information. For example, if the attribute information is color information of a three-dimensional point (such as RGB or YUV), the attribute dimension is set to 3 because the color information is three-dimensional. If the attribute information is reflectance, the attribute dimension is set to 1 because reflectance is one-dimensional. Note that the attribute dimension may also be added to the header of the bitstream's attribute information.
[0429] The coding coefficient (value[j][i]) represents the quantized coding coefficient of the j-th attribute information of the i-th three-dimensional point. For example, if the attribute information is color information, value
[99] [1] represents the coding coefficient of the second-th attribute (e.g., G value) of the 100th three-dimensional point. Also, if the attribute information is reflectance information, value
[0119] [0] represents the coding coefficient of the first-th attribute (e.g., reflectance) of the 120th three-dimensional point.
[0430] Furthermore, if the following conditions are met, the three-dimensional data encoding device may subtract the value 1 from value[j][i] and entropy encode the resulting value. In this case, the three-dimensional data decoding device restores the encoding coefficient by adding the value 1 to the entropy-decoded value[j][i].
[0431] The above conditions apply when (1) attribute_dimension = 1, or (2) attribute_dimension is 1 or greater and the values of all dimensions are equal. For example, if the attribute information is reflectance, then attribute_dimension = 1, so the 3D data encoding device calculates the value by subtracting 1 from the encoding coefficient, and then encodes the calculated value. The 3D data decoding device calculates the encoding coefficient by adding 1 to the decoded value.
[0432] More specifically, for example, if the coding coefficient for reflectance is 10, the three-dimensional data encoding device encodes a value of 9, which is obtained by subtracting 1 from the coding coefficient value of 10. The three-dimensional data decoding device then adds 1 to the decoded value of 9 to calculate the coding coefficient value of 10.
[0433] Furthermore, if the attribute information is color, attribute_dimension=3, so the three-dimensional data encoding device, for example, if the quantized encoding coefficients of each component R, G, and B are the same, subtracts 1 from each encoding coefficient and encodes the resulting value. The three-dimensional data decoding device adds 1 to the decoded value. More specifically, for example, if the encoding coefficients of R, G, and B are (1, 1, 1), the three-dimensional data encoding device encodes (0, 0, 0). The three-dimensional data decoding device adds 1 to each component of (0, 0, 0) to calculate (1, 1, 1). Also, if the encoding coefficients of R, G, and B are (2, 1, 2), the three-dimensional data encoding device encodes (2, 1, 2) as is. The three-dimensional data decoding device uses the decoded (2, 1, 2) as is as the encoding coefficient.
[0434] In this way, by introducing ZeroCnt, patterns where all dimensions are 0 are not generated as values, so it is possible to encode a value obtained by subtracting 1 from the value. Therefore, encoding efficiency can be improved.
[0435] Furthermore, the value[0][i] shown in Figure 48 represents the quantized coding coefficient of the first-dimensional attribute information of the i-th three-dimensional point. As shown in Figure 48, if the layer L (layerL) to which the coding coefficient belongs is smaller than the threshold TH_layer, the amount of code can be reduced by adding the first-dimensional attribute information to the bitstream (without adding the second-dimensional and subsequent attribute information to the bitstream).
[0436] The three-dimensional data encoding device may switch the method of calculating the ZeroCnt value depending on the value of attribute_dimension. For example, if attribute_dimension=3, the three-dimensional data encoding device may count the number of times the coding coefficients of all components (dimensions) are zero. Figure 49 shows an example of coding coefficients and ZeroCnt in this case. For example, in the case of color information shown in Figure 49, the three-dimensional data encoding device counts the number of consecutive coding coefficients where the R, G, and B components are all zero, and adds the counted number as ZeroCnt to the bitstream. This eliminates the need to encode ZeroCnt for each component, reducing overhead. Therefore, coding efficiency can be improved. Note that even if attribute_dimension is 2 or more, the three-dimensional data encoding device may calculate ZeroCnt for each dimension and add the calculated ZeroCnt to the bitstream.
[0437] Figure 50 is a flowchart of the three-dimensional data encoding process according to this embodiment. First, the three-dimensional data encoding device encodes the position information (geometry) (S6601). For example, the three-dimensional data encoding device performs encoding using an octree representation.
[0438] Next, the three-dimensional data encoding device converts the attribute information (S6602). For example, if the position of a three-dimensional point changes due to quantization or the like after encoding the position information, the three-dimensional data encoding device reassigns the attribute information of the original three-dimensional point to the changed three-dimensional point. The three-dimensional data encoding device may also interpolate the attribute information values according to the amount of change in position before reassigning. For example, the three-dimensional data encoding device detects N three-dimensional points that are close to the changed three-dimensional position before the change, weights and averages the attribute information values of the N three-dimensional points based on the distance from the changed three-dimensional position to each of the N three-dimensional points, and sets the obtained value as the attribute information value of the changed three-dimensional point. Furthermore, if two or more three-dimensional points change to the same three-dimensional position due to quantization or the like, the three-dimensional data encoding device may assign the average value of the attribute information of the two or more three-dimensional points before the change as the attribute information value of the changed point.
[0439] Next, the three-dimensional data encoding device encodes attribute information (S6603). For example, if the three-dimensional data encoding device encodes multiple attribute information, it may encode the multiple attribute information sequentially. For example, if the three-dimensional data encoding device encodes color and reflectance as attribute information, it generates a bitstream in which the encoded result of reflectance is appended after the encoded result of color. The order in which the multiple encoded results of attribute information are appended to the bitstream does not matter.
[0440] Furthermore, the three-dimensional data encoding device may add information to the header or elsewhere indicating the starting location of the encoded data for each attribute within the bitstream. This allows the three-dimensional data decoding device to selectively decode the attribute information that needs to be decoded, thus omitting the decoding process for attribute information that does not need to be decoded. Therefore, the processing load on the three-dimensional data decoding device can be reduced. In addition, the three-dimensional data encoding device may encode multiple attribute information in parallel and integrate the encoding results into a single bitstream. This allows the three-dimensional data encoding device to encode multiple attribute information at high speed.
[0441] Figure 51 is a flowchart of the attribute information encoding process (S6603). First, the three-dimensional data encoding device generates encoding coefficients from the attribute information using the Haar transform (S6611). Next, the three-dimensional data encoding device applies quantization to the encoding coefficients (S6612). Then, the three-dimensional data encoding device generates encoded attribute information (bitstream) by encoding the quantized encoding coefficients (S6613).
[0442] Furthermore, the three-dimensional data encoding device applies inverse quantization to the encoded coefficients after quantization (S6614). Next, the three-dimensional data decoding device decodes the attribute information by applying the inverse Haar transform to the encoded coefficients after inverse quantization (S6615). For example, the decoded attribute information is referenced in subsequent encoding.
[0443] Figure 52 is a flowchart of the coding coefficient coding process (S6613). First, the three-dimensional data coding device converts the coding coefficient from a signed integer value to an unsigned integer value (S6621). For example, the three-dimensional data coding device converts a signed integer value to an unsigned integer value as follows: If the signed integer value Ta1q is less than 0, the unsigned integer value is set to -1-(2×Ta1q). If the signed integer value Ta1q is 0 or greater, the unsigned integer value is set to 2×Ta1q. Note that if the coding coefficient does not result in a negative value, the three-dimensional data coding device may encode the coding coefficient directly as an unsigned integer value.
[0444] If not all coding coefficients have been processed (No in S6622), the three-dimensional data encoding device determines whether the value of the coding coefficient to be processed is zero (S6623). If the value of the coding coefficient to be processed is zero (Yes in S6623), the three-dimensional data encoding device increments ZeroCnt by 1 (S6624) and returns to step S6622.
[0445] If the value of the coding coefficient to be processed is not zero (No in S6623), the three-dimensional data encoding device encodes ZeroCnt and resets ZeroCnt to 0 (S6625). The three-dimensional data encoding device also arithmetically encodes the coding coefficient to be processed (S6626) and returns to step S6622. For example, the three-dimensional data encoding device performs binary arithmetic encoding. Alternatively, the three-dimensional data encoding device may subtract the value 1 from the coding coefficient and encode the resulting value.
[0446] Furthermore, the processing in steps S6623 to S6626 is repeated for each coding coefficient. Also, if all coding coefficients have been processed (Yes in S6622), the three-dimensional data encoding device terminates the process.
[0447] Figure 53 is a flowchart of the three-dimensional data decoding process according to this embodiment. First, the three-dimensional data decoding device decodes the position information (geometry) from the bitstream (S6631). For example, the three-dimensional data decoding device performs decoding using an octree representation.
[0448] Next, the three-dimensional data decoder decodes attribute information from the bitstream (S6632). For example, if the three-dimensional data decoder decodes multiple attribute information, it may decode them sequentially. For example, if the three-dimensional data decoder decodes color and reflectance as attribute information, it decodes the color encoding result and the reflectance encoding result in the order in which they are added to the bitstream. For example, if the reflectance encoding result is added after the color encoding result in the bitstream, the three-dimensional data decoder decodes the color encoding result, and then decodes the reflectance encoding result. The three-dimensional data decoder may decode the encoding results of the attribute information added to the bitstream in any order.
[0449] Furthermore, the three-dimensional data decoding device may obtain information indicating the starting location of the encoded data for each attribute information within the bitstream by decoding the header, etc. This allows the three-dimensional data decoding device to selectively decode the attribute information that needs to be decoded, thus omitting the decoding process for attribute information that does not need to be decoded. Therefore, the processing load of the three-dimensional data decoding device can be reduced. In addition, the three-dimensional data decoding device may decode multiple attribute information in parallel and integrate the decoding results into a single three-dimensional point cloud. This allows the three-dimensional data decoding device to decode multiple attribute information at high speed.
[0450] Figure 54 is a flowchart of the attribute information decoding process (S6632). First, the three-dimensional data decoding device decodes the coding coefficients from the bitstream (S6641). Next, the three-dimensional data decoding device applies inverse quantization to the coding coefficients (S6642). Then, the three-dimensional data decoding device decodes the attribute information by applying the inverse Haar transform to the coding coefficients after inverse quantization (S6643).
[0451] Figure 55 is a flowchart of the coding coefficient decoding process (S6641). First, the three-dimensional data decoding device decodes ZeroCnt from the bitstream (S6651). If not all coding coefficients have been processed (No in S6652), the three-dimensional data decoding device determines whether ZeroCnt is greater than 0 (S6653).
[0452] If ZeroCnt is greater than zero (Yes in S6653), the three-dimensional data decoder sets the coding coefficient to be processed to 0 (S6654). Next, the three-dimensional data decoder subtracts 1 from ZeroCnt (S6655) and returns to step S6652.
[0453] If ZeroCnt is zero (No in S6653), the three-dimensional data decoder decodes the coding coefficients to be processed (S6656). For example, the three-dimensional data decoder uses binary arithmetic decoding. Alternatively, the three-dimensional data decoder may add the value 1 to the decoded coding coefficients.
[0454] Next, the three-dimensional data decoding device decodes ZeroCnt, sets the obtained value to ZeroCnt (S6657), and returns to step S6652.
[0455] Furthermore, the processing in steps S6653 to S6657 is repeated for each coding coefficient. Also, if all coding coefficients have been processed (Yes in S6652), the three-dimensional data encoding device converts the decoded coding coefficients from unsigned integer values to signed integer values (S6658). For example, the three-dimensional data decoding device may convert the decoded coding coefficients from unsigned integer values to signed integer values as follows: If the LSB (least significant bit) of the decoded unsigned integer value Ta1u is 1, the signed integer value Ta1q is set to -((Ta1u+1)>>1). If the LSB of the decoded unsigned integer value Ta1u is not 1 (it is 0), the signed integer value Ta1q is set to (Ta1u>>1). Note that if the coding coefficients do not result in negative values, the three-dimensional data decoding device may use the decoded coding coefficients directly as signed integer values.
[0456] Figure 56 is a block diagram of the attribute information coding unit 6600 included in the three-dimensional data coding device. The attribute information coding unit 6600 comprises a sorting unit 6601, a Haar transformation unit 6602, a quantization unit 6603, an inverse quantization unit 6604, an inverse Haar transformation unit 6605, a memory 6606, and an arithmetic coding unit 6607.
[0457] The sorting unit 6601 generates Morton codes using the position information of three-dimensional points and sorts multiple three-dimensional points in Morton code order. The Haar transform unit 6602 generates coding coefficients by applying the Haar transform to the attribute information. The quantization unit 6603 quantizes the coding coefficients of the attribute information.
[0458] The inverse quantization unit 6604 inversely quantizes the encoded coefficients after quantization. The inverse Haar transform unit 6605 applies the inverse Haar transform to the encoded coefficients. The memory 6606 stores the attribute information values of multiple decoded three-dimensional points. For example, the attribute information of decoded three-dimensional points stored in the memory 6606 may be used for predicting unencoded three-dimensional points.
[0459] The arithmetic coding unit 6607 calculates ZeroCnt from the quantized coding coefficients and arithmetically codes ZeroCnt. The arithmetic coding unit 6607 also arithmetically codes the non-zero coding coefficients after quantization. The arithmetic coding unit 6607 may also binarize the coding coefficients before arithmetic coding. Furthermore, the arithmetic coding unit 6607 may generate and code various header information.
[0460] Figure 57 is a block diagram of the attribute information decoding unit 6610 included in the three-dimensional data decoding device. The attribute information decoding unit 6610 comprises an arithmetic decoding unit 6611, an inverse quantization unit 6612, an inverse Haar transform unit 6613, and a memory 6614.
[0461] The arithmetic decoding unit 6611 arithmetically decodes the ZeroCnt and coding coefficients contained in the bitstream. The arithmetic decoding unit 6611 may also decode various header information.
[0462] The inverse quantization unit 6612 inversely quantizes the arithmetic-decoded coding coefficients. The inverse Haar transform unit 6613 applies the inverse Haar transform to the coding coefficients after inverse quantization. The memory 6614 stores the attribute information values of multiple decoded three-dimensional points. For example, the attribute information of decoded three-dimensional points stored in the memory 6614 may be used to predict undecoded three-dimensional points.
[0463] In the above embodiment, an example was shown in which three-dimensional points are encoded in the order from lower layers to upper layers, but this is not necessarily the only way. For example, a method of scanning the encoded coefficients after the Haar transform in the order from upper layers to lower layers may be used. In this case as well, the three-dimensional data encoding device may encode the number of consecutive values of 0 as ZeroCnt.
[0464] Furthermore, the three-dimensional data encoding device may switch whether or not to use the encoding method using ZeroCnt described in this embodiment on a WLD, SPC, or volume basis. In this case, the three-dimensional data encoding device may add information to the header information indicating whether or not the encoding method using ZeroCnt has been applied. This enables the three-dimensional data decoding device to perform decoding appropriately. As an example of a switching method, for example, the three-dimensional data encoding device counts the number of times an encoding coefficient with a value of 0 occurs for one volume. If the count value exceeds a predetermined threshold, the three-dimensional data encoding device applies the method using ZeroCnt to the next volume; if the count value is below the threshold, the method using ZeroCnt is not applied to the next volume. This allows the three-dimensional data encoding device to appropriately switch whether or not to apply the encoding method using ZeroCnt according to the characteristics of the three-dimensional points to be encoded, thereby improving encoding efficiency.
[0465] The following describes another method (modification) of this embodiment. The three-dimensional data encoding device scans and encodes the encoded coefficients (unsigned integer values) after quantization in a certain order. For example, the three-dimensional data encoding device encodes multiple three-dimensional points sequentially from the three-dimensional points in the lower layers toward the upper layers.
[0466] Figure 58 shows examples of the first and second code sequences when this method is applied to the attribute information shown in Figure 43. In this example, the three-dimensional data encoding device encodes multiple coding coefficients in the order of Ta1q, Ta5q, Tb1q, Tb3q, Tc1q, and d0q contained in the lower layer L. Here, the lower the layer, the more likely the quantized coding coefficients are to become zero. The following are some of the factors that contribute to this.
[0467] The coding coefficients of the lower layer L tend to be zero for some three-dimensional points being coded, as they exhibit higher frequency components than the upper layers. Furthermore, due to the switching of quantization scales according to the importance mentioned above, the quantization scale becomes larger in the lower layers, making it easier for the coded coefficients after quantization to become zero.
[0468] Thus, the lower the layer, the more likely the quantized coding coefficients are to be zero, and the more likely consecutive zeros are to occur in the first code sequence. The three-dimensional data encoding device counts the number of times the value zero occurs in the first code sequence and encodes the number of consecutive zeros (ZeroCnt) instead of encoding the consecutive zeros. This improves encoding efficiency when the quantized coding coefficients are consecutive zeros, by encoding the number of consecutive zeros rather than encoding a large number of zeros.
[0469] Furthermore, the three-dimensional data encoding device may encode information indicating the total number of occurrences of the value 0. This reduces the overhead of encoding ZeroCnt and improves encoding efficiency.
[0470] For example, a three-dimensional data encoding device encodes the total number of encoding coefficients with a value of 0 as TotalZeroCnt. In the example shown in Figure 58, when the three-dimensional data decoding device decodes the second ZeroCnt (value 1) in the second code sequence, the total number of decoded ZeroCnts becomes N+1 (=TotalZeroCnt). Therefore, the three-dimensional data decoding device can determine that no more 0s will occur from this point onward. Consequently, the three-dimensional data encoding device no longer needs to encode ZeroCnt for each value, thus reducing the amount of code.
[0471] Furthermore, the three-dimensional data encoding device may entropy encode TotalZeroCnt. For example, the three-dimensional data encoding device may binarize the value of TotalZeroCnt using a trunket unary code with a total of T encoded three-dimensional points, and then arithmetic encode each bit after binarization. In this case, the three-dimensional data encoding device may improve encoding efficiency by using a different encoding table for each bit. For example, the three-dimensional data encoding device may use encoding table 1 for the first bit, encoding table 2 for the second bit, and encoding table 3 for subsequent bits. In this way, the three-dimensional data encoding device can improve encoding efficiency by switching encoding tables for each bit.
[0472] Furthermore, the three-dimensional data encoding device may binarize TotalZeroCnt using exponential golom before performing arithmetic encoding. This improves efficiency compared to binarization and arithmetic encoding using trunket unary coding when the value of TotalZeroCnt tends to be large. The three-dimensional data encoding device may also add a flag to the header to switch between using trunket unary coding and exponential golom coding. This allows the three-dimensional data encoding device to improve encoding efficiency by selecting the optimal binarization method. The three-dimensional data decoding device can then correctly decode the bitstream by referring to the flag included in the header and switching the binarization method.
[0473] Figure 59 shows an example of the syntax for attribute information (attribute_data) in this modified example. The attribute information (attribute_data) shown in Figure 59 includes the total number of zeros (TotalZeroCnt) in addition to the attribute information shown in Figure 48. Other information is the same as in Figure 48. The total number of zeros (TotalZeroCnt) indicates the total number of coding coefficients with a value of 0 after quantization.
[0474] Furthermore, the three-dimensional data encoding device may switch the calculation method for TotalZeroCnt and ZeroCnt depending on the value of attribute_dimension. For example, if attribute_dimension=3, the three-dimensional data encoding device may count the number of times the coding coefficients of all components (dimensions) are zero. Figure 60 shows an example of the coding coefficients, ZeroCnt, and TotalZeroCnt in this case. For example, in the case of color information shown in Figure 60, the three-dimensional data encoding device counts the number of consecutive coding coefficients where the R, G, and B components are all zero, and adds the counted number as TotalZeroCnt and ZeroCnt to the bitstream. This eliminates the need to encode TotalZeroCnt and ZeroCnt for each component, reducing overhead. Therefore, encoding efficiency can be improved. Note that even if attribute_dimension is 2 or more, the three-dimensional data encoding device may calculate TotalZeroCnt and ZeroCnt for each dimension and add the calculated TotalZeroCnt and ZeroCnt to the bitstream.
[0475] Figure 61 is a flowchart of the coding coefficient coding process (S6613) in this modified example. First, the three-dimensional data coding device converts the coding coefficients from signed integer values to unsigned integer values (S6661). Next, the three-dimensional data coding device encodes TotalZeroCnt (S6662).
[0476] If not all coding coefficients have been processed (No in S6663), the three-dimensional data encoding device determines whether the value of the coding coefficient to be processed is zero (S6664). If the value of the coding coefficient to be processed is zero (Yes in S6664), the three-dimensional data encoding device increments ZeroCnt by 1 (S6665) and returns to step S6663.
[0477] If the value of the coding coefficient to be processed is not zero (No in S6664), the three-dimensional data encoding device determines whether TotalZeroCnt is greater than 0 (S6666). If TotalZeroCnt is greater than 0 (Yes in S6666), the three-dimensional data encoding device encodes ZeroCnt and sets TotalZeroCnt to TotalZeroCnt-ZeroCnt (S6667).
[0478] After step S6667, or if TotalZeroCnt is 0 (No in S6666), the three-dimensional data encoding device encodes the encoding coefficients, resets ZeroCnt to 0 (S6668), and returns to step S6663. For example, the three-dimensional data encoding device performs binary arithmetic encoding. Alternatively, the three-dimensional data encoding device may subtract the value 1 from the encoding coefficients and encode the resulting value.
[0479] Furthermore, steps S6664 to S6668 are repeated for each coding coefficient. Also, if all coding coefficients have been processed (Yes in S6663), the three-dimensional data encoding device terminates the process.
[0480] Figure 62 is a flowchart of the coding coefficient decoding process (S6641) in this modified example. First, the three-dimensional data decoding device decodes TotalZeroCnt from the bitstream (S6671). Next, the three-dimensional data decoding device decodes ZeroCnt from the bitstream and sets TotalZeroCnt to TotalZeroCnt-ZeroCnt (S6672).
[0481] If not all coding coefficients have been processed (No in S6673), the three-dimensional data encoding device determines whether ZeroCnt is greater than 0 (S6674).
[0482] If ZeroCnt is greater than zero (Yes in S6674), the three-dimensional data decoder sets the coding coefficient to be processed to 0 (S6675). Next, the three-dimensional data decoder subtracts 1 from ZeroCnt (S6676) and returns to step S6673.
[0483] If ZeroCnt is zero (No in S6674), the three-dimensional data decoder decodes the coding coefficients to be processed (S6677). For example, the three-dimensional data decoder uses binary arithmetic decoding. Alternatively, the three-dimensional data decoder may add the value 1 to the decoded coding coefficients.
[0484] Next, the three-dimensional data decoding device determines whether TotalZeroCnt is greater than 0 (S6678). If TotalZeroCnt is greater than 0 (Yes in S6678), the three-dimensional data decoding device decodes ZeroCnt, sets the obtained value to ZeroCnt, sets TotalZeroCnt to TotalZeroCnt-ZeroCnt (S6679), and returns to step S6673. If TotalZeroCnt is 0 (No in S6678), the three-dimensional data decoding device returns to step S6673.
[0485] Furthermore, the processing in steps S6674 to S6679 is repeated for each coding coefficient. Also, if all coding coefficients have been processed (Yes in S6673), the three-dimensional data encoding device converts the decoded coding coefficients from unsigned integer values to signed integer values (S6680).
[0486] Figure 63 shows an example of another syntax for attribute information (attribute_data). The attribute information (attribute_data) shown in Figure 63 includes value[j][i]_greater_zero_flag, value[j][i]_greater_one_flag, and value[j][i] instead of the coding coefficient (value[j][i]) shown in Figure 48. Other information is the same as in Figure 48.
[0487] The value[j][i]_greater_zero_flag indicates whether the value of the coding coefficient (value[j][i]) is greater than 0 (i.e., 1 or greater). In other words, the value[j][i]_greater_zero_flag indicates whether the value of the coding coefficient (value[j][i]) is 0.
[0488] For example, if the coding coefficient value is greater than 0, value[j][i]_greater_zero_flag is set to value 1, and if the coding coefficient value is 0, value[j][i]_greater_zero_flag is set to value 0. The three-dimensional data encoding device does not have to append value[j][i] to the bitstream if the value of value[j][i]_greater_zero_flag is 0. In this case, the three-dimensional data decoding device may determine that the value of value[j][i] is 0. This reduces the amount of coding required.
[0489] The `value[j][i]_greater_one_flag` indicates whether the value of the coding coefficient (`value[j][i]`) is greater than 1 (i.e., 2 or greater). In other words, `value[j][i]_greater_one_flag` indicates whether the value of the coding coefficient (`value[j][i]`) is 1.
[0490] For example, if the coding coefficient value is greater than 1, value[j][i]_greater_one_flag is set to the value 1. Otherwise (if the coding coefficient value is 1 or less), value[j][i]_greater_one_flag is set to the value 0. The 3D data encoding device does not have to append value[j][i] to the bitstream if the value of value[j][i]_greater_one_flag is 0. In this case, the 3D data decoding device may determine that the value of value[j][i] is 1.
[0491] `value[j][i]` represents the quantized coding coefficient of the j-th attribute information of the i-th three-dimensional point. For example, if the attribute information is color information, `value
[99] [1]` represents the coding coefficient of the second-th attribute (e.g., G value) of the 100th three-dimensional point. Also, if the attribute information is reflectance information, `value
[0119] [0]` represents the coding coefficient of the first-th attribute (e.g., reflectance) of the 120th three-dimensional point.
[0492] The three-dimensional data encoding device may append value[j][i] to the bitstream if value[j][i]_greater_zero_flag=1 and value[j][i]_greater_one_flag=1. Alternatively, the three-dimensional data encoding device may append a value obtained by subtracting 2 from value[j][i] to the bitstream. In this case, the three-dimensional data decoding device calculates the encoding coefficient by adding the value 2 to the decoded value[j][i].
[0493] The three-dimensional data encoding device may entropy encode value[j][i]_greater_zero_flag and value[j][i]_greater_one_flag. For example, binary arithmetic coding and binary arithmetic decoding may be used. This can improve encoding efficiency.
[0494] (Embodiment 9) To achieve high compression, attribute information contained in PCC (Point Cloud Compression) data is transformed using multiple methods such as Lifting, RAHT (Region Adaptive Hierarchical Transform), or other transformation techniques. Here, Lifting is one of the transformation methods that uses LoD (Level of Detail).
[0495] Since important signal information tends to be contained in low-frequency components, the amount of code is reduced by quantizing the high-frequency components. In other words, the conversion process has strong energy compression characteristics. Also, depending on the magnitude of the quantization parameter, precision is lost due to quantization.
[0496] Figure 64 is a block diagram showing the configuration of a three-dimensional data encoding device according to this embodiment. This three-dimensional data encoding device comprises a subtraction unit 7001, a transformation unit 7002, a transformation matrix holding unit 7003, a quantization unit 7004, a quantization control unit 7005, and an entropy encoding unit 7006.
[0497] The subtraction unit 7001 calculates a coefficient value which is the difference between the input data and the reference data. For example, the input data is attribute information contained in the point cloud data, and is a predicted value of the attribute information compared to the reference data.
[0498] The transformation unit 7002 performs a transformation process to convert the values to coefficient values. For example, this transformation process is a process of classifying multiple attribute information into LoDs. This transformation process may also be a Haar transformation, etc. The transformation matrix holding unit 7003 holds the transformation matrix used in the transformation process by the transformation unit 7002. For example, this transformation matrix is a Haar transformation matrix. Here, we show an example in which the three-dimensional data encoding device has the function of performing both transformation processing using LoDs and transformation processing such as Haar transformations, but it may have only one of these functions. The three-dimensional data encoding device may also selectively use these two types of transformation processing. Furthermore, the three-dimensional data encoding device may switch the transformation processing to be used for each predetermined processing unit.
[0499] The quantization unit 7004 generates quantized values by quantizing coefficient values. The quantization control unit 7005 controls the quantization parameters used by the quantization unit 7004 for quantization. For example, the quantization control unit 7005 may switch the quantization parameters (or quantization steps) according to the hierarchical structure of the encoding. This allows the amount of generated code to be controlled for each hierarchical structure by selecting appropriate quantization parameters for each hierarchical structure. The quantization control unit 7005 may also set the quantization parameters below a certain hierarchical level that contain frequency components that have little impact on subjective image quality to their maximum value, and set the quantization coefficients below that hierarchical level to 0. This reduces the amount of generated code while suppressing the degradation of subjective image quality. Furthermore, the quantization control unit 7005 can control subjective image quality and the amount of generated code more precisely. Here, "hierarchical level" refers to the hierarchy (depth in the tree structure) in LoD or RAHT (Haar transform).
[0500] The entropy coding unit 7006 generates a bitstream by entropy coding (e.g., arithmetic coding) the quantization coefficients. The entropy coding unit 7006 also encodes the quantization parameters for each hierarchical level set by the quantization control unit 7005.
[0501] Figure 65 is a block diagram showing the configuration of a three-dimensional data decoding device according to this embodiment. This three-dimensional data decoding device comprises an entropy decoding unit 7011, an inverse quantization unit 7012, a quantization control unit 7013, an inverse transformation unit 7014, a transformation matrix holding unit 7015, and an addition unit 7016.
[0502] The entropy decoding unit 7011 decodes the quantization coefficients and the quantization parameters for each hierarchy from the bitstream. The inverse quantization unit 7012 generates coefficient values by inverse quantization of the quantization coefficients. The quantization control unit 7013 controls the quantization parameters used by the inverse quantization unit 7012 for inverse quantization based on the hierarchy quantization parameters obtained by the entropy decoding unit 7011.
[0503] The inverse transformation unit 7014 performs an inverse transformation on the coefficient values. For example, the inverse transformation unit 7014 performs an inverse Haar transformation on the coefficient values. The transformation matrix holder unit 7015 holds the transformation matrix used in the inverse transformation process performed by the inverse transformation unit 7014. For example, this transformation matrix is an inverse Haar transformation matrix.
[0504] The addition unit 7016 generates output data by adding reference data to the coefficient value. For example, the output data is attribute information included in the point cloud data, and is a predicted value of the attribute information in relation to the reference data.
[0505] Next, we will explain an example of setting quantization parameters for each level. In encoding attribute information such as Predicting / Lifting, different quantization parameters are applied based on the Level of Data (LoD) level. For example, reducing the quantization parameters at lower levels improves the accuracy of those levels. This improves the prediction accuracy at higher levels. Conversely, increasing the quantization parameters at higher levels reduces the amount of data. In this way, quantization tree values (Qt) can be set individually for each LoD according to the user's usage policy. Here, quantization tree values are, for example, quantization parameters.
[0506] Figure 66 shows an example of a Level of Direction (LoD) configuration. For example, as shown in Figure 66, independent Qt0 to Qt2 are set for LoD0 to LoD2.
[0507] Furthermore, in attribute information encoding using RAHT, different quantization parameters are applied based on the depth of the tree structure. For example, reducing the quantization parameters at lower levels improves the accuracy of those levels. This improves the prediction accuracy at higher levels. Conversely, increasing the quantization parameters at higher levels reduces the amount of data. In this way, quantization tree values (Qt) can be set individually for each depth of the tree structure according to the user's usage policy.
[0508] Figure 67 shows an example of the hierarchical structure (tree structure) of RAHT. For example, as shown in Figure 67, independent Qt0 to Qt2 are set for each depth of the tree structure.
[0509] The configuration of the three-dimensional data encoding device according to this embodiment will be described below. Figure 68 is a block diagram showing the configuration of the three-dimensional data encoding device 7020 according to this embodiment. The three-dimensional data encoding device 7020 generates encoded data (encoded stream) by encoding point cloud data. This three-dimensional data encoding device 7020 includes a division unit 7021, a plurality of position information encoding units 7022, a plurality of attribute information encoding units 7023, an additional information encoding unit 7024, and a multiplexing unit 7025.
[0510] The division unit 7021 generates multiple division data by dividing the point cloud data. Specifically, the division unit 7021 generates multiple division data by dividing the space of the point cloud data into multiple subspaces. Here, a subspace is either a tile or a slice, or a combination of a tile and a slice. More specifically, the point cloud data includes location information, attribute information (such as color or reflectance), and additional information. The division unit 7021 generates multiple division location information by dividing the location information and generates multiple division attribute information by dividing the attribute information. The division unit 7021 also generates additional information related to the division.
[0511] For example, the division unit 7021 first divides the point cloud into tiles. Next, the division unit 7021 further divides the resulting tiles into slices.
[0512] Multiple location information encoding units 7022 generate multiple encoded location information by encoding multiple partitioned location information. For example, the location information encoding unit 7022 encodes partitioned location information using an N-tree structure such as an octree. Specifically, in an octree, the target space is divided into 8 nodes (subspaces), and 8 bits of information (occupancy code) are generated to indicate whether or not a point cloud is contained in each node. Furthermore, nodes containing point clouds are further divided into 8 nodes, and 8 bits of information are generated to indicate whether or not a point cloud is contained in each of these 8 nodes. This process is repeated until the number of point clouds contained in a predetermined hierarchy or node falls below a threshold. For example, multiple location information encoding units 7022 process multiple partitioned location information in parallel.
[0513] The attribute information encoding unit 7023 generates encoded attribute information, which is encoded data, by encoding attribute information using the configuration information generated by the location information encoding unit 7022. For example, the attribute information encoding unit 7023 determines the reference point (reference node) to be referenced in encoding the target point (target node) to be processed, based on the octave tree structure generated by the location information encoding unit 7022. For example, the attribute information encoding unit 7023 references a surrounding node or adjacent node whose parent node in the octave tree is the same as the target node. Note that the method for determining the reference relationship is not limited to this.
[0514] Furthermore, the encoding process for location information or attribute information may include at least one of the following: quantization processing, prediction processing, and arithmetic encoding processing. In this case, a reference means using a reference node to calculate the predicted value of attribute information, or using the state of a reference node (for example, occupancy information indicating whether or not a point cloud is included in the reference node) to determine the encoding parameters. For example, encoding parameters are quantization parameters in the quantization processing, or context in arithmetic encoding, etc.
[0515] The additional information encoding unit 7024 generates encoded additional information by encoding the additional information contained in the point cloud data and the additional information related to data division generated by the division unit 7021 during division.
[0516] The multiplexing unit 7025 generates encoded data (encoded stream) by multiplexing multiple encoded position information, multiple encoded attribute information, and encoded additional information, and transmits the generated encoded data. The encoded additional information is used during decoding.
[0517] Figure 69 is a block diagram of the division section 7021. The division section 7021 includes a tile division section 7031 and a slice division section 7032.
[0518] The tile division unit 7031 generates multiple tile position information by dividing position information (Geometry) into tiles. The tile division unit 7031 also generates multiple tile attribute information by dividing attribute information into tiles. Furthermore, the tile division unit 7031 outputs tile metadata information (Tile MetaData) that includes information related to tile division and information generated during tile division.
[0519] The slice division unit 7032 generates multiple division position information (multiple slice position information) by dividing multiple tile position information into slices. The slice division unit 7032 also generates multiple division attribute information (multiple slice attribute information) by dividing multiple tile attribute information into slices. Furthermore, the slice division unit 7032 outputs slice metadata information (Slice MetaData) that includes information related to slice division and information generated during slice division.
[0520] Furthermore, the tile division unit 7031 and the slice division unit 7032 determine quantization tree values (quantization parameters) based on the generated additional information.
[0521] Figure 70 is a block diagram of the attribute information coding unit 7023. The attribute information coding unit 7023 includes a conversion unit 7035, a quantization unit 7036, and an entropy coding unit 7037.
[0522] The conversion unit 7035 classifies the partitioned attribute information into a hierarchy such as LoD, and generates coefficient values (difference values) by calculating the difference between the partitioned attribute information and the predicted values. Alternatively, the conversion unit 7035 may generate coefficient values by performing a Haar transformation on the partitioned attribute information.
[0523] The quantization unit 7036 generates quantized values by quantizing the coefficient values. Specifically, the quantization unit 7036 divides the coefficients in a quantization step based on the quantization parameters. The entropy coding unit 7037 generates coded attribute information by entropy coding the quantized values.
[0524] The configuration of the three-dimensional data decoding device according to this embodiment will be described below. Figure 71 is a block diagram showing the configuration of the three-dimensional data decoding device 7040. The three-dimensional data decoding device 7040 restores point cloud data by decoding encoded data (encoded stream) generated when point cloud data is encoded. This three-dimensional data decoding device 7040 includes a demultiplexing unit 7041, a plurality of position information decoding units 7042, a plurality of attribute information decoding units 7043, an additional information decoding unit 7044, and a coupling unit 7045.
[0525] The demultiplexing unit 7041 generates multiple encoded position information, multiple encoded attribute information, and encoded additional information by demultiplexing the encoded data (encoded stream).
[0526] The multiple location information decoding units 7042 generate multiple segmented location information by decoding multiple encoded location information. For example, the multiple location information decoding units 7042 process multiple encoded location information in parallel.
[0527] The multiple attribute information decoding unit 7043 generates multiple segmented attribute information by decoding multiple encoded attribute information. For example, the multiple attribute information decoding unit 7043 processes multiple encoded attribute information in parallel.
[0528] Multiple additional information decoding units 7044 generate additional information by decoding encoded additional information.
[0529] The coupling unit 7045 generates position information by combining multiple division position information using additional information. The coupling unit 7045 generates attribute information by combining multiple division attribute information using additional information.
[0530] Figure 72 is a block diagram of the attribute information decoding unit 7043. The attribute information decoding unit 7043 includes an entropy decoding unit 7051, an inverse quantization unit 7052, and an inverse transform unit 7053. The entropy decoding unit 7051 generates quantized values by entropy decoding the encoded attribute information. The inverse quantization unit 7052 generates coefficient values by inverse quantization of the quantized values. Specifically, it multiplies the coefficient values by a quantization step based on the quantization tree values (quantization parameters) obtained from the bitstream. The inverse transform unit 7053 generates partitioned attribute information by inverse transforming the coefficient values. Here, inverse transform is, for example, a process of adding a predicted value to the coefficient value. Alternatively, inverse transform is an inverse Haar transform.
[0531] The following describes an example of how to determine quantization parameters. Figure 73 shows an example of setting quantization parameters in tile and slice partitioning.
[0532] When the quantization parameter value is small, the original information is more easily preserved. For example, the default value for the quantization parameter is 1. For instance, in encoding processing using PCC data tiles, the quantization parameter for tiles representing major roads is set to a small value to maintain data quality. On the other hand, the quantization parameter for tiles representing surrounding areas is set to a large value. This reduces the data quality of the surrounding areas, but improves encoding efficiency.
[0533] Similarly, in encoding processes using PCC data slices, sidewalks, trees, and buildings are important for self-localization and mapping, so the quantization parameters for the sidewalk, tree, and building slices are set to small values. On the other hand, since moving objects and other data are less important, the quantization parameters for the moving object and other data slices are set to high values.
[0534] Furthermore, when using ΔQP (DeltaQP), as described later, the three-dimensional data encoding device may set a negative value for ΔQP to reduce the quantization parameter when encoding three-dimensional points belonging to important areas such as major roads, thereby reducing the quantization error. This makes it possible to bring the decoded attribute values of three-dimensional points belonging to important areas closer to their values before encoding. Also, when encoding three-dimensional points belonging to less important areas such as surrounding regions, the three-dimensional data encoding device may set a positive value for ΔQP to increase the quantization parameter, thereby reducing the amount of information. This makes it possible to reduce the overall coding amount while maintaining the amount of information for important areas.
[0535] The following describes an example of information indicating quantization parameters for each layer. When quantizing and encoding attribute information of three-dimensional points, in addition to the quantization parameter QPbase for frames, slices, or tiles, a mechanism is introduced to control quantization parameters at a finer level. For example, when a three-dimensional data encoding device encodes attribute information using Levels of Data (LoD), it sets up a Delta_Layer for each LoD and performs encoding while changing the value of the quantization parameter by adding the Delta_Layer to the value of QPbase for each LoD. The three-dimensional data encoding device also adds the Delta_Layer used for encoding to the bitstream header, etc. This allows the three-dimensional data encoding device to encode the attribute information of three-dimensional points while changing the quantization parameter for each LoD according to, for example, the target code amount and the generated code amount, so that it can ultimately generate a bitstream with a code amount close to the target code amount. Furthermore, the three-dimensional data decoding device can appropriately decode the bitstream by decoding the QPbase and Delta_Layer contained in the header to generate the quantization parameters used by the three-dimensional data encoding device.
[0536] Figure 74 shows an example of encoding the attribute information of all three-dimensional points using the quantization parameter QPbase. Figure 75 shows an example of encoding by switching the quantization parameter for each level of the LoD. In the example shown in Figure 75, the quantization parameter of the first LoD is calculated by adding the Delta_Layer of the first LoD to QPbase. For the second and subsequent LoDs, the quantization parameter of the LoD to be processed is calculated by adding the Delta_Layer of the LoD to the quantization parameter of the previous LoD. For example, the first quantization parameter QP3 of LoD3 is calculated as QP3 = QP2 + Delta_Layer[3].
[0537] Note that Delta_Layer[i] for each LoD may represent the difference from QPbase. That is, the quantization parameter QPi of the i-th LoDi is given by QPi = QPbase + Delta_Layer[i]. For example, QP1 is given by QPbase + Delta_Layer[1] and QP2 is given by QPbase + Delta_Layer[2].
[0538] Figure 76 shows an example of the syntax for an attribute header information. Here, an attribute header is, for example, a header at the frame, slice, or tile level, and is a header for attribute information. As shown in Figure 76, the attribute header includes QPbase (base quantization parameter), NumLayer (number of layers), and Delta_Layer[i] (difference quantization parameter).
[0539] QPbase indicates the value of the quantization parameter that serves as the basis for a frame, slice, or tile. NumLayer indicates the number of layers in the LoD or RAHT. In other words, NumLayer indicates the number of Delta_Layer[i] included in the attribute information header.
[0540] Delta_Layer[i] represents the value of ΔQP for layer i. Here, ΔQP is the value obtained by subtracting the quantization parameter of layer i from the quantization parameter of layer i-1. Alternatively, ΔQP may be the value obtained by subtracting the quantization parameter of layer i from QPbase. Furthermore, ΔQP can take on a positive or negative value. Note that Delta_Layer[0] does not need to be added to the header. In this case, the quantization parameter of layer 0 is equal to QPbase. This reduces the header code size.
[0541] Figure 77 shows another example of the attribute header information syntax. The attribute header shown in Figure 77 includes the delta_Layer_present_flag in addition to the attribute header shown in Figure 76.
[0542] The `delta_Layer_present_flag` flag indicates whether or not the Delta_Layer is included in the bitstream. For example, a value of 1 indicates that the Delta_Layer is included in the bitstream, and a value of 0 indicates that the Delta_Layer is not included in the bitstream. If `delta_Layer_present_flag` is 0, the 3D data decoder will, for example, treat the Delta_Layer as 0 and proceed with the subsequent decoding process.
[0543] Here, we have described an example where the quantization parameters are shown using QPbase and Delta_Layer, but the quantization steps may also be shown using QPbase and Delta_Layer. The quantization steps are calculated from the quantization parameters using a predetermined formula or table. In the quantization process, the three-dimensional data encoding device divides the coefficient values by the quantization steps. In the inverse quantization process, the three-dimensional data decoding device recovers the coefficient values by multiplying the quantized values by the quantization steps.
[0544] Next, we will explain an example of controlling quantization parameters at an even finer level. Figure 78 shows an example of controlling quantization parameters at a finer level than the Level of Dimension (LoD).
[0545] For example, when a three-dimensional data encoding device encodes attribute information using Level of Data (LoD), it defines a Delta_Layer for each LoD hierarchy, as well as an ADelta_QP and a NumPointADelta representing the position information of the three-dimensional point to which the ADelta_QP is added. The three-dimensional data encoding device performs encoding while changing the values of the quantization parameters based on the Delta_Layer, ADelta_QP, and NumPointADelta.
[0546] Furthermore, the three-dimensional data encoding device may add the ADelta and NumPointADelta used for encoding to the bitstream header, etc. This allows the three-dimensional data encoding device to encode attribute information of multiple three-dimensional points while changing the quantization parameters for each three-dimensional point according to the target code amount and the generated code amount. As a result, the three-dimensional data encoding device can ultimately generate a bitstream with a code amount close to the target code amount. In addition, the three-dimensional data decoding device can appropriately decode the bitstream by decoding the QPbase, Delta_Layer, and ADelta contained in the header to generate the quantization parameters used by the three-dimensional data encoding device.
[0547] For example, as shown in Figure 78, the quantized value QP4 of the N0th attribute information is calculated as QP4 = QP3 + ADelta_QP[0].
[0548] Furthermore, an encoding / decoding order reversed from the one shown in Figure 78 may be used. For example, encoding / decoding may be performed in the order of LoD3, LoD2, LoD1, LoD0.
[0549] Figure 79 shows an example of the syntax for the attribute header information when using the example shown in Figure 78. The attribute header information shown in Figure 79 further includes NumADelta, NumPointADelta[i], and ADelta_QP[i] in addition to the attribute header information shown in Figure 76.
[0550] NumADelta indicates the number of ADelta_QPs contained in the bitstream. NumPointADelta[i] indicates the identification number of the 3D point A to which ADelta_QP[i] is applied. For example, NumPointADelta[i] indicates the number of 3D points from the first 3D point to 3D point A in the encoding / decoding order. Alternatively, NumPointADelta[i] may indicate the number of 3D points from the first 3D point in the LoD to which 3D point A belongs to to 3D point A.
[0551] Alternatively, NumPointADelta[i] may represent the difference between the identification number of the three-dimensional point indicated by NumPointADelta[i-1] and the identification number of three-dimensional point A. This allows the value of NumPointADelta[i] to be reduced, thereby reducing the amount of sign.
[0552] ADelta_QP[i] represents the value of ΔQP for the three-dimensional point indicated by NumPointADelta[i]. In other words, ADelta_QP[i] represents the difference between the quantization parameter of the three-dimensional point indicated by NumPointADelta[i] and the quantization parameter of the three-dimensional point immediately preceding that point.
[0553] Figure 80 shows an alternative syntax example of the attribute header information when using the example shown in Figure 78. The attribute header information shown in Figure 80 further includes delta_Layer_present_flag and additional_delta_QP_present_flag, and includes NumADelta_minus1 instead of NumADelta, compared to the attribute header information header shown in Figure 79.
[0554] delta_Layer_present_flag is the same as the one already explained using Figure 77.
[0555] `additional_delta_QP_present_flag` is a flag that indicates whether ADelta_QP is included in the bitstream. For example, a value of 1 indicates that ADelta_QP is included in the bitstream, and a value of 0 indicates that ADelta_QP is not included in the bitstream. If `additional_delta_QP_present_flag` is 0, the 3D data decoder will, for example, treat ADelta_QP as 0 and proceed with the subsequent decoding process.
[0556] NumADelta_minus1 represents the number of ADelta_QPs in the bitstream minus 1. By adding a value obtained by subtracting 1 from the number of ADelta_QPs to the header, the amount of code in the header can be reduced. For example, a three-dimensional data decoder calculates NumADelta = NumADelta_minus1 + 1. ADelta_QP[i] represents the value of the i-th ADelta_QP. Note that ADelta_QP[i] can be set to a negative value as well as a positive value.
[0557] Figure 81 is a flowchart of the three-dimensional data encoding process according to this embodiment. First, the three-dimensional data encoding device encodes the position information (geometry) (S7001). For example, the three-dimensional data encoding device performs encoding using an octree representation.
[0558] Next, the three-dimensional data encoding device transforms the attribute information (S7002). For example, if the position of a three-dimensional point changes due to quantization or the like after encoding the position information, the three-dimensional data encoding device reassigns the attribute information of the original three-dimensional point to the changed three-dimensional point. The three-dimensional data encoding device may also interpolate the attribute information values according to the amount of change in position before reassignment. For example, the three-dimensional data encoding device detects N three-dimensional points that are close to the changed three-dimensional position before the change, and weights and averages the attribute information values of the N three-dimensional points based on the distance from the changed three-dimensional position to each of the N three-dimensional points, determining the resulting value as the attribute information value of the changed three-dimensional point. Furthermore, if two or more three-dimensional points change to the same three-dimensional position due to quantization or the like, the three-dimensional data encoding device may assign the average value of the attribute information of the two or more three-dimensional points before the change as the attribute information value of the changed point.
[0559] Next, the three-dimensional data encoding device encodes attribute information (S7003). For example, if the three-dimensional data encoding device encodes multiple attribute information, it may encode the multiple attribute information sequentially. For example, if the three-dimensional data encoding device encodes color and reflectance as attribute information, it generates a bitstream in which the encoded result of reflectance is appended after the encoded result of color. The order in which the multiple encoded results of attribute information are appended to the bitstream does not matter.
[0560] Furthermore, the three-dimensional data encoding device may add information to the header or elsewhere indicating the starting location of the encoded data for each attribute within the bitstream. This allows the three-dimensional data decoding device to selectively decode the attribute information that needs to be decoded, thus omitting the decoding process for attribute information that does not need to be decoded. Therefore, the processing load on the three-dimensional data decoding device can be reduced. In addition, the three-dimensional data encoding device may encode multiple attribute information in parallel and integrate the encoding results into a single bitstream. This allows the three-dimensional data encoding device to encode multiple attribute information at high speed.
[0561] Figure 82 is a flowchart of the attribute information encoding process (S7003). First, the three-dimensional data encoding device sets the Level of Data (LoD) (S7011). In other words, the three-dimensional data encoding device assigns each three-dimensional point to one of several LoDs.
[0562] Next, the three-dimensional data encoding device starts a loop for each Level of Data (LoD) (S7012). In other words, the three-dimensional data encoding device repeats the process from steps S7013 to S7021 for each LoD.
[0563] Next, the three-dimensional data encoding device starts a loop for each three-dimensional point (S7013). In other words, the three-dimensional data encoding device repeats the process from steps S7014 to S7020 for each three-dimensional point.
[0564] First, the 3D data encoding device searches for multiple surrounding points, which are 3D points that exist around the target 3D point to be processed, and uses them to calculate the predicted value of the target 3D point (S7014). Next, the 3D data encoding device calculates the weighted average of the attribute information values of the multiple surrounding points and sets the obtained value as the predicted value P (S7015). Next, the 3D data encoding device calculates the prediction residual, which is the difference between the attribute information of the target 3D point and the predicted value (S7016). Next, the 3D data encoding device calculates the quantized value by quantizing the prediction residual (S7017). Next, the 3D data encoding device arithmetically encodes the quantized value (S7018). Next, the 3D data encoding device determines ΔQP (S7019). The ΔQP determined here is used to determine the quantization parameter used in the subsequent quantization of the prediction residual.
[0565] Furthermore, the three-dimensional data encoding device calculates the inverse quantized value by inverse quantizing the quantized value (S7020). Next, the three-dimensional data encoding device generates the decoded value by adding the predicted value to the inverse quantized value (S7021). Next, the three-dimensional data encoding device terminates the loop for three-dimensional points (S7022). Furthermore, the three-dimensional data encoding device terminates the loop for LoDs (S7023).
[0566] Figure 83 is a flowchart of the ΔQP determination process (S7019). First, the three-dimensional data encoding device calculates the hierarchy i to which the target three-dimensional point A to be encoded next belongs and the encoding order N (S7031). Hierarchy i represents, for example, the LoD hierarchy or the RAHT hierarchy.
[0567] Next, the three-dimensional data encoding device adds the generated code amount to the cumulative code amount (S7032). Here, the cumulative code amount is the cumulative code amount for one frame, one slice, or one tile of the target three-dimensional point. Note that the cumulative code amount may be the cumulative code amount obtained by adding the code amounts of multiple frames, multiple slices, or multiple tiles. Furthermore, the cumulative code amount of attribute information may be used, or the cumulative code amount obtained by adding both position information and attribute information may be used.
[0568] Next, the three-dimensional data encoding device determines whether the cumulative code amount is greater than the target code amount × TH1 (S7033). Here, the target code amount is the target code amount for one frame, one slice, or one tile of the target three-dimensional point. Note that the target code amount may be the sum of multiple frames, multiple slices, or multiple tiles. Furthermore, the target code amount of attribute information may be used, or the target code amount of both position information and attribute information may be used.
[0569] If the cumulative code amount is less than or equal to the target code amount × TH1 (No in S7033), the three-dimensional data encoding device determines whether the cumulative code amount is greater than the target code amount × TH2 (S7036).
[0570] Here, thresholds TH1 and TH2 are set to values ranging from 0.0 to 1.0, for example. Also, TH1 > TH2. For example, if the cumulative code amount exceeds the value of target code amount × TH1 (Yes in S7033), the three-dimensional data encoding device determines that it is necessary to suppress the code amount immediately and sets ADelta_QP to value α in order to increase the quantization parameter of the next three-dimensional point N. The three-dimensional data encoding device also sets NumPointADelta to value N and increments j by 1 (S7034). Next, the three-dimensional data encoding device adds ADelta_QP=α and NumPointADelta=N to the header (S7035). Note that the value α may be a fixed value or a variable value. For example, the three-dimensional data encoding device may determine the value of α based on the magnitude of the difference between the cumulative code amount and target code amount × TH1. For example, the three-dimensional data encoding device sets a larger value of α the larger the difference between the cumulative code amount and target code amount × TH1. This allows the three-dimensional data encoding device to control the quantization parameters so that the cumulative code amount does not exceed the target code amount.
[0571] Furthermore, if the cumulative code amount exceeds the value of target code amount × TH2 (Yes in S7036), the three-dimensional data encoding device sets Delta_Layer to value β in order to increase the quantization parameter of layer i to which the target three-dimensional point A belongs or the next layer i+1 (S7037). For example, if the target three-dimensional point A is at the beginning of layer i, the three-dimensional data encoding device sets Delta_Layer[i] of layer i to value β, and if the target three-dimensional point A is not at the beginning of layer i, it sets Delta_Layer[i+1] of layer i+1 to value β.
[0572] Furthermore, the three-dimensional data encoding device adds Delta_Layer=β of layer i or layer i+1 to the header (S7038). Note that the value β may be a fixed value or a variable value. For example, the three-dimensional data encoding device may determine the value of β based on the magnitude of the difference between the cumulative code amount and the target code amount × TH2. For example, the three-dimensional data encoding device sets a larger value of β the larger the difference between the cumulative code amount and the target code amount × TH2. This allows the three-dimensional data encoding device to control the quantization parameters so that the cumulative code amount does not exceed the target code amount.
[0573] Furthermore, if the cumulative code amount exceeds or is likely to exceed the target code amount, the three-dimensional data encoding device may set the value of ADelta_QP or Delta_Layer so that the quantization parameter reaches the maximum value supported by the standard. This allows the three-dimensional data encoding device to suppress the increase in generated code amount by setting the quantization coefficients to 0 from three-dimensional point A onwards, or from layer i onwards, thereby controlling the cumulative code amount so that it does not exceed the target code amount.
[0574] Furthermore, the three-dimensional data encoding device may lower the quantization parameters so that the generated code amount increases if the cumulative code amount is less than the target code amount × TH3. For example, the three-dimensional data encoding device may lower the quantization parameters by setting a negative value for Delta_Layer or Adelta_QP according to the difference between the cumulative code amount and the target code amount. This allows the three-dimensional data encoding device to generate a bitstream close to the target code amount.
[0575] Figure 84 is a flowchart of the three-dimensional data decoding process according to this embodiment. First, the three-dimensional data decoding device decodes the position information (geometry) from the bitstream (S7005). For example, the three-dimensional data decoding device performs decoding using an octave tree representation.
[0576] Next, the three-dimensional data decoding device decodes attribute information from the bitstream (S7006). For example, if the three-dimensional data decoding device decodes multiple attribute information, it may decode the multiple attribute information sequentially. For example, if the three-dimensional data decoding device decodes color and reflectance as attribute information, it decodes the color encoding result and the reflectance encoding result in the order in which they are added to the bitstream. For example, if the reflectance encoding result is added after the color encoding result in the bitstream, the three-dimensional data decoding device decodes the color encoding result, and then decodes the reflectance encoding result. The three-dimensional data decoding device may decode the encoding results of the attribute information added to the bitstream in any order.
[0577] Furthermore, the three-dimensional data decoding device may obtain information indicating the starting location of the encoded data for each attribute information within the bitstream by decoding the header, etc. This allows the three-dimensional data decoding device to selectively decode the attribute information that needs to be decoded, thus omitting the decoding process for attribute information that does not need to be decoded. Therefore, the processing load of the three-dimensional data decoding device can be reduced. In addition, the three-dimensional data decoding device may decode multiple attribute information in parallel and integrate the decoding results into a single three-dimensional point cloud. This allows the three-dimensional data decoding device to decode multiple attribute information at high speed.
[0578] Figure 85 is a flowchart of the attribute information decoding process (S7006). First, the three-dimensional data decoding device sets the Level of Direction (LoD) (S7041). That is, the three-dimensional data decoding device assigns each of the multiple three-dimensional points with decoded position information to one of the multiple LoDs. For example, this assignment method is the same as the assignment method used in the three-dimensional data encoding device.
[0579] Next, the three-dimensional data decoding device decodes ΔQP from the bitstream (S7042). Specifically, the three-dimensional data encoding device decodes Delta_Layer, ADelta_QP, and NumPointADelta from the bitstream header.
[0580] Next, the three-dimensional data decoding device starts a loop for each Level of Data (LoD) (S7043). In other words, the three-dimensional data decoding device repeats the process from steps S7044 to S7050 for each LoD.
[0581] Next, the three-dimensional data decoding device starts a loop for each three-dimensional point (S7044). In other words, the three-dimensional data decoding device repeats the process from steps S7045 to S7049 for each three-dimensional point.
[0582] First, the 3D data decoding device searches for multiple surrounding points, which are 3D points that exist around the target 3D point to be processed, in order to calculate the predicted value of the target 3D point (S7045). Next, the 3D data decoding device calculates a weighted average of the attribute information values of the multiple surrounding points and sets the obtained value as the predicted value P (S7046). These processes are the same as those performed in the 3D data encoding device.
[0583] Next, the three-dimensional data decoding device arithmetically decodes the quantized values from the bitstream (S7047). The three-dimensional data decoding device also calculates the inverse quantized values by inverse quantizing the decoded quantized values (S7048). In this inverse quantization, the quantization parameters calculated using ΔQP obtained in step S7042 are used.
[0584] Next, the three-dimensional data decoding device generates a decoded value by adding a predicted value to the inverse quantized value (S7049). Then, the three-dimensional data decoding device terminates the loop for each three-dimensional point (S7050). Finally, the three-dimensional data decoding device terminates the loop for each Line of Data (LoD) (S7051).
[0585] Figure 86 is a block diagram of the attribute information coding unit 7023. The attribute information coding unit 7023 comprises an LoD setting unit 7061, a search unit 7062, a prediction unit 7063, a subtraction unit 7064, a quantization unit 7065, an inverse quantization unit 7066, a reconstruction unit 7067, a memory 7068, and a ΔQP calculation unit 7070.
[0586] The LoD setting unit 7061 generates an LoD using the position information of the three-dimensional points. The search unit 7062 searches for neighboring three-dimensional points of each three-dimensional point using the LoD generation result and the distance information between the three-dimensional points. The prediction unit 7063 generates predicted values for the attribute information of the target three-dimensional point. The prediction unit 7063 also assigns the predicted values to multiple prediction modes from 0 to M-1 and selects the prediction mode to be used from the multiple prediction modes.
[0587] The subtraction unit 7064 generates a prediction residual by subtracting the predicted value from the attribute information. The quantization unit 7065 quantizes the prediction residual of the attribute information. The inverse quantization unit 7066 inverse quantizes the prediction residual after quantization. The reconstruction unit 7067 generates a decoded value by adding the predicted value and the prediction residual after inverse quantization. The memory 7068 stores the attribute information values (decoded values) of each decoded three-dimensional point. The attribute information of the decoded three-dimensional points stored in the memory 7068 is used by the prediction unit 7063 to predict the unencoded three-dimensional points.
[0588] The arithmetic coding unit 7069 calculates ZeroCnt from the quantized prediction residuals and arithmetically codes ZeroCnt. The arithmetic coding unit 7069 also arithmetically codes the non-zero prediction residuals after quantization. The arithmetic coding unit 7069 may binarize the prediction residuals before arithmetic coding. The arithmetic coding unit 7069 may also generate and code various header information. The arithmetic coding unit 7069 may also arithmetically code prediction mode information (PredMode) indicating the prediction mode used for coding by the prediction unit 7063 and add it to the bitstream.
[0589] The ΔQP calculation unit 7070 determines the values of Delta_Layer, ADelta_QP, and NumPointADelta from the generated code amount obtained by the arithmetic coding unit 7069 and a predetermined target code amount. Quantization is performed by the quantization unit 7065 using quantization parameters based on the determined Delta_Layer, ADelta_QP, and NumPointADelta. The arithmetic coding unit 7069 then arithmetically codes Delta_Layer, ADelta_QP, and NumPointADelta and adds them to the bitstream.
[0590] Figure 87 is a block diagram of the attribute information decoding unit 7043. The attribute information decoding unit 7043 comprises an arithmetic decoding unit 7071, an LoD setting unit 7072, a search unit 7073, a prediction unit 7074, an inverse quantization unit 7075, a reconstruction unit 7076, and a memory 7077.
[0591] The arithmetic decoding unit 7071 arithmetically decodes ZeroCnt and prediction residuals contained in the bitstream. The arithmetic decoding unit 7071 also decodes various header information. Furthermore, the arithmetic decoding unit 7071 arithmetically decodes prediction mode information (PredMode) from the bitstream and outputs the obtained prediction mode information to the prediction unit 7074. In addition, the arithmetic decoding unit 7071 decodes Delta_Layer, ADelta_QP, and NumPointADelta from the bitstream header.
[0592] The LoD setting unit 7072 generates an LoD using the decoded position information of the three-dimensional points. The search unit 7073 searches for neighboring three-dimensional points of each three-dimensi...
Claims
1. A three-dimensional data encoding method performed by an encoding device, From the attribute information of multiple three-dimensional points contained in point cloud data, multiple coefficient values belonging to one of multiple hierarchies are generated, The aforementioned multiple coefficient values are subjected to quantization to generate multiple quantized values, The system generates a bitstream containing the aforementioned multiple quantization values, information indicating the total number of layers to which the layer parameters are notified, and the layer parameters. In the aforementioned quantization, In the quantization of coefficient values belonging to the hierarchy to which the aforementioned hierarchy parameters are notified, the quantization step is derived using the aforementioned hierarchy parameters. In the quantization of coefficient values belonging to a hierarchy for which the aforementioned hierarchy parameters are not notified, the quantization step is derived using the hierarchy parameters of the last hierarchy for which the aforementioned hierarchy parameters are notified. Three-dimensional data encoding method.
2. A three-dimensional data decoding method performed by a decoding device, From the bitstream, obtain multiple quantization values, information indicating the total number of hierarchical levels for which hierarchical parameters are notified, and the hierarchical parameters. Inverse quantization is performed on the aforementioned plurality of quantization values to generate a plurality of coefficient values belonging to one of the plurality of hierarchies, From the aforementioned coefficient values, multiple attribute information of multiple three-dimensional points included in the point cloud data is calculated, In the aforementioned inverse quantization, In the inverse quantization of quantized values belonging to the hierarchy to which the aforementioned hierarchy parameters are notified, the quantization step is derived using the aforementioned hierarchy parameters. In the inverse quantization of quantized values belonging to a hierarchy for which the aforementioned hierarchy parameters are not notified, the quantization step is derived using the hierarchy parameters of the last hierarchy for which the aforementioned hierarchy parameters are notified. Three-dimensional data decoding method.
3. Processor and Equipped with memory, The processor uses the memory to: From the attribute information of multiple three-dimensional points contained in point cloud data, multiple coefficient values belonging to one of multiple hierarchies are generated, The aforementioned multiple coefficient values are subjected to quantization to generate multiple quantized values, The system generates a bitstream containing the aforementioned multiple quantization values, information indicating the total number of layers to which the layer parameters are notified, and the layer parameters. In the aforementioned quantization, In the quantization of coefficient values belonging to the hierarchy to which the aforementioned hierarchy parameters are notified, the quantization step is derived using the aforementioned hierarchy parameters. In the quantization of coefficient values belonging to a hierarchy for which the aforementioned hierarchy parameters are not notified, the quantization step is derived using the hierarchy parameters of the last hierarchy for which the aforementioned hierarchy parameters are notified. Three-dimensional data encoding device.
4. Processor and Equipped with memory, The processor uses the memory to: From the bitstream, obtain multiple quantization values, information indicating the total number of hierarchical levels for which hierarchical parameters are notified, and the hierarchical parameters. Inverse quantization is performed on the aforementioned plurality of quantization values to generate a plurality of coefficient values belonging to one of the plurality of hierarchies, From the aforementioned coefficient values, multiple attribute information of multiple three-dimensional points included in the point cloud data is calculated, In the aforementioned inverse quantization, In the inverse quantization of quantized values belonging to the hierarchy to which the aforementioned hierarchy parameters are notified, the quantization step is derived using the aforementioned hierarchy parameters. In the inverse quantization of quantized values belonging to a hierarchy for which the aforementioned hierarchy parameters are not notified, the quantization step is derived using the hierarchy parameters of the last hierarchy for which the aforementioned hierarchy parameters are notified. Three-dimensional data decoding device.
Citation Information
Patent Citations
Coding method and apparatus for image signal
JP2006108792A
Image encoding device, image encoding method, and image encoding program
JP2017076979A
Data coding method, data coding device, and data coding program
JP2018078503A
Scalable point cloud compression with transform, and corresponding decompression
US20170347122A1
Map display device
WO2014020663A1