Three-dimensional data encoding method, three-dimensional data decoding method, three-dimensional data encoding device, and three-dimensional data decoding device
By predictive coding and binary processing of three-dimensional points, the problem of difficult to reduce the encoding amount in three-dimensional data encoding is solved, and more efficient data transmission and storage is achieved.
Patent Information
- Application Number
- CN202510115735.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2018-06-06
- Filing Date
- 2019-05-30
- Publication Date
- 2025-05-06
AI Technical Summary
The prior art is difficult to effectively reduce the encoding amount in three-dimensional data encoding, resulting in inefficient data transmission and storage.
By predictively encoding three-dimensional points, the predicted value of the attribute information of the point cloud and its prediction residuals are calculated, and binarized and arithmetic encoding are performed to reduce the amount of encoded data.
It realizes the reduction of encoding amount in three-dimensional data encoding, thereby improving the efficiency of data transmission and storage.
Smart Images

Figure CN119941881A_ABST
Abstract
Description
[0001] This application was filed on May 30, 2019, with Chinese patent application number 201980037127.1 (international application number PCT / JP2019 / 021636), and is a divisional application of a patent application titled “Three-dimensional data encoding method, three-dimensional data decoding method, three-dimensional data encoding device, and three-dimensional data decoding device”. Technical Field
[0002] The present disclosure relates to a three-dimensional data encoding method, a three-dimensional data decoding method, a three-dimensional data encoding device, and a three-dimensional data decoding device. Background Art
[0003] In the future, devices and services that make use of 3D data will become more common in large fields such as computer vision, map information, monitoring, infrastructure inspection, or image distribution, which are used for autonomous operation of cars or robots. 3D data is obtained by various methods such as distance sensors such as rangefinders, stereo cameras, or a combination of multiple single-lens reflex cameras.
[0004] As a method of expressing three-dimensional data, there is a method called point cloud, which expresses the shape of a three-dimensional structure through a group of points in a three-dimensional space. The position and color of the point group are stored in the point cloud. Although point cloud is expected to become the mainstream method of expressing three-dimensional data, the amount of point group data is very large. Therefore, in the accumulation or transmission of three-dimensional data, it is necessary to compress the data volume through encoding, just like two-dimensional dynamic images (as an example, there are MPEG-4AVC or HEVC standardized by MPEG).
[0005] Furthermore, compression of point clouds is partially supported by a public library (PointCloud Library) that performs point cloud association processing.
[0006] Furthermore, there is a known technique for searching for facilities around a vehicle using three-dimensional map data and displaying the facilities (for example, refer to Patent Document 1).
[0007] Prior art literature
[0008] Patent Literature
[0009] Patent Document 1 International Publication No. 2014 / 020663 Summary of the invention
[0010] Problem that the invention aims to solve
[0011] It is desired to reduce the amount of code in encoding three-dimensional data.
[0012] An object of the present disclosure is to provide a three-dimensional data encoding method, a three-dimensional data decoding method, a three-dimensional data encoding device, or a three-dimensional data decoding device capable of reducing the amount of encoding.
[0013] Means used to solve problems
[0014] A three-dimensional data encoding method according to one embodiment of the present invention is a three-dimensional data encoding method for encoding three-dimensional points, comprising: determining whether to use the first three-dimensional point to predict attribute information of the second three-dimensional point based on a distance between the first three-dimensional point and the second three-dimensional point; and calculating a predicted value of the attribute information of the second three-dimensional point based on the first three-dimensional point when the first three-dimensional point is within a predetermined distance range from the second three-dimensional point.
[0015] A three-dimensional data decoding method according to one embodiment of the present invention is a three-dimensional data decoding method for decoding a three-dimensional point, comprising: determining whether to use the first three-dimensional point to predict attribute information of the second three-dimensional point based on a distance between the first three-dimensional point and the second three-dimensional point; and calculating a predicted value of the attribute information of the second three-dimensional point based on the first three-dimensional point when the first three-dimensional point is within a predetermined distance range from the second three-dimensional point.
[0016] A three-dimensional data encoding device according to one embodiment of the present invention is a three-dimensional data encoding device for encoding three-dimensional points, wherein the device comprises a processor and a memory, wherein the processor uses the memory to determine whether to use the first three-dimensional point to predict attribute information of the second three-dimensional point based on a distance between the first three-dimensional point and the second three-dimensional point; and when the first three-dimensional point is within a predetermined distance range from the second three-dimensional point, the predicted value of the attribute information of the second three-dimensional point is calculated based on the first three-dimensional point.
[0017] A three-dimensional data decoding device according to one embodiment of the present invention is a three-dimensional data decoding device for decoding three-dimensional points, wherein the device comprises a processor and a memory, wherein the processor uses the memory to determine whether to use the first three-dimensional point to predict attribute information of the second three-dimensional point based on a distance between the first three-dimensional point and the second three-dimensional point; and when the first three-dimensional point is within a predetermined distance range from the second three-dimensional point, the predicted value of the attribute information of the second three-dimensional point is calculated based on the first three-dimensional point.
[0018] A three-dimensional data encoding method according to one embodiment of the present invention is a three-dimensional data encoding method for encoding three-dimensional points having attribute information, calculating a predicted value of the attribute information of the three-dimensional point, calculating a difference between the attribute information of the three-dimensional point and the predicted value, i.e., a prediction residual, generating binary data by binarizing the prediction residual, and performing arithmetical encoding on the binary data.
[0019] A three-dimensional data decoding method according to one embodiment of the present invention is a three-dimensional data decoding method for decoding a three-dimensional point having attribute information, calculating a predicted value of the attribute information of the three-dimensional point, generating binary data by arithmetically decoding encoded data contained in a bit stream, generating a prediction residual by multi-valuedizing the binary data, and calculating a decoded value of the attribute information of the three-dimensional point by adding the predicted value to the prediction residual.
[0020] Effects of the Invention
[0021] The present disclosure can provide a three-dimensional data encoding method, a three-dimensional data decoding method, a three-dimensional data encoding device, or a three-dimensional data decoding device capable of reducing the amount of encoding. BRIEF DESCRIPTION OF THE DRAWINGS
[0022] Figure 1 The structure of the encoded three-dimensional data according to the first embodiment is shown.
[0023] Figure 2 An example of a prediction structure between SPCs belonging to the lowest layer of the GOS according to the first embodiment is shown.
[0024] Figure 3 An example of an inter-layer prediction structure in embodiment 1 is shown.
[0025] Figure 4 An example of the encoding order of the GOS according to the first embodiment is shown.
[0026] Figure 5 An example of the encoding order of the GOS according to the first embodiment is shown.
[0027] Figure 6 This is a block diagram of a three-dimensional data encoding device according to Embodiment 1.
[0028] Figure 7 This is a flowchart of the encoding process of implementation mode 1.
[0029] Figure 8 This is a block diagram of a three-dimensional data decoding device according to Embodiment 1.
[0030] Fig. 9 This is a flowchart of the decoding process in implementation mode 1.
[0031] Fig.10 An example of meta-information according to the first embodiment is shown.
[0032] Fig.11 A configuration example of a SWLD according to the second embodiment is shown.
[0033] Fig.12An operation example of the server and the client according to the second embodiment is shown.
[0034] Fig.13 An operation example of the server and the client according to the second embodiment is shown.
[0035] Fig.14 An operation example of the server and the client according to the second embodiment is shown.
[0036] Fig.15 An operation example of the server and the client according to the second embodiment is shown.
[0037] Fig.16 This is a block diagram of a three-dimensional data encoding device according to Embodiment 2.
[0038] Fig.17 This is a flowchart of the encoding process of implementation mode 2.
[0039] Fig.18 This is a block diagram of a three-dimensional data decoding device according to Embodiment 2.
[0040] Fig.19 This is a flowchart of the decoding process of implementation mode 2.
[0041] Fig. 20 A configuration example of a WLD according to the second embodiment is shown.
[0042] Fig.21 An example of the octree structure of the WLD according to the second embodiment is shown.
[0043] Fig. 22 A configuration example of a SWLD according to the second embodiment is shown.
[0044] Fig.23 An example of the octree structure of the SWLD according to the second embodiment is shown.
[0045] Fig.24 This is a block diagram of a three-dimensional data creation device according to the third embodiment.
[0046] Fig.25 This is a block diagram of a three-dimensional data transmitting device according to Embodiment 3.
[0047] Fig.26 This is a block diagram of a three-dimensional information processing device according to a fourth embodiment.
[0048] Fig. 27 This is a block diagram of a three-dimensional data creation device according to a fifth embodiment.
[0049] Fig.28 The configuration of the system according to the sixth embodiment is shown.
[0050] Fig.29This is a block diagram of a client device according to a sixth embodiment.
[0051] Fig.30 This is a block diagram of a server in implementation mode 6.
[0052] Fig.31 This is a flowchart of the three-dimensional data creation process performed by the client device of the sixth embodiment.
[0053] Fig.32 This is a flowchart of the sensor information transmission process performed by the client device according to the sixth embodiment.
[0054] Fig.33 This is a flowchart of the three-dimensional data production processing performed by the server of the sixth embodiment.
[0055] Fig.34 This is a flowchart of the three-dimensional map transmission processing performed by the server in the sixth embodiment.
[0056] Fig.35 The configuration of a modified example of the system of the sixth embodiment is shown.
[0057] Fig.36 The configuration of the server and client device according to the sixth embodiment is shown.
[0058] Fig.37 This is a block diagram of a three-dimensional data encoding device according to embodiment 7.
[0059] Fig.38 An example of the prediction residual in Embodiment 7 is shown.
[0060] Fig.39 An example of volume in Embodiment 7 is shown.
[0061] Fig.40 An example of octree representation of volume in Embodiment 7 is shown.
[0062] Fig.41 An example of a bit string of volume in Implementation Example 7 is shown.
[0063] Fig.42 An example of octree representation of volume in Embodiment 7 is shown.
[0064] Fig.43 An example of volume in Embodiment 7 is shown.
[0065] Fig.44 This is a diagram for explaining the intra-frame prediction processing of embodiment 7.
[0066] Fig.45 This is a diagram used to illustrate the rotation and translation processing of embodiment 7.
[0067] Fig.46 An example of the syntax of the RT application flag and RT information according to the seventh embodiment is shown.
[0068] Fig.47 This is a diagram used to illustrate the inter-frame prediction processing of embodiment 7.
[0069] Fig.48 This is a block diagram of a three-dimensional data decoding device according to the seventh embodiment.
[0070] Fig.49 This is a flowchart of a three-dimensional data encoding process performed by the three-dimensional data encoding device according to the seventh embodiment.
[0071] Fig.50 This is a flowchart of a three-dimensional data decoding process performed by the three-dimensional data decoding device according to the seventh embodiment.
[0072] Fig.51 This is a diagram showing the reference relationship in the octree structure of implementation mode 8.
[0073] Fig.52 This is a diagram showing the reference relationship in the spatial area of Implementation Example 8.
[0074] Fig.53 This is a diagram showing an example of adjacent reference nodes according to the eighth embodiment.
[0075] Fig.54 This is a diagram showing the relationship between a parent node and a node in implementation mode 8.
[0076] Fig.55 This is a diagram showing an example of occupancy coding of a parent node in implementation mode 8.
[0077] Fig.56 This is a block diagram showing a three-dimensional data encoding device according to an eighth embodiment.
[0078] Fig.57 This is a block diagram showing a three-dimensional data decoding device according to an eighth embodiment.
[0079] Fig.58 It is a flowchart showing the three-dimensional data encoding processing of implementation mode 8.
[0080] Fig.59 This is a flowchart showing the three-dimensional data decoding process of the eighth embodiment.
[0081] Fig.60 This is a diagram showing an example of switching the coding table in implementation mode 8.
[0082] Fig.61 This is a diagram showing the reference relationship in the spatial region of Modification 1 of Implementation Example 8.
[0083] Fig.62 This is a diagram showing a syntax example of header information according to variant example 1 of implementation example 8.
[0084] Fig.63 This is a diagram showing a syntax example of header information according to variant example 1 of implementation example 8.
[0085] Fig.64 This is a diagram showing an example of adjacent reference nodes according to variant example 2 of implementation example 8.
[0086] Fig.65 This is a diagram showing an example of a target node and adjacent nodes according to variation 2 of implementation example 8.
[0087] Fig.66 This is a diagram of the reference relationship in the octree structure of variant example 3 of implementation example 8.
[0088] Fig.67 This is a diagram showing the reference relationship in the spatial region of variant example 3 of implementation example 8.
[0089] Fig.68 This is a diagram showing an example of three-dimensional points according to the ninth embodiment.
[0090] Fig.69 This is a diagram showing an example of setting LoD in Implementation Example 9.
[0091] Fig.70 This is a diagram showing an example of a threshold value used in setting LoD in the ninth embodiment.
[0092] Fig.71 This is a diagram showing an example of attribute information used in the prediction value of Implementation 9.
[0093] Fig.72 This is a diagram showing an example of the Exponential Golomb code according to the ninth embodiment.
[0094] Fig.73 This is a diagram showing the processing of the Exponential Golomb code according to the ninth embodiment.
[0095] Fig.74 This is a diagram showing a syntax example of an attribute header according to the ninth embodiment.
[0096] Fig.75 This is a diagram showing a syntax example of attribute data according to the ninth embodiment.
[0097] Fig.76 This is a flowchart of the three-dimensional data encoding process of the ninth embodiment.
[0098] Fig.77 This is a flowchart of the attribute information encoding process of the ninth embodiment.
[0099] Fig.78 This is a diagram showing the processing of the Exponential Golomb code according to the ninth embodiment.
[0100] Fig.79 This is a diagram showing an example of a reverse calculation table showing the relationship between the remaining codes and their values according to Implementation Example 9.
[0101] Fig.80 This is a flowchart of the three-dimensional data decoding process according to the ninth embodiment.
[0102] Fig.81 This is a flowchart of the attribute information decoding process of the ninth embodiment.
[0103] Fig.82 This is a block diagram of a three-dimensional data encoding device according to embodiment 9.
[0104] Fig.83 This is a block diagram of a three-dimensional data decoding device according to the ninth embodiment.
[0105] Fig.84 This is a flowchart of the three-dimensional data encoding process of the ninth embodiment.
[0106] Fig.85 This is a flowchart of the three-dimensional data decoding process according to the ninth embodiment. DETAILED DESCRIPTION
[0107] A three-dimensional data encoding method according to one embodiment of the present invention is a three-dimensional data encoding method for encoding three-dimensional points having attribute information, calculating a predicted value of the attribute information of the three-dimensional point, calculating a difference between the attribute information of the three-dimensional point and the predicted value, i.e., a prediction residual, generating binary data by binarizing the prediction residual, and performing arithmetical encoding on the binary data.
[0108] Thus, the three-dimensional data encoding method calculates the prediction residual of the attribute information, and then binarizes and arithmetic encodes the prediction residual, thereby being able to reduce the encoding amount of the encoded data of the attribute information.
[0109] For example, in the arithmetic coding, a different coding table may be used for each bit of the binary data.
[0110] Therefore, the three-dimensional data encoding method can improve encoding efficiency.
[0111] For example, in the arithmetic coding, a larger number of coding tables may be used for lower order bits of the binary data.
[0112] For example, in the arithmetic coding, a coding table used for arithmetic coding of the target bit may be selected according to the value of a higher-order bit of the target bit included in the binary data.
[0113] Therefore, the three-dimensional data encoding method can select the encoding table according to the value of the upper bit, thereby improving the encoding efficiency.
[0114] For example, it may also be that in the binarization, when the prediction residual is less than a threshold value, the binary data is generated by binarizing the prediction residual with a fixed number of bits, and when the prediction residual is greater than the threshold value, the binary data is generated including a first code of the fixed number of bits representing the threshold value and a second code obtained by binarizing a value obtained by subtracting the threshold value from the prediction residual using Exponential Golomb, and in the arithmetic coding, different arithmetic coding methods are used for the first code and the second code.
[0115] Thus, the three-dimensional data encoding method can, for example, perform arithmetical encoding on the first code and the second code using arithmetic encoding methods suitable for the first code and the second code, respectively, and thus can improve encoding efficiency.
[0116] For example, the three-dimensional data encoding method may further quantize the prediction residual, and in the binarization, the quantized prediction residual may be binarized, and the threshold value may be changed according to a quantization scale in the quantization.
[0117] Thus, the three-dimensional data encoding method can use an appropriate threshold value according to the quantization scale, and thus can improve encoding efficiency.
[0118] For example, the second code may include a prefix part and a suffix part, and in the arithmetic coding, different coding tables are used for the prefix part and the suffix part.
[0119] Therefore, the three-dimensional data encoding method can improve encoding efficiency.
[0120] A three-dimensional data decoding method according to one embodiment of the present invention is a three-dimensional data decoding method for decoding a three-dimensional point having attribute information, calculating a predicted value of the attribute information of the three-dimensional point, generating binary data by arithmetically decoding encoded data contained in a bit stream, generating a prediction residual by multi-valuedizing the binary data, and calculating a decoded value of the attribute information of the three-dimensional point by adding the predicted value to the prediction residual.
[0121] Thus, the three-dimensional data decoding method can calculate the prediction residual of the attribute information, and further appropriately decode the bit stream of the attribute information generated by binarizing and arithmetic coding the prediction residual.
[0122] For example, in the arithmetic decoding, a different coding table may be used for each bit of the binary data.
[0123] Therefore, the three-dimensional data decoding method can appropriately decode the bit stream with improved coding efficiency.
[0124] For example, in the arithmetic decoding, a larger number of coding tables may be used for lower bits of the binary data.
[0125] For example, in the arithmetic decoding, a coding table used in arithmetic decoding of the target bit may be selected according to a value of a higher-order bit of the target bit included in the binary data.
[0126] Therefore, the three-dimensional data decoding method can appropriately decode the bit stream with improved coding efficiency.
[0127] For example, it may also be that, in the multi-valued conversion, the first value is generated by multi-valuedizing a first code having a fixed number of bits contained in the binary data, and when the first value is less than a threshold value, the first value is determined as the prediction residual, and when the first value is greater than the threshold value, the second value is generated by multi-valuedizing an exponential Golomb code, i.e., a second code, contained in the binary data, and the prediction residual is generated by adding the first value to the second value, and in the arithmetic decoding, different arithmetic decoding methods are used for the first code and the second code.
[0128] Therefore, the three-dimensional data decoding method can appropriately decode the bit stream with improved coding efficiency.
[0129] For example, the three-dimensional data decoding method may further inverse quantize the prediction residual, and in the adding, the prediction value is added to the inverse quantized prediction residual, and the threshold value is changed according to a quantization scale in the inverse quantization.
[0130] Therefore, the three-dimensional data decoding method can appropriately decode the bit stream with improved coding efficiency.
[0131] For example, the second code may include a prefix part and a suffix part, and in the arithmetic decoding, different coding tables may be used for the prefix part and the suffix part.
[0132] In addition, a three-dimensional data encoding device of one embodiment of the present invention is a three-dimensional data encoding device that encodes three-dimensional points having attribute information, and is equipped with a processor and a memory. The processor uses the memory to calculate a predicted value of the attribute information of the three-dimensional point, calculates a difference between the attribute information of the three-dimensional point and the predicted value, i.e., a predicted residual, generates binary data by binarizing the predicted residual, and performs arithmetically encoding on the binary data.
[0133] Thus, the three-dimensional data encoding device calculates the prediction residual of the attribute information, and further binarizes and arithmetic encodes the prediction residual, thereby being able to reduce the code amount of the encoded data of the attribute information.
[0134] In addition, a three-dimensional data decoding device of one embodiment of the present invention is a three-dimensional data decoding device that decodes a three-dimensional point having attribute information, and is provided with a processor and a memory. The processor uses the memory to calculate a predicted value of the attribute information of the three-dimensional point, generates binary data by arithmetically decoding the encoded data contained in the bit stream, generates a prediction residual by multi-valuedizing the binary data, and calculates a decoded value of the attribute information of the three-dimensional point by adding the predicted value to the prediction residual.
[0135] Thus, the three-dimensional data decoding apparatus can calculate the prediction residual of the attribute information, and further, can appropriately decode the bit stream of the attribute information generated by binarizing and arithmetic coding the prediction residual.
[0136] In addition, these general or specific forms can be implemented by systems, methods, integrated circuits, computer programs, or computer-readable recording media such as CD-ROMs, and can be implemented by any combination of systems, methods, integrated circuits, computer programs, and recording media.
[0137] The following detailed description of the implementation mode is given with reference to the accompanying drawings. In addition, the implementation modes to be described below are all specific examples of the present disclosure. The numerical values, shapes, materials, constituent elements, configuration positions of constituent elements, connection forms, steps, order of steps, etc. shown in the following implementation modes are all examples, and the main purpose is not to limit the present disclosure. Furthermore, the constituent elements of the following implementation modes that are not recorded in the technical solution showing the highest concept are described as arbitrary constituent elements.
[0138] (Implementation Method 1)
[0139] First, the data structure of encoded three-dimensional data (hereinafter also referred to as encoded data) according to the present embodiment will be described. Figure 1 The structure of the encoded three-dimensional data involved in this embodiment is shown.
[0140] In this embodiment, the three-dimensional space is divided into a space (SPC) equivalent to a picture in the encoding of a dynamic image, and the three-dimensional data is encoded in units of space. The space is further divided into volumes (VLM) equivalent to macroblocks in dynamic image encoding, and prediction and conversion are performed in units of VLM. The volume includes a minimum unit corresponding to a position coordinate, namely a plurality of voxels (VXL). In addition, prediction means that, similar to the prediction performed in a two-dimensional image, predicted three-dimensional data similar to the processing unit of the processing object is generated with reference to other processing units, and the difference between the predicted three-dimensional data and the processing unit of the processing object is encoded. Furthermore, the prediction includes not only spatial prediction with reference to other prediction units at the same time, but also temporal prediction with reference to prediction units at different times.
[0141] For example, when encoding a three-dimensional space represented by point group data such as a point cloud, a three-dimensional data encoding device (hereinafter also referred to as an encoding device) encodes each point of the point group or a plurality of points contained in a voxel according to the size of the voxel. If the voxel is subdivided, the three-dimensional shape of the point group can be expressed with high accuracy, and if the size of the voxel is increased, the three-dimensional shape of the point group can be roughly expressed.
[0142] In addition, although the following description takes the case where the three-dimensional data is a point cloud as an example, the three-dimensional data is not limited to the point cloud, and can also be three-dimensional data in any form.
[0143] Furthermore, voxels of a hierarchical structure may be used. In this case, in the n-order hierarchy, whether or not a sampling point exists in the n-1-order or lower hierarchy (the lower hierarchy of the n-order hierarchy) may be sequentially indicated. For example, when decoding only the n-order hierarchy, if a sampling point exists in the n-1-order or lower hierarchy, decoding may be performed by assuming that a sampling point exists at the center of a voxel in the n-order hierarchy.
[0144] Furthermore, the encoding device obtains point group data through a distance sensor, a stereo camera, a monocular camera, a gyroscope, or an inertial sensor.
[0145] As with the encoding of moving images, the space is classified into at least one of the following three prediction structures: an intra-frame space (I-SPC) that can be decoded independently, a prediction space (P-SPC) that can only be referenced unidirectionally, and a bidirectional space (B-SPC) that can be referenced bidirectionally. In addition, the space has two types of time information: decoding time and display time.
[0146] And, if Figure 1As shown, as a processing unit including a plurality of spaces, there is a GOS (Group Of Space) which is a random access unit. Also, as a processing unit including a plurality of GOS, there is a world space (WLD).
[0147] The spatial area occupied by the world space is associated with an absolute position on the earth through GPS or latitude and longitude information. This position information is stored as meta information. In addition, the meta information can be included in the encoded data or transmitted separately from the encoded data.
[0148] Furthermore, within the GOS, all SPCs may be three-dimensionally adjacent, or there may be SPCs that are not three-dimensionally adjacent to other SPCs.
[0149] In addition, the encoding, decoding or referencing of the three-dimensional data included in the processing unit such as GOS, SPC or VLM is also simply referred to as encoding, decoding or referencing the processing unit. The three-dimensional data included in the processing unit includes at least one set of spatial positions such as three-dimensional coordinates and characteristic values such as color information.
[0150] Next, the prediction structure of the SPC in the GOS will be described. Although multiple SPCs in the same GOS or multiple VLMs in the same SPC occupy different spaces, they have the same time information (decoding time and display time).
[0151] Furthermore, in the GOS, the first SPC in the decoding order is the I-SPC. Furthermore, there are two types of GOS, the closed GOS and the open GOS. The closed GOS is a GOS that can decode all the SPCs in the GOS when decoding starts from the first I-SPC. In the open GOS, in the GOS, some SPCs earlier than the display time of the first I-SPC refer to different GOSs and can only be decoded in the GOS.
[0152] In addition, in the case of coded data such as map information, the WLD may be decoded in the reverse direction of the coding order. If there is a dependency between GOS, it is difficult to reproduce the data in the reverse direction. Therefore, in this case, a closed GOS is basically used.
[0153] Furthermore, the GOS has a layer structure in the height direction, and encoding or decoding is performed sequentially starting from the SPC of the bottom layer.
[0154] Figure 2 An example of a prediction structure between SPCs belonging to the lowest layer of the GOS is shown. Figure 3 An example of an inter-layer prediction structure is shown.
[0155] There are more than one I-SPC in the GOS. Although there are objects such as people, animals, cars, bicycles, traffic lights, or buildings that serve as land landmarks in the three-dimensional space, it is particularly effective to encode small-sized objects as I-SPCs. For example, when a three-dimensional data decoding device (hereinafter also referred to as a decoding device) decodes the GOS at a low processing amount or high speed, it only decodes the I-SPC in the GOS.
[0156] Furthermore, the encoding device may switch the encoding interval or the occurrence frequency of the I-SPC according to the density of the objects in the WLD.
[0157] And, in Figure 3 In the configuration shown, the encoding device or decoding device encodes or decodes the plurality of layers sequentially from the lower layer (layer 1). This allows, for example, autonomous vehicles to prioritize data near the ground with a large amount of information.
[0158] In addition, in the coded data used by a drone or the like, coding or decoding may be performed sequentially starting from the SPC of the upper layer in the height direction within the GOS.
[0159] Furthermore, the encoding device or decoding device may encode or decode multiple layers in such a way that the decoding device roughly grasps the GOS and can gradually increase the resolution. For example, the encoding device or decoding device may encode or decode in the order of layers 3, 8, 1, 9, ...
[0160] Next, the corresponding method of the static object and the dynamic object is described.
[0161] In three-dimensional space, there are static objects or scenes such as buildings and roads (hereinafter collectively referred to as static objects), and dynamic objects such as vehicles and people (hereinafter referred to as dynamic objects). Object detection can be performed by extracting feature points from point cloud data or images captured by stereo cameras. Here, an example of a method for encoding dynamic objects is described.
[0162] The first method is a method of encoding without distinguishing between static objects and dynamic objects. The second method is a method of distinguishing between static objects and dynamic objects using identification information.
[0163] For example, GOS is used as the identification unit. In this case, GOS including SPCs constituting static objects and GOS including SPCs constituting dynamic objects are distinguished within the coded data or by identification information stored separately from the coded data.
[0164] Alternatively, SPC is used as the identification unit. In this case, the SPC including only the VLM constituting the static object and the SPC including the VLM constituting the dynamic object are distinguished by the above-mentioned identification information.
[0165] Alternatively, VLM or VXL may be used as the identification unit. In this case, the VLM or VXL including the static object and the VLM or VXL including the dynamic object are distinguished by the above-mentioned identification information.
[0166] Furthermore, the encoding device may encode the dynamic object as one or more VLMs or SPCs, and encode the VLM or SPC including the static object and the SPC including the dynamic object as different GOSs. Furthermore, when the size of the GOS becomes variable according to the size of the dynamic object, the encoding device may store the size of the GOS separately as meta-information.
[0167] Furthermore, the encoding device encodes the static object and the dynamic object independently of each other, and the dynamic object can be overlapped with respect to the world space composed of the static object. In this case, the dynamic object is composed of one or more SPCs, and each SPC corresponds to one or more SPCs constituting the static object overlapped with the SPC. In addition, the dynamic object may not be represented by an SPC, but may be represented by one or more VLMs or VXLs.
[0168] Furthermore, the encoding device may encode static objects and dynamic objects as different streams.
[0169] Furthermore, the encoding device may generate a GOS including one or more SPCs constituting a dynamic object. Furthermore, the encoding device may set the GOS (GOS_M) including the dynamic object and the GOS of the static object corresponding to the spatial region of the GOS_M to be of the same size (occupying the same spatial region). In this way, overlapping processing can be performed in units of GOS.
[0170] The P-SPC or B-SPC constituting the dynamic object may also refer to the SPC included in the encoded different GOS. When the position of the dynamic object changes over time and the same dynamic object is encoded as the GOS at different times, cross-GOS reference is effective from the perspective of compression rate.
[0171] Furthermore, the first method and the second method may be switched according to the purpose of the encoded data. For example, when the encoded three-dimensional data is used as a map, it is desirable to separate the three-dimensional data from the dynamic objects, so the encoding device uses the second method. In addition, when the encoding device encodes the three-dimensional data of an event such as a concert or sports, if it is not necessary to separate the dynamic objects, the first method is used.
[0172] Furthermore, the decoding time and display time of GOS or SPC can be stored in the encoded data or as meta-information. Furthermore, the time information of static objects can all be the same. At this time, the actual decoding time and display time can be determined by the decoding device. Alternatively, different values can be assigned to each GOS or SPC as the decoding time, and the same value can be assigned to all the display times. Moreover, as shown in the decoder mode in dynamic image encoding such as HEVC's HRD (Hypothetical Reference Decoder), the decoder has a buffer of a specified size. As long as the bit stream is read at a specified bit rate according to the decoding time, a model that will not be destroyed and is guaranteed to be decodable can be imported.
[0173] Next, the configuration of the GOS in the world space is described. The coordinates of the three-dimensional space in the world space are represented by three mutually orthogonal coordinate axes (x-axis, y-axis, and z-axis). By setting a prescribed rule in the encoding order of the GOS, spatially adjacent GOS can be encoded continuously in the encoded data. For example, Figure 4 In the example shown, the GOS in the xz plane is continuously encoded. After the encoding of all GOS in an xz plane is completed, the value of the y axis is updated. That is, as the encoding continues, the world space extends in the y axis direction. And the index number of the GOS is set as the encoding order.
[0174] Here, the three-dimensional space of the world space corresponds one-to-one to absolute geographical coordinates such as GPS, latitude and longitude. Alternatively, the three-dimensional space can be represented by a relative position relative to a preset reference position. The directions of the x-axis, y-axis, and z-axis of the three-dimensional space are represented as direction vectors determined based on latitude and longitude, etc., and the direction vectors are stored together with the encoded data as meta information.
[0175] Furthermore, the size of the GOS is set to be fixed, and the encoding device stores the size as meta-information. Furthermore, the size of the GOS can be switched, for example, depending on whether it is in the city, indoors, or outdoors. That is, the size of the GOS can be switched according to the amount or nature of objects that have value as information. Alternatively, the encoding device can appropriately switch the size of the GOS or the interval of the I-SPC in the GOS according to the density of the object, etc. in the same world space. For example, the encoding device sets the size of the GOS to be smaller and the interval of the I-SPC in the GOS to be shorter when the density of the object is higher.
[0176] exist Figure 5In the example, in the area from the 3rd to the 10th GOS, since the density of objects is high, the GOS is subdivided to achieve random access of fine granularity. And, the 7th to the 10th GOS exist on the back of the 3rd to the 6th GOS, respectively.
[0177] Next, the configuration and operation flow of the three-dimensional data encoding device according to this embodiment will be described. Figure 6 It is a block diagram of the three-dimensional data encoding device 100 involved in this embodiment. Figure 7 : is a flowchart showing an operation example of the three-dimensional data encoding device 100 .
[0178] Figure 6 The three-dimensional data encoding device 100 shown generates encoded three-dimensional data 112 by encoding three-dimensional data 111. The three-dimensional data encoding device 100 includes an acquisition unit 101, an encoding region determination unit 102, a division unit 103, and an encoding unit 104.
[0179] like Figure 7 As shown, first, the acquisition unit 101 acquires three-dimensional data 111 as point group data ( S101 ).
[0180] Next, the coding region determination unit 102 determines a coding target region from the spatial region corresponding to the obtained point cloud data (S102). For example, the coding region determination unit 102 determines a spatial region around the position as the coding target region according to the position of the user or vehicle.
[0181] Next, the segmentation unit 103 segments the point group data included in the area of the encoding object into respective processing units. Here, the processing units are the above-mentioned GOS and SPC, etc. And, the area of the encoding object corresponds to the above-mentioned world space, for example. Specifically, the segmentation unit 103 segments the point group data into processing units according to the size of the pre-set GOS, the presence or size of the dynamic object (S103). And, the segmentation unit 103 determines the starting position of the SPC that becomes the beginning in the encoding order in each GOS.
[0182] Next, the encoding unit 104 generates the encoded three-dimensional data 112 by sequentially encoding a plurality of SPCs in each GOS ( S104 ).
[0183] In addition, here, after the area to be coded is divided into GOS and SPC, although an example of coding each GOS is shown, the order of processing is not limited to the above. For example, after the composition of a GOS is determined, the GOS may be coded, and then the composition of the GOS may be determined.
[0184] In this way, the three-dimensional data encoding device 100 generates the encoded three-dimensional data 112 by encoding the three-dimensional data 111. Specifically, the three-dimensional data encoding device 100 divides the three-dimensional data into random access units, that is, divides it into first processing units (GOS) corresponding to three-dimensional coordinates, divides the first processing unit (GOS) into a plurality of second processing units (SPC), and divides the second processing unit (SPC) into a plurality of third processing units (VLM). In addition, the third processing unit (VLM) includes more than one voxel (VXL), and the voxel (VXL) is the smallest unit corresponding to the position information.
[0185] Next, the three-dimensional data encoding device 100 generates the encoded three-dimensional data 112 by encoding each of the plurality of first processing units (GOS). Specifically, the three-dimensional data encoding device 100 encodes each of the plurality of second processing units (SPC) in each first processing unit (GOS). Furthermore, the three-dimensional data encoding device 100 encodes each of the plurality of third processing units (VLM) in each second processing unit (SPC).
[0186] For example, when the first processing unit (GOS) of the processing object is a closed GOS, the three-dimensional data encoding device 100 performs encoding with reference to other second processing units (SPCs) included in the first processing unit (GOS) of the processing object for the second processing unit (SPC) of the processing object included in the first processing unit (GOS) of the processing object. That is, the three-dimensional data encoding device 100 does not refer to the second processing unit (SPC) included in the first processing unit (GOS) different from the first processing unit (GOS) of the processing object.
[0187] Furthermore, when the first processing unit (GOS) of the processing object is an open GOS, the second processing unit (SPC) of the processing object contained in the first processing unit (GOS) of the processing object is encoded with reference to other second processing units (SPCs) contained in the first processing unit (GOS) of the processing object, or a second processing unit (SPC) contained in a first processing unit (GOS) different from the first processing unit (GOS) of the processing object.
[0188] Furthermore, the three-dimensional data encoding device 100 selects one as the type of the second processing unit (SPC) of the processing object from among the first type (I-SPC) that does not refer to other second processing units (SPC), the second type (P-SPC) that refers to one other second processing unit (SPC), and the third type that refers to two other second processing units (SPC), and encodes the second processing unit (SPC) of the processing object according to the selected type.
[0189] Next, the configuration and operation flow of the three-dimensional data decoding device according to the present embodiment will be described. Figure 8 It is a block diagram of the three-dimensional data decoding device 200 according to this embodiment. Fig. 9 : is a flowchart showing an example of the operation of the three-dimensional data decoding device 200 .
[0190] Figure 8 The three-dimensional data decoding device 200 shown generates decoded three-dimensional data 212 by decoding the encoded three-dimensional data 211. Here, the encoded three-dimensional data 211 is, for example, the encoded three-dimensional data 112 generated by the three-dimensional data encoding device 100. The three-dimensional data decoding device 200 includes an acquisition unit 201, a decoding start GOS determination unit 202, a decoding SPC determination unit 203, and a decoding unit 204.
[0191] First, the acquisition unit 201 acquires the encoded three-dimensional data 211 (S201). Next, the decoding start GOS determination unit 202 determines the GOS to be decoded (S202). Specifically, the decoding start GOS determination unit 202 refers to the meta information stored in the encoded three-dimensional data 211 or separately from the encoded three-dimensional data, and determines the GOS including the spatial position, object, or SPC corresponding to the time at which decoding is started as the GOS to be decoded.
[0192] Next, the decoded SPC determination unit 203 determines the type (I, P, B) of the SPC to be decoded in the GOS (S203). For example, the decoded SPC determination unit 203 determines (1) whether to decode only the I-SPC, (2) whether to decode the I-SPC and the P-SPC, or (3) whether to decode all types. In addition, when the type of the SPC to be decoded is predetermined, such as when all SPCs are decoded, this step may not be performed.
[0193] Next, the decoding unit 204 obtains the SPC that is the first in the decoding order (the same as the encoding order) in the GOS, and the address position that starts in the encoded three-dimensional data 211, obtains the encoded data of the first SPC from the address position, and decodes each SPC in sequence from the first SPC (S204). And, the above-mentioned address position is stored in meta information, etc.
[0194] In this way, the three-dimensional data decoding device 200 decodes the decoded three-dimensional data 212. Specifically, the three-dimensional data decoding device 200 generates the decoded three-dimensional data 212 of the first processing unit (GOS) as a random access unit by decoding each of the encoded three-dimensional data 211 of the first processing unit (GOS) corresponding to the three-dimensional coordinates. More specifically, the three-dimensional data decoding device 200 decodes each of the plurality of second processing units (SPC) in each first processing unit (GOS). And, the three-dimensional data decoding device 200 decodes each of the plurality of third processing units (VLM) in each second processing unit (SPC).
[0195] The meta information for random access is described below. This meta information is generated by the three-dimensional data encoding device 100 and included in the encoded three-dimensional data 112 (211).
[0196] In conventional random access of two-dimensional moving images, decoding is started from the head frame of a random access unit near a designated time. However, in the world space, random access to (coordinates or objects, etc.) is also envisioned in addition to time.
[0197] Therefore, in order to realize random access to at least three elements, namely coordinates, objects, and time, a table is prepared in which each element is associated with the index number of the GOS. Furthermore, the index number of the GOS is associated with the address of the I-SPC at the beginning of the GOS. Fig.10 An example of a table included in the meta information is shown. Fig.10 Of all the tables shown, at least one table may be used.
[0198] As an example, random access starting from a coordinate is described below. When accessing the coordinates (x2, y2, z2), the coordinate-GOS table is first referenced, and it can be known that the location with the coordinates (x2, y2, z2) is included in the second GOS. Next, the GOS address table is referenced, and since it can be known that the address of the first I-SPC in the second GOS is addr(2), the decoding unit 204 obtains data from this address and starts decoding.
[0199] In addition, the address may be an address in a logical format or a physical address of a HDD or a memory. Furthermore, information for identifying a file segment may be used instead of an address. For example, a file segment is a unit after segmenting one or more GOSs.
[0200] Furthermore, when the object spans across multiple GOSs, the GOSs to which the multiple objects belong may also be shown in the object GOS table. If the multiple GOSs are closed GOSs, the encoding device and the decoding device can perform encoding or decoding in parallel. In addition, if the multiple GOSs are open GOSs, the compression efficiency can be further improved by having the multiple GOSs refer to each other.
[0201] Examples of objects include people, animals, cars, bicycles, traffic lights, or buildings that serve as landmarks on land, etc. For example, when encoding in world space, the three-dimensional data encoding device 100 extracts feature points unique to the object from a three-dimensional point cloud, etc., detects the object based on the feature points, and can set the detected object as a random access point.
[0202] In this way, the three-dimensional data encoding device 100 generates the first information, which indicates the plurality of first processing units (GOS) and the three-dimensional coordinates corresponding to each of the plurality of first processing units (GOS). And the encoded three-dimensional data 112 (211) includes the first information. And the first information further indicates at least one of the object, time, and data storage destination corresponding to each of the plurality of first processing units (GOS).
[0203] The three-dimensional data decoding device 200 obtains the first information from the encoded three-dimensional data 211 , uses the first information to determine the first processing unit of the encoded three-dimensional data 211 corresponding to the specified three-dimensional coordinates, object or time, and decodes the encoded three-dimensional data 211 .
[0204] Other examples of meta information are described below. In addition to the meta information for random access, the three-dimensional data encoding device 100 can also generate and store the following meta information. Furthermore, the three-dimensional data decoding device 200 can also use this meta information during decoding.
[0205] In the case where three-dimensional data is used as map information, a profile is specified according to the purpose, and information indicating the profile may be included in the meta-information. For example, a profile for urban areas or suburbs is specified, or a profile for flying objects is specified, and the maximum or minimum size of the world space, SPC or VLM is defined respectively. For example, in the profile for urban areas, more detailed information is required than in the suburbs, so the minimum size of VLM is set smaller.
[0206] The meta information may also include a tag value indicating the type of the object. The tag value corresponds to the VLM, SPC, or GOS that constitutes the object. The tag value may be set according to the type of the object, for example, a tag value of "0" indicates a "person", a tag value of "1" indicates a "car", and a tag value of "2" indicates a "traffic light". Alternatively, when the type of the object is difficult to determine or does not need to be determined, a tag value indicating the size, or whether it is a dynamic object or a static object, etc. may be used.
[0207] Furthermore, the meta-information may include information indicating the range of the spatial region occupied by the world space.
[0208] Furthermore, the meta-information may be used as header information common to the entire stream of coded data or to a plurality of SPCs such as an SPC in a GOS, and may store the size of the SPC or VXL.
[0209] Furthermore, the meta-information may include identification information of a distance sensor or a camera used in generating the point cloud, or information indicating the positional accuracy of a point group in the point cloud.
[0210] Also, the meta information may include information showing whether the world space is composed of only static objects or contains dynamic objects.
[0211] Modifications of this embodiment will be described below.
[0212] The encoding device or decoding device may encode or decode two or more different SPCs or GOSs in parallel. The GOSs to be encoded or decoded in parallel can be determined based on meta information indicating the spatial position of the GOSs.
[0213] In the case where three-dimensional data is used as a spatial map for the movement of a vehicle or a flying object, or such a spatial map is generated, the encoding device or decoding device can encode or decode the GOS or SPC contained in the space determined based on GPS, path information, or a zoom factor, etc.
[0214] Furthermore, the decoding device may also start decoding from the space close to its own position or walking path. The encoding device or decoding device may also make the priority of the space far from its own position or walking path lower than the priority of the space close to it to perform encoding or decoding. Here, lowering the priority means lowering the processing order, lowering the resolution (screening post-processing), or lowering the image quality (improving the encoding efficiency. For example, increasing the quantization step size), etc.
[0215] Furthermore, when decoding the coded data which is hierarchically coded in space, the decoding device may decode only the lower layer.
[0216] Furthermore, the decoding device may start decoding from a lower layer according to the zoom ratio or purpose of the map.
[0217] Furthermore, in applications such as estimating the self-position of a car or robot during automatic driving or identifying an object, the encoding device or decoding device may also reduce the resolution of an area outside the area within a specified height from the road surface (the area to be identified) for encoding or decoding.
[0218] Furthermore, the encoding device may also encode the point clouds representing the indoor and outdoor spatial shapes separately, for example, by separating the GOS representing the indoor space (indoor GOS) from the GOS representing the outdoor space (outdoor GOS), so that when the decoding device uses the encoded data, it can select the GOS to be decoded according to the viewpoint position.
[0219] Furthermore, the encoding device may encode the indoor GOS and the outdoor GOS whose coordinates are close to each other by placing them adjacent to each other in the coded stream. For example, the encoding device may correspond the identifiers of the two and store information indicating that the corresponding identifiers are established in the coded stream or in meta-information stored separately. Accordingly, the decoding device can refer to the information in the meta-information to identify the indoor GOS and the outdoor GOS whose coordinates are close to each other.
[0220] Furthermore, the encoding device may switch the size of the GOS or SPC between the indoor GOS and the outdoor GOS. For example, the encoding device sets the size of the GOS to be smaller indoors than outdoors. Furthermore, the encoding device may change the accuracy of extracting feature points from the point cloud or the accuracy of object detection between the indoor GOS and the outdoor GOS.
[0221] Furthermore, the encoding device may add information for the decoding device to distinguish and display dynamic objects from static objects to the encoded data. Accordingly, the decoding device can represent the dynamic objects by combining them with red frames or explanatory texts. In addition, the decoding device may replace the dynamic objects and represent them with only red frames or explanatory texts. Furthermore, the decoding device may represent more detailed object categories. For example, a car may be represented by a red frame, and a person may be represented by a yellow frame.
[0222] Furthermore, the encoding device or decoding device may determine whether to perform encoding or decoding by treating the dynamic object and the static object as different SPCs or GOSs according to the frequency of occurrence of the dynamic object or the ratio of the static object to the dynamic object. For example, when the frequency of occurrence or the ratio of the dynamic object exceeds a threshold, the SPC or GOS in which the dynamic object and the static object are mixed is allowed, and when the frequency of occurrence or the ratio of the dynamic object does not exceed the threshold, the SPC or GOS in which the dynamic object and the static object are mixed is not allowed.
[0223] When the dynamic object is detected from the two-dimensional image information of the camera instead of the point cloud, the encoding device can obtain the information (frame or text, etc.) for identifying the detection result and the object position separately, and encode the information as part of the three-dimensional encoded data. In this case, the decoding device overlaps the auxiliary information (frame or text) representing the dynamic object with the decoding result of the static object.
[0224] Furthermore, the coding device may change the density of VXL or VLM according to the complexity of the shape of the static object, etc. For example, the coding device sets VXL or VLM to be denser when the shape of the static object is more complex. Furthermore, the coding device may determine the quantization step size when quantizing the spatial position or color information according to the density of VXL or VLM. For example, the coding device sets the quantization step size to be smaller when the VXL or VLM is denser.
[0225] As described above, the encoding device or decoding device according to the present embodiment performs spatial encoding or decoding in a spatial unit having coordinate information.
[0226] Furthermore, the encoding device and the decoding device perform encoding or decoding in volume units in space. The volume includes voxels which are the minimum units corresponding to the position information.
[0227] Furthermore, the encoding device and the decoding device establish a table of correspondence between each element of spatial information including coordinates, objects, and time and GOP, or a table of correspondence between each element, so as to establish correspondence between arbitrary elements to perform encoding or decoding. Furthermore, the decoding device determines the coordinates using the value of the selected element, determines the volume, voxel, or space based on the coordinates, and decodes the space including the volume or voxel, or the determined space.
[0228] Then, the encoding device determines a volume, voxel, or space that can be selected by an element by extracting feature points or recognizing an object, and encodes it as a volume, voxel, or space that can be randomly accessed.
[0229] Spaces are divided into three types: I-SPC which can be encoded or decoded as a single space, P-SPC which can be encoded or decoded with reference to any one of the processed spaces, and B-SPC which can be encoded or decoded with reference to any two of the processed spaces.
[0230] One or more volumes correspond to static objects or dynamic objects. Spaces containing static objects and spaces containing dynamic objects are encoded or decoded as different GOSs. That is, SPCs containing static objects and SPCs containing dynamic objects are assigned to different GOSs.
[0231] Dynamic objects are encoded or decoded for each object and correspond to one or more spaces containing only static objects. That is, multiple dynamic objects are encoded separately, and the encoded data of multiple dynamic objects obtained correspond to SPCs containing only static objects.
[0232] The encoding device and the decoding device increase the priority of the I-SPC in the GOS to perform encoding or decoding. For example, the encoding device performs encoding in a manner that reduces the degradation of the I-SPC (after decoding, the original three-dimensional data can be reproduced more faithfully). And, the decoding device, for example, only decodes the I-SPC.
[0233] The encoding device may change the frequency of using the I-SPC according to the density or value (number) of objects in the world space to perform encoding. That is, the encoding device changes the frequency of selecting the I-SPC according to the number or density of objects included in the three-dimensional data. For example, the encoding device increases the frequency of using the I space when the density of objects in the world space is greater.
[0234] Furthermore, the encoding device sets the random access point in units of GOS, and stores information indicating the spatial area corresponding to the GOS in the header information.
[0235] The coding device, for example, adopts a default value as the spatial size of the GOS. In addition, the coding device may also change the size of the GOS according to the value (number) or density of the objects or dynamic objects. For example, the coding device sets the spatial size of the GOS to be smaller when the objects or dynamic objects are denser or more in number.
[0236] Furthermore, the space or volume includes a group of feature points derived from information obtained by sensors such as a depth sensor, a gyroscope, or a camera. The coordinates of the feature points are set to the center position of the voxel. Furthermore, by subdividing the voxels, the position information can be highly accurate.
[0237] The feature point group is derived using a plurality of pictures. The plurality of pictures have at least the following two types of time information: actual time information and the same time information in the plurality of pictures corresponding to the space (for example, encoding time for rate control, etc.).
[0238] Furthermore, encoding or decoding is performed in units of GOS including one or more spaces.
[0239] The encoding device and the decoding device refer to the space in the processed GOS to predict the P space or the B space in the GOS to be processed.
[0240] Alternatively, the encoding device and the decoding device predict the P space or B space in the GOS to be processed by using the processed space in the GOS to be processed without referring to different GOSs.
[0241] Furthermore, the encoding device and the decoding device transmit or receive the encoded stream in units of a world space including one or more GOSs.
[0242] Furthermore, the GOS has a layer structure in at least one direction in the world space, and the encoding device and the decoding device perform encoding or decoding from the lower layer. For example, the GOS that can be randomly accessed belongs to the lowest layer. The GOS belonging to the upper layer only refers to the GOS belonging to the layer below the same layer. That is, the GOS is spatially divided in a predetermined direction and includes a plurality of layers each having one or more SPCs. The encoding device and the decoding device perform encoding or decoding for each SPC by referring to the SPC included in the layer that is the same layer as the SPC or the layer below the SPC.
[0243] Furthermore, the encoding device and the decoding device continuously encode or decode the GOS in the world space unit including the plurality of GOS. The encoding device and the decoding device write or read information indicating the order (direction) of encoding or decoding as metadata. That is, the encoded data includes information indicating the encoding order of the plurality of GOS.
[0244] Furthermore, the encoding device and the decoding device encode or decode two or more different spaces or GOS in parallel.
[0245] Furthermore, the encoding device and the decoding device encode or decode the space information (coordinates, size, etc.) of the space or the GOS.
[0246] Furthermore, the encoding device and the decoding device encode or decode the space or GOS included in the specific space determined based on external information related to the own position and / or area size, such as GPS, path information, or magnification.
[0247] The encoding device or decoding device performs encoding or decoding by giving a lower priority to a space far from the own position than to a space close to the own position.
[0248] The encoding device sets a direction in the world space according to the magnification or purpose, and encodes the GOS having a layer structure in the direction. And the decoding device preferentially decodes the GOS having a layer structure in the direction of the world space set according to the magnification or purpose from the lower layer.
[0249] The encoding device changes the feature point extraction, object recognition accuracy, or space area size contained in the indoor and outdoor spaces. However, the encoding device and the decoding device encode or decode the indoor GOS and the outdoor GOS with close coordinates adjacent to each other in the world space, and also encode or decode these identifiers in correspondence.
[0250] (Implementation Method 2)
[0251] When point cloud coded data is used in actual devices or services, it is desirable to transmit and receive required information according to the application in order to reduce network bandwidth. However, such a function does not exist in the existing 3D data coding structure, and therefore there is no corresponding coding method.
[0252] What will be described in this embodiment is a three-dimensional data encoding method and a three-dimensional data encoding device for providing the function of sending and receiving required information according to the purpose in the encoded data of a three-dimensional point cloud, as well as a three-dimensional data decoding method and a three-dimensional data decoding device for decoding the encoded data.
[0253] A voxel (VXL) having a certain feature value or more is defined as a feature voxel (FVXL), and a world space (WLD) formed by the FVXL is defined as a sparse world space (SWLD). Fig.11 The sparse world space and the configuration example of the world space are shown. SWLD includes: FGOS, which is a GOS composed of FVXL; FSPC, which is an SPC composed of FVXL; and FVLM, which is a VLM composed of FVXL. The data structure and prediction structure of FGOS, FSPC and FVLM can be the same as those of GOS, SPC and VLM.
[0254] The feature quantity refers to a feature quantity that expresses the three-dimensional position information of VXL or the visible light information of the position of VXL, and in particular, a feature quantity that can detect more corners and edges of a three-dimensional object. Specifically, the feature quantity is a three-dimensional feature quantity or a visible light feature quantity described below, but other than this, any feature quantity can be used as long as it expresses the position, brightness, or color information of VXL.
[0255] As the three-dimensional feature, a SHOT feature (Signature of Histograms of Orientations), a PFH feature (Point Feature Histograms), or a PPF feature (Point Pair Feature) is used.
[0256] The SHOT feature is obtained by segmenting the area around the VXL, calculating the inner product between the reference point and the normal vector of the segmented area, and converting it into a histogram. The SHOT feature has a high dimension and high feature expression.
[0257] The PFH feature is obtained by selecting a plurality of two-point groups near VXL, calculating the normal vector etc. from the two points, and histogramming them. Since the PFH feature is a histogram feature, it is robust to small amounts of interference and has a high feature expression.
[0258] The PPF feature is a feature calculated according to the VXL of two points using a normal vector, etc. Since all VXLs are used in this PPF feature, it is robust to occlusion.
[0259] Furthermore, as feature quantities of visible light, SIFT (Scale-Invariant Feature Transform), SURF (Speeded Up Robust Features), or HOG (Histogram of Oriented Gradients) that utilize information such as brightness gradient information of an image can be used.
[0260] SWLD is generated by calculating the above-mentioned feature amount from each VXL of WLD and extracting FVXL. Here, SWLD may be updated every time WLD is updated, or may be updated regularly after a certain period of time has passed regardless of the update timing of WLD.
[0261] SWLD can be generated for each feature quantity. For example, as shown in SWLD1 based on SHOT feature quantity and SWLD2 based on SIFT feature quantity, SWLD can be generated for each feature quantity and used according to the purpose. In addition, the feature quantity of each FVXL calculated can be stored as feature quantity information in each FVXL.
[0262] Next, a method of using the sparse world space (SWLD) will be described. Since the SWLD only includes feature voxels (FVXL), the data size is generally smaller than that of the WLD including all VXLs.
[0263] In applications that use feature quantities to achieve a certain purpose, by using SWLD information instead of WLD, it is possible to reduce the time required to read from the hard disk, and to reduce the bandwidth and transmission time during network transmission. For example, as map information, WLD and SWLD are stored in the server in advance, and the map information to be sent is switched to WLD or SWLD according to the needs of the client, thereby reducing the network bandwidth and transmission time. A specific example is shown below.
[0264] Fig.12 as well as Fig.13 The following shows the use cases of SWLD and WLD. Fig.12 As shown, when the client 1 as a vehicle-mounted device needs map information for the purpose of determining its own position, the client 1 sends a request for obtaining map data for estimating its own position to the server (S301). The server sends SWLD to the client 1 according to the acquisition request (S302). The client 1 uses the received SWLD to determine its own position (S303). At this time, the client 1 obtains the VXL information of the surrounding area of the client 1 through various methods such as a distance sensor such as a rangefinder, a stereo camera, or a combination of multiple monocular cameras, and estimates its own position information based on the obtained VXL information and SWLD. Here, the own position information includes the three-dimensional position information and orientation of the client 1.
[0265] like Fig.13 As shown, when the client 2 as a vehicle-mounted device needs map information for the purpose of drawing a three-dimensional map or the like, the client 2 sends a request for obtaining map data for drawing a map to the server (S311). The server sends a WLD to the client 2 in accordance with the acquisition request (S312). The client 2 uses the received WLD to draw a map (S313). At this time, the client 2 uses, for example, an image taken by itself with a visible light camera and the WLD obtained from the server to create a conceptual image, and draws the created image on a screen such as a car navigation system.
[0266] As described above, the server sends SWLD to the client when the feature amount of each VXL is mainly needed, such as estimating the own position, and sends WLD to the client when detailed VXL information is needed, such as drawing a map. This makes it possible to efficiently send and receive map data.
[0267] In addition, the client can determine which one of SWLD and WLD it needs and request the server to send SWLD or WLD. And the server can determine which one of SWLD and WLD should be sent according to the status of the client or the network.
[0268] Next, a method of switching the transmission and reception between the sparse world space (SWLD) and the world space (WLD) will be described.
[0269] The reception of WLD or SWLD can be switched according to the network bandwidth. Fig.14 An example of operation in this case is shown. For example, when a low-speed network that can use the network bandwidth in an LTE (Long Term Evolution) environment is used, when the client accesses the server via the low-speed network (S321), the SWLD as map information is obtained from the server (S322). In addition, when a high-speed network with sufficient network bandwidth is used in a WiFi environment, the client accesses the server via the high-speed network (S323) and obtains the WLD from the server (S324). Accordingly, the client can obtain appropriate map information according to the network bandwidth of the client.
[0270] Specifically, the client receives SWLD via LTE outdoors, and acquires WLD via WiFi when entering a facility or the like indoors. This allows the client to acquire more detailed map information indoors.
[0271] In this way, the client can request WLD or SWLD from the server according to the frequency band of the network it uses. Alternatively, the client can send information showing the frequency band of the network it uses to the server, and the server can send appropriate data (WLD or SWLD) to the client according to the information. Alternatively, the server can determine the network bandwidth of the client and send appropriate data (WLD or SWLD) to the client.
[0272] Furthermore, the reception of WLD or SWLD can be switched according to the moving speed. Fig.15 An example of operation in this case is shown. For example, when the client moves at high speed (S331), the client receives SWLD from the server (S332). In addition, when the client moves at low speed (S333), the client receives WLD from the server (S334). Accordingly, the client can suppress the network bandwidth and obtain map information according to the speed. Specifically, when the client is driving on a highway, by receiving SWLD with a small amount of data, the map information can be updated at a roughly appropriate speed. In addition, when the client is driving on a general road, by receiving WLD, more detailed map information can be obtained.
[0273] In this way, the client can request WLD or SWLD from the server according to its own moving speed. Alternatively, the client can send information showing its own moving speed to the server, and the server can send appropriate data (WLD or SWLD) to the client according to the information. Alternatively, the server can determine the moving speed of the client and send appropriate data (WLD or SWLD) to the client.
[0274] Furthermore, the client may first obtain the SWLD from the server, and then obtain the WLD of the important area. For example, when the client obtains map data, it first obtains the general map information with the SWLD, selects the area where the features such as buildings, signs, or people appear more, and then obtains the WLD of the selected area. In this way, the client can suppress the amount of data received from the server and obtain the detailed information of the required area.
[0275] Furthermore, the server may create SWLDs for each object based on the WLD, and the client may receive them separately according to the purpose. In this way, the network bandwidth can be suppressed. For example, the server pre-identifies a person or a car from the WLD, and creates a SWLD for the person and a SWLD for the car. When the client wants to obtain information about the people around it, it receives the SWLD for the person, and when it wants to obtain information about the car, it receives the SWLD for the car. Furthermore, the type of such SWLD can be distinguished based on the information (flag or type, etc.) attached to the header.
[0276] Next, the configuration and operation flow of the three-dimensional data encoding device (for example, a server) according to the present embodiment will be described. Fig.16 It is a block diagram of the three-dimensional data encoding device 400 involved in this embodiment. Fig.17 This is a flowchart of the three-dimensional data encoding process performed by the three-dimensional data encoding device 400.
[0277] Fig.16 The three-dimensional data encoding device 400 shown in the figure generates encoded three-dimensional data 413 and 414 as encoded streams by encoding input three-dimensional data 411. Here, the encoded three-dimensional data 413 is encoded three-dimensional data corresponding to WLD, and the encoded three-dimensional data 414 is encoded three-dimensional data corresponding to SWLD. The three-dimensional data encoding device 400 includes: an acquisition unit 401, a coding area determination unit 402, a SWLD extraction unit 403, a WLD encoding unit 404, and a SWLD encoding unit 405.
[0278] like Fig.17 As shown, first, the acquisition unit 401 acquires input three-dimensional data 411 which is point group data in a three-dimensional space (S401).
[0279] Next, the encoding region determination unit 402 determines the encoding target spatial region based on the spatial region where the point cloud data exists ( S402 ).
[0280] Next, the SWLD extraction unit 403 defines the spatial region of the encoding object as WLD, and calculates the feature value from each VXL included in the WLD. In addition, the SWLD extraction unit 403 extracts the VXL whose feature value is greater than a predetermined threshold value, defines the extracted VXL as FVXL, and adds the FVXL to the SWLD to generate the extracted three-dimensional data 412 (S403). That is, the extracted three-dimensional data 412 whose feature value is greater than the threshold value is extracted from the input three-dimensional data 411.
[0281] Next, the WLD encoding unit 404 generates encoded three-dimensional data 413 corresponding to the WLD by encoding the input three-dimensional data 411 corresponding to the WLD (S404). At this time, the WLD encoding unit 404 adds information for distinguishing that the encoded three-dimensional data 413 is a stream containing the WLD to the header of the encoded three-dimensional data 413.
[0282] Then, the SWLD encoding unit 405 generates encoded three-dimensional data 414 corresponding to the SWLD by encoding the extracted three-dimensional data 412 corresponding to the SWLD (S405). At this time, the SWLD encoding unit 405 adds information for distinguishing that the encoded three-dimensional data 414 is a stream including the SWLD to the header of the encoded three-dimensional data 414.
[0283] Furthermore, the processing order of the process of generating the encoded three-dimensional data 413 and the process of generating the encoded three-dimensional data 414 may be reversed from the above. Furthermore, part or all of the above-mentioned processes may be executed in parallel.
[0284] Information assigned to the header of the encoded three-dimensional data 413 and 414 is defined as a parameter such as "world_type". When world_type = 0, it means that the stream contains WLD, and when world_type = 1, it means that the stream contains SWLD. In the case of defining more categories, the assigned value can be increased, such as world_type = 2. In addition, a specific flag can be included in one of the encoded three-dimensional data 413 and 414. For example, the encoded three-dimensional data 414 can be assigned a flag indicating that the stream contains SWLD. In this case, the decoding device can determine whether it is a stream containing WLD or a stream containing SWLD based on the presence or absence of the flag.
[0285] Furthermore, the encoding method used by the WLD encoding unit 404 when encoding the WLD may be different from the encoding method used by the SWLD encoding unit 405 when encoding the SWLD.
[0286] For example, since data is decimated in SWLD, the correlation with surrounding data may be lower than that in WLD. Therefore, in the encoding method for SWLD, inter-frame prediction is prioritized over intra-frame prediction and inter-frame prediction in comparison with the encoding method for WLD.
[0287] Furthermore, the encoding method used for SWLD and the encoding method used for WLD may have different methods of expressing the three-dimensional position. For example, the three-dimensional position of FVXL may be expressed by three-dimensional coordinates in FWLD, and the three-dimensional position may be expressed by an octree described later in WLD, or vice versa.
[0288] Furthermore, the SWLD coding unit 405 performs coding in such a manner that the data size of the coded three-dimensional data 414 of SWLD is smaller than the data size of the coded three-dimensional data 413 of WLD. For example, as described above, the correlation between data in SWLD may be reduced compared to that in WLD. Accordingly, the coding efficiency is reduced, and the data size of the coded three-dimensional data 414 may be larger than the data size of the coded three-dimensional data 413 of WLD. Therefore, when the data size of the obtained coded three-dimensional data 414 is larger than the data size of the coded three-dimensional data 413 of WLD, the SWLD coding unit 405 performs re-coding to generate the coded three-dimensional data 414 with a reduced data size again.
[0289] For example, the SWLD extraction unit 403 generates the extracted three-dimensional data 412 again with the number of extracted feature points reduced, and the SWLD encoding unit 405 encodes the extracted three-dimensional data 412. Alternatively, the degree of quantization in the SWLD encoding unit 405 may be made coarser. For example, in the octree structure described later, the degree of quantization may be made coarser by rounding the data of the lowest layer.
[0290] Furthermore, when the data size of the SWLD encoded three-dimensional data 414 cannot be made smaller than the data size of the WLD encoded three-dimensional data 413, the SWLD encoding unit 405 may not generate the SWLD encoded three-dimensional data 414. Alternatively, the WLD encoded three-dimensional data 413 may be copied to the SWLD encoded three-dimensional data 414. That is, the WLD encoded three-dimensional data 413 may be used directly as the SWLD encoded three-dimensional data 414.
[0291] Next, the configuration and operation flow of the three-dimensional data decoding device (eg, client) according to the present embodiment will be described. Fig.18 It is a block diagram of a three-dimensional data decoding device 500 according to this embodiment. Fig.19 3D data decoding processing performed by the 3D data decoding apparatus 500 is shown in FIG.
[0292] Fig.18 The three-dimensional data decoding device 500 shown generates decoded three-dimensional data 512 or 513 by decoding the encoded three-dimensional data 511. Here, the encoded three-dimensional data 511 is, for example, the encoded three-dimensional data 413 or 414 generated by the three-dimensional data encoding device 400.
[0293] The three-dimensional data decoding device 500 includes an acquisition unit 501 , a header analysis unit 502 , a WLD decoding unit 503 , and a SWLD decoding unit 504 .
[0294] like Fig.19 As shown, first, the acquisition unit 501 acquires the encoded three-dimensional data 511 (S501). Then, the header analysis unit 502 analyzes the header of the encoded three-dimensional data 511 to determine whether the encoded three-dimensional data 511 is a stream containing WLD or a stream containing SWLD (S502). For example, the determination is made by referring to the above-mentioned world_type parameter.
[0295] When the coded three-dimensional data 511 is a stream including WLD ("Yes" in S503), the WLD decoding unit 503 decodes the coded three-dimensional data 511 to generate decoded three-dimensional data 512 of WLD (S504). In addition, when the coded three-dimensional data 511 is a stream including SWLD ("No" in S503), the SWLD decoding unit 504 decodes the coded three-dimensional data 511 to generate decoded three-dimensional data 513 of SWLD (S505).
[0296] Also, similar to the encoding device, the decoding method used by the WLD decoding unit 503 when decoding the WLD may be different from the decoding method used by the SWLD decoding unit 504 when decoding the SWLD. For example, in the decoding method for SWLD, priority may be given to the inter-frame prediction of intra-frame prediction and inter-frame prediction over the decoding method for WLD.
[0297] Furthermore, the decoding method for SWLD and the decoding method for WLD may have different representation methods for the three-dimensional position. For example, the three-dimensional position of FVXL may be represented by three-dimensional coordinates in SWLD and by an octree described later in WLD, or vice versa.
[0298] Next, octree expression as a method of expressing three-dimensional positions will be described. VXL data included in three-dimensional data is converted into an octree structure and encoded. Fig. 20 An example of a VXL of a WLD is shown. Fig.21 Shows Fig. 20 The octree structure of WLD is shown in Figure 1. Fig. 20 In the example shown, there are three VXLs 1 to 3 that are VXLs (hereinafter, effective VXLs) that include point groups. Fig.21 As shown, the octree structure consists of nodes and leaf nodes. Each node has a maximum of 8 nodes or leaf nodes. Each leaf node has VXL information. Here, Fig.21 Among the leaf nodes shown, leaf nodes 1, 2, and 3 represent Fig. 20 VXL1, VXL2, VXL3 shown.
[0299] Specifically, each node and leaf node corresponds to a three-dimensional position. Fig. 20 The block corresponding to node 1 is divided into 8 blocks, and among the 8 blocks, the block including the valid VXL is set as a node, and the other blocks are set as leaf nodes. The block corresponding to the node is further divided into 8 nodes or leaf nodes, and this process is repeated the same number of times as the number of levels in the tree structure. And all the blocks at the bottom are set as leaf nodes.
[0300] and, Fig. 22 Shows from Fig. 20 An example of SWLD generated by WLD is shown. Fig. 20 The feature extraction results of VXL1 and VXL2 shown are determined to be FVXL1 and FVXL2 and are included in SWLD. In addition, VXL3 is not determined to be FVXL and is therefore not included in SWLD. Fig.23 Shows Fig. 22 The octree structure of SWLD is shown in Figure 1. Fig.23 In the octree structure shown, Fig.21 The leaf node 3 shown, which is equivalent to VXL3, is deleted. Fig.21 The node 3 shown has no valid VXL and is changed to a leaf node. In this way, generally speaking, the number of leaf nodes of SWLD is smaller than that of WLD, and the encoded three-dimensional data of SWLD is also smaller than that of WLD.
[0301] Modifications of this embodiment will be described below.
[0302] For example, when a client such as a vehicle-mounted device estimates its own position, it receives SWLD from a server, uses SWLD to estimate its own position, and performs obstacle detection. It then uses various methods such as distance sensors such as rangefinders, stereo cameras, or a combination of multiple monocular cameras to perform obstacle detection based on the three-dimensional information of the surrounding area obtained by itself.
[0303] In general, it is difficult to include VXL data of flat areas in SWLD. For this reason, the server maintains a subsampled world space (SubWLD) that is a downsampled version of WLD for detecting stationary obstacles, and can send SWLD and SubWLD to the client. This can suppress network bandwidth while enabling the client to estimate its own position and detect obstacles.
[0304] Furthermore, when the client quickly depicts three-dimensional map data, it is convenient if the map information is in a grid structure. Therefore, the server can generate a grid based on the WLD and store it in advance as a grid world space (MWLD). For example, when the client needs to perform a rough three-dimensional depiction, the MWLD is received, and when a detailed three-dimensional depiction is required, the WLD is received. In this way, the network bandwidth can be suppressed.
[0305] Furthermore, although the server sets the VXL with a feature value above the threshold value as FVXL from each VXL, FVXL can also be calculated by different methods. For example, if the server determines that VXL, VLM, SPC, or GOS constituting a signal or intersection is required for self-position estimation, driving assistance, or automatic driving, it can be included in SWLD as FVXL, FVLM, FSPC, FGOS. Furthermore, the above judgment can be performed manually. In addition, FVXL obtained by the above method can be added to FVXL set based on the feature value. That is, the SWLD extraction unit 403 can further extract data corresponding to an object with predetermined attributes from the input three-dimensional data 411 as extracted three-dimensional data 412.
[0306] Furthermore, labels different from feature quantities may be assigned to situations that require these purposes. The server may separately maintain FVXL required for self-position estimation such as signals or intersections, driving assistance, or autonomous driving as an upper layer of SWLD (e.g., lane world space).
[0307] Furthermore, the server may also attach attributes to the VXL in the WLD in random access units or specified units. Attributes include, for example, information indicating whether the location is required or not required for estimating the location, or information indicating whether traffic information such as signals or intersections is important. Furthermore, attributes may also include the correspondence between lane information (GDF: Geographic Data Files, etc.) and features (intersections or roads, etc.).
[0308] Furthermore, as a method of updating WLD or SWLD, the following method can be adopted.
[0309] Update information showing changes in people, construction, or street trees (trajectory orientation) is uploaded to the server as a point group or metadata. The server updates the WLD based on the upload, and then updates the SWLD using the updated WLD.
[0310] Furthermore, when the client detects a mismatch between the three-dimensional information generated by itself and the three-dimensional information received from the server when estimating its own position, the client can send the three-dimensional information generated by itself to the server together with the update notification. In this case, the server updates the SWLD using the WLD. If the SWLD is not updated, the server determines that the WLD itself is old.
[0311] Furthermore, although information for distinguishing between WLD and SWLD is added as the header information of the coded stream, when there are multiple world spaces such as a grid world space or a lane world space, information for distinguishing them may be added to the header information. Furthermore, when there are multiple SWLDs with different feature quantities, information for distinguishing them may also be added to the header information.
[0312] Furthermore, although SWLD is composed of FVXL, it may also include VXL that is not determined to be FVXL. For example, SWLD may include adjacent VXL used when calculating the feature quantity of FVXL. Accordingly, even if each FVXL of SWLD does not have feature quantity information attached, the client can calculate the feature quantity of FVXL when receiving SWLD. In addition, at this time, SWLD may include information for distinguishing whether each VXL is FVXL or VXL.
[0313] As described above, the three-dimensional data encoding device 400 extracts extracted three-dimensional data 412 (second three-dimensional data) whose feature value is above a threshold value from the input three-dimensional data 411 (first three-dimensional data), and generates encoded three-dimensional data 414 (first encoded three-dimensional data) by encoding the extracted three-dimensional data 412.
[0314] According to this, the three-dimensional data encoding device 400 generates the encoded three-dimensional data 414 obtained by encoding the data whose feature quantity is greater than the threshold value. In this way, the amount of data can be reduced compared to the case where the input three-dimensional data 411 is directly encoded. Therefore, the three-dimensional data encoding device 400 can reduce the amount of data during transmission.
[0315] Furthermore, the three-dimensional data encoding device 400 further encodes the input three-dimensional data 411 to generate encoded three-dimensional data 413 (second encoded three-dimensional data).
[0316] According to this, the three-dimensional data encoding device 400 can selectively transmit the encoded three-dimensional data 413 and the encoded three-dimensional data 414 according to the purpose of use, for example.
[0317] Furthermore, the extracted three-dimensional data 412 is encoded by a first encoding method, and the input three-dimensional data 411 is encoded by a second encoding method different from the first encoding method.
[0318] Accordingly, the three-dimensional data encoding device 400 can adopt appropriate encoding methods for the input three-dimensional data 411 and the extracted three-dimensional data 412 respectively.
[0319] Furthermore, in the first encoding method, compared with the second encoding method, priority is given to the inter-frame prediction between the intra-frame prediction and the inter-frame prediction.
[0320] According to this, the three-dimensional data encoding device 400 can increase the priority of inter-frame prediction for the extracted three-dimensional data 412 where the correlation between adjacent data is likely to be low.
[0321] Furthermore, the first encoding method and the second encoding method have different methods of expressing the three-dimensional position. For example, in the second encoding method, the three-dimensional position is expressed by an octree, while in the first encoding method, the three-dimensional position is expressed by three-dimensional coordinates.
[0322] According to this, the three-dimensional data encoding device 400 can adopt a more appropriate three-dimensional position expression method for three-dimensional data having different data numbers (number of VXLs or FVXLs).
[0323] Furthermore, at least one of the coded three-dimensional data 413 and 414 includes an identifier indicating whether the coded three-dimensional data is coded three-dimensional data obtained by encoding the input three-dimensional data 411 or coded three-dimensional data obtained by encoding a part of the input three-dimensional data 411. That is, the identifier indicates whether the coded three-dimensional data is the coded three-dimensional data 413 of the WLD or the coded three-dimensional data 414 of the SWLD.
[0324] Based on this, the decoding device can easily determine whether the acquired encoded three-dimensional data is the encoded three-dimensional data 413 or the encoded three-dimensional data 414.
[0325] Furthermore, the three-dimensional data encoding device 400 encodes the extracted three-dimensional data 412 in such a manner that the data amount of the encoded three-dimensional data 414 is smaller than the data amount of the encoded three-dimensional data 413 .
[0326] According to this, the three-dimensional data encoding device 400 can make the data amount of the encoded three-dimensional data 414 smaller than the data amount of the encoded three-dimensional data 413 .
[0327] Furthermore, the three-dimensional data encoding device 400 further extracts data corresponding to an object having a predetermined attribute from the input three-dimensional data 411 as extracted three-dimensional data 412. For example, the object having a predetermined attribute is an object required for self-position estimation, driving assistance, or automatic driving, such as a signal or an intersection.
[0328] Accordingly, the three-dimensional data encoding device 400 can generate the encoded three-dimensional data 414 including the data required by the decoding device.
[0329] Furthermore, the three-dimensional data encoding device 400 (server) further transmits one of the encoded three-dimensional data 413 and 414 to the client according to the status of the client.
[0330] Accordingly, the three-dimensional data encoding device 400 can send appropriate data according to the status of the client.
[0331] Furthermore, the status of the client includes the communication status of the client (eg, network bandwidth) or the moving speed of the client.
[0332] Furthermore, the three-dimensional data encoding device 400 further transmits one of the encoded three-dimensional data 413 and 414 to the client according to the request of the client.
[0333] Thereby, the three-dimensional data encoding device 400 can send appropriate data according to the request of the client.
[0334] Furthermore, the three-dimensional data decoding device 500 according to the present embodiment decodes the encoded three-dimensional data 413 or 414 generated by the three-dimensional data encoding device 400 .
[0335] That is, the three-dimensional data decoding device 500 decodes the encoded three-dimensional data 414 obtained by encoding the extracted three-dimensional data 412 whose feature amount extracted from the input three-dimensional data 411 is greater than the threshold value by the first decoding method. Furthermore, the three-dimensional data decoding device 500 decodes the encoded three-dimensional data 413 obtained by encoding the input three-dimensional data 411 by using a second decoding method different from the first decoding method.
[0336] According to this, the three-dimensional data decoding device 500 can selectively receive the encoded three-dimensional data 414 and the encoded three-dimensional data 413 obtained by encoding data having a feature value greater than a threshold value, for example, according to the purpose of use. According to this, the three-dimensional data decoding device 500 can reduce the amount of data during transmission. Furthermore, the three-dimensional data decoding device 500 can adopt appropriate decoding methods for the input three-dimensional data 411 and the extracted three-dimensional data 412, respectively.
[0337] Furthermore, in the first decoding method, compared with the second decoding method, priority is given to the inter prediction between the intra prediction and the inter prediction.
[0338] According to this, the three-dimensional data decoding apparatus 500 can increase the priority of inter-frame prediction for extracting three-dimensional data in which the correlation between adjacent data is likely to be low.
[0339] Furthermore, the first decoding method and the second decoding method use different methods for expressing the three-dimensional position. For example, the second decoding method expresses the three-dimensional position by an octree, while the first decoding method expresses the three-dimensional position by three-dimensional coordinates.
[0340] According to this, the three-dimensional data decoding apparatus 500 can adopt a more appropriate three-dimensional position expression method for three-dimensional data having different data numbers (number of VXLs or FVXLs).
[0341] Furthermore, at least one of the encoded three-dimensional data 413 and 414 includes an identifier indicating whether the encoded three-dimensional data is encoded three-dimensional data obtained by encoding the input three-dimensional data 411 or encoded three-dimensional data obtained by encoding a part of the input three-dimensional data 411. The three-dimensional data decoding device 500 identifies the encoded three-dimensional data 413 and 414 by referring to the identifier.
[0342] Based on this, the three-dimensional data decoding device 500 can easily determine whether the obtained encoded three-dimensional data is the encoded three-dimensional data 413 or the encoded three-dimensional data 414 .
[0343] Furthermore, the 3D data decoding device 500 notifies the server of the state of the client (the 3D data decoding device 500). The 3D data decoding device 500 receives one of the encoded 3D data 413 and 414 transmitted from the server according to the state of the client.
[0344] Accordingly, the three-dimensional data decoding apparatus 500 can receive appropriate data according to the status of the client.
[0345] Furthermore, the status of the client includes the communication status of the client (eg, network bandwidth) or the moving speed of the client.
[0346] Furthermore, the three-dimensional data decoding apparatus 500 further requests one of the encoded three-dimensional data 413 and 414 from the server, and receives one of the encoded three-dimensional data 413 and 414 transmitted from the server in accordance with the request.
[0347] Thereby, the three-dimensional data decoding device 500 can receive appropriate data according to the application.
[0348] (Implementation method 3)
[0349] In this embodiment, a method for transmitting and receiving three-dimensional data between vehicles will be described. For example, the three-dimensional data is transmitted and received between the own vehicle and surrounding vehicles.
[0350] Fig.24 This is a block diagram of a three-dimensional data production device 620 according to this embodiment. The three-dimensional data production device 620 is included in the vehicle, for example, and produces denser third three-dimensional data 636 by synthesizing the received second three-dimensional data 635 with the first three-dimensional data 632 produced by the three-dimensional data production device 620.
[0351] The three-dimensional data creation device 620 includes a three-dimensional data creation unit 621 , a request range determination unit 622 , a search unit 623 , a reception unit 624 , a decoding unit 625 , and a synthesis unit 626 .
[0352] First, the three-dimensional data creation unit 621 uses sensor information 631 detected by a sensor of the own vehicle to create first three-dimensional data 632. Next, the request range determination unit 622 determines a request range, which is a three-dimensional space range where the created first three-dimensional data 632 does not have enough data.
[0353] Next, the search unit 623 searches for surrounding vehicles that have three-dimensional data of the requested range, and sends request range information 633 showing the requested range to the surrounding vehicles determined by the search. Next, the receiving unit 624 receives the encoded three-dimensional data 634 (S624) as the encoded stream of the requested range from the surrounding vehicles. In addition, the search unit 623 can indiscriminately issue requests to all vehicles existing in the determined range, and receive the encoded three-dimensional data 634 from the responding party. In addition, the search unit 623 is not limited to vehicles, and can also issue requests to objects such as traffic lights or signs, and receive the encoded three-dimensional data 634 from the object.
[0354] Next, the decoder 625 decodes the received encoded three-dimensional data 634 to obtain second three-dimensional data 635. Next, the synthesizer 626 synthesizes the first three-dimensional data 632 and the second three-dimensional data 635 to create denser third three-dimensional data 636.
[0355] Next, the configuration and operation of the three-dimensional data transmitting device 640 according to this embodiment will be described. Fig.25 It is a block diagram of the three-dimensional data transmitting device 640.
[0356] The three-dimensional data sending device 640 is included in the above-mentioned surrounding vehicles, for example, and processes the fifth three-dimensional data 652 produced by the surrounding vehicles into the sixth three-dimensional data 654 requested by the own vehicle, and generates encoded three-dimensional data 634 by encoding the sixth three-dimensional data 654, and sends the encoded three-dimensional data 634 to the own vehicle.
[0357] The three-dimensional data transmitting device 640 includes a three-dimensional data creating unit 641 , a receiving unit 642 , an extracting unit 643 , a coding unit 644 , and a transmitting unit 645 .
[0358] First, the three-dimensional data creation unit 641 creates fifth three-dimensional data 652 using sensor information 651 detected by sensors provided in surrounding vehicles. Next, the receiving unit 642 receives the requested range information 633 transmitted from the own vehicle.
[0359] Next, the extraction unit 643 extracts the three-dimensional data of the requested range indicated by the requested range information 633 from the fifth three-dimensional data 652, and processes the fifth three-dimensional data 652 into the sixth three-dimensional data 654. Next, the encoding unit 644 encodes the sixth three-dimensional data 654, thereby generating the encoded three-dimensional data 634 as an encoded stream. Then, the transmission unit 645 transmits the encoded three-dimensional data 634 to the own vehicle.
[0360] In addition, although an example is described here in which the own vehicle has the three-dimensional data creation device 620 and the surrounding vehicles have the three-dimensional data transmission device 640, each vehicle may also have the functions of the three-dimensional data creation device 620 and the three-dimensional data transmission device 640.
[0361] (Implementation 4)
[0362] In this embodiment, the operation related to abnormal situations in the estimation of the own position based on the three-dimensional map will be described.
[0363] The use of autonomous movement of moving objects such as automatic driving of automobiles, robots, or flying objects such as drones will expand in the future. An example of a method for achieving such autonomous movement is a method in which a moving object estimates its position in a three-dimensional map (self-position estimation) and drives according to the map.
[0364] The self-position estimation is achieved by matching the three-dimensional map with the three-dimensional information around the own vehicle obtained by sensors such as the rangefinder (LIDAR, etc.) or stereo camera mounted on the own vehicle (hereinafter referred to as the own vehicle detection three-dimensional data), and estimating the own vehicle position within the three-dimensional map.
[0365] As shown in the HD map proposed by HERE, a 3D map is not only a 3D point cloud, but also includes 2D map data such as road and intersection shape information, and information such as congestion and accidents that change in real time. The 3D map is composed of multiple layers such as 3D data, 2D data, and metadata that changes in real time, and the device can obtain only the required data or refer to the required data.
[0366] The point cloud data may be the above-mentioned SWLD, or may include point group data that is not a feature point. Furthermore, the transmission and reception of the point cloud data is basically performed in one or more random access units.
[0367] As a matching method for the three-dimensional map and the three-dimensional data detected by the own vehicle, the following method can be adopted. For example, the device compares the shapes of the point groups in each other's point clouds and determines the parts with high similarity between the feature points as the same position. In addition, when the three-dimensional map is composed of SWLD, the device compares and matches the feature points constituting the SWLD with the three-dimensional feature points extracted from the three-dimensional data detected by the own vehicle.
[0368] Here, in order to estimate the own position with high accuracy, the following (A) and (B) need to be met: (A) a three-dimensional map and three-dimensional data for detecting the own vehicle are available, and (B) their accuracy meets a predetermined benchmark. However, in the following abnormal situation, (A) or (B) cannot be met.
[0369] (1) A three-dimensional map cannot be obtained through the communication path.
[0370] (2) There is no three-dimensional map or the obtained three-dimensional map is damaged.
[0371] (3) The sensor of the own vehicle fails, or due to bad weather, the accuracy of the three-dimensional data generated by the own vehicle is insufficient.
[0372] The following describes the operation for dealing with these abnormal situations. Although the operation is described below using a vehicle as an example, the following method can also be applied to all moving objects that move autonomously, such as robots and drones.
[0373] The following describes the configuration and operation of the three-dimensional information processing device according to the present embodiment for detecting abnormalities in three-dimensional data corresponding to a three-dimensional map or a vehicle. Fig.26 It is a block diagram showing a configuration example of a three-dimensional information processing device 700 according to this embodiment.
[0374] The three-dimensional information processing device 700 is mounted on a mobile object such as a motor vehicle. Fig.26As shown, the three-dimensional information processing device 700 includes a three-dimensional map acquisition unit 701 , a vehicle detection data acquisition unit 702 , an abnormality determination unit 703 , a countermeasure operation determination unit 704 , and an operation control unit 705 .
[0375] In addition, the three-dimensional information processing device 700 may also include a camera for obtaining a two-dimensional image, or may include a sensor for one-dimensional data using ultrasonic waves or lasers, which is used to detect structural objects or moving objects around the vehicle, which is not shown in the figure. In addition, the three-dimensional information processing device 700 may also include a communication unit (not shown) for obtaining a three-dimensional map through a mobile communication network such as 4G or 5G, or communication between vehicles, or communication between roads and vehicles.
[0376] The three-dimensional map acquisition unit 701 acquires a three-dimensional map 711 near the driving route. For example, the three-dimensional map acquisition unit 701 acquires the three-dimensional map 711 through a mobile communication network, communication between vehicles, or communication between roads and vehicles.
[0377] Next, the own vehicle detection data acquisition unit 702 acquires the own vehicle detection three-dimensional data 712 based on the sensor information. For example, the own vehicle detection data acquisition unit 702 generates the own vehicle detection three-dimensional data 712 based on the sensor information acquired by the sensor included in the own vehicle.
[0378] Next, the abnormality determination unit 703 detects an abnormality by performing a predetermined check on at least one of the obtained three-dimensional map 711 and the vehicle detection three-dimensional data 712. That is, the abnormality determination unit 703 determines whether at least one of the obtained three-dimensional map 711 and the vehicle detection three-dimensional data 712 is abnormal.
[0379] When an abnormal situation is detected, the countermeasure action determination unit 704 determines a countermeasure action for the abnormal situation. Next, the action control unit 705 controls the operation of each processing unit required for executing the countermeasure action, such as the three-dimensional map acquisition unit 701.
[0380] In addition, when no abnormality is detected, the three-dimensional information processing device 700 ends the processing.
[0381] The three-dimensional information processing device 700 estimates the own position of the vehicle having the three-dimensional information processing device 700 using the three-dimensional map 711 and the own vehicle detection three-dimensional data 712. Then, the three-dimensional information processing device 700 uses the result of the own position estimation to make the vehicle automatically drive.
[0382] Accordingly, the three-dimensional information processing device 700 obtains map data (three-dimensional map 711) including the first three-dimensional position information via the channel. For example, the first three-dimensional position information is encoded in units of partial spaces having three-dimensional coordinate information, and the first three-dimensional position information includes a plurality of random access units, each of which is a collection of more than one partial space and can be independently decoded. For example, the first three-dimensional position information is data (SWLD) encoded with a feature point whose three-dimensional feature quantity is above a predetermined threshold.
[0383] Furthermore, the three-dimensional information processing device 700 generates the second three-dimensional position information (the vehicle detection three-dimensional data 712) based on the information detected by the sensor. Next, the three-dimensional information processing device 700 performs an abnormality determination process on the first three-dimensional position information or the second three-dimensional position information to determine whether the first three-dimensional position information or the second three-dimensional position information is abnormal.
[0384] When determining that the first three-dimensional position information or the second three-dimensional position information is abnormal, the three-dimensional information processing device 700 determines a countermeasure for the abnormality. Next, the three-dimensional information processing device 700 executes control required for executing the countermeasure.
[0385] According to this, the three-dimensional information processing device 700 can detect abnormality in the first three-dimensional position information or the second three-dimensional position information and perform a corresponding operation.
[0386] (Implementation method 5)
[0387] In this embodiment, a method of transmitting three-dimensional data to a rear vehicle and the like will be described.
[0388] Fig. 27 1 is a block diagram showing a configuration example of a three-dimensional data production device 810 according to the present embodiment. The three-dimensional data production device 810 is mounted on a vehicle, for example. The three-dimensional data production device 810 transmits and receives three-dimensional data with external traffic cloud monitoring, a vehicle in front, or a vehicle behind, and produces and accumulates the three-dimensional data.
[0389] The three-dimensional data production device 810 includes: a data receiving unit 811, a communication unit 812, a receiving control unit 813, a format conversion unit 814, multiple sensors 815, a three-dimensional data production unit 816, a three-dimensional data synthesis unit 817, a three-dimensional data storage unit 818, a communication unit 819, a sending control unit 820, a format conversion unit 821, and a data sending unit 822.
[0390] The data receiving unit 811 receives three-dimensional data 831 from traffic cloud monitoring or the vehicle ahead. The three-dimensional data 831 includes, for example, a point cloud including information on areas that cannot be detected by the vehicle's sensor 815, visible light images, depth information, sensor position information, or speed information.
[0391] The communication unit 812 communicates with the traffic cloud monitoring or the vehicle ahead, and sends a data transmission request or the like to the traffic cloud monitoring or the vehicle ahead.
[0392] The reception control unit 813 exchanges information such as a corresponding format with the communication partner via the communication unit 812, and establishes communication with the communication partner.
[0393] The format conversion unit 814 generates three-dimensional data 832 by performing format conversion or the like on the three-dimensional data 831 received by the data reception unit 811. Furthermore, when the three-dimensional data 831 is compressed or encoded, the format conversion unit 814 performs decompression or decoding processing.
[0394] The plurality of sensors 815 are a group of sensors such as LiDAR, a visible light camera, or an infrared camera that obtains information outside the vehicle and generates sensor information 833. For example, when the sensor 815 is a laser sensor such as LiDAR, the sensor information 833 is three-dimensional data such as a point cloud (point group data). In addition, the number of sensors 815 may not be multiple.
[0395] The three-dimensional data generating unit 816 generates three-dimensional data 834 based on the sensor information 833. The three-dimensional data 834 includes, for example, point cloud, visible light image, depth information, sensor position information, or speed information.
[0396] The three-dimensional data synthesis unit 817 synthesizes the three-dimensional data 832 produced by traffic cloud monitoring or the front vehicle, etc., into the three-dimensional data 834 produced based on the sensor information 833 of the own vehicle, thereby constructing three-dimensional data 835 that also includes the space in front of the front vehicle that cannot be detected by the sensor 815 of the own vehicle.
[0397] The three-dimensional data accumulation unit 818 accumulates the generated three-dimensional data 835 and the like.
[0398] The communication unit 819 communicates with the traffic cloud monitoring or the vehicle behind, and sends a data transmission request and the like to the traffic cloud monitoring or the vehicle behind.
[0399] The transmission control unit 820 exchanges information such as the corresponding format with the communication partner via the communication unit 819 to establish communication with the communication partner. In addition, the transmission control unit 820 determines the transmission area of the space of the three-dimensional data to be transmitted based on the three-dimensional data construction information of the three-dimensional data 832 generated by the three-dimensional data synthesis unit 817 and the data transmission request from the communication partner.
[0400] Specifically, the transmission control unit 820 determines the transmission area including the space in front of the vehicle that cannot be detected by the sensor of the rear vehicle according to the data transmission request from the traffic cloud monitoring or the rear vehicle. In addition, the transmission control unit 820 determines the transmission area by judging whether the space that can be transmitted or the transmitted space has been updated based on the three-dimensional data construction information. For example, the transmission control unit 820 determines the area that is both specified by the data transmission request and the area where the corresponding three-dimensional data 835 exists as the transmission area. In addition, the transmission control unit 820 notifies the format corresponding to the communication partner and the transmission area to the format conversion unit 821.
[0401] The format conversion unit 821 generates three-dimensional data 837 by converting the three-dimensional data 836 of the transmission area in the three-dimensional data 835 accumulated in the three-dimensional data accumulation unit 818 into a format corresponding to the receiving side. In addition, the format conversion unit 821 may compress or encode the three-dimensional data 837 to reduce the data amount.
[0402] The data transmission unit 822 transmits the three-dimensional data 837 to the traffic cloud monitoring or the rear vehicle. The three-dimensional data 837 includes, for example, a point cloud in front of the vehicle including information on the area that becomes a blind spot for the rear vehicle, a visible light image, depth information, or sensor position information.
[0403] In addition, although the example in which the format conversion units 814 and 821 perform format conversion is described here, format conversion may not be performed.
[0404] With this configuration, the three-dimensional data production device 810 obtains three-dimensional data 831 of an area that cannot be detected by the sensor 815 of the own vehicle from the outside, and generates three-dimensional data 835 by synthesizing the three-dimensional data 831 with three-dimensional data 834 based on sensor information 833 detected by the sensor 815 of the own vehicle. In this way, the three-dimensional data production device 810 can generate three-dimensional data of a range that cannot be detected by the sensor 815 of the own vehicle.
[0405] In addition, the three-dimensional data production device 810 can send three-dimensional data of the space in front of its own vehicle that cannot be detected by the sensors of the rear vehicle to the traffic cloud monitoring or the rear vehicle, etc. according to the data sending request from the traffic cloud monitoring or the rear vehicle.
[0406] (Implementation 6)
[0407] In the fifth embodiment, a client device such as a vehicle sends three-dimensional data to another vehicle or a server such as a traffic cloud monitoring device. In this embodiment, the client device sends sensor information obtained by the sensor to the server or other client devices.
[0408] First, the configuration of a system according to this embodiment will be described. Fig.28 1 shows the structure of the system for transmitting and receiving three-dimensional maps and sensor information according to the present embodiment. The system includes a server 901 and client devices 902A and 902B. In addition, when the client devices 902A and 902B are not specifically distinguished, they are also referred to as client devices 902.
[0409] The client device 902 is, for example, an in-vehicle device mounted on a mobile body such as a vehicle. The server 901 is, for example, a traffic cloud monitoring system, and can communicate with a plurality of client devices 902 .
[0410] The server 901 transmits the three-dimensional map composed of the point cloud to the client device 902. In addition, the composition of the three-dimensional map is not limited to the point cloud, and can also be expressed by other three-dimensional data such as a mesh structure.
[0411] The client device 902 sends the sensor information obtained by the client device 902 to the server 901. The sensor information includes, for example, at least one of LiDAR information, visible light images, infrared images, depth images, sensor position information, and speed information.
[0412] The data sent and received between the server 901 and the client device 902 may be compressed when it is desired to reduce the data, and may not be compressed when it is desired to maintain the accuracy of the data. When compressing the data, a three-dimensional compression method based on an octree may be used in the point cloud, for example. In addition, a two-dimensional image compression method may be used in the visible light image, the infrared image, and the depth image. The two-dimensional image compression method is, for example, MPEG-4AVC or HEVC standardized by MPEG.
[0413] Furthermore, the server 901 transmits the three-dimensional map managed by the server 901 to the client device 902 in accordance with the three-dimensional map transmission request from the client device 902. In addition, the server 901 may transmit the three-dimensional map without waiting for the three-dimensional map transmission request from the client device 902. For example, the server 901 may broadcast the three-dimensional map to one or more client devices 902 in a predetermined space. Furthermore, the server 901 may transmit the three-dimensional map suitable for the location of the client device 902 at regular intervals to the client device 902 that has received a transmission request once. Furthermore, the server 901 may transmit the three-dimensional map to the client device 902 every time the three-dimensional map managed by the server 901 is updated.
[0414] The client device 902 sends a three-dimensional map transmission request to the server 901. For example, when the client device 902 wants to estimate its own position while driving, the client device 902 sends a three-dimensional map transmission request to the server 901.
[0415] In addition, in the following cases, the client device 902 may also issue a three-dimensional map transmission request to the server 901. In the case where the three-dimensional map held by the client device 902 is relatively old, the client device 902 may also issue a three-dimensional map transmission request to the server 901. For example, in the case where the client device 902 obtains the three-dimensional map and a certain period of time has passed, the client device 902 may also issue a three-dimensional map transmission request to the server 901.
[0416] Alternatively, the client device 902 may send a three-dimensional map transmission request to the server 901 before a certain time when the client device 902 is about to leave the space shown on the three-dimensional map held by the client device 902. For example, the client device 902 may send a three-dimensional map transmission request to the server 901 when the client device 902 is within a predetermined distance from the boundary of the space shown on the three-dimensional map held by the client device 902. Furthermore, when the moving path and moving speed of the client device 902 are known, the time when the client device 902 leaves the space shown on the three-dimensional map held by the client device 902 may be predicted based on the known moving path and moving speed.
[0417] When the error between the three-dimensional data generated by the client device 902 based on the sensor information and the position of the three-dimensional map is greater than a certain range, the client device 902 may send a request to the server 901 to send the three-dimensional map.
[0418] The client device 902 transmits the sensor information to the server 901 in accordance with the sensor information transmission request transmitted from the server 901. In addition, the client device 902 may transmit the sensor information to the server 901 without waiting for the sensor information transmission request from the server 901. For example, when the client device 902 has received a sensor information transmission request from the server 901 once, the client device 902 may periodically transmit the sensor information to the server 901 within a certain period. Furthermore, when the error between the three-dimensional data generated by the client device 902 based on the sensor information and the position of the three-dimensional map obtained from the server 901 is greater than a certain range, the client device 902 may determine that there is a possibility that the three-dimensional map around the client device 902 has changed, and transmit the determination result together with the sensor information to the server 901.
[0419] The server 901 issues a request to send sensor information to the client device 902. For example, the server 901 receives location information of the client device 902 such as GPS from the client device 902. When the server 901 determines that the client device 902 is close to a space with little information in the three-dimensional map managed by the server 901 based on the location information of the client device 902, the server 901 issues a request to send sensor information to the client device 902 in order to regenerate the three-dimensional map. In addition, the server 901 may issue a request to send sensor information when it is desired to update the three-dimensional map, when it is desired to check the road conditions such as when snow is accumulated or when a disaster occurs, when it is desired to check the congestion conditions or the accident conditions.
[0420] Furthermore, the client device 902 may set the data volume of the sensor information to be sent to the server 901 according to the communication state or the frequency band when receiving the sensor information sending request received from the server 901. Setting the data volume of the sensor information to be sent to the server 901 means, for example, increasing or decreasing the data itself or selecting an appropriate compression method.
[0421] Fig.29 is a block diagram showing an example of the configuration of the client device 902. The client device 902 receives a three-dimensional map composed of a point cloud or the like from the server 901, and estimates the position of the client device 902 itself based on the three-dimensional data produced based on the sensor information of the client device 902. The client device 902 then transmits the obtained sensor information to the server 901.
[0422] The client device 902 includes: a data receiving unit 1011, a communication unit 1012, a receiving control unit 1013, a format conversion unit 1014, multiple sensors 1015, a three-dimensional data production unit 1016, a three-dimensional image processing unit 1017, a three-dimensional data storage unit 1018, a format conversion unit 1019, a communication unit 1020, a sending control unit 1021, and a data sending unit 1022.
[0423] The data receiving unit 1011 receives a three-dimensional map 1031 from the server 901. The three-dimensional map 1031 is data including a point cloud such as WLD or SWLD. The three-dimensional map 1031 may include either compressed data or uncompressed data.
[0424] The communication unit 1012 communicates with the server 901 and transmits a data transmission request (for example, a transmission request of a three-dimensional map) and the like to the server 901 .
[0425] The reception control unit 1013 exchanges information such as the corresponding format with the communication partner via the communication unit 1012, and establishes communication with the communication partner.
[0426] The format conversion unit 1014 generates a three-dimensional map 1032 by performing format conversion on the three-dimensional map 1031 received by the data receiving unit 1011. Furthermore, the format conversion unit 1014 performs decompression or decoding processing when the three-dimensional map 1031 is compressed or encoded. In addition, the format conversion unit 1014 does not perform decompression or decoding processing when the three-dimensional map 1031 is non-compressed data.
[0427] The plurality of sensors 1015 are a group of sensors mounted on the client device 902 such as LiDAR, a visible light camera, an infrared camera, or a depth sensor for obtaining information outside the vehicle, and generate sensor information 1033. For example, when the sensor 1015 is a laser sensor such as LiDAR, the sensor information 1033 is three-dimensional data such as a point cloud (point group data). In addition, the number of sensors 1015 may not be multiple.
[0428] The three-dimensional data generating unit 1016 generates three-dimensional data 1034 of the surroundings of the own vehicle based on the sensor information 1033. For example, the three-dimensional data generating unit 1016 generates point cloud data having color information of the surroundings of the own vehicle using information obtained by LiDAR and visible light images obtained by a visible light camera.
[0429] The three-dimensional image processing unit 1017 uses the received three-dimensional map 1032 such as the point cloud and the three-dimensional data 1034 of the surroundings of the own vehicle generated based on the sensor information 1033 to perform the own vehicle's own position estimation processing, etc. Alternatively, the three-dimensional image processing unit 1017 may synthesize the three-dimensional map 1032 and the three-dimensional data 1034 to produce three-dimensional data 1035 of the surroundings of the own vehicle, and use the produced three-dimensional data 1035 to perform the own position estimation processing.
[0430] The three-dimensional data storage unit 1018 stores the three-dimensional map 1032 , the three-dimensional data 1034 , the three-dimensional data 1035 , and the like.
[0431] The format conversion unit 1019 generates sensor information 1037 by converting the sensor information 1033 into a format corresponding to the receiving side. In addition, the format conversion unit 1019 can reduce the amount of data by compressing or encoding the sensor information 1037. In addition, when format conversion is not required, the format conversion unit 1019 can omit the processing. In addition, the format conversion unit 1019 can control the amount of data to be transmitted according to the designation of the transmission range.
[0432] The communication unit 1020 communicates with the server 901 , and receives a data transmission request (a sensor information transmission request) and the like from the server 901 .
[0433] The transmission control unit 1021 exchanges information such as a corresponding format with the communication partner via the communication unit 1020, thereby establishing communication.
[0434] The data transmission unit 1022 transmits the sensor information 1037 to the server 901. The sensor information 1037 includes information obtained by multiple sensors 1015, such as information obtained by LiDAR, a brightness image (visible light image) obtained by a visible light camera, an infrared image obtained by an infrared camera, a depth image obtained by a depth sensor, sensor position information, and speed information.
[0435] Next, the configuration of the server 901 will be described. Fig.30 1 is a block diagram showing an example of the configuration of the server 901. The server 901 receives sensor information sent from the client device 902, and creates three-dimensional data based on the received sensor information. The server 901 updates the three-dimensional map managed by the server 901 using the created three-dimensional data. Furthermore, the server 901 sends the updated three-dimensional map to the client device 902 in accordance with the three-dimensional map sending request from the client device 902.
[0436] The server 901 includes: a data receiving unit 1111, a communication unit 1112, a receiving control unit 1113, a format conversion unit 1114, a three-dimensional data production unit 1116, a three-dimensional data synthesis unit 1117, a three-dimensional data storage unit 1118, a format conversion unit 1119, a communication unit 1120, a sending control unit 1121, and a data sending unit 1122.
[0437] The data receiving unit 1111 receives sensor information 1037 from the client device 902. The sensor information 1037 includes, for example, information obtained by LiDAR, a brightness image (visible light image) obtained by a visible light camera, an infrared image obtained by an infrared camera, a depth image obtained by a depth sensor, sensor position information, and speed information.
[0438] The communication unit 1112 communicates with the client device 902 and transmits a data transmission request (for example, a sensor information transmission request) and the like to the client device 902 .
[0439] The reception control unit 1113 exchanges information such as the corresponding format with the communication partner via the communication unit 1112, thereby establishing communication.
[0440] When the received sensor information 1037 is compressed or encoded, the format conversion unit 1114 performs decompression or decoding processing to generate the sensor information 1132. When the sensor information 1037 is uncompressed data, the format conversion unit 1114 does not perform decompression or decoding processing.
[0441] The three-dimensional data generating unit 1116 generates three-dimensional data 1134 of the surroundings of the client device 902 based on the sensor information 1132. For example, the three-dimensional data generating unit 1116 generates point cloud data having color information of the surroundings of the client device 902 using information obtained by LiDAR and visible light images obtained by a visible light camera.
[0442] The three-dimensional data synthesis unit 1117 synthesizes the three-dimensional data 1134 generated based on the sensor information 1132 with the three-dimensional map 1135 managed by the server 901 , thereby updating the three-dimensional map 1135 .
[0443] The three-dimensional data storage unit 1118 stores the three-dimensional map 1135 and the like.
[0444] The format conversion unit 1119 generates the three-dimensional map 1031 by converting the three-dimensional map 1135 into a format corresponding to the receiving side. In addition, the format conversion unit 1119 can also reduce the amount of data by compressing or encoding the three-dimensional map 1135. In addition, when format conversion is not required, the format conversion unit 1119 can also omit the processing. In addition, the format conversion unit 1119 can control the amount of data sent according to the designation of the sending range.
[0445] The communication unit 1120 communicates with the client device 902 , and receives a data transmission request (a transmission request of a three-dimensional map) and the like from the client device 902 .
[0446] The transmission control unit 1121 exchanges information such as a corresponding format with the communication partner via the communication unit 1120 to establish communication.
[0447] The data transmission unit 1122 transmits the three-dimensional map 1031 to the client device 902. The three-dimensional map 1031 is data including a point cloud such as WLD or SWLD. The three-dimensional map 1031 may include either compressed data or uncompressed data.
[0448] Next, the workflow of the client device 902 will be described. Fig.31 1 is a flowchart showing the operation performed by the client device 902 when obtaining a three-dimensional map.
[0449] First, the client device 902 requests the server 901 to send a three-dimensional map (point cloud, etc.) (S1001). At this time, the client device 902 also sends the location information of the client device 902 obtained by GPS, etc., and accordingly, the server 901 can be requested to send a three-dimensional map related to the location information.
[0450] Next, the client device 902 receives the three-dimensional map from the server 901 (S1002). If the received three-dimensional map is compressed data, the client device 902 decodes the received three-dimensional map to generate a non-compressed three-dimensional map (S1003).
[0451] Next, the client device 902 creates three-dimensional data 1034 of the surroundings of the client device 902 based on the sensor information 1033 obtained from the plurality of sensors 1015 (S1004). Next, the client device 902 estimates its own position using the three-dimensional map 1032 received from the server 901 and the three-dimensional data 1034 created based on the sensor information 1033 (S1005).
[0452] Fig.32101 is a flowchart showing the operation of the client device 902 when transmitting sensor information. First, the client device 902 receives a request to transmit sensor information from the server 901 (S1011). The client device 902 that has received the transmission request transmits the sensor information 1037 to the server 901 (S1012). In addition, when the sensor information 1033 includes a plurality of information obtained by a plurality of sensors 1015, the client device 902 compresses each information in a compression method suitable for each information, thereby generating the sensor information 1037.
[0453] Next, the workflow of the server 901 is described. Fig.33 1 is a flowchart showing the operation of the server 901 when acquiring sensor information. First, the server 901 requests the client device 902 to send sensor information (S1021). Next, the server 901 receives the sensor information 1037 sent from the client device 902 in accordance with the request (S1022). Next, the server 901 uses the received sensor information 1037 to create three-dimensional data 1134 (S1023). Next, the server 901 reflects the created three-dimensional data 1134 on the three-dimensional map 1135 (S1024).
[0454] Fig.34 1031 is a flowchart showing the work of the server 901 when sending a three-dimensional map. First, the server 901 receives a request to send a three-dimensional map from the client device 902 (S1031). The server 901 that has received the request to send a three-dimensional map sends the three-dimensional map 1031 to the client device 902 (S1032). At this time, the server 901 can extract a three-dimensional map in the vicinity corresponding to the location information of the client device 902, and send the extracted three-dimensional map. Furthermore, the server 901 can compress the three-dimensional map composed of the point cloud, for example, using an octree compression method, and send the compressed three-dimensional map.
[0455] Hereinafter, modified examples of the present embodiment will be described.
[0456] The server 901 uses the sensor information 1037 received from the client device 902 to create three-dimensional data 1134 near the location of the client device 902. Next, the server 901 matches the created three-dimensional data 1134 with a three-dimensional map 1135 of the same area managed by the server 901, and calculates the difference between the three-dimensional data 1134 and the three-dimensional map 1135. When the difference is greater than a predetermined threshold, the server 901 determines that some abnormality has occurred around the client device 902. For example, when the ground subsidence occurs due to a natural disaster such as an earthquake, it can be considered that there will be a large difference between the three-dimensional map 1135 managed by the server 901 and the three-dimensional data 1134 created based on the sensor information 1037.
[0457] The sensor information 1037 may also include at least one of the type of sensor, the performance of the sensor, and the model of the sensor. In addition, a category ID corresponding to the performance of the sensor may be added to the sensor information 1037. For example, when the sensor information 1037 is information obtained by LiDAR, an identifier may be assigned in consideration of the performance of the sensor, for example, category 1 may be assigned to a sensor that can obtain information with an accuracy of several mm, category 2 may be assigned to a sensor that can obtain information with an accuracy of several cm, and category 3 may be assigned to a sensor that can obtain information with an accuracy of several m. In addition, the server 901 may estimate the performance information of the sensor from the model of the client device 902. For example, when the client device 902 is mounted on a vehicle, the server 901 may determine the specification information of the sensor according to the model of the vehicle. In this case, the server 901 may obtain the information of the model of the vehicle in advance, or include the information in the sensor information. In addition, the server 901 may switch the degree of correction of the three-dimensional data 1134 produced using the sensor information 1037 using the obtained sensor information 1037. For example, when the sensor performance is high accuracy (category 1), the server 901 does not perform correction on the three-dimensional data 1134. When the sensor performance is low accuracy (category 3), the server 901 applies correction suitable for the accuracy of the sensor to the three-dimensional data 1134. For example, the server 901 increases the degree (intensity) of correction as the accuracy of the sensor is lower.
[0458] The server 901 may also simultaneously send a request to send sensor information to multiple client devices 902 existing in a certain space. When the server 901 receives multiple sensor information from multiple client devices 902, it is not necessary to use all the sensor information to create the three-dimensional data 1134. For example, the server 901 may select the sensor information to be used according to the performance of the sensor. For example, when the server 901 updates the three-dimensional map 1135, it may select high-precision sensor information (category 1) from the received multiple sensor information and use the selected sensor information to create the three-dimensional data 1134.
[0459] The server 901 is not limited to servers such as traffic cloud monitoring, but can also be other client devices (car-mounted). Fig.35 The system configuration in this case is shown.
[0460] For example, the client device 902C sends a request to send sensor information to the client device 902A that is nearby, and obtains the sensor information from the client device 902A. Then, the client device 902C uses the obtained sensor information of the client device 902A to create three-dimensional data, and updates the three-dimensional map of the client device 902C. In this way, the client device 902C can utilize the performance of the client device 902C to generate a three-dimensional map of the space that can be obtained from the client device 902A. For example, when the performance of the client device 902C is high, this situation can be considered to occur.
[0461] In this case, the client device 902A that has provided the sensor information is given the right to obtain the high-precision three-dimensional map generated by the client device 902C. The client device 902A receives the high-precision three-dimensional map from the client device 902C in accordance with the right.
[0462] Furthermore, the client device 902C may send a request to send sensor information to multiple client devices 902 (client device 902A and client device 902B) in the vicinity. If the sensor of the client device 902A or the client device 902B is high-performance, the client device 902C can use the sensor information obtained by the high-performance sensor to create three-dimensional data.
[0463] Fig.36 1 is a block diagram showing the functional configuration of the server 901 and the client device 902. The server 901 includes, for example, a 3D map compression / decoding processing unit 1201 that compresses and decodes a 3D map and a sensor information compression / decoding processing unit 1202 that compresses and decodes sensor information.
[0464] The client device 902 includes: a three-dimensional map decoding processing unit 1211, and a sensor information compression processing unit 1212. The three-dimensional map decoding processing unit 1211 receives the coded data of the compressed three-dimensional map, decodes the coded data and obtains the three-dimensional map. The sensor information compression processing unit 1212 does not compress the three-dimensional data produced by the obtained sensor information, but compresses the sensor information itself, and sends the compressed coded data of the sensor information to the server 901. According to this structure, the client device 902 can keep the processing unit (device or LSI) for decoding the three-dimensional map (point cloud, etc.) inside, without having to keep the processing unit for compressing the three-dimensional data of the three-dimensional map (point cloud, etc.) inside. In this way, the cost and power consumption of the client device 902 can be suppressed.
[0465] As described above, the client device 902 involved in this embodiment is mounted on the mobile body, and generates three-dimensional data 1034 of the surroundings of the mobile body based on the sensor information 1033 indicating the surrounding conditions of the mobile body obtained by the sensor 1015 mounted on the mobile body. The client device 902 estimates the own position of the mobile body using the generated three-dimensional data 1034. The client device 902 transmits the obtained sensor information 1033 to the server 901 or another mobile body 902.
[0466] Based on this, the client device 902 transmits the sensor information 1033 to the server 901 or the like. In this way, the amount of data to be transmitted may be reduced compared to the case of transmitting three-dimensional data. Furthermore, since it is not necessary to perform processing such as compression or encoding of three-dimensional data on the client device 902, the amount of processing on the client device 902 can be reduced. Therefore, the client device 902 can reduce the amount of data to be transmitted or simplify the configuration of the device.
[0467] Furthermore, the client device 902 further sends a request to send a three-dimensional map to the server 901, and receives the three-dimensional map 1031 from the server 901. The client device 902 estimates its own position using the three-dimensional data 1034 and the three-dimensional map 1032.
[0468] Furthermore, the sensor information 1033 includes at least one of information obtained by a laser sensor, a brightness image (visible light image), an infrared image, a depth image, position information of the sensor, and speed information of the sensor.
[0469] Also, the sensor information 1033 includes information showing the performance of the sensor.
[0470] Furthermore, the client device 902 encodes or compresses the sensor information 1033, and transmits the encoded or compressed sensor information 1037 to the server 901 or another mobile body 902. Thus, the client device 902 can reduce the amount of data to be transmitted.
[0471] For example, the client device 902 includes a processor and a memory, and the processor performs the above-mentioned processing using the memory.
[0472] Furthermore, the server 901 according to the present embodiment can communicate with the client device 902 mounted on the mobile body, and receive sensor information 1037 indicating the surrounding conditions of the mobile body obtained by the sensor 1015 mounted on the mobile body from the client device 902. The server 901 generates three-dimensional data 1134 of the surroundings of the mobile body based on the received sensor information 1037.
[0473] Accordingly, the server 901 uses the sensor information 1037 sent from the client device 902 to create the three-dimensional data 1134. In this way, compared with the case where the client device 902 sends the three-dimensional data, it is possible to reduce the amount of data to be sent. In addition, since it is not necessary to perform processing such as compression or encoding of the three-dimensional data on the client device 902, the processing amount of the client device 902 can be reduced. In this way, the server 901 can reduce the amount of data to be transmitted or simplify the configuration of the device.
[0474] Furthermore, the server 901 further transmits a request for transmitting sensor information to the client device 902 .
[0475] Furthermore, the server 901 further updates the three-dimensional map 1135 using the created three-dimensional data 1134 , and transmits the three-dimensional map 1135 to the client device 902 in response to a transmission request for the three-dimensional map 1135 from the client device 902 .
[0476] Furthermore, the sensor information 1037 includes at least one of information obtained by a laser sensor, a brightness image (visible light image), an infrared image, a depth image, position information of the sensor, and speed information of the sensor.
[0477] Also, the sensor information 1037 includes information showing the performance of the sensor.
[0478] Furthermore, the server 901 further calibrates the three-dimensional data according to the performance of the sensor. Accordingly, the three-dimensional data production method can improve the quality of the three-dimensional data.
[0479] Furthermore, when receiving sensor information, the server 901 receives a plurality of sensor information 1037 from a plurality of client devices 902, and selects the sensor information 1037 used in the production of the three-dimensional data 1134 based on a plurality of information indicating the performance of the sensors included in the plurality of sensor information 1037. In this way, the server 901 can improve the quality of the three-dimensional data 1134.
[0480] Furthermore, the server 901 decodes or decompresses the received sensor information 1137, and creates three-dimensional data 1134 based on the decoded or decompressed sensor information 1132. In this way, the server 901 can reduce the amount of data to be transmitted.
[0481] For example, the server 901 includes a processor and a memory, and the processor performs the above-mentioned processing using the memory.
[0482] (Implementation 7)
[0483] In this embodiment, a method for encoding and a method for decoding three-dimensional data using an inter-frame prediction process will be described.
[0484] Fig.37 3D data encoding device 1300 according to the present embodiment is a block diagram. The 3D data encoding device 1300 generates a coded bit stream (hereinafter also simply referred to as a bit stream) as a coded signal by encoding 3D data. Fig.37 As shown, the three-dimensional data encoding device 1300 includes: a segmentation unit 1301, a subtraction unit 1302, a transformation unit 1303, a quantization unit 1304, an inverse quantization unit 1305, an inverse transformation unit 1306, an addition unit 1307, a reference volume memory 1308, an intra-frame prediction unit 1309, a reference space memory 1310, an inter-frame prediction unit 1311, a prediction control unit 1312, and an entropy coding unit 1313.
[0485] The segmentation unit 1301 segments each space (SPC) included in the three-dimensional data into a plurality of volumes (VLM) as coding units. Furthermore, the segmentation unit 1301 performs octree representation (Octree) on the voxels in each volume. In addition, the segmentation unit 1301 may make the space and the volume the same size and perform octree representation on the space. Furthermore, the segmentation unit 1301 may also attach information required for octree representation (depth information, etc.) to the header of the bitstream, etc.
[0486] The subtraction unit 1302 calculates the difference between the volume (coding target volume) output from the division unit 1301 and the prediction volume generated by intra prediction or inter prediction described later, and outputs the calculated difference as a prediction residual to the transformation unit 1303 . Fig.38An example of calculating the prediction residual is shown. The bit strings of the encoding target volume and the prediction volume shown here are, for example, position information indicating the positions of three-dimensional points (for example, point clouds) included in the volume.
[0487] The following describes the octree representation and the scanning order of voxels. The volume is transformed into an octree structure (octreeization) and then encoded. The octree structure consists of nodes and leaf nodes. Each node has 8 nodes or leaf nodes, and each leaf node has voxel (VXL) information. Fig.39 An example of the configuration of a volume including a plurality of voxels is shown. Fig.40 Shows the Fig.39 The volume shown is transformed into an example of an octree structure. Here, Fig.40 The leaf nodes 1, 2, and 3 in the leaf nodes shown represent Fig.39 The voxels VXL1, VXL2, and VXL3 shown express the VXL (hereinafter referred to as effective VXL) including the point group.
[0488] The octree is represented by a binary sequence of 0 and 1. For example, when a node or valid VXL is set to a value of 1 and the others are set to a value of 0, each node and leaf node is assigned Fig.40 The binary sequence shown in FIG. Then, the binary sequence is scanned in a width-first or depth-first scanning order. For example, in the case of a width-first scan, Fig.41 The binary sequence shown in A. When the depth-first scan is performed, we get Fig.41 The binary sequence shown in B. The binary sequence obtained by this scanning is encoded by entropy coding, so that the amount of information is reduced.
[0489] Next, the depth information in the octree representation is explained. The depth in the octree representation is used to control the granularity of the point cloud information contained in the volume. If the depth is set to a large value, the point cloud information can be reproduced at a finer level, but the amount of data used to represent the nodes and leaf nodes will increase. On the contrary, if the depth is set to a small value, although the amount of data can be reduced, multiple point cloud information at different positions and colors will be regarded as the same position and the same color, so the original information of the point cloud information will be lost.
[0490] For example, Fig.42 Shows the Fig.40 The octree with depth = 2 is shown as an example of expressing an octree with depth = 1. Fig.42 The octree shown is Fig.40 The octree shown has a small amount of data. Fig.42 The octree shown is similar to Fig.42Compared with the octree shown, the number of bits after binary serialization is less. Fig.40 The leaf nodes 1 and 2 shown in the figure become Fig.41 The leaf node 1 shown is shown. That is, Fig.40 The leaf node 1 and the leaf node 2 shown are information on different locations.
[0491] Fig.43 Shown with Fig.42 The volume corresponding to the octree shown. Fig.39 The VXL1 and VXL2 shown are Fig.43 In this case, the three-dimensional data encoding device 1300 corresponds to VXL12. Fig.39 The color information of VXL1 and VXL2 shown in the figure generates Fig.43 For example, the three-dimensional data encoding device 1300 calculates the color information of VXL1 and VXL2 as the color information of VXL12 using the average value, median value, or weighted average value. In this way, the three-dimensional data encoding device 1300 can control the reduction of the data amount by changing the depth of the octree.
[0492] The three-dimensional data encoding device 1300 may also use any one of the world space units, space units, and volume units to set the depth information of the octree. In addition, at this time, the three-dimensional data encoding device 1300 may also attach the depth information to the header information of the world space, the header information of the space, or the header information of the volume. In addition, the same value may be used as the depth information in all world spaces, spaces, and volumes at different times. In this case, the three-dimensional data encoding device 1300 may also attach the depth information to the header information that manages the world space at all times.
[0493] When the voxel contains color information, the transformation unit 1303 applies a frequency transformation such as an orthogonal transformation to the prediction residual of the color information of the voxels in the volume. For example, the transformation unit 1303 scans the prediction residual in a certain scanning order to produce a one-dimensional arrangement. Thereafter, the transformation unit 1303 transforms the one-dimensional arrangement into the frequency domain by applying a one-dimensional orthogonal transformation to the produced one-dimensional arrangement. Accordingly, when the value of the prediction residual in the volume is close, the value of the frequency component of the low frequency band becomes larger, and the value of the frequency component of the high frequency band becomes smaller. Therefore, the quantization unit 1304 can more effectively reduce the amount of coding.
[0494] Furthermore, the transformation unit 1303 may use orthogonal transformation of two or more dimensions instead of one-dimensional orthogonal transformation. For example, the transformation unit 1303 maps the prediction residual into a two-dimensional arrangement in a certain scanning order, and applies a two-dimensional orthogonal transformation to the obtained two-dimensional arrangement. Furthermore, the transformation unit 1303 may select the orthogonal transformation method to be used from a plurality of orthogonal transformation methods. In this case, the three-dimensional data encoding device 1300 attaches information indicating which orthogonal transformation method is used to the bit stream. Furthermore, the transformation unit 1303 may select the orthogonal transformation method to be used from a plurality of orthogonal transformation methods with different dimensions. In this case, the three-dimensional data encoding device 1300 attaches information indicating which dimensional orthogonal transformation method is used to the bit stream.
[0495] For example, the transformation unit 1303 matches the scanning order of the prediction residual with the scanning order (width-first or depth-first, etc.) in the octree within the volume. Accordingly, since it is not necessary to attach information showing the scanning order of the prediction residual to the bitstream, the additional overhead can be reduced. In addition, the transformation unit 1303 may also apply a scanning order different from the scanning order of the octree. In this case, the three-dimensional data encoding device 1300 attaches information showing the scanning order of the prediction residual to the bitstream. Accordingly, the three-dimensional data encoding device 1300 can efficiently encode the prediction residual. In addition, the three-dimensional data encoding device 1300 may attach information (flag, etc.) indicating whether the scanning order of the octree is applicable to the bitstream, and when the scanning order of the octree is not applicable, the information showing the scanning order of the prediction residual is attached to the bitstream.
[0496] The conversion unit 1303 may convert not only the prediction residual of the color information but also other attribute information of the voxel. For example, the conversion unit 1303 may convert and encode information such as reflectivity obtained when a point cloud is obtained by LiDAR or the like.
[0497] The transform unit 1303 may skip the process when the space does not have attribute information such as color information. Furthermore, the three-dimensional data encoding device 1300 may add information (flag) indicating whether to skip the process of the transform unit 1303 to the bit stream.
[0498] The quantization unit 1304 quantizes the frequency components of the prediction residual generated by the transformation unit 1303 using the quantization control parameters to generate quantization coefficients. The amount of information is reduced accordingly. The generated quantization coefficients are output to the entropy coding unit 1313. The quantization unit 1304 can control the quantization control parameters according to world space units, space units, or volume units. At this time, the three-dimensional data encoding device 1300 attaches the quantization control parameters to the respective header information, etc. In addition, the quantization unit 1304 can also change the weight according to the frequency components of each prediction residual to perform quantization control. For example, the quantization unit 1304 can perform fine quantization on the low-frequency components and coarse quantization on the high-frequency components. In this case, the three-dimensional data encoding device 1300 can attach parameters representing the weights of each frequency component to the header.
[0499] The quantization unit 1304 may skip the process when the space does not have attribute information such as color information. In addition, the three-dimensional data encoding device 1300 may add information (flag) indicating whether the process of the quantization unit 1304 is skipped to the bit stream.
[0500] The inverse quantization unit 1305 inversely quantizes the quantization coefficients generated by the quantization unit 1304 using the quantization control parameters, thereby generating inverse quantization coefficients of the prediction residual, and outputs the generated inverse quantization coefficients to the inverse transformation unit 1306 .
[0501] The inverse transform unit 1306 generates an inverse transform applied prediction residual by applying inverse transform to the inverse quantized coefficients generated by the inverse quantization unit 1305. Since the inverse transform applied prediction residual is a prediction residual generated after quantization, it may not be completely consistent with the prediction residual output by the transform unit 1303.
[0502] The adding unit 1307 adds the prediction residual after inverse transformation generated by the inverse transform unit 1306 and the prediction volume generated by intra-frame prediction or inter-frame prediction described later, which is used to generate the prediction residual before quantization, to generate a reconstructed volume. The reconstructed volume is stored in the reference volume memory 1308 or the reference space memory 1310.
[0503] The intra prediction unit 1309 generates a predicted volume of the encoding target volume using the attribute information of the adjacent volume stored in the reference volume memory 1308. The attribute information includes the color information or reflectance of the voxel. The intra prediction unit 1309 generates a predicted value of the color information or reflectance of the encoding target volume.
[0504] Fig.44 1309 is a diagram for explaining the operation of the intra prediction unit 1309. For example, Fig.44As shown, the intra-frame prediction unit 1309 generates a predicted volume of the encoding target volume (volume idx=3) based on the adjacent volume (volume idx=0). Here, volume idx is identifier information added to the volume in the space, and different values are assigned to each volume. The order of assigning volume idx can be the same as the encoding order or different from the encoding order. For example, as Fig.44 The intra prediction unit 1309 uses the average value of the color information of the voxels contained in the volume idx=0 which is the adjacent volume as the predicted value of the color information of the encoding object volume shown. In this case, the prediction residual is generated by subtracting the predicted value of the color information from the color information of each voxel contained in the encoding object volume. The processing after the transformation unit 1303 is performed on the prediction residual. And, in this case, the three-dimensional data encoding device 1300 adds the adjacent volume information and the prediction mode information to the bit stream. Here, the adjacent volume information is information showing the adjacent volume used in the prediction, for example, the volume idx showing the adjacent volume used in the prediction. And the prediction mode information shows the mode used in the generation of the prediction volume. The mode refers to, for example, an average value mode for generating a prediction value based on the average value of the voxels in the adjacent volume, or an intermediate value mode for generating a prediction value based on the intermediate value of the voxels in the adjacent volume.
[0505] The intra-frame prediction unit 1309 may also generate a prediction volume based on a plurality of adjacent volumes. Fig.44 In the illustrated configuration, the intra prediction unit 1309 generates prediction volume 0 based on volume idx=0, and generates prediction volume 1 based on volume idx=1. Then, the intra prediction unit 1309 generates the final prediction volume by averaging the prediction volume 0 and the prediction volume 1. In this case, the three-dimensional data encoding device 1300 may also attach multiple volume idxs of the multiple volumes used in generating the prediction volume to the bitstream.
[0506] Fig.45 The inter-frame prediction process involved in this embodiment is shown in the mode. The inter-frame prediction unit 1311 uses the coded space at a different time T_LX to encode (inter-frame prediction) for the space (SPC) at a certain time T_Cur. In this case, the inter-frame prediction unit 1311 applies rotation and translation processing to the coded space at different time T_LX to perform encoding processing.
[0507] Furthermore, the three-dimensional data encoding device 1300 adds RT information related to the rotation and translation processing of the space at the different time T_LX to the bitstream. The different time T_LX is, for example, the time T_L0 before the certain time T_Cur. At this time, the three-dimensional data encoding device 1300 may also add RT information RT_L0 related to the rotation and translation processing of the space at the time T_L0 to the bitstream.
[0508] Alternatively, the different time T_LX is, for example, time T_L1 after the certain time T_Cur. In this case, the three-dimensional data encoding apparatus 1300 may add RT information RT_L1 about the spatial rotation and translation processing applied at time T_L1 to the bit stream.
[0509] Alternatively, the inter prediction unit 1311 performs encoding by referring to spaces at different times T_L0 and T_L1 (bi-prediction). In this case, the three-dimensional data encoding device 1300 may add both RT information RT_L0 and RT_L1 about rotation and translation applied to the space to the bitstream.
[0510] In addition, although T_L0 is set as the time before T_Cur and T_L1 is set as the time after T_Cur, it is not limited to this. For example, T_L0 and T_L1 can both be the time before T_Cur. Or, T_L0 and T_L1 can both be the time after T_Cur.
[0511] Furthermore, the three-dimensional data encoding device 1300 may add RT information related to the rotation and translation applied to each space to the bitstream when encoding with reference to multiple spaces at different times. For example, the three-dimensional data encoding device 1300 manages the multiple encoded spaces referred to by two reference lists (L0 list and L1 list). When the first reference space in the L0 list is set to L0R0, the second reference space in the L0 list is set to L0R1, the first reference space in the L1 list is set to L1R0, and the second reference space in the L1 list is set to L1R1, the three-dimensional data encoding device 1300 adds RT information RT_L0R0 of L0R0, RT information RT_L0R1 of L0R1, RT information RT_L1R0 of L1R0, and RT information RT_L1R1 of L1R1 to the bitstream. For example, the three-dimensional data encoding device 1300 adds these RT information to the header of the bitstream.
[0512] Furthermore, the three-dimensional data encoding device 1300 may determine whether rotation and translation are applicable to each reference space when encoding with reference to a plurality of reference spaces at different times. In this case, the three-dimensional data encoding device 1300 may attach information (RT applicable flag, etc.) indicating whether rotation and translation are applicable to each reference space to header information of the bitstream, etc. For example, the three-dimensional data encoding device 1300 calculates RT information and ICP error value using an ICP (Interactive Closest Point) algorithm according to the encoding object space and each reference space to be referenced. When the ICP error value is below a predetermined certain value, the three-dimensional data encoding device 1300 determines that rotation and translation are not required and sets the RT applicable flag to OFF (invalid). In addition, when the ICP error value is larger than the above-mentioned certain value, the three-dimensional data encoding device 1300 sets the RT applicable flag to ON (valid) and attaches the RT information to the bitstream.
[0513] Fig.46 An example of a syntax in which RT information and an RT applicable flag are attached to a header is shown. In addition, the number of bits allocated to each syntax can be determined based on the range that the syntax can take. For example, when the number of reference spaces included in the reference list L0 is 8, 3 bits can be allocated in MaxRefSpc_l0. The number of allocated bits can be changed according to the values that each syntax can take, or the number of allocated bits can be fixed without being affected by the possible values. In the case of fixing the number of allocated bits, the three-dimensional data encoding device 1300 can attach the fixed number of bits to other header information.
[0514] Here, Fig.46 MaxRefSpc_10 shown shows the number of reference spaces included in reference list L0. RT_flag_10[i] is the RT application flag for reference space i in reference list L0. When RT_flag_10[i] is 1, rotation and translation are applied to reference space i. When RT_flag_10[i] is 0, rotation and translation are not applied to reference space i.
[0515] R_l0[i] and T_l0[i] are RT information of reference space i in reference list L0. R_l0[i] is rotation information of reference space i in reference list L0. The rotation information indicates the content of the applied rotation processing, such as a rotation matrix or quaternion. T_l0[i] is translation information of reference space i in reference list L0. The translation information indicates the content of the applied translation processing, such as a translation vector.
[0516] MaxRefSpc_l1 indicates the number of reference spaces included in reference list L1. RT_flag_l1[i] is an RT application flag for reference space i in reference list L1. When RT_flag_l1[i] is 1, rotation and translation are applied to reference space i. When RT_flag_l1[i] is 0, rotation and translation are not applied to reference space i.
[0517] R_l1[i] and T_l1[i] are RT information of reference space i in reference list L1. R_l1[i] is rotation information of reference space i in reference list L1. The rotation information indicates the content of the applied rotation processing, such as a rotation matrix or quaternion. T_l1[i] is translation information of reference space i in reference list L1. The translation information indicates the content of the applied translation processing, such as a translation vector.
[0518] The inter-frame prediction unit 1311 generates a predicted volume of the encoding target volume using information of the encoded reference space stored in the reference space memory 1310. As described above, before generating the predicted volume of the encoding target volume, the inter-frame prediction unit 1311 uses the ICP (Interactive Closest Point) algorithm to obtain RT information in the encoding target space and the reference space in order to make the positional relationship between the encoding target space and the entire reference space close. Then, the inter-frame prediction unit 1311 applies rotation and translation processing to the reference space using the obtained RT information, thereby obtaining the reference space B. Thereafter, the inter-frame prediction unit 1311 generates a predicted volume of the encoding target volume in the encoding target space using information in the reference space B. Here, the three-dimensional data encoding device 1300 adds the RT information used to obtain the reference space B to the header information of the encoding target space, etc.
[0519] In this way, the inter-frame prediction unit 1311 applies rotation and translation processing to the reference space, so as to make the overall positional relationship between the encoding object space and the reference space close, and then uses the information of the reference space to generate a prediction volume, thereby improving the accuracy of the prediction volume. In addition, since the prediction residual can be suppressed, the amount of encoding can be reduced. In addition, although an example of performing ICP using the encoding object space and the reference space is shown here, it is not limited to this. For example, in order to reduce the amount of processing, the inter-frame prediction unit 1311 can also perform ICP using at least one of the encoding object space from which the number of voxels or point clouds is extracted, and the reference space from which the number of voxels or point clouds is extracted, so as to obtain RT information.
[0520] Furthermore, when the ICP error value obtained from the ICP result is smaller than a predetermined first threshold value, that is, when the positional relationship between the encoding object space and the reference space is close, the inter-frame prediction unit 1311 may determine that rotation and translation processing are not required, and does not perform rotation and translation. In this case, the three-dimensional data encoding device 1300 may not add RT information to the bitstream, thereby suppressing additional overhead.
[0521] Furthermore, when the ICP error value is greater than a predetermined second threshold, the inter-frame prediction unit 1311 determines that the shape change in space is large, and intra-frame prediction can be applied to all volumes of the encoding object space. Hereinafter, the space to which intra-frame prediction is applied is referred to as intra-frame space. Furthermore, the second threshold is a value greater than the above-mentioned first threshold. Furthermore, it is not limited to ICP, and any method can be applied as long as the method of obtaining RT information from two voxel sets or two point cloud sets.
[0522] Furthermore, when the three-dimensional data contains attribute information such as shape or color, the inter-frame prediction unit 1311 searches, for example, a volume in the reference space that is closest to the shape or color attribute information of the encoding target volume as a prediction volume of the encoding target volume in the encoding target space. Furthermore, the reference space is, for example, a reference space after the above-mentioned rotation and translation processing. The inter-frame prediction unit 1311 generates a prediction volume based on the volume (reference volume) obtained by the search. Fig.47 is a diagram for explaining the generation of the prediction volume. Fig.47 When the encoding target volume (volume idx=0) shown in the figure is encoded by using inter-frame prediction, the reference volumes in the reference space are scanned in sequence while searching for the volume with the smallest prediction residual, that is, the difference between the encoding target volume and the reference volume. The inter-frame prediction unit 1311 selects the volume with the smallest prediction residual as the prediction volume. The prediction residual between the encoding target volume and the prediction volume is encoded by the processing after the transformation unit 1303. Here, the prediction residual refers to the difference between the attribute information of the encoding target volume and the attribute information of the prediction volume. In addition, the three-dimensional data encoding device 1300 adds the volume idx of the reference volume in the reference space referred to as the prediction volume to the header of the bit stream.
[0523] exist Fig.47 In the example shown, the reference volume idx=4 of the reference space L0R0 is selected as the prediction volume of the encoding target volume. Then, the prediction residual between the encoding target volume and the reference volume and the reference volume idx=4 are encoded and added to the bit stream.
[0524] In addition, although the description here is made by taking the generation of the predicted volume of the attribute information as an example, the same processing can be performed on the predicted volume of the position information.
[0525] The prediction control unit 1312 controls whether to use intra-frame prediction or inter-frame prediction to encode the encoding target volume. Here, a mode including intra-frame prediction and inter-frame prediction is referred to as a prediction mode. For example, the prediction control unit 1312 calculates the prediction residual when the encoding target volume is predicted by intra-frame prediction and the prediction residual when it is predicted by inter-frame prediction as an evaluation value, and selects the prediction mode with the smaller evaluation value. Alternatively, the prediction control unit 1312 may apply orthogonal transformation, quantization, and entropy coding to the prediction residual of intra-frame prediction and the prediction residual of inter-frame prediction, respectively, to calculate the actual amount of code, and select the prediction mode using the calculated amount of code as the evaluation value. Furthermore, overhead information other than the prediction residual (reference volume idx information, etc.) may be added to the evaluation value. Furthermore, the prediction control unit 1312 may also generally select intra-frame prediction when the encoding target space is predetermined to be encoded in the intra-frame space.
[0526] The entropy coding unit 1313 generates a coded signal (coded bit stream) by performing variable length coding on the quantized coefficients input from the quantization unit 1304. Specifically, the entropy coding unit 1313 binarizes the quantized coefficients and performs arithmetic coding on the obtained binary signal, for example.
[0527] Next, a three-dimensional data decoding device that decodes the encoded signal generated by the three-dimensional data encoding device 1300 will be described. Fig.48 14 is a block diagram of a three-dimensional data decoding device 1400 according to the present embodiment. The three-dimensional data decoding device 1400 includes an entropy decoding unit 1401, an inverse quantization unit 1402, an inverse transformation unit 1403, an addition unit 1404, a reference volume memory 1405, an intra-frame prediction unit 1406, a reference space memory 1407, an inter-frame prediction unit 1408, and a prediction control unit 1409.
[0528] The entropy decoding unit 1401 performs variable length decoding on the coded signal (coded bit stream). For example, the entropy decoding unit 1401 performs arithmetic decoding on the coded signal to generate a binary signal, and generates a quantized coefficient based on the generated binary signal.
[0529] The inverse quantization unit 1402 performs inverse quantization on the quantized coefficients input from the entropy decoding unit 1401 using a quantization parameter added to a bit stream or the like, thereby generating inverse quantized coefficients.
[0530] The inverse transform unit 1403 generates a prediction residual by performing an inverse transform on the inverse quantized coefficient input from the inverse quantization unit 1402. For example, the inverse transform unit 1403 generates a prediction residual by performing an inverse orthogonal transform on the inverse quantized coefficient based on information added to the bit stream.
[0531] The adder 1404 adds the prediction residual generated by the inverse transform unit 1403 to the prediction volume generated by intra prediction or inter prediction to generate a reconstructed volume. The reconstructed volume is output as decoded three-dimensional data and stored in the reference volume memory 1405 or the reference space memory 1407.
[0532] The intra prediction unit 1406 generates a prediction volume by intra prediction using the reference volume in the reference volume memory 1405 and the information added to the bitstream. Specifically, the intra prediction unit 1406 obtains the prediction mode information and the adjacent volume information (e.g., volume idx) added to the bitstream, and generates a prediction volume using the adjacent volume indicated by the adjacent volume information in the mode indicated by the prediction mode information. In addition, the details of these processes are the same as the processes of the intra prediction unit 1309 described above, except that the information added to the bitstream is used.
[0533] The inter-frame prediction unit 1408 generates a prediction volume through inter-frame prediction using the reference space in the reference space memory 1407 and the information attached to the bitstream. Specifically, the inter-frame prediction unit 1408 uses the RT information of each reference space attached to the bitstream, applies rotation and translation processing to the reference space, and generates a prediction volume using the reference space after application. In addition, when the RT application flag of each reference space exists in the bitstream, the inter-frame prediction unit 1408 applies rotation and translation processing to the reference space according to the RT application flag. In addition, the details of the above-mentioned processing are the same as the processing of the above-mentioned inter-frame prediction unit 1311, except that the information attached to the bitstream is used.
[0534] Whether to decode the decoding target volume by intra prediction or inter prediction is controlled by the prediction control unit 1409. For example, the prediction control unit 1409 selects intra prediction or inter prediction according to information indicating the prediction mode to be used, which is added to the bit stream. In addition, when it is predetermined that the decoding target space is decoded as the intra space, the prediction control unit 1409 may normally select the intra prediction.
[0535] The following is a description of a variation of the present embodiment. In the present embodiment, although the application of rotation and translation in spatial units is described as an example, rotation and translation in finer units may also be applied. For example, the three-dimensional data encoding device 1300 may divide the space into subspaces and apply rotation and translation in subspace units. In this case, the three-dimensional data encoding device 1300 generates RT information according to each subspace, and attaches the generated RT information to the header of the bit stream. In addition, the three-dimensional data encoding device 1300 may use volume units as encoding units to apply rotation and translation. In this case, the three-dimensional data encoding device 1300 generates RT information in encoding volume units, and attaches the generated RT information to the header of the bit stream. Moreover, the above may be combined. That is, the three-dimensional data encoding device 1300 may apply rotation and translation in finer units after applying rotation and translation in large units. For example, the three-dimensional data encoding device 1300 may apply rotation and translation in spatial units, and apply different rotations and translations to each of the multiple volumes contained in the obtained space.
[0536] Furthermore, although the present embodiment is described by taking the application of rotation and translation to the reference space as an example, the present invention is not limited thereto. For example, the three-dimensional data encoding device 1300 may apply scaling processing to change the size of the three-dimensional data. Furthermore, the three-dimensional data encoding device 1300 may also apply any one or two of rotation, translation, and scaling. Furthermore, as described above, when the processing is applied in different units in multiple stages, the type of processing applied in each unit may be different. For example, rotation and translation may be applied in the spatial unit, and translation may be applied in the volume unit.
[0537] Note that these modifications are also applicable to the three-dimensional data decoding device 1400 .
[0538] As described above, the three-dimensional data encoding device 1300 according to this embodiment performs the following processing. Fig.48 3D data encoding apparatus 1300 performs an inter-frame prediction process.
[0539] First, the three-dimensional data encoding device 1300 generates predicted position information (e.g., predicted volume) using the position information of three-dimensional points included in the target three-dimensional data (e.g., encoding target space) and the reference three-dimensional data (e.g., reference space) at different times (S1301). Specifically, the three-dimensional data encoding device 1300 generates the predicted position information by applying rotation and translation processing to the position information of the three-dimensional points included in the reference three-dimensional data.
[0540] In addition, the three-dimensional data encoding device 1300 performs rotation and translation processing with a first unit (e.g., space), and generates predicted position information with a second unit (e.g., volume) that is smaller than the first unit. For example, the three-dimensional data encoding device 1300 searches for a volume in which the difference between the encoding object volume and the position information contained in the encoding object space is the smallest from among a plurality of volumes contained in the reference space after the rotation and translation processing, and uses the obtained volume as the predicted volume. In addition, the three-dimensional data encoding device 1300 may perform the rotation and translation processing and the generation of the predicted position information with the same unit.
[0541] Furthermore, the three-dimensional data encoding device 1300 may apply a first rotation and translation process to the position information of the three-dimensional points contained in the reference three-dimensional data using a first unit (e.g., space), and may apply a second rotation and translation process to the position information of the three-dimensional points obtained by the first rotation and translation process using a second unit (e.g., volume) that is smaller than the first unit, thereby generating predicted position information.
[0542] Here, the position information of the three-dimensional point and the predicted position information are as follows: Fig.41 As shown, the octree structure is used for representation. For example, the position information and predicted position information of the three-dimensional point are represented in a scanning order that prioritizes the width among the depth and width in the octree structure. Alternatively, the position information and predicted position information of the three-dimensional point are represented in a scanning order that prioritizes the depth and width among the depth and width in the octree structure.
[0543] And, if Fig.46 As shown, the three-dimensional data encoding device 1300 encodes the RT applicable flag indicating whether the rotation and translation processing is applied to the position information of the three-dimensional point contained in the reference three-dimensional data. That is, the three-dimensional data encoding device 1300 generates a coded signal (coded bit stream) including the RT applicable flag. In addition, the three-dimensional data encoding device 1300 encodes the RT information indicating the content of the rotation and translation processing. That is, the three-dimensional data encoding device 1300 generates a coded signal (coded bit stream) including the RT information. Alternatively, the three-dimensional data encoding device 1300 may encode the RT information when the RT applicable flag indicates that the rotation and translation processing is applicable, and may not encode the RT information when the RT applicable flag indicates that the rotation and translation processing is not applicable.
[0544] The three-dimensional data includes, for example, position information of three-dimensional points and attribute information (color information, etc.) of each three-dimensional point. The three-dimensional data encoding device 1300 generates predicted attribute information by using the attribute information of the three-dimensional points included in the reference three-dimensional data (S1302).
[0545] Next, the three-dimensional data encoding device 1300 uses the predicted position information to encode the position information of the three-dimensional points included in the object three-dimensional data. Fig.38 As shown, differential position information which is a difference between the position information of the three-dimensional point included in the target three-dimensional data and the predicted position information is calculated (S1303).
[0546] Furthermore, the three-dimensional data encoding device 1300 uses the predicted attribute information to encode the attribute information of the three-dimensional points included in the target three-dimensional data. For example, the three-dimensional data encoding device 1300 calculates the difference between the attribute information of the three-dimensional points included in the target three-dimensional data and the predicted attribute information, that is, differential attribute information (S1304). Next, the three-dimensional data encoding device 1300 transforms and quantizes the calculated differential attribute information (S1305).
[0547] Finally, the 3D data encoding device 1300 encodes (for example, entropy encoding) the differential position information and the quantized differential attribute information (S1306). That is, the 3D data encoding device 1300 generates an encoded signal (encoded bit stream) including the differential position information and the differential attribute information.
[0548] In addition, when the three-dimensional data does not include attribute information, the three-dimensional data encoding device 1300 may not perform steps S1302, S1304, and S1305. In addition, the three-dimensional data encoding device 1300 may only encode the position information of the three-dimensional point or encode the attribute information of the three-dimensional point.
[0549] and, Fig.49 The processing sequence shown is only an example and is not limited thereto. For example, since the processing of location information (S1301, S1303) and the processing of attribute information (S1302, S1304, S1305) are independent of each other, they can be executed in any order, or some of them can be processed in parallel.
[0550] As described above, in this embodiment, the three-dimensional data encoding device 1300 generates predicted position information using the position information of the three-dimensional points included in the target three-dimensional data and the reference three-dimensional data at different times, and encodes the difference between the position information of the three-dimensional points included in the target three-dimensional data and the predicted position information, that is, the differential position information. Accordingly, since the amount of data of the encoded signal can be reduced, the encoding efficiency can be improved.
[0551] Furthermore, in this embodiment, the three-dimensional data encoding device 1300 generates predicted attribute information by using attribute information of three-dimensional points included in the reference three-dimensional data, and encodes the difference between the attribute information of the three-dimensional points included in the target three-dimensional data and the predicted attribute information, that is, the differential attribute information. Accordingly, since the data amount of the encoded signal can be reduced, the encoding efficiency can be improved.
[0552] For example, the three-dimensional data encoding device 1300 includes a processor and a memory, and the processor performs the above-mentioned processing using the memory.
[0553] Fig.48 3D data decoding apparatus 1400 performs an inter-frame prediction process.
[0554] First, the three-dimensional data decoding apparatus 1400 decodes (eg, performs entropy decoding) the difference position information and the difference attribute information according to the coded signal (coded bit stream) ( S1401 ).
[0555] Furthermore, the three-dimensional data decoding device 1400 decodes the RT application flag indicating whether the rotation and translation processing is applied to the position information of the three-dimensional point included in the reference three-dimensional data based on the coded signal. Furthermore, the three-dimensional data decoding device 1400 decodes the RT information indicating the content of the rotation and translation processing. In addition, the three-dimensional data decoding device 1400 decodes the RT information when the RT application flag indicates that the rotation and translation processing is applied, and does not decode the RT information when the RT application flag indicates that the rotation and translation processing is not applied.
[0556] Next, the three-dimensional data decoding apparatus 1400 performs inverse quantization and inverse transformation on the decoded differential attribute information ( S1402 ).
[0557] Next, the three-dimensional data decoding device 1400 generates predicted position information (e.g., predicted volume) using the position information of the three-dimensional points included in the target three-dimensional data (e.g., decoding target space) and the reference three-dimensional data (e.g., reference space) at different times (S1403). Specifically, the three-dimensional data decoding device 1400 generates the predicted position information by applying rotation and translation processing to the position information of the three-dimensional points included in the reference three-dimensional data.
[0558] More specifically, when the RT application flag indicates that the rotation and translation processing is applied, the three-dimensional data decoding device 1400 applies the rotation and translation processing to the position information of the three-dimensional point included in the reference three-dimensional data indicated by the RT information. Also, when the RT application flag indicates that the rotation and translation processing is not applied, the three-dimensional data decoding device 1400 does not apply the rotation and translation processing to the position information of the three-dimensional point included in the reference three-dimensional data.
[0559] In addition, the three-dimensional data decoding device 1400 may perform the rotation and translation processing in a first unit (e.g., space), and may generate the predicted position information in a second unit (e.g., volume) that is smaller than the first unit. In addition, the three-dimensional data decoding device 1400 may also perform the rotation and translation processing, and the generation of the predicted position information in the same unit.
[0560] Furthermore, the three-dimensional data decoding device 1400 may apply a first rotation and translation process with a first unit (e.g., space) to the position information of the three-dimensional points included in the reference three-dimensional data, and may apply a second rotation and translation process with a second unit (e.g., volume) that is smaller than the first unit to the position information of the three-dimensional points obtained by the first rotation and translation process, thereby generating predicted position information.
[0561] Here, the position information of the three-dimensional point and the predicted position information are as follows: Fig.41 As shown, the octree structure is used for representation. For example, the position information and predicted position information of the three-dimensional point are represented in a scanning order that prioritizes the width among the depth and width in the octree structure. Alternatively, the position information and predicted position information of the three-dimensional point are represented in a scanning order that prioritizes the depth among the depth and width in the octree structure.
[0562] The three-dimensional data decoding apparatus 1400 generates predicted attribute information by using the attribute information of the three-dimensional points included in the reference three-dimensional data ( S1404 ).
[0563] Next, the three-dimensional data decoding device 1400 decodes the coded position information contained in the coded signal by using the predicted position information, thereby restoring the position information of the three-dimensional point contained in the target three-dimensional data. Here, the coded position information is, for example, differential position information, and the three-dimensional data decoding device 1400 restores the position information of the three-dimensional point contained in the target three-dimensional data by adding the differential position information to the predicted position information (S1405).
[0564] Furthermore, the three-dimensional data decoding device 1400 decodes the coded attribute information included in the coded signal by using the predicted attribute information, thereby restoring the attribute information of the three-dimensional point included in the target three-dimensional data. Here, the coded attribute information is, for example, differential attribute information, and the three-dimensional data decoding device 1400 restores the attribute information of the three-dimensional point included in the target three-dimensional data by adding the differential attribute information to the predicted attribute information (S1406).
[0565] Alternatively, if the three-dimensional data does not include attribute information, the three-dimensional data decoding device 1400 may not perform steps S1402, S1404, and S1406. Furthermore, the three-dimensional data decoding device 1400 may only decode the position information of the three-dimensional point or decode the attribute information of the three-dimensional point.
[0566] and, Fig.50 The order of processing shown is an example and is not limited to this. For example, since the processing of location information (S1403, S1405) and the processing of attribute information (S1402, S1404, S1406) are independent of each other, they can be performed in any order, and some of them can be processed in parallel.
[0567] (Implementation 8)
[0568] A method of controlling reference during encoding of the occupancy rate encoding in this embodiment will be described. In addition, the following mainly describes the operation of the three-dimensional data encoding device, but the same processing can be performed in the three-dimensional data decoding device.
[0569] Fig.51 as well as Fig.52 is a diagram showing a reference relationship involved in this embodiment, Fig.51 It is a graph that represents the reference relationship on the octree structure. Fig.52 It is a graph that represents reference relationships in a spatial area.
[0570] In this embodiment, when the three-dimensional data encoding device encodes the encoding information of the node of the encoding object (hereinafter referred to as the object node), the encoding information of each node in the parent node to which the object node belongs is referenced. However, the encoding information of each node in other nodes (hereinafter referred to as parent adjacent nodes) in the same layer as the parent node is not referenced. In other words, the three-dimensional data encoding device sets the parent adjacent node to be unable to be referenced, or prohibits reference.
[0571] In addition, the three-dimensional data encoding device may also allow reference to the encoding information in the parent node (hereinafter referred to as the grandparent node) to which the parent node belongs. That is, the three-dimensional data encoding device may also refer to the encoding information of the parent node and the grandparent node to which the object node belongs, and encode the encoding information of the object node.
[0572] Here, the coding information is, for example, an occupancy code. When encoding the occupancy code of the object node, the three-dimensional data coding device refers to information indicating whether each node in the parent node to which the object node belongs contains a point group (hereinafter referred to as occupancy information). In other words, when encoding the occupancy code of the object node, the three-dimensional data coding device refers to the occupancy code of the parent node. On the other hand, the three-dimensional data coding device does not refer to the occupancy information of each node in the parent adjacent node. That is, the three-dimensional data coding device does not refer to the occupancy code of the parent adjacent node. In addition, the three-dimensional data coding device may also refer to the occupancy information of each node in the grandparent node. That is, the three-dimensional data coding device may also refer to the occupancy information of the parent node and the parent adjacent node.
[0573] For example, when encoding the occupancy code of the object node, the three-dimensional data encoding device uses the occupancy code of the parent node or the grandparent node to which the object node belongs, and switches the coding table used when entropy coding the occupancy code of the object node. In addition, the details are described later. At this time, the three-dimensional data encoding device may not refer to the occupancy code of the parent adjacent node. Thus, when encoding the occupancy code of the object node, the three-dimensional data encoding device can appropriately switch the coding table according to the information of the occupancy code of the parent node or the grandparent node, thereby improving the coding efficiency. In addition, the three-dimensional data encoding device does not refer to the parent adjacent node, thereby suppressing the confirmation processing of the information of the parent adjacent node and the memory capacity used to store the processing. In addition, it becomes easy to scan the occupancy code of each node of the octree in depth-first order and encode it.
[0574] Hereinafter, an example of switching the coding table using the occupancy rate coding of the parent node will be described. Fig.53 is a diagram showing an example of a target node and adjacent reference nodes. Fig.54 It is a graph representing the relationship between parent nodes and nodes. Fig.55 is a diagram showing an example of encoding the occupancy rate of a parent node. Here, the adjacent reference node refers to a node that is spatially adjacent to the target node and is referenced when encoding the target node. Fig.53 In the example shown, the adjacent nodes are nodes belonging to the same layer as the target node. In addition, as reference adjacent nodes, the node X adjacent to the target block in the x direction, the node Y adjacent to the y direction, and the node Z adjacent to the z direction are used. That is, one adjacent block is set as the reference adjacent block in each of the x, y, and z directions.
[0575] also, Fig.54 The node numbers shown are examples, and the relationship between the node numbers and the positions of the nodes is not limited to this. Fig.55In , the node 0 is allocated to the lower bit and the node 7 is allocated to the upper bit, but the allocation may be in the reverse order. In addition, each node may be allocated to any bit.
[0576] The three-dimensional data encoding device determines a coding table when entropy coding is performed on the occupancy coding of the target node, for example, by the following equation.
[0577] CodingTable=(FlagX<<2)+(FlagY<<1)+(FlagZ)
[0578] Here, CodingTable represents a coding table for encoding the occupancy of the object node, and represents any value from 0 to 7. FlagX is the occupancy information of the adjacent node X, and represents 1 if the adjacent node X contains (occupies) the point group, and represents 0 if not. FlagY is the occupancy information of the adjacent node Y, and represents 1 if the adjacent node Y contains (occupies) the point group, and represents 0 if not. FlagZ is the occupancy information of the adjacent node Z, and represents 1 if the adjacent node Z contains (occupies) the point group, and represents 0 if not.
[0579] Furthermore, since information indicating whether an adjacent node is occupied is included in the occupancy code of the parent node, the three-dimensional data encoding device may select the encoding table using the value indicated in the occupancy code of the parent node.
[0580] As can be seen from the above, the three-dimensional data encoding device can improve the encoding efficiency by switching the encoding table using information indicating whether the neighboring nodes of the target node include a point group.
[0581] In addition, if Fig.53 As shown, the three-dimensional data encoding device can also switch the adjacent reference node according to the spatial position of the object node in the parent node. That is, the three-dimensional data encoding device can also switch the adjacent node for reference among multiple adjacent nodes according to the spatial position of the parent node of the object node.
[0582] Next, configuration examples of a three-dimensional data encoding device and a three-dimensional data decoding device will be described. Fig.56 This is a block diagram of a three-dimensional data encoding device 2100 according to this embodiment. Fig.56 The three-dimensional data encoding device 2100 shown includes an octree generation unit 2101 , a geometric information calculation unit 2102 , a coding table selection unit 2103 , and an entropy encoding unit 2104 .
[0583] The octree generation unit 2101 generates, for example, an octree based on the input three-dimensional points (point cloud), and generates an occupancy code for each node included in the octree. The geometry information calculation unit 2102 obtains occupancy information indicating whether the adjacent reference node of the object node is occupied. For example, the geometry information calculation unit 2102 obtains the occupancy information of the adjacent reference node from the occupancy code of the parent node to which the object node belongs. In addition, Fig.53 As shown, the geometric information calculation unit 2102 may switch the adjacent reference node according to the position of the target node in the parent node. In addition, the geometric information calculation unit 2102 does not refer to the occupancy information of each node in the parent adjacent node.
[0584] The coding table selection unit 2103 selects a coding table used in entropy coding of the occupancy coding of the target node using the occupancy information of the adjacent reference nodes calculated by the geometric information calculation unit 2102. The entropy coding unit 2104 generates a bit stream by entropy coding the occupancy coding using the selected coding table. In addition, the entropy coding unit 2104 may also attach information indicating the selected coding table to the bit stream.
[0585] Fig.57 This is a block diagram of a three-dimensional data decoding device 2110 according to this embodiment. Fig.57 The three-dimensional data decoding device 2110 shown includes an octree generation unit 2111 , a geometric information calculation unit 2112 , a coding table selection unit 2113 , and an entropy decoding unit 2114 .
[0586] The octree generation unit 2111 generates an octree of a certain space (node) using the header information of the bitstream, etc. The octree generation unit 2111 generates a large space (root node) using the size of the x-axis, y-axis, and z-axis directions of the certain space added to the header information, and generates an octree by dividing the space into two in the x-axis, y-axis, and z-axis directions, respectively, to generate eight small spaces A (nodes A0 to A7). In addition, nodes A0 to A7 are sequentially set as object nodes.
[0587] The geometry information calculation unit 2112 obtains the occupancy information indicating whether the adjacent reference node of the object node is occupied. For example, the geometry information calculation unit 2112 obtains the occupancy information of the adjacent reference node from the occupancy rate code of the parent node to which the object node belongs. Fig.53 As shown, the geometric information calculation unit 2112 may switch the adjacent reference node according to the position of the target node in the parent node. In addition, the geometric information calculation unit 2112 does not refer to the occupancy information of each node in the parent adjacent node.
[0588] The coding table selection unit 2113 selects a coding table (decoding table) used for entropy decoding of the occupancy code of the target node using the occupancy information of the adjacent reference nodes calculated by the geometric information calculation unit 2112. The entropy decoding unit 2114 generates a three-dimensional point by entropy decoding the occupancy code using the selected coding table. In addition, the coding table selection unit 2113 decodes and obtains the information of the selected coding table attached to the bit stream, and the entropy decoding unit 2114 may also use the coding table indicated by the obtained information.
[0589] Each bit of the occupancy code (8 bits) included in the bit stream indicates whether a point group is included in each of the eight small spaces A (nodes A0 to A7). Furthermore, the three-dimensional data decoding device divides the small space node A0 into eight small spaces B (nodes B0 to B7) and generates an octree, decodes the occupancy code, and obtains information indicating whether each node of the small space B contains a point group. In this way, the three-dimensional data decoding device decodes the occupancy code of each node while generating an octree from a large space to a small space.
[0590] The following describes the flow of processing by the three-dimensional data encoding device and the three-dimensional data decoding device. Fig.58 3D data encoding processing in a 3D data encoding device. First, the 3D data encoding device determines (defines) a space (object node) containing part or all of the input 3D point group (S2101). Next, the 3D data encoding device divides the object node 8 into 8 small spaces (nodes) (S2102). Next, the 3D data encoding device generates an occupancy code of the object node according to whether each node contains a point group (S2103).
[0591] Next, the three-dimensional data encoding device calculates (obtains) the occupancy information of the adjacent reference nodes of the object node from the occupancy code of the parent node of the object node (S2104). Next, the three-dimensional data encoding device selects a coding table used in entropy coding based on the occupancy information of the adjacent reference nodes of the determined object node (S2105). Next, the three-dimensional data encoding device performs entropy coding on the occupancy code of the object node using the selected coding table (S2106).
[0592] Furthermore, the three-dimensional data encoding device repeatedly divides each node 8 and encodes the occupancy rate code of each node until the node cannot be divided (S2107). That is, the processing of steps S2102 to S2106 is recursively repeated.
[0593] Fig.59This is a flowchart of a three-dimensional data decoding method in a three-dimensional data decoding device. First, the three-dimensional data decoding device uses the header information of the bit stream to determine (define) the space (object node) to be decoded (S2111). Next, the three-dimensional data decoding device divides the object node 8 to generate 8 small spaces (nodes) (S2112). Next, the three-dimensional data decoding device calculates (obtains) the occupancy information of the adjacent reference nodes of the object node from the occupancy code of the parent node of the object node (S2113).
[0594] Next, the three-dimensional data decoding device selects a coding table used for entropy decoding based on the occupancy information of the adjacent reference nodes (S2114). Next, the three-dimensional data decoding device entropy decodes the occupancy code of the target node using the selected coding table (S2115).
[0595] Furthermore, the three-dimensional data decoding device repeatedly divides each node 8 and decodes the occupancy code of each node until the node cannot be divided (S2116). That is, the processing of steps S2112 to S2115 is recursively repeated.
[0596] Next, an example of switching the coding table will be described. Fig.60 is a diagram showing an example of switching of the coding table. Fig.60 As shown in the coding table 0, the same context model can be applied to multiple occupancy codes. In addition, each occupancy code can also be assigned a different context model. Thus, the context model can be assigned according to the probability of occurrence of the occupancy code, thereby improving the coding efficiency. In addition, a context model that updates the probability table according to the frequency of occurrence of the occupancy code can also be used. In addition, a context model that fixes the probability table can also be used.
[0597] Hereinafter, Modification 1 of the present embodiment will be described. Fig.61 In the above embodiment, the three-dimensional data encoding device does not refer to the occupancy rate encoding of the parent adjacent node, but it is also possible to switch whether to refer to the occupancy rate encoding of the parent adjacent node according to a specific condition.
[0598] For example, when the three-dimensional data encoding device encodes the occupancy rate of the object node by referring to the occupancy information of the node in the parent adjacent node while encoding the octree with a width-first scan. On the other hand, when the three-dimensional data encoding device encodes the occupancy rate of the object node by referring to the occupancy information of the node in the parent adjacent node while encoding the octree with a depth-first scan. In this way, according to the scanning order (encoding order) of the nodes of the octree, the nodes that can be referenced are appropriately switched, so that the encoding efficiency can be improved and the processing load can be suppressed.
[0599] Furthermore, the three-dimensional data encoding device may add information such as whether the octree is encoded with width first or depth first to the header of the bit stream. Fig.62 This is a diagram showing an example of the syntax of header information in this case. Fig.62 The octree_scan_order shown is encoding order information (encoding order flag) indicating the encoding order of the octree. For example, when octree_scan_order is 0, it indicates width priority, and when it is 1, it indicates depth priority. Thus, by referring to octree_scan_order, the three-dimensional data decoding device can know whether the bit stream is encoded in width priority or depth priority, and can appropriately decode the bit stream.
[0600] Furthermore, the three-dimensional data encoding device may add information indicating whether or not to prohibit reference to a parent adjacent node to header information of the bit stream. Fig.63 The figure shows a syntax example of the header information in this case. limit_refer_flag is the prohibition switching information (prohibition switching flag) indicating whether to prohibit the reference to the parent adjacent node. For example, when limit_refer_flag is 1, it prohibits the reference to the parent adjacent node, and when it is 0, it indicates that there is no reference restriction (the reference to the parent adjacent node is permitted).
[0601] That is, the three-dimensional data encoding device determines whether to prohibit reference to the parent adjacent node, and switches whether to prohibit or allow reference to the parent adjacent node based on the result of the above determination. In addition, the three-dimensional data encoding device generates a bit stream including prohibition switching information, the prohibition switching information is the result of the above determination, and indicates whether to prohibit reference to the parent adjacent node.
[0602] Furthermore, the three-dimensional data decoding device obtains prohibition switching information indicating whether to prohibit referring to the parent adjacent node from the bit stream, and switches whether to prohibit or permit referring to the parent adjacent node based on the prohibition switching information.
[0603] Thus, the three-dimensional data encoding device can control the reference of the parent adjacent node and generate a bit stream. In addition, the three-dimensional data decoding device can obtain information indicating whether to prohibit the reference of the parent adjacent node from the header of the bit stream.
[0604] In addition, in this embodiment, as an example of prohibiting the coding process of referring to the parent adjacent node, the coding process of occupancy coding is recorded as an example, but it is not necessarily limited to this. For example, the same method can also be applied when encoding other information of the node of the octree. For example, when encoding other attribute information such as color, normal vector, or reflectivity attached to the node, the method of this embodiment can also be applied. In addition, the same method can also be applied when encoding the coding table or the predicted value.
[0605] Next, a second variation of the present embodiment will be described. Fig.53 Although an example using three reference adjacent nodes is shown, four or more reference adjacent nodes may be used. Fig.64 This is a diagram showing an example of a target node and reference adjacent nodes.
[0606] For example, the three-dimensional data encoding device calculates the Fig.64 The occupancy coding of the object node shown is a coding table when entropy coding is performed.
[0607] CodingTable=(FlagX0<<3)+(FlagX1<<2)+(FlagY<<1)+(FlagZ)
[0608] Here, CodingTable represents a coding table for encoding the occupancy of the object node, and represents any value from 0 to 15. FlagXN is the occupancy information of the adjacent node XN (N=0…1), and represents 1 if the adjacent node XN contains (occupies) a point group, and represents 0 if not. FlagY is the occupancy information of the adjacent node Y, and represents 1 if the adjacent node Y contains (occupies) a point group, and represents 0 if not. FlagZ is the occupancy information of the adjacent node Z, and represents 1 if the adjacent node Z contains (occupies) a point group, and represents 0 if not.
[0609] At this time, if the adjacent node is, for example Fig.64 In the case where the adjacent node X0 is not referenceable (prohibited from reference), the three-dimensional data encoding device may use a fixed value such as 1 (occupied) or 0 (non-occupied) as a substitute value.
[0610] Fig.65 is a graph showing examples of object nodes and adjacent nodes. Fig.65 As shown, when it is impossible to refer to (prohibited to refer to) the adjacent nodes, the occupancy rate code of the grandparent node of the object node can also be referred to to calculate the occupancy information of the adjacent nodes. Fig.65 The three-dimensional data encoding device can also use the occupancy information of the adjacent node G0 to calculate the FlagX0 of the above formula, and use the calculated FlagX0 to determine the value of the encoding table. Fig.65 The neighboring node G0 shown is a neighboring node that can be determined by the occupancy rate coding of the grandparent node. The neighboring node X1 is a neighboring node that can be determined by the occupancy rate coding of the parent node.
[0611] Hereinafter, Modification 3 of the present embodiment will be described. Fig.66 as well as Fig.67 is a diagram showing a reference relationship involved in this modification example, Fig.66 It is a graph that represents the reference relationship on the octree structure. Fig.67 It is a graph that represents reference relationships in a spatial area.
[0612] In this variant, when encoding the encoding information of the node of the encoding object (hereinafter referred to as the object node 2), the three-dimensional data encoding device refers to the encoding information of each node in the parent node to which the object node 2 belongs. That is, the three-dimensional data encoding device allows reference to the information (for example, occupancy information) of the child nodes of the first node whose parent node is the same as the parent node of the object node among the plurality of adjacent nodes. Fig.66 When encoding the occupancy rate code of the object node 2 shown in FIG. 1 , reference is made to the nodes existing in the parent node to which the object node 2 belongs, for example, Fig.66 The occupancy rate of the object node is shown in . Fig.67 As shown, Fig.66 The occupancy code of the object node shown indicates whether each node in the object node adjacent to the object node 2 is occupied. Therefore, the three-dimensional data encoding device can switch the encoding table of the occupancy code of the object node 2 according to the finer shape of the object node, thereby improving the encoding efficiency.
[0613] The three-dimensional data encoding device may calculate a coding table for entropy coding the occupancy coding of the target node 2 by, for example, the following equation.
[0614] CodingTable=(FlagX1<<5)+(FlagX2<<4)+(FlagX3<<3)+(FlagX4<<2)+(FlagY<<1)+(FlagZ)
[0615] Here, CodingTable represents a coding table for encoding the occupancy of the object node 2, and represents any value from 0 to 63. FlagXN is the occupancy information of the adjacent node XN (N=1…4), and represents 1 if the adjacent node XN contains (occupies) a point group, and represents 0 if not. FlagY is the occupancy information of the adjacent node Y, and represents 1 if the adjacent node Y contains (occupies) a point group, and represents 0 if not. FlagZ is the occupancy information of the adjacent node Y, and represents 1 if the adjacent node Z contains (occupies) a point group, and represents 0 if not.
[0616] Furthermore, the three-dimensional data encoding device may change the method of calculating the encoding table according to the node position of the object node 2 within the parent node.
[0617] In addition, if it is not prohibited to refer to the parent adjacent node, the three-dimensional data encoding device can refer to the encoding information of each node in the parent adjacent node. For example, if it is not prohibited to refer to the parent adjacent node, it is allowed to refer to the information (such as occupancy information) of the child node of the third node whose parent node is different from the parent node of the object node. Fig.65 In the example shown, the three-dimensional data encoding device refers to the occupancy code of the adjacent node X0 whose parent node is different from the parent node of the object node to obtain the occupancy information of the child node of the adjacent node X0. The three-dimensional data encoding device switches the encoding table used in the entropy encoding of the occupancy code of the object node based on the obtained occupancy information of the child node of the adjacent node X0.
[0618] As described above, the three-dimensional data encoding device according to the present embodiment encodes information (such as occupancy rate encoding) of object nodes included in an N-way tree structure (N is an integer greater than or equal to 2) of a plurality of three-dimensional points included in three-dimensional data. Fig.51 as well as Fig.52 As shown, in the above encoding, the three-dimensional data encoding device allows reference to information (e.g., occupancy information) of a first node whose parent node is the same as the parent node of the object node among a plurality of adjacent nodes spatially adjacent to the object node, and prohibits reference to information (e.g., occupancy information) of a second node whose parent node is different from the parent node of the object node. In other words, in the above encoding, the three-dimensional data encoding device allows reference to information (e.g., occupancy rate encoding) of the parent node, and prohibits reference to information (e.g., occupancy rate encoding) of other nodes (parent adjacent nodes) at the same layer as the parent node.
[0619] Thus, the three-dimensional data encoding device can improve the encoding efficiency by referring to the information of the first node whose parent node is the same as the parent node of the object node among the plurality of adjacent nodes that are spatially adjacent to the object node. In addition, the three-dimensional data encoding device can reduce the processing amount because it does not refer to the information of the second node whose parent node is different from the parent node of the object node among the plurality of adjacent nodes. In this way, the three-dimensional data encoding device can improve the encoding efficiency and reduce the processing amount.
[0620] For example, the three-dimensional data encoding device further determines whether to prohibit reference to the second node information, and in the above encoding, based on the result of the above determination, switches whether to prohibit or permit reference to the second node information. The three-dimensional data encoding device further generates a switching prohibition information (for example, Fig.63 The limit_refer_flag shown in the bit stream is configured such that the prohibition switching information is the result of the above determination and indicates whether reference to the second node is prohibited.
[0621] This allows the three-dimensional data encoding device to switch whether to prohibit reference to the information of the second node. In addition, the three-dimensional data decoding device can appropriately perform decoding processing using the prohibition switching information.
[0622] For example, the information of the object node is information indicating whether there is a three-dimensional point in each of the child nodes belonging to the object node (for example, occupancy coding), the information of the first node is information indicating whether there is a three-dimensional point in the first node (occupancy information of the first node), and the information of the second node is information indicating whether there is a three-dimensional point in the second node (occupancy information of the second node).
[0623] For example, in the above encoding, the three-dimensional data encoding device selects a coding table based on whether a three-dimensional point exists at the first node, and uses the selected coding table to entropy encode the information of the target node (eg, occupancy code).
[0624] For example, Fig.66 as well as Fig.67 As shown, the three-dimensional data encoding device allows reference to information (for example, occupancy information) of child nodes of a first node among a plurality of adjacent nodes during the encoding.
[0625] Therefore, the three-dimensional data encoding device can improve encoding efficiency because it can refer to more detailed information of adjacent nodes.
[0626] For example, Fig.53 As shown, in the above encoding, the three-dimensional data encoding device switches the adjacent node to be referenced among the plurality of adjacent nodes according to the spatial position of the parent node of the object node.
[0627] Thus, the three-dimensional data encoding device can refer to appropriate adjacent nodes according to the spatial position of the target node in the parent node.
[0628] For example, the three-dimensional data encoding device includes a processor and a memory, and the processor performs the above-mentioned processing using the memory.
[0629] In addition, the three-dimensional data decoding device according to the present embodiment decodes information (such as occupancy code) of the object node included in the N (N is an integer greater than or equal to 2) tree structure of a plurality of three-dimensional points included in the three-dimensional data. Fig.51 as well as Fig.52As shown, in the above decoding, the three-dimensional data decoding device allows reference to information (e.g., occupancy information) of a first node whose parent node is the same as the parent node of the object node among a plurality of adjacent nodes spatially adjacent to the object node, and prohibits reference to information (e.g., occupancy information) of a second node whose parent node is different from the parent node of the object node. In other words, in the above decoding, the three-dimensional data decoding device allows reference to information (e.g., occupancy code) of the parent node, and prohibits reference to information (e.g., occupancy code) of other nodes (parent adjacent nodes) in the same layer as the parent node.
[0630] Thus, the three-dimensional data decoding device can improve the coding efficiency by referring to the information of the first node whose parent node is the same as the parent node of the object node among the plurality of adjacent nodes that are spatially adjacent to the object node. In addition, the three-dimensional data decoding device can reduce the processing amount because it does not refer to the information of the second node whose parent node is different from the parent node of the object node among the plurality of adjacent nodes. In this way, the three-dimensional data decoding device can improve the coding efficiency and reduce the processing amount.
[0631] For example, the three-dimensional data decoding device further obtains, from the bit stream, prohibition switching information indicating whether to prohibit reference to the information of the second node (for example, Fig.63 limit_refer_flag) is shown in the above decoding. Based on the prohibition switching information, whether to prohibit or permit reference to the information of the second node is switched.
[0632] Thus, the three-dimensional data decoding device can appropriately perform decoding processing using the switching inhibition information.
[0633] For example, the information of the object node is information indicating whether there is a three-dimensional point in each of the child nodes belonging to the object node (for example, occupancy coding), the information of the first node is information indicating whether there is a three-dimensional point in the first node (occupancy information of the first node), and the information of the second node is information indicating whether there is a three-dimensional point in the second node (occupancy information of the second node).
[0634] For example, in the above decoding, the three-dimensional data decoding device selects a coding table based on whether a three-dimensional point exists at the first node, and uses the selected coding table to entropy decode the information of the target node (eg, occupancy code).
[0635] For example, Fig.66 as well as Fig.67 As shown, the three-dimensional data decoding device allows reference to information (for example, occupancy information) of child nodes of a first node among a plurality of adjacent nodes during the above decoding.
[0636] As a result, the three-dimensional data decoding device can refer to more detailed information of adjacent nodes, thereby improving encoding efficiency.
[0637] For example, Fig.53 As shown, in the above decoding, the three-dimensional data decoding device switches the adjacent node to be referenced among the plurality of adjacent nodes according to the spatial position of the target node in the parent node.
[0638] Thus, the three-dimensional data decoding device can refer to appropriate adjacent nodes according to the spatial position of the target node in the parent node.
[0639] For example, the three-dimensional data decoding device includes a processor and a memory, and the processor performs the above-mentioned processing using the memory.
[0640] (Implementation method 9)
[0641] The information of the three-dimensional point group includes position information (geometry) and attribute information (attribute). The position information includes coordinates (x coordinate, y coordinate, z coordinate) based on a certain point. When encoding the position information, instead of directly encoding the coordinates of each three-dimensional point, a method is used to reduce the amount of encoding by using an octree to represent the position of each three-dimensional point and encoding the information of the octree.
[0642] On the other hand, the attribute information includes information indicating color information (RGB, YUV, etc.), reflectivity, normal vector, etc. of each 3D point. For example, the 3D data encoding device can encode the attribute information using a different encoding method from that of the position information.
[0643] In this embodiment, a method for encoding attribute information is described. In addition, in this embodiment, an integer value is used as the value of the attribute information for description. For example, when each color component of the color information RGB or YUV has an 8-bit precision, each color component takes an integer value of 0 to 255. When the reflectivity value has a 10-bit precision, the reflectivity value takes an integer value of 0 to 1023. In addition, when the bit precision of the attribute information is a decimal precision, the three-dimensional data encoding device may also multiply the value by a scaling value and round it to an integer value so that the value of the attribute information becomes an integer value. In addition, the three-dimensional data encoding device may also attach the scaling value to the header of the bit stream, etc.
[0644] As a method for encoding attribute information of a three-dimensional point, it is considered to calculate the predicted value of the attribute information of the three-dimensional point and encode the difference (prediction residual) between the original attribute information value and the predicted value. For example, when the value of the attribute information of the three-dimensional point p is Ap and the predicted value is Pp, the three-dimensional data encoding device encodes its differential absolute value Diffp = |Ap-Pp|. In this case, if the predicted value Pp can be generated with high precision, the value of the differential absolute value Diffp becomes smaller. Therefore, for example, by entropy encoding the differential absolute value Diffp using a coding table that generates a smaller number of bits as the value is smaller, the amount of coding can be reduced.
[0645] As a method for generating a predicted value of attribute information, it is considered to use attribute information of other three-dimensional points located around the object three-dimensional point of the encoding object, namely, reference three-dimensional points. Here, the reference three-dimensional point refers to a three-dimensional point within a predetermined distance range from the object three-dimensional point. For example, when there is an object three-dimensional point p = (x1, y1, z1) and a three-dimensional point q = (x2, y2, z2), the three-dimensional data encoding device calculates the Euclidean distance d (p, q) between the three-dimensional point p and the three-dimensional point q shown in (Formula A1).
[0646]
Formula 1
[0647]
[0648] When the Euclidean distance d(p, q) is less than a predetermined threshold value THd, the three-dimensional data encoding device determines that the position of the three-dimensional point q is close to the position of the object three-dimensional point p, and determines that the value of the attribute information of the three-dimensional point q is used in the generation of the predicted value of the attribute information of the object three-dimensional point p. In addition, the distance calculation method may also be other methods, for example, the Mahalanobis distance may also be used. In addition, the three-dimensional data encoding device may also determine not to use the three-dimensional points outside the predetermined distance range from the object three-dimensional point for prediction processing. For example, when there is a three-dimensional point r and the distance d(p, r) between the object three-dimensional point p and the three-dimensional point r is greater than the threshold value THd, the three-dimensional data encoding device may also determine not to use the three-dimensional point r for prediction. In addition, the three-dimensional data encoding device may also attach information indicating the threshold value THd to the header of the bit stream, etc.
[0649] Fig.68 3D points are shown in the figure. In this example, the distance d(p, q) between the target 3D point p and the 3D point q is less than the threshold value THd. Therefore, the 3D data encoding device determines that the 3D point q is a reference 3D point of the target 3D point p, and determines that the value of the attribute information Aq of the 3D point q is used in the generation of the predicted value Pp of the attribute information Ap of the target 3D point p.
[0650] On the other hand, the distance d(p, r) between the object 3D point p and the 3D point r is greater than the threshold value THd. Therefore, the 3D data encoding device determines that the 3D point r is not a reference 3D point of the object 3D point p, and determines that the value of the attribute information Ar of the 3D point r is not used in the generation of the predicted value Pp of the attribute information Ap of the object 3D point p.
[0651] Furthermore, when encoding the attribute information of the target 3D point using the prediction value, the 3D data encoding device uses the 3D point whose attribute information has been encoded and decoded as a reference 3D point. Similarly, when decoding the attribute information of the target 3D point of the decoding target using the prediction value, the 3D data decoding device uses the 3D point whose attribute information has been decoded as a reference 3D point. Thus, the same prediction value can be generated during encoding and decoding, so that the bit stream of the 3D point generated by encoding can be correctly decoded on the decoding side.
[0652] In addition, when encoding the attribute information of the 3D points, it is considered to classify each 3D point into multiple levels using the position information of the 3D points and then encode them. Here, each level after classification is called LoD (Level of Detail). Fig.69 The method of generating LoD is described.
[0653] First, the three-dimensional data encoding device selects an initial point a0 and assigns it to LoD0. Next, the three-dimensional data encoding device extracts a point a1 whose distance from point a0 is greater than the threshold Thres_LoD[0] of LoD0 and assigns it to LoD0. Next, the three-dimensional data encoding device extracts a point a2 whose distance from point a1 is greater than the threshold Thres_LoD[0] of LoD0 and assigns it to LoD0. In this way, the three-dimensional data encoding device constructs LoD0 in such a way that the distance between each point in LoD0 is greater than the threshold Thres_LoD[0].
[0654] Next, the 3D data encoding device selects a point b0 that has not been assigned a LoD, and assigns it to LoD1. Next, the 3D data encoding device extracts a point b1 that has not been assigned a LoD and whose distance from point b0 is greater than the threshold Thres_LoD[1] of LoD1, and which has not been assigned a LoD, and assigns it to LoD1. Next, the 3D data encoding device extracts a point b2 that has not been assigned a LoD and whose distance from point b1 is greater than the threshold Thres_LoD[1] of LoD1, and which has not been assigned a LoD, and assigns it to LoD1. In this way, the 3D data encoding device constructs LoD1 in such a way that the distance between each point in LoD1 is greater than the threshold Thres_LoD[1].
[0655] Next, the three-dimensional data encoding device selects point c0 that has not been assigned a LoD, and assigns it to LoD2. Next, the three-dimensional data encoding device extracts point c1 that has not been assigned a LoD and whose distance from point c0 is greater than the threshold Thres_LoD[2] of LoD2, and which has not been assigned a LoD, and assigns it to LoD2. Next, the three-dimensional data encoding device extracts point c2 that has not been assigned a LoD and whose distance from point c1 is greater than the threshold Thres_LoD[2] of LoD2, and which has not been assigned a LoD, and assigns it to LoD2. In this way, the three-dimensional data encoding device constructs LoD2 in such a way that the distance between each point in LoD2 is greater than the threshold Thres_LoD[2]. For example, Fig.70 As shown, the threshold values of each LoD, Thres_LoD[0], Thres_LoD[1], and Thres_LoD[2] are set.
[0656] In addition, the three-dimensional data encoding device may also add information indicating the threshold value of each LoD to the header of the bit stream. Fig.70 In the illustrated example, the three-dimensional data encoding device may add threshold values Thres_LoD[0], Thres_LoD[1], and Thres_LoD[2] to the header.
[0657] Alternatively, the 3D data encoding device may also assign all 3D points that are not assigned LoD to the lowest layer of LoD. In this case, the 3D data encoding device can reduce the amount of coding in the head by not adding the threshold of the lowest layer of LoD to the head. Fig.70 In the example shown, the 3D data encoding device adds the thresholds Thres_LoD[0] and Thres_LoD[1] to the header, and does not add Thres_LoD[2] to the header. In this case, the 3D data decoding device may also estimate the value of Thres_LoD[2] to be 0. In addition, the 3D data encoding device may also add the number of layers of LoD to the header. Thus, the 3D data decoding device can use the number of layers of LoD to determine the LoD of the lowest layer.
[0658] In addition, if Fig.70 As shown in FIG. 1 , the threshold value of each LoD layer is set to be larger as it is closer to the upper layer, so that the upper layer (the layer closer to LoD0) becomes a sparse point group with a long distance between three-dimensional points, and the lower layer becomes a dense point group with a short distance between three-dimensional points. Fig.70 In the example shown, LoD0 is the highest layer.
[0659] In addition, the method of selecting the initial three-dimensional point when setting each LoD may also depend on the encoding order when encoding the position information. For example, the three-dimensional data encoding device selects the three-dimensional point that is first encoded when encoding the position information as the initial point a0 of LoD0, and selects points a1 and a2 to constitute LoD0 based on the initial point a0. Moreover, the three-dimensional data encoding device may also select the three-dimensional point whose position information is encoded earliest among the three-dimensional points that do not belong to LoD0 as the initial point b0 of LoD1. That is, the three-dimensional data encoding device may also select the three-dimensional point whose position information is encoded earliest among the three-dimensional points that do not belong to the upper layer (LoD0 to LoDn-1) of LoDn as the initial point n0 of LoDn. Thus, the three-dimensional data decoding device can construct the same LoD as that during encoding by using the same initial point selection method during decoding, and thus can properly decode the bit stream. Specifically, the three-dimensional data decoding device selects the three-dimensional point whose position information is decoded earliest among the three-dimensional points that do not belong to the upper layer of LoDn as the initial point n0 of LoDn.
[0660] The following describes a method for generating predicted values of attribute information of three-dimensional points using LoD information. For example, when encoding three-dimensional points included in LoD0 in sequence, the three-dimensional data encoding device uses the encoded and decoded (hereinafter, also referred to as "encoded") attribute information included in LoD0 and LoD1 to generate the object three-dimensional point included in LoD1. In this way, the three-dimensional data encoding device uses the encoded attribute information included in LoDn' (n'<=n) to generate predicted values of attribute information of three-dimensional points included in LoDn. That is, the three-dimensional data encoding device does not use the attribute information of three-dimensional points included in the lower layer of LoDn in the calculation of the predicted values of the attribute information of the three-dimensional points included in LoDn.
[0661] For example, the three-dimensional data encoding device generates a predicted value of the attribute information of the three-dimensional point by calculating the average value of the attribute values of less than N three-dimensional points among the encoded three-dimensional points around the object three-dimensional point of the encoding object. In addition, the three-dimensional data encoding device can add the value of N to the header of the bit stream, etc. In addition, the three-dimensional data encoding device can also change the value of N for each three-dimensional point and add the value of N to each three-dimensional point. Thus, it is possible to select an appropriate N for each three-dimensional point, thereby improving the accuracy of the predicted value. Therefore, the prediction residual can be reduced. In addition, the three-dimensional data encoding device can also add the value of N to the header of the bit stream and fix the value of N in the bit stream. Thus, it is not necessary to encode or decode the value of N for each three-dimensional point, thereby reducing the amount of processing. In addition, the three-dimensional data encoding device can also encode the value of N for each LoD separately. Thus, by selecting an appropriate N for each LoD, the coding efficiency can be improved.
[0662] Alternatively, the 3D data encoding device may also calculate the predicted value of the attribute information of the 3D point by taking the weighted average of the attribute information of the surrounding N 3D points that have been encoded. For example, the 3D data encoding device calculates the weight using the distance information of the target 3D point and the surrounding N 3D points.
[0663] When the three-dimensional data encoding device encodes the value of N for each LoD, for example, the higher the LoD layer, the larger the value of N is set, and the lower the LoD layer, the smaller the value of N is set. In the upper layer of the LoD, the distance between the three-dimensional points belonging to the layer is far, so it is possible to set the value of N to a large value and select a plurality of surrounding three-dimensional points for averaging, thereby improving the prediction accuracy. In addition, since the distance between the three-dimensional points belonging to the layer in the lower layer of the LoD is close, it is possible to set the value of N to a small value to suppress the processing amount of averaging while performing efficient prediction.
[0664] Fig.71 is a diagram showing an example of attribute information used in the prediction value. As described above, the predicted value of the point P included in LoDN' (N' <= N) is generated using the encoded surrounding points P' included in LoDN'. Here, the surrounding points P' are selected based on the distance from the point P. For example, the attribute information of points a0, a1, a2, b0, and b1 is used to generate Fig.71 The predicted value of the attribute information of point b2 is shown.
[0665] The selected surrounding points change according to the value of N. For example, when N=5, a0, a1, a2, b0, and b1 are selected as the surrounding points of point b2. When N=4, points a0, a1, a2, and b1 are selected based on the distance information.
[0666] The prediction is calculated by weighted averaging depending on the distance. For example, Fig.71 In the example shown, the predicted value a2p of point a2 is calculated by taking the weighted average of the attribute information of points a0 and a1 as shown in (Formula A2) and (Formula A3). i It is the value of the attribute information of point ai.
[0667]
Formula 2
[0668]
[0669] In addition, the predicted value b2p of point b2 is calculated by weighted average of the attribute information of points a0, a1, a2, b0, and b1 as shown in (Formula A4) to (Formula A6). i It is the value of the attribute information of point bi.
[0670]
Formula 3
[0671]
[0672] In addition, the three-dimensional data encoding device can also calculate the difference between the value of the attribute information of the three-dimensional point and the predicted value generated from the surrounding points (prediction residual), and quantize the calculated prediction residual. For example, the three-dimensional data encoding device quantizes the prediction residual by dividing it by a quantization scale (also called a quantization step size). In this case, the smaller the quantization scale, the smaller the error (quantization error) that may be generated due to quantization. On the contrary, the larger the quantization scale, the larger the quantization error.
[0673] In addition, the three-dimensional data encoding device may also change the quantization scale used for each LoD. For example, the higher the level of the three-dimensional data encoding device, the smaller the quantization scale, and the lower the level of the three-dimensional data encoding device, the larger the quantization scale. The value of the attribute information of the three-dimensional point belonging to the upper layer may be used as the predicted value of the attribute information of the three-dimensional point belonging to the lower layer. Therefore, the quantization scale of the upper layer can be reduced to suppress the quantization error generated in the upper layer, and the coding efficiency can be improved by improving the accuracy of the predicted value. In addition, the three-dimensional data encoding device may also attach the quantization scale used for each LoD to the header, etc. As a result, the three-dimensional data decoding device can correctly decode the quantization scale, and thus can properly decode the bit stream.
[0674] In addition, the three-dimensional data encoding device may also transform the signed integer value (signed quantized value) of the quantized prediction residual into an unsigned integer value (unsigned quantized value). Thus, when entropy encoding is performed on the prediction residual, there is no need to consider the generation of negative integers. In addition, the three-dimensional data encoding device does not necessarily need to transform the signed integer value into an unsigned integer value, and for example, the sign bit may be entropy encoded separately.
[0675] The prediction residual is calculated by subtracting the predicted value from the original value. For example, as shown in (Formula A7), the prediction residual a2r of point a2 is calculated by subtracting the predicted value a2p of point a2 from the value A2 of the attribute information of point a2. As shown in (Formula A8), the prediction residual b2r of point b2 is calculated by subtracting the predicted value b2p of point b2 from the value B2 of the attribute information of point b2.
[0676] a2r=A2-a2p…(Formula A7)
[0677] b2r=B2-b2p...(Formula A8)
[0678] In addition, the prediction residual is quantized by dividing by QS (Quantization Step). For example, the quantization value a2q of point a2 is calculated by (Formula A9). The quantization value b2q of point b2 is calculated by (Formula A10). Here, QS_LoD0 is the QS for LoD0, and QS_LoD1 is the QS for LoD1. That is, QS can be changed according to LoD.
[0679] a2q=a2r / QS_LoD0…(Formula A9)
[0680] b2q=b2r / QS_LoD1…(Formula A10)
[0681] In addition, as described below, the three-dimensional data encoding device converts the signed integer value as the quantized value into an unsigned integer value. When the signed integer value a2q is less than 0, the three-dimensional data encoding device sets the unsigned integer value a2u to -1-(2×a2q). When the signed integer value a2q is greater than or equal to 0, the three-dimensional data encoding device sets the unsigned integer value a2u to 2×a2q.
[0682] Similarly, when the signed integer value b2q is less than 0, the three-dimensional data encoding device sets the unsigned integer value b2u to -1-(2×b2q). When the signed integer value b2q is greater than or equal to 0, the three-dimensional data encoding device sets the unsigned integer value b2u to 2×b2q.
[0683] Furthermore, the three-dimensional data encoding device may encode the quantized prediction residual (unsigned integer value) by entropy encoding. For example, after binarizing the unsigned integer value, binary arithmetic coding may be applied.
[0684] In addition, in this case, the three-dimensional data encoding device can also switch the binarization method according to the value of the prediction residual. For example, when the prediction residual pu is less than the threshold R_TH, the three-dimensional data encoding device binarizes the prediction residual pu with the required fixed number of bits to represent the threshold R_TH. In addition, when the prediction residual pu is greater than the threshold R_TH, the three-dimensional data encoding device binarizes the binarized data of the threshold R_TH and the value of (pu-R_TH) using Exponential-Golomb or the like.
[0685] For example, when the threshold R_TH is 63 and the prediction residual pu is less than 63, the three-dimensional data encoding device binarizes the prediction residual pu with 6 bits. In addition, when the prediction residual pu is greater than 63, the three-dimensional data encoding device binarizes the binary data (111111) and (pu-63) of the threshold R_TH using Exponential Golomb, thereby performing arithmetic coding.
[0686] In a more specific example, when the prediction residual pu is 32, the three-dimensional data encoding device generates 6-bit binary data (100000) and performs arithmetic coding on the bit string. In addition, when the prediction residual pu is 66, the three-dimensional data encoding device generates binary data (111111) representing the threshold value R_TH using Exponential Golomb and a bit string (00100) of value 3 (66-63), and performs arithmetic coding on the bit string (111111+00100).
[0687] Thus, the 3D data encoding device switches the binarization method according to the size of the prediction residual, thereby being able to perform encoding while suppressing a sharp increase in the number of binarization bits when the prediction residual becomes larger. In addition, the 3D data encoding device may also add the threshold R_TH to the header of the bitstream, etc.
[0688] For example, in the case of encoding at a high bit rate, that is, when the quantization scale is small, the quantization error becomes smaller, the prediction accuracy becomes higher, and the prediction residual may not become larger as a result. Therefore, in this case, the three-dimensional data encoding device sets the threshold R_TH to be large. As a result, the possibility of encoding the binary data of the threshold R_TH becomes lower, and the encoding efficiency is improved. On the contrary, in the case of encoding at a low bit rate, that is, when the quantization scale is large, the quantization error becomes larger, the prediction accuracy becomes worse, and the prediction residual may become larger as a result. Therefore, in this case, the three-dimensional data encoding device sets the threshold R_TH to be small. As a result, it is possible to prevent the sharp increase in the bit length of the binary data.
[0689] In addition, the three-dimensional data encoding device may also switch the threshold R_TH for each LoD, and attach the threshold R_TH of each LoD to the header, etc. That is, the three-dimensional data encoding device may also switch the binarization method for each LoD. For example, in the upper layer, due to the long distance between the three-dimensional points, the prediction accuracy deteriorates, and the prediction residual may become larger as a result. Therefore, the three-dimensional data encoding device prevents the sharp increase in the bit length of the binary data by setting the threshold R_TH to be small for the upper layer. In addition, in the lower layer, due to the short distance between the three-dimensional points, the prediction accuracy becomes high, and the prediction residual may become smaller as a result. Therefore, the three-dimensional data encoding device improves the encoding efficiency by setting the threshold R_TH to be large for the hierarchy.
[0690] Fig.72 is a diagram showing an example of an Exp Golomb code, and is a diagram showing the relationship between the value before binarization (multi-value) and the bit after binarization (code). Fig.72 The 0s and 1s shown are reversed.
[0691] In addition, the three-dimensional data encoding device applies arithmetic coding to the binary data of the prediction residual. Thus, the coding efficiency can be improved. In addition, when arithmetic coding is applied, in the binary data, the tendency of the occurrence probability of 0 and 1 of each bit may be different in the part binarized with n bits, that is, the n-bit code (n-bit code) and the part binarized using the exponential Golomb, that is, the remaining code (remaining code). Therefore, the three-dimensional data encoding device can also switch the application method of arithmetic coding through n-bit coding and remaining coding.
[0692] For example, for n-bit coding, the three-dimensional data coding device uses a different coding table (probability table) to perform arithmetic coding on each bit. At this time, the three-dimensional data coding device can also change the number of coding tables used for each bit. For example, the three-dimensional data coding device uses 1 coding table to perform arithmetic coding on the leading bit b0 of the n-bit coding. In addition, the three-dimensional data coding device uses 2 coding tables for the next bit b1. In addition, the three-dimensional data coding device switches the coding table used in the arithmetic coding of the bit b1 according to the value of b0 (0 or 1). Similarly, the three-dimensional data coding device also uses 4 coding tables for the next bit b2. In addition, the three-dimensional data coding device switches the coding table used in the arithmetic coding of the bit b2 according to the values of b0 and b1 (0 to 3).
[0693] Thus, the three-dimensional data encoding device uses 2 when performing arithmetic coding on each bit bn-1 of the n-bit code. n-1 In addition, the three-dimensional data encoding device switches the encoding table to be used according to the value (occurrence pattern) of the bit before bn-1. As a result, the three-dimensional data encoding device can use an appropriate encoding table for each bit, thereby improving encoding efficiency.
[0694] In addition, the three-dimensional data encoding device may also reduce the number of encoding tables used for each bit. For example, when performing arithmetic coding on each bit bn-1, the three-dimensional data encoding device may also switch between two encoding tables according to the value (generation pattern) of the m bits (m<n-1) before bn-1. m The three-dimensional data encoding device can also update the probability of occurrence of 0 and 1 in each coding table according to the value of the binary data actually generated. In addition, the three-dimensional data encoding device can also fix the probability of occurrence of 0 and 1 in the coding table of a part of the bits. Thus, the number of updates of the probability of occurrence can be suppressed, so the amount of processing can be reduced.
[0695] For example, when the n-bit code is b0b1b2…bn-1, there is one coding table for b0 (CTb0). There are two coding tables for b1 (CTb10, CTb11). In addition, the coding table used is switched according to the value of b0 (0 to 1). There are four coding tables for b2 (CTb20, CTb21, CTb22, CTb23). In addition, the coding table used is switched according to the values of b0 and b1 (0 to 3). There are 2 coding tables for bn-1. n-1 (CTbn0, CTbn1, ..., CTbn(2 n-1 -1)). In addition, according to the value of b0b1…bn-2 (0~2 n-1 -1) to switch the encoding table used.
[0696] In addition, the three-dimensional data encoding device may also set 0 to 2 instead of binarizing the n-bit code. n m-ary arithmetic coding of the value -1 (m=2 n ). In addition, when the three-dimensional data encoding device performs arithmetic encoding on the n-bit code using m-ary, the three-dimensional data decoding device can also restore the n-bit code by arithmetic decoding of m-ary.
[0697] Fig.73 2 is a diagram for explaining the processing when the residual code is an exponential Golomb code, for example. Fig.73 As shown, the portion binarized using the exponential Golomb, i.e., the remaining code, includes the prefix part and the suffix part. For example, the three-dimensional data encoding device switches the coding table between the prefix part and the suffix part. That is, the three-dimensional data encoding device uses the coding table for the prefix to perform arithmetic coding on each bit included in the prefix part, and uses the coding table for the suffix to perform arithmetic coding on each bit included in the suffix part.
[0698] In addition, the three-dimensional data encoding device may also update the occurrence probabilities of 0 and 1 in each coding table according to the value of the binary data actually generated. Alternatively, the three-dimensional data encoding device may also fix the occurrence probabilities of 0 and 1 in a certain coding table. Thus, the number of updates of the occurrence probability can be suppressed, thereby reducing the amount of processing. For example, the three-dimensional data encoding device may also update the occurrence probability for the prefix part and fix the occurrence probability for the suffix part.
[0699] In addition, the three-dimensional data encoding device decodes the quantized prediction residual by inverse quantization and reconstruction, and uses the decoded prediction residual, i.e., the decoded value, for subsequent prediction of the three-dimensional point of the encoding object. Specifically, the three-dimensional data encoding device calculates the inverse quantization value by multiplying the quantized prediction residual (quantization value) by the quantization scale, and adds the inverse quantization value and the prediction value to obtain the decoded value (reconstruction value).
[0700] For example, the inverse quantization value a2iq of point a2 is calculated by (Formula A11) using the quantization value a2q of point a2. The inverse quantization value b2iq of point b2 is calculated by (Formula A12) using the quantization value b2q of point b2. Here, QS_LoD0 is the QS for LoD0, and QS_LoD1 is the QS for LoD1. That is, QS can be changed according to LoD.
[0701] a2iq=a2q×QS_LoD0…(Formula A11)
[0702] b2iq=b2q×QS_LoD1…(Formula A12)
[0703] For example, as shown in (Formula A13), the decoded value a2rec of point a2 is calculated by adding the predicted value a2p of point a2 to the inverse quantized value a2iq of point a2. As shown in (Formula A14), the decoded value b2rec of point b2 is calculated by adding the predicted value b2p of point b2 to the inverse quantized value b2iq of point b2.
[0704] a2rec=a2iq+a2p…(Formula A13)
[0705] b2rec=b2iq+b2p…(Formula A14)
[0706] Hereinafter, an example of the syntax of the bit stream according to the present embodiment will be described. Fig.74 1 is a diagram showing an example of the syntax of the attribute header (attribute_header) of this embodiment. The attribute header is the header information of the attribute information. Fig.74 As shown, the attribute header includes the number of layers information (NumLoD), three-dimensional point information (NumOfPoint[i]), layer threshold (Thres_Lod[i]), surrounding point information (NumNeighorPoint[i]), prediction threshold (THd[i]), quantization scale (QS[i]), and binarization threshold (R_TH[i]).
[0707] The number of layers information (NumLoD) indicates the number of layers of LoD used.
[0708] The three-dimensional point number information (NumOfPoint[i]) represents the number of three-dimensional points belonging to level i. In addition, the three-dimensional data encoding device may also attach the three-dimensional point total number information (AllNumOfPoint) representing the total number of three-dimensional points to another header. In this case, the three-dimensional data encoding device may not attach NumOfPoint[NumLoD - 1], which represents the number of three-dimensional points belonging to the bottommost level, to the header. In this case, the three-dimensional data decoding device can calculate NumOfPoint[NumLoD - 1] by (Equation A15). Thereby, the encoding amount of the header can be reduced.
[0709]
Equation 4
[0710]
[0711] The level threshold (Thres_Lod[i]) is a threshold for setting level i. The three-dimensional data encoding device and the three-dimensional data decoding device form LoDi such that the distance between each point within LoDi is greater than the threshold Thres_LoD[i]. In addition, the three-dimensional data encoding device may not attach the value of Thres_Lod[NumLoD - 1] (the bottommost level) to the header. In this case, the three-dimensional data decoding device estimates the value of Thres_Lod[NumLoD - 1] as 0. Thereby, the encoding amount of the header can be reduced.
[0712] The surrounding point number information (NumNeighorPoint[i]) represents the upper limit value of the number of surrounding points used in the generation of the predicted value of the three-dimensional points belonging to level i. When the number of surrounding points M is less than NumNeighorPoint[i] (M < NumNeighorPoint[i]), the three-dimensional data encoding device may also use M surrounding points to calculate the predicted value. In addition, when it is not necessary to separate the value of NumNeighorPoint[i] in each LoD, the three-dimensional data encoding device may attach one surrounding point number information (NumNeighorPoint) used in all LoDs to the header.
[0713] The prediction threshold (THd[i]) represents the upper limit value of the distance between the surrounding three-dimensional points and the object three-dimensional point used in the prediction of the object three-dimensional point to be encoded or decoded at level i. The three-dimensional data encoding device and the three-dimensional data decoding device do not use the three-dimensional points whose distance from the object three-dimensional point is farther than THd[i] for prediction. In addition, when it is not necessary to separate the value of THd[i] in each LoD, the three-dimensional data encoding device may attach one prediction threshold (THd) used in all LoDs to the header.
[0714] The quantization scale (QS[i]) represents the quantization scale used in quantization and inverse quantization of level i.
[0715] The binarization threshold (R_TH[i]) is a threshold for switching the binarization method of the prediction residual of the three-dimensional point belonging to the layer i. For example, when the prediction residual is less than the threshold R_TH, the three-dimensional data encoding device binarizes the prediction residual pu with a fixed number of bits, and when the prediction residual is greater than the threshold R_TH, the binarization data of the threshold R_TH and the value of (pu-R_TH) are binarized using the exponential Golomb. In addition, when there is no need to switch the value of R_TH[i] in each LoD, the three-dimensional data encoding device can also attach a binarization threshold (R_TH) used in all LoDs to the header.
[0716] In addition, R_TH[i] may be a maximum value represented by nbit. For example, in 6bit, R_TH is 63, and in 8bit, R_TH is 255. In addition, the three-dimensional data encoding device may encode the number of bits instead of encoding the maximum value represented by nbit as the binarization threshold. For example, the three-dimensional data encoding device may append the value 6 to the header when R_TH[i]=63, and append the value 8 to the header when R_TH[i]=255. In addition, the three-dimensional data encoding device may define the minimum value (minimum number of bits) of the number of bits representing R_TH[i], and append the relative number of bits relative to the minimum value to the header. For example, the three-dimensional data encoding device may append the value 0 to the header when R_TH[i]=63 and the minimum number of bits is 6, and append the value 2 to the header when R_TH[i]=255 and the minimum number of bits is 6.
[0717] In addition, the three-dimensional data encoding device may entropy encode at least one of NumLoD, Thres_Lod[i], NumNeighborPoint[i], THd[i], QS[i], and R_TH[i] and append it to the header. For example, the three-dimensional data encoding device may binarize each value and perform arithmetic encoding. In addition, in order to reduce the amount of processing, the three-dimensional data encoding device may encode each value with a fixed length.
[0718] In addition, the three-dimensional data encoding device may not add at least one of NumLoD, Thres_Lod[i], NumNeighborPoint[i], THd[i], QS[i], and R_TH[i] to the header. For example, the value of at least one of them may also be specified by a profile or level of a standard. In this way, the bit amount of the header can be reduced.
[0719] Fig.75 1 is a diagram showing a syntax example of attribute data (attribute_data) according to the present embodiment. The attribute data includes encoded data of attribute information of a plurality of three-dimensional points. Fig.75 As shown, the attribute data includes an n-bit code and a remaining code.
[0720] The n-bit code is the coded data of the prediction residual of the value of the attribute information or a part thereof. The bit length of the n-bit code depends on the value of R_TH[i]. For example, when the value shown in R_TH[i] is 63, the n-bit code is 6 bits, and when the value shown in R_TH[i] is 255, the n-bit code is 8 bits.
[0721] The remaining code is the coded data after exponential Golomb coding in the coded data of the prediction residual of the value of the attribute information. When the n-bit code is the same as R_TH[i], the remaining code is encoded or decoded. In addition, the three-dimensional data decoding device adds the value of the n-bit code and the value of the remaining code to decode the prediction residual. In addition, when the n-bit code is not the same value as R_TH[i], the remaining code may not be encoded or decoded.
[0722] The following describes the flow of processing in the three-dimensional data encoding device. Fig.76 This is a flowchart of a three-dimensional data encoding process performed by a three-dimensional data encoding device.
[0723] First, the three-dimensional data encoding device encodes position information (geometry) (S3001). For example, the three-dimensional data encoding is performed using an octree representation.
[0724] After encoding the position information, the three-dimensional data encoding device reallocates the original three-dimensional point's attribute information to the changed three-dimensional point when the position of the three-dimensional point changes due to quantization or the like (S3002). For example, the three-dimensional data encoding device reallocates the attribute information by interpolating the value of the attribute information according to the amount of change in the position. For example, the three-dimensional data encoding device detects N three-dimensional points before the change that are close to the changed three-dimensional position, and performs weighted averaging on the values of the attribute information of the N three-dimensional points. For example, in the weighted averaging, the three-dimensional data encoding device determines the weight based on the distance from the changed three-dimensional position to each of the N three-dimensional points. Then, the three-dimensional data encoding device determines the value obtained by the weighted averaging as the value of the attribute information of the changed three-dimensional point. In addition, when two or more three-dimensional points change to the same three-dimensional position due to quantization or the like, the three-dimensional data encoding device may also allocate the average value of the attribute information of the two or more three-dimensional points before the change as the value of the attribute information of the changed three-dimensional point.
[0725] Next, the three-dimensional data encoding device encodes the reallocated attribute information (Attribute) (S3003). For example, in the case of encoding multiple attribute information, the three-dimensional data encoding device may also encode the multiple attribute information in sequence. For example, in the case of encoding color and reflectivity as attribute information, the three-dimensional data encoding device may also generate a bit stream in which the encoding result of reflectivity is attached after the encoding result of color. In addition, the order of the multiple encoding results of the attribute information attached to the bit stream is not limited to this order, and can be any order.
[0726] In addition, the three-dimensional data encoding device may also attach information indicating the starting position of the encoded data of each attribute information in the bit stream to the header, etc. Thus, the three-dimensional data decoding device can selectively decode the attribute information that needs to be decoded, and thus can omit the decoding process of the attribute information that does not need to be decoded. Therefore, the processing amount of the three-dimensional data decoding device can be reduced. In addition, the three-dimensional data encoding device may also encode multiple attribute information in parallel and merge the encoding results into one bit stream. Thus, the three-dimensional data encoding device can encode multiple attribute information at high speed.
[0727] Fig.77 3003 is a flowchart of the attribute information encoding process. First, the three-dimensional data encoding device sets the LoD (S3011). That is, the three-dimensional data encoding device assigns each three-dimensional point to any one of a plurality of LoDs.
[0728] Next, the three-dimensional data encoding device starts a loop in units of LoDs (S3012). That is, the three-dimensional data encoding device repeatedly performs the processing of steps S3013 to S3021 for each LoD.
[0729] Next, the three-dimensional data encoding device starts a loop in three-dimensional point units (S3013). That is, the three-dimensional data encoding device repeatedly performs the processing of steps S3014 to S3020 for each three-dimensional point.
[0730] First, the three-dimensional data encoding device searches for three-dimensional points that are used in calculating the predicted value of the object three-dimensional point of the processing object, that is, multiple surrounding points (S3014). Next, the three-dimensional data encoding device calculates the weighted average of the values of the attribute information of the multiple surrounding points, and sets the obtained value as the predicted value P (S3015). Next, the three-dimensional data encoding device calculates the difference between the attribute information of the object three-dimensional point and the predicted value, that is, the prediction residual (S3016). Next, the three-dimensional data encoding device calculates a quantized value by quantizing the prediction residual (S3017). Next, the three-dimensional data encoding device performs arithmetic coding on the quantized value (S3018).
[0731] In addition, the three-dimensional data encoding device calculates an inverse quantization value by inverse quantizing the quantization value (S3019). Next, the three-dimensional data encoding device generates a decoded value by adding a prediction value to the inverse quantization value (S3020). Next, the three-dimensional data encoding device ends the loop of the three-dimensional point unit (S3021). In addition, the three-dimensional data encoding device ends the loop of the LoD unit (S3022).
[0732] Hereinafter, a three-dimensional data decoding process in a three-dimensional data decoding device for decoding a bit stream generated by the three-dimensional data encoding device described above will be described.
[0733] The three-dimensional data decoding device generates decoded binary data by performing arithmetic decoding on the binary data of the attribute information in the bit stream generated by the three-dimensional data encoding device in the same method as the three-dimensional data encoding device. In addition, in the three-dimensional data encoding device, when the application method of arithmetic coding is switched between the part binarized by n bits (n-bit coding) and the part binarized by exponential Golomb (residual coding), the three-dimensional data decoding device performs decoding in accordance with the arithmetic decoding when the arithmetic decoding is applied.
[0734] For example, in an arithmetic decoding method for n-bit coding, a three-dimensional data decoding device uses a different coding table (decoding table) to perform arithmetic decoding on each bit. At this time, the three-dimensional data decoding device may also change the number of coding tables used for each bit. For example, one coding table is used to perform arithmetic decoding on the leading bit b0 of the n-bit coding. In addition, the three-dimensional data decoding device uses two coding tables for the next bit b1. In addition, the three-dimensional data decoding device switches the coding table used in the arithmetic decoding of the bit b1 according to the value of b0 (0 or 1). Similarly, the three-dimensional data decoding device further uses four coding tables for the next bit b2. In addition, the three-dimensional data decoding device switches the coding table used in the arithmetic decoding of the bit b2 according to the values of b0 and b1 (0 to 3).
[0735] Thus, the three-dimensional data decoding device uses 2 when performing arithmetic decoding on each bit bn-1 of the n-bit code. n-1 In addition, the three-dimensional data decoding device switches the coding table to be used according to the value (occurrence pattern) of the bit before bn-1. As a result, the three-dimensional data decoding device can use an appropriate coding table for each bit to appropriately decode the bit stream with improved coding efficiency.
[0736] In addition, the three-dimensional data decoding device may also reduce the number of coding tables used for each bit. For example, when performing arithmetic decoding on each bit bn-1, the three-dimensional data decoding device may switch between two encoding tables according to the value (occurrence pattern) of the m bits (m<n-1) before bn-1. m The three-dimensional data decoding device can appropriately decode the bit stream with improved coding efficiency while suppressing the number of coding tables used in each bit. In addition, the three-dimensional data decoding device can also update the occurrence probability of 0 and 1 in each coding table according to the value of the binary data actually generated. In addition, the three-dimensional data decoding device can also fix the occurrence probability of 0 and 1 in the coding table of a part of the bits. In this way, the number of updates of the occurrence probability can be suppressed, so the processing amount can be reduced.
[0737] For example, when the n-bit code is b0b1b2…bn-1, there is one coding table for b0 (CTb0). There are two coding tables for b1 (CTb10, CTb11). In addition, the coding table is switched according to the value of b0 (0 to 1). There are four coding tables for b2 (CTb20, CTb21, CTb22, CTb23). In addition, the coding table is switched according to the values of b0 and b1 (0 to 3). There are 2 coding tables for bn-1. n-1 (CTbn0, CTbn1, ..., CTbn(2 n-1 -1)). In addition, according to the value of b0b1…bn-2 (0~2 n-1-1) to switch the encoding table.
[0738] For example, Fig.78 2 is a diagram for explaining the processing when the residual code is an exponential Golomb code. Fig.78 As shown, the part (remaining code) that the three-dimensional data encoding device binarizes using the exponential Golomb and encodes includes the prefix part and the suffix part. For example, the three-dimensional data decoding device switches the encoding table between the prefix part and the suffix part. That is, the three-dimensional data decoding device uses the encoding table for the prefix to perform arithmetic decoding on each bit included in the prefix part, and uses the encoding table for the suffix to perform arithmetic decoding on each bit included in the suffix part.
[0739] In addition, the three-dimensional data decoding device may also update the occurrence probabilities of 0 and 1 in each coding table according to the value of the binary data generated during decoding. Alternatively, the three-dimensional data decoding device may also fix the occurrence probabilities of 0 and 1 in a certain coding table. Thus, the number of updates of the occurrence probabilities can be suppressed, thereby reducing the amount of processing. For example, the three-dimensional data decoding device may also update the occurrence probability for the prefix part and fix the occurrence probability for the suffix part.
[0740] In addition, the three-dimensional data decoding device converts the binary data of the prediction residual obtained by arithmetic decoding into multiple values in accordance with the encoding method used in the three-dimensional data encoding device, thereby decoding the quantized prediction residual (unsigned integer value). The three-dimensional data decoding device first calculates the value of the n-bit code decoded by arithmetic decoding the binary data of the n-bit code. Then, the three-dimensional data decoding device compares the value of the n-bit code with the value of R_TH.
[0741] When the value of the n-bit code is consistent with the value of R_TH, the three-dimensional data decoding device determines that there is a bit encoded by Exponential Golomb next, and performs arithmetic decoding on the binary data encoded by Exponential Golomb, that is, the residual code. Then, the three-dimensional data decoding device calculates the value of the residual code based on the decoded residual code using a back-calculation table indicating the relationship between the residual code and the value. Fig.79 : is a diagram showing an example of a back-estimation table showing the relationship between the residual code and its value. Next, the three-dimensional data decoding apparatus adds the obtained residual code value to R_TH to obtain a multi-valued quantized prediction residual.
[0742] On the other hand, when the value of the n-bit code is inconsistent with the value of R_TH (the value is smaller than R_TH), the three-dimensional data decoding device directly determines the value of the n-bit code as the prediction residual after quantization. Thus, the three-dimensional data decoding device can appropriately decode the bit stream generated by switching the binarization method according to the value of the prediction residual in the three-dimensional data encoding device.
[0743] In addition, when the threshold R_TH is added to the header of the bit stream, the three-dimensional data decoding device may decode the value of the threshold R_TH from the header and use the decoded value of the threshold R_TH to switch the decoding method. In addition, when the threshold R_TH is added to the header for each LoD, the three-dimensional data decoding device uses the decoded threshold R_TH to switch the decoding method for each LoD.
[0744] For example, when the threshold R_TH is 63 and the decoded n-bit code value is 63, the three-dimensional data decoding device decodes the remaining code using the Exponential Golomb method to obtain the value of the remaining code. Fig.79 In the example shown, the remaining code is 00100, and the remaining code value is 3. Next, the three-dimensional data decoding device adds the value 63 of the threshold value R_TH and the value 3 of the remaining code to obtain a prediction residual value 66.
[0745] Furthermore, when the decoded n-bit code value is 32, the three-dimensional data decoding apparatus sets the n-bit code value 32 as the value of the prediction residual.
[0746] In addition, the three-dimensional data decoding device converts the decoded quantized prediction residual from an unsigned integer value to a signed integer value by, for example, processing opposite to the processing in the three-dimensional data encoding device. Thus, the three-dimensional data decoding device can appropriately decode a bit stream generated without considering the generation of negative integers when entropy encoding the prediction residual. In addition, the three-dimensional data decoding device does not necessarily need to convert the unsigned integer value to a signed integer value, and can also decode the sign bit when decoding a bit stream generated by separately entropy encoding the sign bit.
[0747] The three-dimensional data decoding device decodes the quantized prediction residual transformed into a signed integer value through inverse quantization and reconstruction, thereby generating a decoded value. In addition, the three-dimensional data decoding device uses the generated decoded value for subsequent prediction of the three-dimensional point of the decoding object. Specifically, the three-dimensional data decoding device calculates the inverse quantization value by multiplying the quantized prediction residual by the decoded quantization scale, and adds the inverse quantization value and the prediction value to obtain the decoded value.
[0748] The decoded unsigned integer value (unsigned quantized value) is converted into a signed integer value by the following processing. When the LSB (least significant bit) of the decoded unsigned integer value a2u is 1, the three-dimensional data decoding device sets the signed integer value a2q to -((a2u+1)>>1). When the LSB of the unsigned integer value a2u is not 1, the three-dimensional data decoding device sets the signed integer value a2q to (a2u>>1).
[0749] Similarly, when the LSB of the decoded unsigned integer value b2u is 1, the three-dimensional data decoding device sets the signed integer value b2q to -((b2u+1)>>1). When the LSB of the unsigned integer value n2u is not 1, the three-dimensional data decoding device sets the signed integer value b2q to (b2u>>1).
[0750] Note that details of the inverse quantization and reconstruction processing performed by the three-dimensional data decoding device are the same as those of the inverse quantization and reconstruction processing in the three-dimensional data encoding device.
[0751] The following describes the flow of processing in the three-dimensional data decoding device. Fig.80 3D data decoding processing performed by a 3D data decoding device. First, the 3D data decoding device decodes position information (geometry) from a bit stream (S3031). For example, the 3D data decoding device uses an octree representation for decoding.
[0752] Next, the three-dimensional data decoding device decodes the attribute information (Attribute) from the bit stream (S3032). For example, in the case of decoding multiple types of attribute information, the three-dimensional data decoding device may also decode the multiple types of attribute information in sequence. For example, in the case of decoding color and reflectivity as attribute information, the three-dimensional data decoding device decodes the encoding result of color and the encoding result of reflectivity in the order in which they are attached to the bit stream. For example, in the bit stream, in the case where the encoding result of reflectivity is attached after the encoding result of color, the three-dimensional data decoding device decodes the encoding result of color and then decodes the encoding result of reflectivity. In addition, the three-dimensional data decoding device may decode the encoding results of the attribute information attached to the bit stream in any order.
[0753] In addition, the three-dimensional data decoding device can also obtain information indicating the starting position of the coded data of each attribute information in the bit stream by decoding the header. As a result, the three-dimensional data decoding device can selectively decode the attribute information that needs to be decoded, so that the decoding process of the attribute information that does not need to be decoded can be omitted. Therefore, the processing amount of the three-dimensional data decoding device can be reduced. In addition, the three-dimensional data decoding device can also decode multiple attribute information in parallel and merge the decoding results into one three-dimensional point group. As a result, the three-dimensional data decoding device can decode multiple attribute information at high speed.
[0754] Fig.81 3041). That is, the 3D data decoding device assigns a plurality of 3D points having decoded position information to any one of a plurality of LoDs. For example, the assignment method is the same as the assignment method used in the 3D data encoding device.
[0755] Next, the three-dimensional data decoding device starts a loop in units of LoDs (S3042). That is, the three-dimensional data decoding device repeatedly performs the processing of steps S3043 to S3049 for each LoD.
[0756] Next, the three-dimensional data decoding device starts a three-dimensional point unit loop (S3043). That is, the three-dimensional data decoding device repeatedly performs the processing of steps S3044 to S3048 for each three-dimensional point.
[0757] First, the three-dimensional data decoding device searches for three-dimensional points that are used in calculating the predicted value of the object three-dimensional point of the processing object, that is, multiple surrounding points (S3044). Next, the three-dimensional data decoding device calculates the weighted average of the values of the attribute information of the mult...
Claims
1. A three-dimensional data encoding method is a three-dimensional data encoding method for encoding three-dimensional points, comprising: determining whether to use the first three-dimensional point to predict attribute information of the second three-dimensional point according to a distance between the first three-dimensional point and the second three-dimensional point; as well as When the first three-dimensional point is within a predetermined distance range from the second three-dimensional point, a predicted value of the attribute information of the second three-dimensional point is calculated based on the first three-dimensional point.
2. The three-dimensional data encoding method according to claim 1, wherein: When the first three-dimensional point is not within a predetermined distance range from the second three-dimensional point, it is determined that the first three-dimensional point is not to be used.
3. The three-dimensional data encoding method according to claim 1, wherein: Also includes: Calculating a prediction residual which is a difference between the attribute information of the second three-dimensional point and the predicted value; Generate binary data by binarizing the prediction residual; as well as The binary data is arithmetically encoded.
4. A three-dimensional data decoding method is a three-dimensional data decoding method for decoding three-dimensional points, comprising: determining whether to use the first three-dimensional point to predict attribute information of the second three-dimensional point according to a distance between the first three-dimensional point and the second three-dimensional point; as well as When the first three-dimensional point is within a predetermined distance range from the second three-dimensional point, a predicted value of the attribute information of the second three-dimensional point is calculated based on the first three-dimensional point.
5. The three-dimensional data decoding method according to claim 4, wherein: When the first three-dimensional point is not within a predetermined distance range from the second three-dimensional point, it is determined that the first three-dimensional point is not to be used.
6. The three-dimensional data decoding method according to claim 4, wherein: Also includes: generating binary data by arithmetically decoding the encoded data; Generate prediction residuals by multi-valuedizing the binary data; as well as The attribute information of the second three-dimensional point is calculated by adding the predicted value to the prediction residual.
7. A three-dimensional data encoding device is a three-dimensional data encoding device for encoding three-dimensional points, wherein: have: processor; as well as Memory, The processor uses the memory, determining whether to use the first three-dimensional point to predict attribute information of the second three-dimensional point according to a distance between the first three-dimensional point and the second three-dimensional point; as well as When the first three-dimensional point is within a predetermined distance range from the second three-dimensional point, a predicted value of the attribute information of the second three-dimensional point is calculated based on the first three-dimensional point.
8. A three-dimensional data decoding device is a three-dimensional data decoding device for decoding three-dimensional points, wherein: have: Processor; and Memory, The processor uses the memory, determining whether to use the first three-dimensional point to predict attribute information of the second three-dimensional point according to a distance between the first three-dimensional point and the second three-dimensional point; as well as When the first three-dimensional point is within a predetermined distance range from the second three-dimensional point, a predicted value of the attribute information of the second three-dimensional point is calculated based on the first three-dimensional point.
Citation Information
Patent Citations
Map display device
WO2014020663A1