Encoding method, decoding method, encoding device, decoding device and computer program product

By selecting three-dimensional points used to predict the attribute information of three-dimensional point and using Morton code to optimize the predicted value generation, the problem of low encoding efficiency of three-dimensional data is solved, and more efficient data compression is achieved.

CN120223904APending Publication Date: 2025-06-27PANASONIC INTELLECTUAL PROPERTY CORP OF AMERICA
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510372317.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2019-02-28
Filing Date
2020-02-28
Publication Date
2025-06-27

AI Technical Summary

Technical Problem

The prior art has low encoding efficiency in three-dimensional data encoding, making it difficult to effectively compress large-scale point cloud data.

Method used

By selecting a three-dimensional point for predicting the prediction value generation of three-dimensional point attribute information, a plurality of three-dimensional point candidates are evaluated, and a three-dimensional point set for the prediction value generation is determined whether to include it in the set of three-dimensional points for predicting value generation, and a three-dimensional point for predicting value generation is selected using the Morton code.

Benefits of technology

The efficiency of three-dimensional data encoding is improved, and by selecting three-dimensional points close to the encoding object as candidates for predicted value generation, the increase in the amount of data is reduced, thereby improving the encoding efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120223904A_ABST
    Figure CN120223904A_ABST
Patent Text Reader

Abstract

The invention relates to an encoding method, a decoding method, an encoding device, a decoding device and a computer program product. The encoding method is executed by an encoding device that selects a three-dimensional point for prediction value generation for predicting attribute information of a three-dimensional point to be predicted, and includes evaluating at least one of a plurality of three-dimensional point candidates, and generating a prediction value based on a result of the evaluation. It is determined whether or not the evaluated three-dimensional point candidate is included in a set of three-dimensional points for prediction value generation comprising N three-dimensional points, and at least one three-dimensional point in the set of three-dimensional points for prediction value generation is selected using a Morton code assigned to the three-dimensional point candidate.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] This application is a divisional of the patent application with the application date of February 28, 2020, application number 202080016358.7, and invention title "3D Data Encoding Method, 3D Data Decoding Method, 3D Data Encoding Apparatus, and 3D Data Decoding Apparatus". Technical Field

[0002] The present disclosure relates to a 3D data encoding method, a 3D data decoding method, a 3D data encoding apparatus, and a 3D data decoding apparatus. Background Art

[0003] In large fields such as computer vision, map information, monitoring, infrastructure inspection, or video distribution for autonomous operation of automobiles or robots, devices or services that make flexible use of 3D data will be popularized in the future. 3D data is obtained by various methods such as distance sensors such as rangefinders, stereo cameras, or combinations of multiple monocular cameras.

[0004] As a representation method of 3D data, there is a representation method called point cloud, which represents the shape of a 3D structure by a point group in 3D space. The position and color of the point group are stored in the point cloud. Although it is expected that the point cloud will become the mainstream as a representation method of 3D data, the data volume of the point group is very large. Therefore, in the storage or transmission of 3D data, like 2D moving images (as an example, MPEG-4 AVC or HEVC standardized by MPEG), data volume compression needs to be performed by encoding.

[0005] In addition, for the compression of point clouds, some are supported by publicly available libraries (PointCloud Library) that perform point cloud correlation processing.

[0006] In addition, there is a well-known technique that uses 3D map data to retrieve facilities around a vehicle and display them (for example, refer to Patent Document 1).

[0007] Prior Art Documents

[0008] Patent Documents

[0009] Patent Document 1 International Publication No. 2014 / 020663 Summary of the Invention

[0010] Problems to be Solved by the Invention

[0011] It is desired to improve the encoding efficiency in the encoding of 3D data.

[0012] An object of the present disclosure is to provide a three-dimensional data encoding method, a three-dimensional data decoding method, a three-dimensional data encoding apparatus, or a three-dimensional data decoding apparatus that can improve encoding efficiency.

[0013] Means for solving the problem

[0014] An encoding method according to one aspect of the present disclosure is an encoding method executed by an encoding apparatus that selects three-dimensional points for generating a predicted value of attribute information of a three-dimensional point to be predicted. At least one of a plurality of three-dimensional point candidates is evaluated, and based on the evaluation result, it is determined whether to include the evaluated three-dimensional point candidate in a set of three-dimensional points for generating a predicted value composed of N three-dimensional points. At least one three-dimensional point in the set of three-dimensional points for generating a predicted value is selected using the Morton code assigned to the three-dimensional point candidate.

[0015] In addition, a decoding method according to one aspect of the present disclosure is a decoding method executed by a decoding apparatus that selects three-dimensional points for generating a predicted value of attribute information of a three-dimensional point to be predicted. At least one of a plurality of three-dimensional point candidates is evaluated, and based on the evaluation result, it is determined whether to include the evaluated three-dimensional point candidate in a set of three-dimensional points for generating a predicted value composed of N three-dimensional points. At least one three-dimensional point in the set of three-dimensional points for generating a predicted value is selected using the Morton code assigned to the three-dimensional point candidate.

[0016] In addition, an encoding apparatus according to one aspect of the present disclosure selects three-dimensional points for generating a predicted value of attribute information of a three-dimensional point to be predicted, and includes: a processor; and a memory. The processor uses the memory to evaluate at least one of a plurality of three-dimensional point candidates, and based on the evaluation result, determines whether to include the evaluated three-dimensional point candidate in a set of three-dimensional points for generating a predicted value composed of N three-dimensional points, and selects at least one three-dimensional point in the set of three-dimensional points for generating a predicted value using the Morton code assigned to the three-dimensional point candidate.

[0017] In addition, a decoding apparatus according to one aspect of the present disclosure selects three-dimensional points for generating a predicted value of attribute information of a three-dimensional point to be predicted, and includes: a processor; and a memory. The processor uses the memory to evaluate at least one of a plurality of three-dimensional point candidates, and based on the evaluation result, determines whether to include the evaluated three-dimensional point candidate in a set of three-dimensional points for generating a predicted value composed of N three-dimensional points, and selects at least one three-dimensional point in the set of three-dimensional points for generating a predicted value using the Morton code assigned to the three-dimensional point candidate.

[0018] In addition, a computer program product according to one aspect of the present disclosure is used to execute the above encoding method or decoding method.

[0019] In addition, these general or specific forms can be implemented by a system, a method, an integrated circuit, a computer program, or a recording medium such as a computer-readable CD-ROM, and can be implemented by any combination of a system, a method, an integrated circuit, a computer program, and a recording medium.

[0020] Advantages of the Invention

[0021] The present disclosure can provide a three-dimensional data encoding method, a three-dimensional data decoding method, a three-dimensional data encoding device, or a three-dimensional data decoding device that can improve encoding efficiency. BRIEF DESCRIPTION OF THE DRAWINGS

[0022] Figure 1 Shows the configuration of encoded three-dimensional data in Embodiment 1.

[0023] Figure 2 Shows an example of the prediction structure between SPCs belonging to the lowest layer of the GOS in Embodiment 1.

[0024] Figure 3 Shows an example of the inter-layer prediction structure in Embodiment 1.

[0025] Figure 4 Shows an example of the encoding order of the GOS in Embodiment 1.

[0026] Figure 5 Shows an example of the encoding order of the GOS in Embodiment 1.

[0027] Figure 6 Is a block diagram of the three-dimensional data encoding device in Embodiment 1.

[0028] Figure 7 Is a flowchart of the encoding process in Embodiment 1.

[0029] Figure 8 Is a block diagram of the three-dimensional data decoding device in Embodiment 1.

[0030] Fig. 9 Is a flowchart of the decoding process in Embodiment 1.

[0031] Fig.10 Shows an example of the meta-information in Embodiment 1.

[0032] Fig.11 Shows a configuration example of the SWLD in Embodiment 2.

[0033] Fig.12 Shows a working example of the server and the client in Embodiment 2.

[0034] Fig.13 Shows a working example of the server and the client in Embodiment 2.

[0035] Fig.14 Shows an operation example of the server and the client in Embodiment 2.

[0036] Fig.15 Shows an operation example of the server and the client in Embodiment 2.

[0037] Fig.16 Is a block diagram of the three-dimensional data encoding device in Embodiment 2.

[0038] Fig.17 Is a flowchart of the encoding process in Embodiment 2.

[0039] Fig.18 Is a block diagram of the three-dimensional data decoding device in Embodiment 2.

[0040] Fig.19 Is a flowchart of the decoding process in Embodiment 2.

[0041] Fig. 20 Shows a configuration example of the WLD in Embodiment 2.

[0042] Fig.21 Shows an example of the octree structure of the WLD in Embodiment 2.

[0043] Fig. 22 Shows a configuration example of the SWLD in Embodiment 2.

[0044] Fig.23 Shows an example of the octree structure of the SWLD in Embodiment 2.

[0045] Fig.24 Is a block diagram of the three-dimensional data production device in Embodiment 3.

[0046] Fig.25 Is a block diagram of the three-dimensional data transmission device in Embodiment 3.

[0047] Fig.26 Is a block diagram of the three-dimensional information processing device in Embodiment 4.

[0048] Fig. 27 Is a block diagram of the three-dimensional data production device in Embodiment 5.

[0049] Fig.28 Shows the configuration of the system in Embodiment 6.

[0050] Fig.29 Is a block diagram of the client device in Embodiment 6.

[0051] Fig.30 Is a block diagram of the server in Embodiment 6.

[0052] Fig.31 It is a flowchart of three-dimensional data creation processing performed by the client device of Embodiment 6.

[0053] Fig.32 It is a flowchart of sensor information transmission processing performed by the client device of Embodiment 6.

[0054] Fig.33 It is a flowchart of three-dimensional data creation processing performed by the server of Embodiment 6.

[0055] Fig.34 It is a flowchart of three-dimensional map transmission processing performed by the server of Embodiment 6.

[0056] Fig.35 It shows the configuration of a modified example of the system of Embodiment 6.

[0057] Fig.36 It shows the configurations of the server and the client device of Embodiment 6.

[0058] Fig.37 It is a block diagram of the three-dimensional data encoding device of Embodiment 7.

[0059] Fig.38 It shows an example of the prediction residual of Embodiment 7.

[0060] Fig.39 It shows an example of the volume of Embodiment 7.

[0061] Fig.40 It shows an example of the octree representation of the volume of Embodiment 7.

[0062] Fig.41 It shows an example of the bit string of the volume of Embodiment 7.

[0063] Fig.42 It shows an example of the octree representation of the volume of Embodiment 7.

[0064] Fig.43 It shows an example of the volume of Embodiment 7.

[0065] Fig.44 It is a diagram for explaining the intra prediction processing of Embodiment 7.

[0066] Fig.45 It is a diagram for explaining the rotation and translation processing of Embodiment 7.

[0067] Fig.46 It shows an example of the syntax of the RT application flag and RT information of Embodiment 7.

[0068] Fig.47It is a diagram for explaining the inter-frame prediction processing of Embodiment 7.

[0069] Fig.48 It is a block diagram of the three-dimensional data decoding device of Embodiment 7.

[0070] Fig.49 It is a flowchart of the three-dimensional data encoding processing performed by the three-dimensional data encoding device of Embodiment 7.

[0071] Fig.50 It is a flowchart of the three-dimensional data decoding processing performed by the three-dimensional data decoding device of Embodiment 7.

[0072] Fig.51 It is a diagram showing an example of three-dimensional points of Embodiment 8.

[0073] Fig.52 It is a diagram showing a setting example of LoD of Embodiment 8.

[0074] Fig.53 It is a diagram showing an example of a threshold value used in the setting of LoD of Embodiment 8.

[0075] Fig.54 It is a diagram showing an example of attribute information used in the predicted value of Embodiment 8.

[0076] Fig.55 It is a diagram showing an example of the exponential Golomb code of Embodiment 8.

[0077] Fig.56 It is a diagram showing the processing for the exponential Golomb code of Embodiment 8.

[0078] Fig.57 It is a diagram showing a syntax example of the attribute header of Embodiment 8.

[0079] Fig.58 It is a diagram showing a syntax example of the attribute data of Embodiment 8.

[0080] Fig.59 It is a flowchart of the three-dimensional data encoding processing of Embodiment 8.

[0081] Fig.60 It is a flowchart of the attribute information encoding processing of Embodiment 8.

[0082] Fig.61 It is a diagram showing the processing for the exponential Golomb code of Embodiment 8.

[0083] Fig.62 It is a diagram showing an example of a reverse table representing the relationship between the residual coding and its value of Embodiment 8.

[0084] Fig.63It is a flowchart of the three-dimensional data decoding process of Embodiment 8.

[0085] Fig.64 It is a flowchart of the attribute information decoding process of Embodiment 8.

[0086] Fig.65 It is a block diagram of the three-dimensional data encoding device of Embodiment 8.

[0087] Fig.66 It is a block diagram of the three-dimensional data decoding device of Embodiment 8.

[0088] Fig.67 It is a flowchart of the three-dimensional data encoding process of Embodiment 8.

[0089] Fig.68 It is a flowchart of the three-dimensional data decoding process of Embodiment 8.

[0090] Fig.69 It is a diagram showing the first example of a table of predicted values calculated in each prediction mode of Embodiment 9.

[0091] Fig.70 It is a diagram showing an example of the attribute information used in the predicted values of Embodiment 9.

[0092] Fig.71 It is a diagram showing the second example of a table of predicted values calculated in each prediction mode of Embodiment 9.

[0093] Fig.72 It is a diagram showing the third example of a table of predicted values calculated in each prediction mode of Embodiment 9.

[0094] Fig.73 It is a diagram showing the fourth example of a table of predicted values calculated in each prediction mode of Embodiment 9.

[0095] Fig.74 It is a diagram showing an example of the reference relationship of Embodiment 10.

[0096] Fig.75 It is a diagram showing an example of the reference relationship of Embodiment 10.

[0097] Fig.76 It is a diagram showing an example of setting the number of search times for each LoD of Embodiment 10.

[0098] Fig.77 It is a diagram showing an example of the reference relationship of Embodiment 10.

[0099] Fig.78 It is a diagram showing an example of the reference relationship of Embodiment 10.

[0100] Fig.79 It is a diagram showing an example of the reference relationship of Embodiment 10.

[0101] Fig.80 It is a diagram showing a syntactic example of the attribute information header of Embodiment 10.

[0102] Fig.81 It is a diagram showing a syntactic example of the attribute information header of Embodiment 10.

[0103] Fig.82 It is a flowchart of the three-dimensional data encoding process of Embodiment 10.

[0104] Fig.83 It is a flowchart of the attribute information encoding process of Embodiment 10.

[0105] Fig.84 It is a flowchart of the three-dimensional data decoding process of Embodiment 10.

[0106] Fig.85 It is a flowchart of the attribute information decoding process of Embodiment 10.

[0107] Fig.86 It is a flowchart of the surrounding point search process of Embodiment 10.

[0108] Fig.87 It is a flowchart of the surrounding point search process of Embodiment 10.

[0109] Fig.88 It is a flowchart of the surrounding point search process of Embodiment 10.

[0110] Fig.89 It is a flowchart of the three-dimensional data encoding process of Embodiment 10.

[0111] Fig.90 It is a flowchart of the three-dimensional data decoding process of Embodiment 10.

[0112] Fig.91 It is a diagram for explaining a method of selecting N three-dimensional points of Embodiment 11.

[0113] Fig.92 It is a diagram showing an example of the bounding box of group Gk of Embodiment 11.

[0114] Fig.93 It is a diagram for explaining the process of selecting candidates for N three-dimensional points when the first three-dimensional point and multiple second three-dimensional points of Embodiment 11 belong to the same group.

[0115] Fig.94 It is a diagram for explaining the process of selecting candidates for N three-dimensional points when the first three-dimensional point and multiple second three-dimensional points of Embodiment 11 belong to different groups.

[0116] Fig.95 This is a diagram for explaining the process of selecting a 3D point candidate from multiple second 3D points belonging to different hierarchies in Embodiment 11.

[0117] Fig.96 This is a diagram for explaining the process of selecting or updating a 3D point candidate from the groups before and after the initial group in Embodiment 11.

[0118] Fig.97 This is a diagram for explaining an example of a group with a smaller priority bounding box in Embodiment 11.

[0119] Fig.98 This is a flowchart of the 3D data encoding process of the 3D data encoding device in Embodiment 11.

[0120] Fig.99 This is a flowchart of the attribute information encoding process in Embodiment 11.

[0121] Fig.100 This is a flowchart of the 3D data decoding process of the 3D data decoding device in Embodiment 11.

[0122] Fig.101 This is a flowchart of the attribute information decoding process in Embodiment 11.

[0123] Fig.102 This is a flowchart of the search process for surrounding points in Embodiment 11.

[0124] Fig.103 This is a flowchart of the search process for surrounding points in Embodiment 11.

[0125] Fig.104 This is a block diagram showing the structure of the attribute information encoding unit included in the 3D data encoding device in Embodiment 11.

[0126] Fig.105 This is a block diagram showing the structure of the attribute information decoding unit included in the 3D data decoding device in Embodiment 11.

[0127] Fig.106 This is a flowchart of the 3D data encoding process in Embodiment 11.

[0128] Fig.107 This is a flowchart of the 3D data decoding process in Embodiment 11. Detailed implementation mode

[0129] A three-dimensional data encoding method according to an aspect of the present disclosure, wherein, among a plurality of three-dimensional points, a three-dimensional point closest to a first three-dimensional point is selected as a candidate for calculating a predicted value of the attribute information of the first three-dimensional point, the attribute information of the three-dimensional point selected as the candidate is used to calculate the predicted value, a difference between the attribute information of the first three-dimensional point and the calculated predicted value, i.e., a prediction residual, is calculated, a bitstream is generated based on the prediction residual, and in the selection of the candidate, when there are a plurality of three-dimensional points having the same distance from the first three-dimensional point among the plurality of three-dimensional points, the candidate is selected based on the first Morton code of the first three-dimensional point.

[0130] Thereby, a second three-dimensional point close to the first three-dimensional point to be encoded can be selected as a candidate for calculating the predicted value, and thus the encoding efficiency can be improved.

[0131] For example, it may also be that, in the selection of the candidate, when there are a plurality of second three-dimensional points having the same distance from the first three-dimensional point, a three-dimensional point having a Morton code close to the first Morton code is selected from the plurality of second three-dimensional points as the candidate.

[0132] For example, it may also be that the plurality of three-dimensional points are arranged in one dimension in Morton code order, the plurality of three-dimensional points have a first group and a second group, the first group includes a third three-dimensional point belonging to a higher level than the first three-dimensional point and having a Morton code close to the first Morton code, the second group includes a fourth three-dimensional point belonging to the higher level and having a Morton code larger than the Morton code of the third three-dimensional point, the second group is adjacent to the first group, and the search for the candidate of the second three-dimensional point starts from the second group after the first group.

[0133] For example, it may also be that, when the third three-dimensional point is selected as the candidate, if the distance between the first three-dimensional point and the third three-dimensional point is equal to the distance between the first three-dimensional point and the fourth three-dimensional point, the three-dimensional point with the smaller Morton code among the third three-dimensional point and the fourth three-dimensional point is maintained as the candidate.

[0134] For example, it may also be that, when the third three-dimensional point is selected as the candidate, if the distance between the first three-dimensional point and the third three-dimensional point is equal to the distance between the first three-dimensional point and the fourth three-dimensional point, the candidate is not updated.

[0135] For example, it may also be that the plurality of three-dimensional points have a third group, the third group is adjacent to the first group and includes a fifth three-dimensional point having a Morton code smaller than the Morton code of the third three-dimensional point, and the search for the candidate starts from the third group after the second group.

[0136] For example, it may also be that the plurality of three-dimensional points have a fourth group and a fifth group. The fourth group is adjacent to the third group and includes three-dimensional points having a Morton code larger than the Morton code of the third three-dimensional point. The fifth group is adjacent to the fourth group and includes three-dimensional points having a Morton code smaller than the Morton code of the fourth three-dimensional point. The candidates are searched in the order of the third group, the fourth group, and the fifth group.

[0137] For example, it may also be that at least one of the first group, the second group, the third group, the fourth group, and the fifth group includes only one three-dimensional point.

[0138] A three-dimensional data decoding method according to an aspect of the present disclosure, wherein a prediction residual of a first three-dimensional point among a plurality of three-dimensional points is obtained by obtaining a bitstream. Among a plurality of three-dimensional points around the first three-dimensional point, a three-dimensional point having the closest distance to the first three-dimensional point is selected as a candidate for calculating a predicted value of the attribute information of the first three-dimensional point. The predicted value is calculated using the attribute information of the three-dimensional point selected as the candidate. The attribute information of the first three-dimensional point is calculated by adding the predicted value and the prediction residual. In the selection of the candidate, when there are a plurality of three-dimensional points having the same distance from the first three-dimensional point among the plurality of three-dimensional points, the candidate is selected based on the first Morton code of the first three-dimensional point.

[0139] Thereby, the attribute information of the first three-dimensional point of the processing target can be appropriately decoded.

[0140] A three-dimensional data encoding apparatus according to an aspect of the present disclosure, which includes: a processor; and a memory. The processor uses the memory to select, from a plurality of three-dimensional points, a three-dimensional point having the closest distance to the first three-dimensional point as a candidate for calculating a predicted value of the attribute information of the first three-dimensional point. The predicted value is calculated using the attribute information of the three-dimensional point selected as the candidate. A difference between the attribute information of the first three-dimensional point and the calculated predicted value, that is, a prediction residual, is calculated. A bitstream is generated based on the prediction residual. In the selection of the candidate, when there are a plurality of three-dimensional points having the same distance from the first three-dimensional point among the plurality of three-dimensional points, the candidate is selected based on the first Morton code of the first three-dimensional point.

[0141] Thereby, a second three-dimensional point close to the first three-dimensional point to be encoded can be selected as a candidate for calculating the predicted value, and thus the encoding efficiency can be improved.

[0142] A three-dimensional data decoding device according to an aspect of the present disclosure includes: a processor; and a memory. The processor uses the memory to obtain a prediction residual of a first three-dimensional point among a plurality of three-dimensional points by obtaining a bitstream, selects a three-dimensional point closest to the first three-dimensional point from among the plurality of three-dimensional points around the first three-dimensional point as a candidate for calculating a predicted value of attribute information of the first three-dimensional point, calculates the predicted value using the attribute information of the three-dimensional point selected as the candidate, calculates the attribute information of the first three-dimensional point by adding the predicted value and the prediction residual, and in the selection of the candidate, when there are a plurality of three-dimensional points having the same distance from the first three-dimensional point among the plurality of three-dimensional points, selects the candidate based on a first Morton code of the first three-dimensional point.

[0143] Thereby, it is possible to appropriately decode the attribute information of the first three-dimensional point to be processed.

[0144] In addition, these general or specific aspects can be implemented by a system, a device, an integrated circuit, a computer program, or a recording medium such as a computer-readable CD-ROM, and can be implemented by any combination of a system, a device, an integrated circuit, a computer program, and a recording medium.

[0145] Hereinafter, embodiments will be specifically described with reference to the drawings. In addition, all the embodiments to be described below are specific examples showing the present disclosure. The numerical values, shapes, materials, constituent elements, arrangement positions and connection forms of the constituent elements, steps, order of steps, etc. shown in the following embodiments are all examples, and the gist thereof is not to limit the present disclosure. And, among the constituent elements of the following embodiments, the constituent elements not described in the independent technical solution representing the uppermost concept are described as arbitrary constituent elements.

[0146] (Embodiment 1)

[0147] First, a data structure of the encoded three-dimensional data (hereinafter also referred to as encoded data) according to the present embodiment will be described. Figure 1 The configuration of the encoded three-dimensional data according to the present embodiment is shown.

[0148] In the present embodiment, a three-dimensional space is divided into spaces (SPCs) corresponding to pictures in the encoding of a moving picture, and three-dimensional data is encoded in units of space. The space is further divided into volumes (VLMs) corresponding to macroblocks and the like in the moving picture encoding, and prediction and transformation are performed in units of VLM. A volume includes a plurality of voxels (VXLs), which are the smallest units corresponding to position coordinates. In addition, prediction means, similar to the prediction performed in a two-dimensional image, generating predicted three-dimensional data similar to the processing unit to be processed by referring to other processing units, and encoding the difference between the predicted three-dimensional data and the processing unit to be processed. And this prediction includes not only spatial prediction by referring to other prediction units at the same time but also temporal prediction by referring to prediction units at different times.

[0149] For example, when a three-dimensional data encoding device (hereinafter also referred to as an encoding device) encodes a three-dimensional space represented by point cloud data such as a point cloud, it encodes each point of the point cloud or a plurality of points included in a voxel together according to the size of the voxel. If the voxel is subdivided, the three-dimensional shape of the point cloud can be represented with high precision, and if the size of the voxel is increased, the three-dimensional shape of the point cloud can be represented roughly.

[0150] In addition, although the case where the three-dimensional data is a point cloud is described as an example below, the three-dimensional data is not limited to the point cloud and can be three-dimensional data in any form.

[0151] Also, hierarchical voxels can be used. In this case, in the nth hierarchy, it is possible to sequentially show whether there are sampling points in the hierarchies below the (n - 1)th hierarchy (the lower layer of the nth hierarchy). For example, when only the nth hierarchy is decoded, when there are sampling points in the hierarchies below the (n - 1)th hierarchy, it is possible to perform decoding by regarding that there are sampling points at the center of the voxels in the nth hierarchy.

[0152] And the encoding device obtains point cloud data through a distance sensor, a stereo camera, a monocular camera, a gyroscope, or an inertial sensor, etc.

[0153] Regarding the space, similar to the encoding of a moving picture, it is at least classified into any one of the following three prediction structures: an intra-frame space (I-SPC) that can be decoded independently, a predictive space (P-SPC) that can only be referred to unidirectionally, and a bi-directional space (B-SPC) that can be referred to bidirectionally. And the space has two types of time information: a decoding time and a display time.

[0154] And, as Figure 1As shown, as a processing unit including multiple spaces, there is a GOS (Group Of Space) which is a random access unit. Moreover, as a processing unit including multiple GOSs, there is a world space (WLD).

[0155] The space area occupied by the world space is corresponded to the absolute position on the earth through GPS or latitude and longitude information, etc. This position information is stored as meta information. In addition, the meta information can be included in the encoded data or transmitted separately from the encoded data.

[0156] Also, within a GOS, all SPCs can be three-dimensionally adjacent, or there can be SPCs that are not three-dimensionally adjacent to other SPCs.

[0157] In addition, hereinafter, the processes such as encoding, decoding, or referring to the three-dimensional data included in processing units such as GOS, SPC, or VLM are also simply referred to as encoding, decoding, or referring to the processing unit, etc. And the three-dimensional data included in the processing unit includes at least one group of spatial positions such as three-dimensional coordinates and characteristic values such as color information, for example.

[0158] Next, the prediction structure of SPCs in a GOS will be described. Multiple SPCs within the same GOS, or multiple VLMs within the same SPC, although occupying different spaces from each other, hold the same time information (decoding time and display time).

[0159] Also, within a GOS, the SPC that is the first in the decoding order is an I-SPC. And there are two types of GOSs in a GOS: a closed GOS and an open GOS. A closed GOS is a GOS that can decode all SPCs within the GOS when starting to decode from the first I-SPC. In an open GOS, within the GOS, a part of the SPCs whose display time is earlier than that of the first I-SPC refer to different GOSs and can only be decoded in that GOS.

[0160] In addition, in the encoded data such as map information, there is a case of decoding the WLD in the direction opposite to the encoding order. If there is a dependency between GOSs, it is difficult to perform reverse regeneration. Therefore, in this case, a closed GOS is basically adopted.

[0161] Also, a GOS has a layer structure in the height direction, and encoding or decoding is performed sequentially starting from the SPCs in the bottom layer.

[0162] Figure 2 An example of the prediction structure between SPCs belonging to the bottommost layer of a GOS is shown. Figure 3 An example of the inter-layer prediction structure is shown.

[0163] There is more than one I-SPC in the GOS. Although there are objects such as people, animals, cars, bicycles, traffic lights, or buildings that serve as land marks in the three-dimensional space, it is effective especially when encoding small-sized objects as I-SPCs. For example, when a three-dimensional data decoding device (hereinafter also referred to as the decoding device) decodes the GOS with a low processing volume or at high speed, it only decodes the I-SPCs in the GOS.

[0164] Moreover, the encoding device can switch the encoding interval or the occurrence frequency of the I-SPCs according to the density of the objects in the WLD.

[0165] Moreover, in Figure 3 In the configuration shown, the encoding device or the decoding device encodes or decodes multiple layers sequentially starting from the lower layer (layer 1). Accordingly, for example, for an automatically moving vehicle or the like, it is possible to increase the priority of the data near the ground with a large amount of information.

[0166] In addition, in the encoded data used in a drone or the like, in the GOS, encoding or decoding can be performed sequentially starting from the SPC of the upper layer in the height direction.

[0167] Moreover, the encoding device or the decoding device can also encode or decode multiple layers in such a way that the decoding device generally grasps the GOS and can gradually increase the resolution. For example, the encoding device or the decoding device can perform encoding or decoding in the order of layer 3, 8, 1, 9...

[0168] Next, a method for corresponding static objects and dynamic objects will be described.

[0169] In the three-dimensional space, there are static objects or scenes such as buildings or roads (hereinafter collectively referred to as static objects), and dynamic objects such as vehicles or people (hereinafter referred to as dynamic objects). Detection of the objects can be performed separately by extracting feature points from the data of the point cloud or the images captured by a stereo camera or the like. Here, an example of the encoding method for dynamic objects will be described.

[0170] The first method is a method of encoding without distinguishing between static objects and dynamic objects. The second method is a method of distinguishing between static objects and dynamic objects by identification information.

[0171] For example, the GOS is used as an identification unit. In this case, the GOS including the SPCs constituting the static object and the GOS including the SPCs constituting the dynamic object are distinguished in the encoded data or by the identification information stored separately from the encoded data.

[0172] Alternatively, an SPC is used as an identification unit. In this case, only the SPCs that constitute the VLMs of the static object and the SPCs that include the VLMs that constitute the dynamic object are distinguished by the above-described identification information.

[0173] Alternatively, a VLM or a VXL can be used as an identification unit. In this case, the VLMs or VXLs that include static objects and the VLMs or VXLs that include dynamic objects are distinguished by the above-described identification information.

[0174] Moreover, the encoding device can encode a dynamic object as one or more VLMs or SPCs, and encode the VLMs or SPCs that include static objects and the SPCs that include dynamic objects as different GOSs. Moreover, when the size of the GOS becomes variable according to the size of the dynamic object, the encoding device separately stores the size of the GOS as meta information.

[0175] Moreover, the encoding device encodes the static object and the dynamic object independently of each other, and for the world space constituted by the static object, the dynamic object can be overlapped. At this time, the dynamic object is constituted by one or more SPCs, and each SPC corresponds to one or more SPCs of the static object that overlaps the SPC. In addition, the dynamic object may not be represented by an SPC, and may be represented by one or more VLMs or VXLs.

[0176] Moreover, the encoding device can encode the static object and the dynamic object as different streams.

[0177] Moreover, the encoding device can also generate a GOS that includes one or more SPCs that constitute a dynamic object. Moreover, the encoding device can set the GOS (GOS_M) that includes the dynamic object and the GOS of the static object corresponding to the spatial region of GOS_M to have the same size (occupy the same spatial region). In this way, overlapping processing can be performed in units of GOS.

[0178] The P-SPC or B-SPC that constitutes the dynamic object can also refer to the SPCs included in different encoded GOSs. The position of the dynamic object changes over time. In the case where the same dynamic object is encoded as GOSs at different times, the cross-GOS reference is effective from the viewpoint of the compression ratio.

[0179] Moreover, the above-described first method and second method can also be switched according to the use of the encoded data. For example, when encoding three-dimensional data for use as a map, since it is desired to separate from the dynamic object, the encoding device adopts the second method. In addition, when the encoding device encodes three-dimensional data of an event such as a concert or a sports event, if it is not necessary to separate the dynamic object, the first method is adopted.

[0180] Moreover, the decoding time and the display time of GOS or SPC can be stored in the encoded data or stored as meta-information. Also, the time information of static objects can all be the same. In this case, the actual decoding time and display time can be determined by the decoding device. Alternatively, as the decoding time, different values can be assigned for each GOS or SPC, and as the display time, the same value can be assigned to all. Also, as shown in the decoder mode in video coding such as the HRD (Hypothetical Reference Decoder) of HEVC, the decoder has a buffer of a specified size. As long as the bitstream is read at a specified bit rate according to the decoding time, a model that will not be damaged and can be guaranteed to be decoded can be imported.

[0181] Next, the configuration of GOS in the world space will be described. The coordinates of the three-dimensional space in the world space are represented by three mutually orthogonal coordinate axes (x-axis, y-axis, z-axis). By setting a specified rule in the encoding order of GOS, GOS that are adjacent in space can be encoded continuously in the encoded data. For example, in Figure 4 the example shown, the GOS in the xz plane are encoded continuously. After the encoding of all the GOS in one xz plane is completed, the value of the y-axis is updated. That is, as the encoding progresses, the world space extends in the y-axis direction. Also, the index number of GOS is set as the encoding order.

[0182] Here, the three-dimensional space of the world space corresponds one-to-one with GPS or geographical absolute coordinates such as latitude and longitude. Alternatively, the three-dimensional space can be represented by the relative position with respect to a preset reference position. The directions of the x-axis, y-axis, and z-axis of the three-dimensional space are represented as direction vectors determined based on latitude and longitude, etc., and this direction vector is stored together with the encoded data as meta-information.

[0183] Also, the size of GOS is set to be fixed, and the encoding device stores this size as meta-information. Also, the size of GOS can be switched according to, for example, whether it is in the city or indoors or outdoors. That is, the size of GOS can be switched according to the quantity or nature of the object having the value as information. Alternatively, the encoding device can appropriately switch the size of GOS or the interval of I-SPC within GOS in the same world space according to the density of the object, etc. For example, the encoding device sets the size of GOS to be smaller and the interval of I-SPC within GOS to be shorter when the density of the object is higher.

[0184] In Figure 5In the example, in the region from the 3rd to the 10th GOS, due to the high density of objects, in order to achieve random access with a fine granularity, the GOS is subdivided. Also, the 7th to 10th GOSs are respectively located on the back side of the 3rd to 6th GOSs.

[0185] Next, the configuration and operation process of the three-dimensional data encoding device according to this embodiment will be described. Figure 6 FIG. 5 is a block diagram of a three-dimensional data encoding device 100 according to this embodiment. Figure 7 FIG. 7 is a flowchart showing an operation example of the three-dimensional data encoding device 100.

[0186] Figure 6 The three-dimensional data encoding device 100 shown generates encoded three-dimensional data 112 by encoding three-dimensional data 111. The three-dimensional data encoding device 100 includes: an acquisition unit 101, an encoding region determination unit 102, a division unit 103, and an encoding unit 104.

[0187] As Figure 7 shown, first, the acquisition unit 101 acquires three-dimensional data 111 as point cloud data (S101).

[0188] Next, the encoding region determination unit 102 determines the region to be encoded from the spatial region corresponding to the acquired point cloud data (S102). For example, the encoding region determination unit 102 determines the spatial region around the position of the user or the vehicle as the region to be encoded according to the position.

[0189] Next, the division unit 103 divides the point cloud data included in the region to be encoded into respective processing units. Here, the processing units are the above-mentioned GOS and SPC, etc. And the region to be encoded corresponds to the above-mentioned world space, for example. Specifically, the division unit 103 divides the point cloud data into processing units according to the size of the GOS set in advance, the presence or absence or size of dynamic objects (S103). And the division unit 103 determines the start position of the SPC that becomes the head in the encoding order in each GOS.

[0190] Next, the encoding unit 104 generates encoded three-dimensional data 112 by sequentially encoding a plurality of SPCs in each GOS (S104).

[0191] In addition, here, after dividing the region to be encoded into GOS and SPC, although an example of encoding each GOS is shown, the order of processing is not limited to the above. For example, after determining the configuration of one GOS, the GOS can be encoded, and then the configuration of the GOS can be determined, etc.

[0192] In this way, the three-dimensional data encoding device 100 generates the encoded three-dimensional data 112 by encoding the three-dimensional data 111. Specifically, the three-dimensional data encoding device 100 divides the three-dimensional data into random access units, that is, divides it into first processing units (GOS) corresponding to three-dimensional coordinates respectively, divides the first processing units (GOS) into a plurality of second processing units (SPC), and divides the second processing units (SPC) into a plurality of third processing units (VLM). Moreover, the third processing unit (VLM) includes one or more voxels (VXL), and the voxel (VXL) is the smallest unit corresponding to the position information.

[0193] Next, the three-dimensional data encoding device 100 generates the encoded three-dimensional data 112 by encoding each of the plurality of first processing units (GOS). Specifically, the three-dimensional data encoding device 100 encodes each of the plurality of second processing units (SPC) in each first processing unit (GOS). Moreover, the three-dimensional data encoding device 100 encodes each of the plurality of third processing units (VLM) in each second processing unit (SPC).

[0194] For example, when the first processing unit (GOS) of the processing object is a closed GOS, for the second processing unit (SPC) of the processing object included in the first processing unit (GOS) of the processing object, encoding is performed with reference to other second processing units (SPC) included in the first processing unit (GOS) of the processing object. That is, the three-dimensional data encoding device 100 does not refer to the second processing units (SPC) included in the first processing unit (GOS) different from the first processing unit (GOS) of the processing object.

[0195] Moreover, when the first processing unit (GOS) of the processing object is an open GOS, for the second processing unit (SPC) of the processing object included in the first processing unit (GOS) of the processing object, encoding is performed with reference to other second processing units (SPC) included in the first processing unit (GOS) of the processing object or the second processing units (SPC) included in the first processing unit (GOS) different from the first processing unit (GOS) of the processing object.

[0196] Moreover, the three-dimensional data encoding device 100 selects one from the first type (I-SPC) that never refers to other second processing units (SPC), the second type (P-SPC) that refers to one other second processing unit (SPC), and the third type that refers to two other second processing units (SPC) as the type of the second processing unit (SPC) of the processing object, and encodes the second processing unit (SPC) of the processing object according to the selected type.

[0197] Next, the configuration and operation process of the three-dimensional data decoding device according to this embodiment will be described. Figure 8 is a block diagram of the three-dimensional data decoding device 200 according to this embodiment. Fig. 9 is a flowchart showing an operation example of the three-dimensional data decoding device 200.

[0198] Figure 8 The three-dimensional data decoding device 200 shown generates decoded three-dimensional data 212 by decoding the encoded three-dimensional data 211. Here, the encoded three-dimensional data 211 is, for example, the encoded three-dimensional data 112 generated by the three-dimensional data encoding device 100. The three-dimensional data decoding device 200 includes: an acquisition unit 201, a decoding start GOS determination unit 202, a decoding SPC determination unit 203, and a decoding unit 204.

[0199] First, the acquisition unit 201 acquires the encoded three-dimensional data 211 (S201). Next, the decoding start GOS determination unit 202 determines the GOS to be decoded (S202). Specifically, the decoding start GOS determination unit 202 refers to the meta information in the encoded three-dimensional data 211 or stored separately from the encoded three-dimensional data, and determines the GOS including the spatial position, object, or SPC corresponding to the time at which decoding starts as the GOS to be decoded.

[0200] Next, the decoding SPC determination unit 203 determines the type (I, P, B) of the SPC to be decoded within the GOS (S203). For example, the decoding SPC determination unit 203 determines (1) whether to decode only I-SPC, (2) whether to decode I-SPC and P-SPC, (3) whether to decode all types. In addition, when the type of SPC to be decoded, such as all SPCs, is specified in advance, this step may not be performed.

[0201] Next, the decoding unit 204 acquires the address position in the encoded three-dimensional data 211 where the SPC at the beginning in the decoding order (the same as the encoding order) within the GOS starts, acquires the encoded data of the beginning SPC from this address position, and sequentially decodes each SPC from this beginning SPC (S204). And the above address position is stored in meta information or the like.

[0202] In this way, the three-dimensional data decoding device 200 decodes the decoded three-dimensional data 212. Specifically, the three-dimensional data decoding device 200 generates the decoded three-dimensional data 212 of the first processing unit (GOS) as a random access unit by decoding each of the encoded three-dimensional data 211 of the first processing unit (GOS) corresponding to the three-dimensional coordinates respectively. More specifically, the three-dimensional data decoding device 200 decodes each of the plurality of second processing units (SPC) in each of the first processing units (GOS). Further, the three-dimensional data decoding device 200 decodes each of the plurality of third processing units (VLM) in each of the second processing units (SPC).

[0203] The meta information for random access will be described below. This meta information is generated by the three-dimensional data encoding device 100 and is included in the encoded three-dimensional data 112 (211).

[0204] In the random access of conventional two-dimensional moving images, decoding starts from the first frame of the random access unit near the specified time. However, in the world space, random access is envisioned not only for time but also for (coordinates or objects, etc.).

[0205] Therefore, in order to achieve random access to at least the three elements of coordinates, objects, and time, a table in which the index numbers of each element are associated with the GOS is prepared. Moreover, the index number of the GOS is associated with the address of the I-SPC that is the start of the GOS. Fig.10 An example of the table included in the meta information is shown. Additionally, it is not necessary to use Fig.10 all of the shown tables, and at least one table can be used.

[0206] Hereinafter, as an example, random access starting from coordinates will be described. When accessing the coordinates (x2, y2, z2), first referring to the coordinate-GOS table, it can be known that the location with coordinates (x2, y2, z2) is included in the second GOS. Then, referring to the GOS address table, since it can be known that the address of the I-SPC at the start of the second GOS is addr(2), the decoding unit 204 obtains data from this address and starts decoding.

[0207] In addition, the address can be an address in the logical format or a physical address of an HDD or a memory. Also, information for determining a file segment can be used instead of the address. For example, a file segment is a unit obtained by segmenting one or more GOSs, etc.

[0208] Also, in the case where the object spans multiple GOSs, the GOSs to which the multiple objects belong can also be shown in the object GOS table. If the multiple GOSs are closed GOSs, the encoding device and the decoding device can perform encoding or decoding in parallel. In addition, if the multiple GOSs are open GOSs, by cross-referencing the multiple GOSs with each other, the compression efficiency can be further improved.

[0209] Examples of the object include a person, an animal, a car, a bicycle, a traffic signal, or a building serving as a land mark. For example, when encoding in the world space, the three-dimensional data encoding device 100 extracts the feature points unique to the object from a three-dimensional point cloud or the like, detects the object based on the feature points, and can set the detected object as a random access point.

[0210] In this way, the three-dimensional data encoding device 100 generates the first information, which shows multiple first processing units (GOSs) and the three-dimensional coordinates corresponding to each of the multiple first processing units (GOSs). And the encoded three-dimensional data 112(211) includes this first information. And the first information further shows at least one of the object, the time, and the data storage destination corresponding to each of the multiple first processing units (GOSs).

[0211] The three-dimensional data decoding device 200 obtains the first information from the encoded three-dimensional data 211, uses the first information to determine the encoded three-dimensional data 211 of the first processing unit corresponding to the specified three-dimensional coordinates, object, or time, and decodes the encoded three-dimensional data 211.

[0212] Examples of other meta-information will be described below. In addition to the meta-information for random access, the three-dimensional data encoding device 100 can also generate and store the following meta-information. And the three-dimensional data decoding device 200 can also use this meta-information during decoding.

[0213] In the case of using the three-dimensional data as map information, etc., a profile is specified according to the use, and the information showing the profile can be included in the meta-information. For example, a profile for urban areas or suburbs is specified, or a profile for flying objects is specified, and the maximum or minimum size of the world space, SPC, or VLM is defined respectively. For example, in the profile for urban areas, more detailed information is required than in the suburbs, so the minimum size of the VLM is set smaller.

[0214] The meta-information may also include a tag value indicating the type of the object. This tag value corresponds to the VLM, SPC, or GOS that constitutes the object. The tag value can be set according to the type of the object, etc. For example, the tag value "0" represents "person", the tag value "1" represents "car", and the tag value "2" represents "traffic signal". Alternatively, in a case where it is difficult to determine or unnecessary to determine the type of the object, a tag value indicating properties such as size, or whether it is a dynamic object or a static object may also be used.

[0215] Furthermore, the meta-information may also include information indicating the range of the spatial region occupied by the world space.

[0216] Furthermore, the meta-information may store the size of the SPC or VXL as the entire stream of encoded data or as header information shared by multiple SPCs such as SPCs within the GOS.

[0217] Furthermore, the meta-information may also include identification information such as a distance sensor or a camera used in the generation of the point cloud, or may include information indicating the position accuracy of the point group within the point cloud.

[0218] Furthermore, the meta-information may include information indicating whether the world space is composed only of static objects or contains dynamic objects.

[0219] A modification example of the present embodiment will be described below.

[0220] The encoding device or the decoding device may encode or decode two or more different SPCs or GOSs in parallel. The GOSs to be encoded or decoded in parallel can be determined based on the meta-information indicating the spatial position of the GOSs, etc.

[0221] In a case where the three-dimensional data is used as a spatial map when a vehicle or a flying object moves, or in a case where such a spatial map is generated, etc., the encoding device or the decoding device may encode or decode the GOS or SPC included in the space determined based on GPS, path information, or magnification ratio, etc.

[0222] Furthermore, the decoding device may also start decoding sequentially from the space close to its own position or the walking path. The encoding device or the decoding device may also perform encoding or decoding by making the priority of the space far from its own position or the walking path lower than that of the close space. Here, lowering the priority means lowering the processing order, lowering the resolution (post-filtering), or lowering the image quality (improving the encoding efficiency. For example, increasing the quantization step size), etc.

[0223] Furthermore, when decoding the encoded data hierarchically encoded within the space, the decoding device may also decode only the lower hierarchy.

[0224] Also, the decoding device can also start decoding from the lower level according to the zoom ratio or usage of the map.

[0225] Also, in applications such as self-position estimation or object recognition performed during the automatic driving of a vehicle or a robot, the encoding device or the decoding device can also reduce the resolution of areas outside the area within a specified height from the road surface (the area to be recognized) for encoding or decoding.

[0226] Also, the encoding device can also encode the point clouds representing the spatial shapes of the indoor and outdoor spaces independently. For example, by separating the GOS representing the indoor (indoor GOS) from the GOS representing the outdoor (outdoor GOS), the decoding device can select the GOS to be decoded according to the viewpoint position when using the encoded data.

[0227] Also, the encoding device can make the indoor GOS and the outdoor GOS with close coordinates adjacent in the encoding stream for encoding. For example, the encoding device corresponds their identifiers and stores the information showing the identifiers corresponding in the encoding stream or in the meta-information stored separately. Accordingly, the decoding device can identify the indoor GOS and the outdoor GOS with close coordinates by referring to the information in the meta-information.

[0228] Also, the encoding device can also switch the size of the GOS or SPC between the indoor GOS and the outdoor GOS. For example, the encoding device sets the size of the GOS to be smaller indoors than outdoors. Also, the encoding device can also change the accuracy when extracting feature points from the point cloud or the accuracy of object detection, etc., between the indoor GOS and the outdoor GOS.

[0229] Also, the encoding device can attach information for the decoding device to distinguish and display dynamic objects from static objects to the encoded data. Accordingly, the decoding device can combine the dynamic objects with a red frame or explanatory text, etc., for display. In addition, the decoding device can also represent only with a red frame or explanatory text instead of the dynamic objects. And the decoding device can represent more detailed object categories. For example, a car can use a red frame and a person can use a yellow frame.

[0230] Also, the encoding device or the decoding device can determine whether to encode or decode the dynamic objects and the static objects as different SPCs or GOSs according to the appearance frequency of the dynamic objects, or the ratio of the static objects to the dynamic objects, etc. For example, when the appearance frequency or ratio of the dynamic objects exceeds the threshold, the SPC or GOS in which the dynamic objects and the static objects are mixed is allowed, and when the appearance frequency or ratio of the dynamic objects does not exceed the threshold, the SPC or GOS in which the dynamic objects and the static objects are mixed is not allowed.

[0231] When the dynamic object is detected from the two-dimensional image information of the camera instead of the point cloud, the encoding device can obtain the information (such as a box or text) for identifying the detection result and the object position separately, and encode these information as part of the three-dimensional encoded data. In this case, the decoding device overlays and displays the auxiliary information (box or text) representing the dynamic object on the decoding result of the static object.

[0232] Moreover, the encoding device can change the density of VXL or VLM according to the complexity of the shape of the static object, etc. For example, the encoding device sets VXL or VLM to be denser when the shape of the static object is more complex. Also, the encoding device can determine the quantization step size, etc. when quantifying the spatial position or color information according to the density of VXL or VLM. For example, the encoding device sets the quantization step size to be smaller when VXL or VLM is denser.

[0233] As described above, the encoding device or decoding device according to this embodiment performs spatial encoding or decoding in a spatial unit having coordinate information.

[0234] Moreover, the encoding device and the decoding device perform encoding or decoding in volume units within the space. The volume includes voxels, which are the smallest units corresponding to the position information.

[0235] Moreover, the encoding device and the decoding device establish correspondences between arbitrary elements by using a table in which each element of the spatial information including coordinates, objects, and time, etc. is associated with a GOP, or a table corresponding between each element, and perform encoding or decoding. And the decoding device judges the coordinates by using the value of the selected element, determines the volume, voxel, or space according to the coordinates, and decodes the space including the volume or voxel, or the determined space.

[0236] Moreover, the encoding device judges the volume, voxel, or space that can be selected by the element through feature point extraction or object recognition, and encodes it as a volume, voxel, or space that can be randomly accessed.

[0237] The space is divided into three types, namely: I-SPC that can be encoded or decoded by the space alone, P-SPC that encodes or decodes with reference to any one processed space, and B-SPC that encodes or decodes with reference to any two processed spaces.

[0238] One or more volumes correspond to static objects or dynamic objects. The space containing the static object and the space containing the dynamic object are encoded or decoded as different GOSs respectively. That is, the SPC containing the static object and the SPC containing the dynamic object are assigned to different GOSs.

[0239] Dynamic objects are encoded or decoded on a per-object basis, corresponding to more than one space that contains only static objects. That is, multiple dynamic objects are encoded separately, and the encoded data of the multiple dynamic objects corresponds to the SPC that contains only static objects.

[0240] The encoding device and the decoding device increase the priority of the I-SPC in the GOS to perform encoding or decoding. For example, the encoding device performs encoding in a manner that reduces the degradation of the I-SPC (after decoding, the original three-dimensional data can be reproduced more faithfully). Also, the decoding device decodes only the I-SPC, for example.

[0241] The encoding device can change the frequency of using the I-SPC according to the density or value (quantity) of the objects in the world space to perform encoding. That is, the encoding device changes the frequency of selecting the I-SPC according to the quantity or density of the objects included in the three-dimensional data. For example, the encoding device increases the usage frequency of the I-space when the density of the objects in the world space is greater.

[0242] Also, the encoding device sets random access points in units of GOS, and stores the information showing the spatial region corresponding to the GOS in the header information.

[0243] The encoding device uses a default value as the spatial size of the GOS, for example. In addition, the encoding device can also change the size of the GOS according to the value (quantity) or density of the objects or dynamic objects. For example, the encoding device sets the spatial size of the GOS to be smaller when the objects or dynamic objects are denser or the quantity is larger.

[0244] Also, the space or volume includes a feature point group derived using information obtained by sensors such as depth sensors, gyroscopes, or cameras. The coordinates of the feature points are set as the center positions of the voxels. And through the subdivision of the voxels, high-precision position information can be achieved.

[0245] The feature point group is derived using multiple pictures. The multiple pictures have at least the following two types of time information, namely: actual time information, and the same time information in the multiple pictures corresponding to the space (for example, the encoding time for rate control, etc.).

[0246] Also, encoding or decoding is performed in units of GOS that includes more than one space.

[0247] The encoding device and the decoding device predict the P space or B space in the GOS to be processed with reference to the space in the processed GOS.

[0248] Alternatively, the encoding device and the decoding device do not refer to different GOSs, and use the processed space within the GOS of the object to be processed to predict the P space or B space within the GOS of the object to be processed.

[0249] Moreover, the encoding device and the decoding device send or receive an encoded stream in units of a world space including one or more GOSs.

[0250] Furthermore, the GOS has a layer structure at least in one direction within the world space, and the encoding device and the decoding device perform encoding or decoding starting from the lower layer. For example, a GOS that enables random access belongs to the lowest layer. A GOS belonging to an upper layer only refers to GOSs belonging to layers below the same layer. That is, the GOS is spatially divided in a predefined direction and includes multiple layers each having one or more SPCs. The encoding device and the decoding device perform encoding or decoding for each SPC by referring to the SPCs included in the same layer as or a lower layer than the SPC.

[0251] In addition, the encoding device and the decoding device continuously perform encoding or decoding on GOSs within a world space unit including multiple GOSs. The encoding device and the decoding device write or read information indicating the order (direction) of encoding or decoding as metadata. That is, the encoded data includes information indicating the encoding order of multiple GOSs.

[0252] Also, the encoding device and the decoding device perform encoding or decoding on two or more different spaces or GOSs in parallel.

[0253] Moreover, the encoding device and the decoding device encode or decode the spatial information (coordinates, size, etc.) of a space or GOS.

[0254] Furthermore, the encoding device and the decoding device encode or decode the space or GOS included in a specific space determined according to external information such as GPS, path information, or magnification related to its own position or / and area size.

[0255] The encoding device or the decoding device performs encoding or decoding by setting a lower priority for a space farther from its own position than for a space closer to its own position.

[0256] The encoding device sets one direction in the world space according to magnification or use, and encodes the GOS having a layer structure in that direction. And the decoding device preferentially performs decoding starting from the lower layer for the GOS having a layer structure in one direction of the world space set according to magnification or use.

[0257] The encoding device changes the extraction of feature points, the accuracy of object recognition, or the size of the spatial region included in the indoor and outdoor spaces. However, the encoding device and the decoding device encode or decode by making the indoor GOS and the outdoor GOS with adjacent coordinates adjacent in the world space, and also encode or decode by corresponding these identifiers.

[0258] (Embodiment 2)

[0259] When using the encoded data of the point cloud for an actual device or service, in order to suppress the network bandwidth, it is desired to transmit and receive the required information according to the usage. However, such a function does not exist in the encoding structure of the three-dimensional data so far, and thus there is no encoding method corresponding thereto.

[0260] In this embodiment, a three-dimensional data encoding method and a three-dimensional data encoding device for providing a function of transmitting and receiving the required information according to the usage in the encoded data of the three-dimensional point cloud, a three-dimensional data decoding method for decoding the encoded data, and a three-dimensional data decoding device will be described.

[0261] A voxel (VXL) having a feature amount equal to or more than a certain value is defined as a feature voxel (FVXL), and a world space (WLD) composed of FVXL is defined as a sparse world space (SWLD). Fig.11 A configuration example of the sparse world space and the world space is shown. In the SWLD, there are included: FGOS, which is a GOS composed of FVXL; FSPC, which is an SPC composed of FVXL; and FVLM, which is a VLM composed of FVXL. The data structure and the prediction structure of FGOS, FSPC, and FVLM can be the same as those of GOS, SPC, and VLM.

[0262] The feature amount refers to a feature amount representing the three-dimensional position information of the VXL or the visible light information at the VXL position, and in particular, a feature amount capable of detecting more features such as the corners and edges of a three-dimensional object. Specifically, although the feature amount is the three-dimensional feature amount or the visible light feature amount described below, as long as it is a feature amount representing the position, brightness, or color information of the VXL, it can be any feature amount.

[0263] As the three-dimensional feature amount, a SHOT feature amount (Signature of Histograms of OrienTations), a PFH feature amount (Point Feature Histograms), or a PPF feature amount (Point Pair Feature) is adopted.

[0264] The SHOT feature quantity is obtained by segmenting the periphery of the VXL, calculating the inner product of the normal vector of the reference point and the segmented region, and performing histogramming. This SHOT feature quantity has the characteristics of high dimensionality and high feature expressiveness.

[0265] The PFH feature quantity is obtained by selecting multiple two-point groups near the VXL, calculating the normal vector, etc. based on these two points, and performing histogramming. Since this PFH feature quantity is a histogram feature, it is robust against a small amount of interference and has the characteristic of high feature expressiveness.

[0266] The PPF feature quantity is a feature quantity calculated using the normal vector, etc. according to two VXLs. In this PPF feature quantity, since all VXLs are used, it is robust against occlusion.

[0267] Moreover, as feature quantities of visible light, it is possible to use SIFT (Scale-Invariant Feature Transform), SURF (Speeded Up Robust Features), or HOG (Histogram of Oriented Gradients), etc. which adopt information such as the luminance gradient information of an image.

[0268] The SWLD is generated by calculating the above-mentioned feature quantities from each VXL of the WLD and extracting the FVXL. Here, the SWLD can be updated each time the WLD is updated, or it can be updated regularly after a certain period of time regardless of the update timing of the WLD.

[0269] The SWLD can be generated for each feature quantity. For example, as shown by the SWLD1 based on the SHOT feature quantity and the SWLD2 based on the SIFT feature quantity, the SWLD can be generated separately for each feature quantity, and the SWLD can be distinguished and used according to the usage. Also, the feature quantities of each calculated FVXL can be held as feature quantity information in each FVXL.

[0270] Next, the method of using the sparse world space (SWLD) will be described. Since the SWLD only contains feature voxels (FVXL), generally, the data size is smaller compared to the WLD that includes all VXLs.

[0271] In an application that uses feature quantities to achieve a certain purpose, by replacing WLD with SWLD information, the read time from the hard disk can be suppressed, and the bandwidth and transmission time during network transmission can be suppressed. For example, as map information, WLD and SWLD are stored in the server in advance, and by switching the transmitted map information to WLD or SWLD according to the requirements from the client, the network bandwidth and transmission time can be suppressed. The following shows specific examples.

[0272] Fig.12 and Fig.13 shows examples of the use of SWLD and WLD. As Fig.12 shown, when the client 1 as a vehicle-mounted device needs map information for its own position judgment, the client 1 sends a request for obtaining map data for its own position estimation to the server (S301). The server sends SWLD to the client 1 according to this acquisition request (S302). The client 1 uses the received SWLD to judge its own position (S303). At this time, the client 1 obtains VXL information around the client 1 by various methods such as a distance sensor such as a rangefinder, a stereo camera, or a combination of multiple monocular cameras, and estimates its own position information based on the obtained VXL information and SWLD. Here, the own position information includes the three-dimensional position information and orientation of the client 1.

[0273] As Fig.13 shown, when the client 2 as a vehicle-mounted device needs map information for map drawing such as a three-dimensional map, the client 2 sends a request for obtaining map data for map drawing to the server (S311). The server sends WLD to the client 2 according to this acquisition request (S312). The client 2 uses the received WLD to perform map drawing (S313). At this time, the client 2, for example, uses an image taken by its own visible light camera and the WLD obtained from the server to create a concept image, and depicts the created image on a screen such as a car navigation.

[0274] As shown above, the server sends SWLD to the client in applications that mainly require the feature quantities of each VXL for its own position estimation, and sends WLD to the client when detailed VXL information is required, such as map drawing. Accordingly, the map data can be efficiently transmitted and received.

[0275] In addition, the client can judge which of SWLD and WLD it needs and request the server to send SWLD or WLD. And the server can judge which of SWLD or WLD should be sent according to the condition of the client or the network.

[0276] Next, a method for switching the transmission and reception between the Sparse World Space (SWLD) and the World Space (WLD) will be described.

[0277] The reception of the WLD or SWLD can be switched according to the network bandwidth. Fig.14 A working example in this case is shown. For example, when a low-speed network with a network bandwidth such as in an LTE (Long Term Evolution) environment is used, when the client accesses the server via the low-speed network (S321), the client obtains the SWLD as map information from the server (S322). Additionally, when a high-speed network with a surplus network bandwidth such as in a WiFi environment is used, the client accesses the server via the high-speed network (S323) and obtains the WLD from the server (S324). Accordingly, the client can obtain appropriate map information according to the network bandwidth of the client.

[0278] Specifically, the client receives the SWLD via LTE outdoors, and when entering indoors such as in a facility, obtains the WLD via WiFi. Accordingly, the client can obtain more detailed map information of the interior.

[0279] In this way, the client can request the WLD or SWLD from the server according to the frequency band of the network it uses. Alternatively, the client can send information indicating the frequency band of the network it uses to the server, and the server sends appropriate data (WLD or SWLD) to the client according to this information. Or, the server can determine the network bandwidth of the client and send appropriate data (WLD or SWLD) to the client.

[0280] Moreover, the reception of the WLD or SWLD can be switched according to the moving speed. Fig.15 A working example in this case is shown. For example, when the client is moving at high speed (S331), the client receives the SWLD from the server (S332). Additionally, when the client is moving at low speed (S333), the client receives the WLD from the server (S334). Accordingly, the client can both suppress the network bandwidth and obtain map information according to the speed. Specifically, when the client is driving on a highway, by receiving the SWLD with a small data volume, the map information can be updated at an appropriate speed approximately. Additionally, when the client is driving on an ordinary road, by receiving the WLD, more detailed map information can be obtained.

[0281] In this way, the client can request the WLD or SWLD from the server according to its own moving speed. Alternatively, the client can send the information indicating its own moving speed to the server, and the server sends appropriate data (WLD or SWLD) to the client according to this information. Or, the server can determine the moving speed of the client and send appropriate data (WLD or SWLD) to the client.

[0282] Also, it can be that the client first obtains the SWLD from the server and then obtains the WLD of the important areas therein. For example, when the client obtains map data, it first obtains the general map information in the form of SWLD, screens out the areas where features such as buildings, signs, or people appear more frequently from it, and then obtains the WLD of the screened areas. Accordingly, the client can both suppress the amount of received data from the server and obtain the detailed information of the required areas.

[0283] Also, it can be that the server separately creates the SWLD for each object according to the WLD, and the client receives them separately according to the usage. Accordingly, the network bandwidth can be suppressed. For example, the server pre-identifies people or vehicles from the WLD and creates the SWLD for people and the SWLD for vehicles. When the client wants to obtain information about the people around, it receives the SWLD for people, and when it wants to obtain information about vehicles, it receives the SWLD for vehicles. And the types of such SWLD can be distinguished according to the information (flags or types, etc.) attached to the header, etc.

[0284] Next, the configuration and the working process of the three-dimensional data encoding device (such as the server) according to the present embodiment will be described. Fig.16 FIG. is a block diagram of the three-dimensional data encoding device 400 according to the present embodiment. Fig.17 FIG. is a flowchart of the three-dimensional data encoding process performed by the three-dimensional data encoding device 400.

[0285] Fig.16 The three-dimensional data encoding device 400 shown generates encoded three-dimensional data 413 and 414 as an encoded stream by encoding the input three-dimensional data 411. Here, the encoded three-dimensional data 413 is the encoded three-dimensional data corresponding to the WLD, and the encoded three-dimensional data 414 is the encoded three-dimensional data corresponding to the SWLD. The three-dimensional data encoding device 400 includes: an acquisition unit 401, an encoding area determination unit 402, an SWLD extraction unit 403, a WLD encoding unit 404, and an SWLD encoding unit 405.

[0286] As Fig.17 shown, first, the acquisition unit 401 acquires the input three-dimensional data 411 as point cloud data in a three-dimensional space (S401).

[0287] Next, the coding area determination unit 402 determines the spatial area of the coding object based on the spatial area where the point cloud data exists (S402).

[0288] Next, the SWLD extraction unit 403 defines the spatial area of the coding object as the WLD, and calculates the feature amount based on each VXL included in the WLD. Further, the SWLD extraction unit 403 extracts the VXL whose feature amount is equal to or greater than a preset threshold value, defines the extracted VXL as the FVXL, and generates the extracted three-dimensional data 412 by adding the FVXL to the SWLD (S403). That is, the extracted three-dimensional data 412 whose feature amount is equal to or greater than the threshold value is extracted from the input three-dimensional data 411.

[0289] Next, the WLD coding unit 404 generates the coded three-dimensional data 413 corresponding to the WLD by coding the input three-dimensional data 411 corresponding to the WLD (S404). At this time, the WLD coding unit 404 attaches information for distinguishing that the coded three-dimensional data 413 is a stream including the WLD to the head of the coded three-dimensional data 413.

[0290] Further, the SWLD coding unit 405 generates the coded three-dimensional data 414 corresponding to the SWLD by coding the extracted three-dimensional data 412 corresponding to the SWLD (S405). At this time, the SWLD coding unit 405 attaches information for distinguishing that the coded three-dimensional data 414 is a stream including the SWLD to the head of the coded three-dimensional data 414.

[0291] Further, the processing order of the process of generating the coded three-dimensional data 413 and the process of generating the coded three-dimensional data 414 may be opposite to the above. Further, a part or all of the above processes may be executed in parallel.

[0292] The information given to the heads of the coded three-dimensional data 413 and 414 is defined as a parameter such as "world_type", for example. When world_type = 0, it indicates that the stream includes the WLD, and when world_type = 1, it indicates that the stream includes the SWLD. When defining more other categories, the assigned value can be increased as in world_type = 2. Further, a specific flag may be included in one of the coded three-dimensional data 413 and 414. For example, the coded three-dimensional data 414 may be given a flag indicating that the stream includes the SWLD. In this case, the decoding device can determine whether the stream includes the WLD or the SWLD based on the presence or absence of the flag.

[0293] Further, the coding method used by the WLD coding unit 404 when coding the WLD may be different from the coding method used by the SWLD coding unit 405 when coding the SWLD.

[0294] For example, since SWLD data is selected, its correlation with surrounding data may be lower compared to WLD. Therefore, in the encoding method for SWLD, among intra-frame prediction and inter-frame prediction, inter-frame prediction is prioritized compared to the encoding method for WLD.

[0295] Also, the representation method of three-dimensional positions may be different between the encoding method for SWLD and the encoding method for WLD. For example, it may be that in FWLD, the three-dimensional position of FVXL is represented by three-dimensional coordinates, and in WLD, the three-dimensional position is represented by an octree described later, and vice versa.

[0296] Moreover, the SWLD encoding unit 405 encodes in such a way that the data size of the encoded three-dimensional data 414 of SWLD is smaller than the data size of the encoded three-dimensional data 413 of WLD. As described above, for example, the correlation between data in SWLD may be lower than that in WLD. Accordingly, the encoding efficiency decreases, and the data size of the encoded three-dimensional data 414 may be larger than the data size of the encoded three-dimensional data 413 of WLD. Therefore, when the data size of the obtained encoded three-dimensional data 414 of SWLD is larger than the data size of the encoded three-dimensional data 413 of WLD, the SWLD encoding unit 405 re-encodes to regenerate the encoded three-dimensional data 414 with a reduced data size.

[0297] For example, the SWLD extraction unit 403 regenerates the extracted three-dimensional data 412 with a reduced number of extracted feature points, and the SWLD encoding unit 405 encodes this extracted three-dimensional data 412. Alternatively, the quantization level in the SWLD encoding unit 405 can be made coarser. For example, in the octree structure described later, by rounding the data in the bottom layer, the quantization level can be made coarser.

[0298] Also, when the SWLD encoding unit 405 cannot make the data size of the encoded three-dimensional data 414 of SWLD smaller than the data size of the encoded three-dimensional data 413 of WLD, it may not generate the encoded three-dimensional data 414 of SWLD. Alternatively, the encoded three-dimensional data 413 of WLD can be copied to the encoded three-dimensional data 414 of SWLD. That is, as the encoded three-dimensional data 414 of SWLD, the encoded three-dimensional data 413 of WLD can be directly used.

[0299] Next, the configuration and operation flow of the three-dimensional data decoding device (such as a client) according to the present embodiment will be described. Fig.18 It is a block diagram of the three-dimensional data decoding device 500 according to the present embodiment. Fig.19 It is a flowchart of the three-dimensional data decoding process performed by the three-dimensional data decoding device 500.

[0300] Fig.18 The three-dimensional data decoding device 500 shown generates decoded three-dimensional data 512 or 513 by decoding the encoded three-dimensional data 511. Here, the encoded three-dimensional data 511 is, for example, the encoded three-dimensional data 413 or 414 generated by the three-dimensional data encoding device 400.

[0301] The three-dimensional data decoding device 500 includes: an acquisition unit 501, a header analysis unit 502, a WLD decoding unit 503, and a SWLD decoding unit 504.

[0302] As Fig.19 shown, first, the acquisition unit 501 acquires the encoded three-dimensional data 511 (S501). Next, the header analysis unit 502 analyzes the header of the encoded three-dimensional data 511 to determine whether the encoded three-dimensional data 511 is a stream containing WLD or a stream containing SWLD (S502). For example, the determination is made with reference to the above-mentioned world_type parameter.

[0303] In the case where the encoded three-dimensional data 511 is a stream containing WLD ("Yes" in S503), the WLD decoding unit 503 decodes the encoded three-dimensional data 511 to generate the decoded three-dimensional data 512 of WLD (S504). In addition, in the case where the encoded three-dimensional data 511 is a stream containing SWLD ("No" in S503), the SWLD decoding unit 504 decodes the encoded three-dimensional data 511 to generate the decoded three-dimensional data 513 of SWLD (S505).

[0304] And, similar to the encoding device, the decoding method used by the WLD decoding unit 503 when decoding WLD and the decoding method used by the SWLD decoding unit 504 when decoding SWLD can be different. For example, in the decoding method for SWLD, inter-frame prediction in intra-frame prediction and inter-frame prediction can be given priority compared to the decoding method for WLD.

[0305] And, in the decoding method for SWLD and the decoding method for WLD, the expression methods of three-dimensional positions can be different. For example, in SWLD, the three-dimensional position of FVXL can be expressed by three-dimensional coordinates, and in WLD, the three-dimensional position can be expressed by an octree described later, and vice versa.

[0306] Next, the octree representation as the expression method of the three-dimensional position will be described. The VXL data included in the three-dimensional data is converted into an octree structure and then encoded. Fig. 20 An example of the VXL of WLD is shown. Fig.21 Shown is Fig. 20 the octree structure of the WLD shown. In Fig. 20 In the example shown, there are three VXLs (hereinafter, effective VXLs) that include point groups, namely VXL1 to VXL3. As Fig.21 shown, the octree structure is composed of nodes and leaf nodes. Each node has a maximum of eight nodes or leaf nodes. Each leaf node has VXL information. Here, Fig.21 among the leaf nodes shown, leaf nodes 1, 2, and 3 respectively represent Fig. 20 the VXL1, VXL2, and VXL3 shown.

[0307] Specifically, each node and leaf node correspond to a three-dimensional position. Node 1 corresponds to Fig. 20 all the blocks shown. The block corresponding to Node 1 is divided into eight blocks. Among the eight blocks, the blocks including the effective VXL are set as nodes, and the other blocks are set as leaf nodes. The blocks corresponding to the nodes are further divided into eight nodes or leaf nodes, and the number of times this process is repeated is the same as the number of levels in the tree structure. And all the blocks in the bottom layer are set as leaf nodes.

[0308] And, Fig. 22 an example of the SWLD generated from the Fig. 20 shown WLD is shown. Fig. 20 The results of feature quantity extraction of the VXL1 and VXL2 shown are judged as FVXL1 and FVXL2 and are added to the SWLD. In addition, since VXL3 is not judged as FVXL, it is not included in the SWLD. Fig.23 An example of the octree structure of the Fig. 22 shown SWLD is shown. In the Fig.23 octree structure shown, Fig.21 the leaf node 3 corresponding to VXL3 shown is deleted. Accordingly, Fig.21 the node 3 shown has no effective VXL and is changed to a leaf node. Thus, generally, the number of leaf nodes in the SWLD is smaller than that in the WLD, and the encoded three-dimensional data of the SWLD is also smaller than that of the WLD.

[0309] The following describes a modification example of the present embodiment.

[0310] For example, it may also be the case where, when a client such as an in-vehicle device estimates its own position, it receives the SWLD from the server, uses the SWLD for its own position estimation, and performs obstacle detection. In this case, various methods such as a distance sensor such as a rangefinder, a stereo camera, or a combination of multiple monocular cameras are used to perform obstacle detection based on the three-dimensional information of the surrounding area obtained by itself.

[0311] Also, generally speaking, it is difficult to include VXL data of flat areas in the SWLD. For this reason, the server maintains a downsampled world space (SubWLD) obtained by downsampling the WLD for detecting stationary obstacles, and can send the SWLD and the SubWLD to the client. Accordingly, both the network bandwidth can be suppressed, and the self-position estimation and obstacle detection can be performed on the client side.

[0312] Also, when the client quickly depicts three-dimensional map data, it may be convenient if the map information has a grid structure. Thus, the server can generate a grid based on the WLD and maintain it in advance as a grid world space (MWLD). For example, when the client needs to perform rough three-dimensional depiction, it receives the MWLD, and when it needs to perform detailed three-dimensional depiction, it receives the WLD. Accordingly, the network bandwidth can be suppressed.

[0313] Also, although the server sets the VXL whose feature amount is above the threshold as the FVXL from each VXL, the FVXL can also be calculated by different methods. For example, if the server determines that VXL, VLM, SPC, or GOS constituting a signal or an intersection is required for self-position estimation, driving assistance, or autonomous driving, etc., it can be included in the SWLD as FVXL, FVLM, FSPC, FGOS. And the above determination can be made manually. In addition, the FVXL obtained by the above method can be added to the FVXL set based on the feature amount. That is, the SWLD extraction unit 403 can further extract data corresponding to an object having a predetermined attribute from the input three-dimensional data 411 as the extracted three-dimensional data 412.

[0314] Also, different labels can be assigned to the situations that are required for these uses, different from the feature amount. The server can separately maintain the FVXL required for self-position estimation, driving assistance, or autonomous driving, such as signals or intersections, as an upper layer of the SWLD (for example, lane world space).

[0315] Also, the server can attach attributes to the VXL in the WLD in units of random access or specified units. The attributes include, for example, information indicating whether it is required or not required for self-position estimation, or information indicating whether it is important as traffic information such as signals or intersections. And the attributes can also include the correspondence relationship with Features (intersections or roads, etc.) in lane information (GDF: Geographic DataFiles, etc.).

[0316] Also, as a method for updating the WLD or SWLD, the following method can be adopted.

[0317] Update information such as changes in people, construction, or street trees (facing the trajectory) is loaded into the server as a point cloud or metadata. The server updates the WLD based on this load, and then updates the SWLD using the updated WLD.

[0318] Moreover, when the client detects a mismatch between the three-dimensional information generated by itself during self-position estimation and the three-dimensional information received from the server, it can send the three-dimensional information generated by itself to the server together with an update notification. In this case, the server updates the SWLD using the WLD. If the SWLD is not updated, the server determines that the WLD itself is old.

[0319] In addition, as header information of the encoded stream, information for distinguishing between WLD and SWLD is attached. For example, in a case where there are multiple world spaces such as a grid world space or a lane world space, information for distinguishing them can be attached to the header information. Also, in a case where there are multiple SWLDs with different feature amounts, information for distinguishing them separately can be attached to the header information.

[0320] Furthermore, although the SWLD is composed of FVXLs, it can also include VXLs that are not determined to be FVXLs. For example, the SWLD can include adjacent VXLs used when calculating the feature amounts of FVXLs. Accordingly, even when no feature amount information is attached to each FVXL of the SWLD, the client can calculate the feature amounts of the FVXLs when receiving the SWLD. Also, at this time, the SWLD can include information for distinguishing whether each VXL is an FVXL or a VXL.

[0321] As described above, the three-dimensional data encoding device 400 extracts the extracted three-dimensional data 412 (second three-dimensional data) whose feature amount is equal to or greater than the threshold from the input three-dimensional data 411 (first three-dimensional data), and generates the encoded three-dimensional data 414 (first encoded three-dimensional data) by encoding the extracted three-dimensional data 412.

[0322] Accordingly, the three-dimensional data encoding device 400 generates the encoded three-dimensional data 414 obtained by encoding data whose feature amount is equal to or greater than the threshold. In this way, compared with the case of directly encoding the input three-dimensional data 411, the data amount can be reduced. Therefore, the three-dimensional data encoding device 400 can reduce the data amount during transmission.

[0323] Moreover, the three-dimensional data encoding device 400 further generates the encoded three-dimensional data 413 (second encoded three-dimensional data) by encoding the input three-dimensional data 411.

[0324] Accordingly, the three-dimensional data encoding device 400 can selectively transmit the encoded three-dimensional data 413 and the encoded three-dimensional data 414 according to the usage purpose or the like.

[0325] Moreover, the extracted three-dimensional data 412 is encoded by the first encoding method, and the input three-dimensional data 411 is encoded by the second encoding method different from the first encoding method.

[0326] Accordingly, the three-dimensional data encoding device 400 can adopt appropriate encoding methods for the input three-dimensional data 411 and the extracted three-dimensional data 412 respectively.

[0327] Moreover, in the first encoding method, among intra prediction and inter prediction, inter prediction is prioritized compared with the second encoding method.

[0328] Accordingly, the three-dimensional data encoding device 400 can increase the priority of inter prediction for the extracted three-dimensional data 412 where the correlation between adjacent data is likely to decrease.

[0329] Moreover, in the first encoding method and the second encoding method, the expression methods of three-dimensional positions are different. For example, in the second encoding method, the three-dimensional position is expressed by an octree, and in the first encoding method, the three-dimensional position is expressed by three-dimensional coordinates.

[0330] Accordingly, the three-dimensional data encoding device 400 can adopt a more appropriate expression method of three-dimensional positions for three-dimensional data with different numbers of data (the number of VXL or FVXL).

[0331] Moreover, at least one of the encoded three-dimensional data 413 and 414 includes an identifier indicating whether the encoded three-dimensional data is the encoded three-dimensional data obtained by encoding the input three-dimensional data 411 or the encoded three-dimensional data obtained by encoding a part of the input three-dimensional data 411. That is, this identifier indicates whether the encoded three-dimensional data is the encoded three-dimensional data 413 of WLD or the encoded three-dimensional data 414 of SWLD.

[0332] Accordingly, the decoding device can easily determine whether the acquired encoded three-dimensional data is the encoded three-dimensional data 413 or the encoded three-dimensional data 414.

[0333] Moreover, the three-dimensional data encoding device 400 encodes the extracted three-dimensional data 412 in such a way that the data amount of the encoded three-dimensional data 414 is less than the data amount of the encoded three-dimensional data 413.

[0334] Accordingly, the three-dimensional data encoding device 400 can make the data amount of the encoded three-dimensional data 414 less than the data amount of the encoded three-dimensional data 413.

[0335] Furthermore, the three-dimensional data encoding device 400 further extracts data corresponding to an object having a pre-specified attribute from the input three-dimensional data 411 as the extracted three-dimensional data 412. For example, an object having a pre-specified attribute refers to an object required for self-position estimation, driving assistance, or autonomous driving, etc., such as a signal or an intersection.

[0336] Accordingly, the three-dimensional data encoding device 400 can generate the encoded three-dimensional data 414 including the data required by the decoding device.

[0337] Furthermore, the three-dimensional data encoding device 400 (server) further sends one of the encoded three-dimensional data 413 and 414 to the client according to the state of the client.

[0338] Accordingly, the three-dimensional data encoding device 400 can send appropriate data according to the state of the client.

[0339] Furthermore, the state of the client includes the communication status of the client (e.g., network bandwidth) or the moving speed of the client.

[0340] Furthermore, the three-dimensional data encoding device 400 further sends one of the encoded three-dimensional data 413 and 414 to the client according to the request of the client.

[0341] Accordingly, the three-dimensional data encoding device 400 can send appropriate data according to the request of the client.

[0342] Furthermore, the three-dimensional data decoding device 500 according to the present embodiment decodes the encoded three-dimensional data 413 or 414 generated by the above three-dimensional data encoding device 400.

[0343] That is, the three-dimensional data decoding device 500 decodes the encoded three-dimensional data 414 obtained by encoding the extracted three-dimensional data 412 whose feature amount extracted from the input three-dimensional data 411 is above the threshold value by the first decoding method. And the three-dimensional data decoding device 500 decodes the encoded three-dimensional data 413 obtained by encoding the input three-dimensional data 411 by using a second decoding method different from the first decoding method.

[0344] Accordingly, the three-dimensional data decoding device 500 can selectively receive, for example, according to the usage purpose, etc., the encoded three-dimensional data 414 and the encoded three-dimensional data 413 obtained by encoding the data whose feature amount is above the threshold value. Accordingly, the three-dimensional data decoding device 500 can reduce the amount of data during transmission. Moreover, the three-dimensional data decoding device 500 can adopt appropriate decoding methods for the input three-dimensional data 411 and the extracted three-dimensional data 412 respectively.

[0345] Also, in the first decoding method, among intra prediction and inter prediction, inter prediction is prioritized compared to the second decoding method.

[0346] Accordingly, the three-dimensional data decoding device 500 can increase the priority of inter prediction for the extracted three-dimensional data where the correlation between adjacent data is likely to be low.

[0347] Also, the representation methods of three-dimensional positions are different between the first decoding method and the second decoding method. For example, in the second decoding method, the three-dimensional position is represented by an octree, and in the first decoding method, the three-dimensional position is represented by three-dimensional coordinates.

[0348] Accordingly, the three-dimensional data decoding device 500 can adopt a more appropriate representation method of three-dimensional positions for three-dimensional data with different numbers of data (the number of VXL or FVXL).

[0349] Also, at least one of the encoded three-dimensional data 413 and 414 includes an identifier indicating whether the encoded three-dimensional data is obtained by encoding the input three-dimensional data 411 or by encoding a part of the input three-dimensional data 411. The three-dimensional data decoding device 500 identifies the encoded three-dimensional data 413 and 414 with reference to this identifier.

[0350] Accordingly, the three-dimensional data decoding device 500 can easily determine whether the obtained encoded three-dimensional data is the encoded three-dimensional data 413 or the encoded three-dimensional data 414.

[0351] Also, the three-dimensional data decoding device 500 further notifies the server of the state of the client (the three-dimensional data decoding device 500). The three-dimensional data decoding device 500 receives one of the encoded three-dimensional data 413 and 414 sent from the server according to the state of the client.

[0352] Accordingly, the three-dimensional data decoding device 500 can receive appropriate data according to the state of the client.

[0353] Also, the state of the client includes the communication status of the client (such as network bandwidth) or the moving speed of the client.

[0354] Also, the three-dimensional data decoding device 500 further requests one of the encoded three-dimensional data 413 and 414 from the server and receives one of the encoded three-dimensional data 413 and 414 sent from the server according to this request.

[0355] Accordingly, the three-dimensional data decoding device 500 can receive appropriate data corresponding to the usage.

[0356] (Embodiment 3)

[0357] In this embodiment, a method for transmitting and receiving three-dimensional data between vehicles will be described. For example, three-dimensional data is transmitted and received between the host vehicle and surrounding vehicles.

[0358] Fig.24 It is a block diagram of the three-dimensional data creation device 620 according to this embodiment. The three-dimensional data creation device 620 is included in the host vehicle, for example, and creates a denser third three-dimensional data 636 by synthesizing the received second three-dimensional data 635 with the first three-dimensional data 632 created by the three-dimensional data creation device 620.

[0359] The three-dimensional data creation device 620 includes: a three-dimensional data creation unit 621, a request range determination unit 622, a search unit 623, a reception unit 624, a decoding unit 625, and a synthesis unit 626.

[0360] First, the three-dimensional data creation unit 621 creates the first three-dimensional data 632 using the sensor information 631 detected by the sensors provided in the host vehicle. Next, the request range determination unit 622 determines a request range, which is a three-dimensional space range where the data in the created first three-dimensional data 632 is insufficient.

[0361] Next, the search unit 623 searches for surrounding vehicles that hold three-dimensional data within the request range, and transmits request range information 633 indicating the request range to the surrounding vehicles identified through the search. Next, the reception unit 624 receives the encoded three-dimensional data 634 as an encoded stream within the request range from the surrounding vehicles (S624). In addition, the search unit 623 can issue requests indiscriminately to all vehicles existing within the determined range, and receive the encoded three-dimensional data 634 from the responding party. Also, the search unit 623 is not limited to vehicles, and can also issue requests to objects such as traffic lights or signs, and receive the encoded three-dimensional data 634 from the object.

[0362] Next, the received encoded three-dimensional data 634 is decoded by the decoding unit 625 to obtain the second three-dimensional data 635. Next, the first three-dimensional data 632 and the second three-dimensional data 635 are synthesized by the synthesis unit 626 to create a denser third three-dimensional data 636.

[0363] Next, the configuration and operation of the three-dimensional data transmission device 640 according to this embodiment will be described. Fig.25 It is a block diagram of the three-dimensional data transmission device 640.

[0364] The three-dimensional data transmission device 640 is included, for example, in the surrounding vehicles described above. It processes the fifth three-dimensional data 652 created by the surrounding vehicles into the sixth three-dimensional data 654 requested by its own vehicle, generates encoded three-dimensional data 634 by encoding the sixth three-dimensional data 654, and transmits the encoded three-dimensional data 634 to its own vehicle.

[0365] The three-dimensional data transmission device 640 includes: a three-dimensional data creation unit 641, a reception unit 642, an extraction unit 643, an encoding unit 644, and a transmission unit 645.

[0366] First, the three-dimensional data creation unit 641 creates the fifth three-dimensional data 652 using the sensor information 651 detected by the sensors provided in the surrounding vehicles. Next, the reception unit 642 receives the requested range information 633 transmitted from its own vehicle.

[0367] Next, the extraction unit 643 extracts the three-dimensional data of the requested range indicated by the requested range information 633 from the fifth three-dimensional data 652, and processes the fifth three-dimensional data 652 into the sixth three-dimensional data 654. Then, the encoding unit 644 encodes the sixth three-dimensional data 654, thereby generating the encoded three-dimensional data 634 as an encoded stream. Thus, the transmission unit 645 transmits the encoded three-dimensional data 634 to its own vehicle.

[0368] In addition, here, although an example in which its own vehicle is equipped with the three-dimensional data creation device 620 and the surrounding vehicles are equipped with the three-dimensional data transmission device 640 has been described, each vehicle may also have the functions of the three-dimensional data creation device 620 and the three-dimensional data transmission device 640.

[0369] (Embodiment 4)

[0370] In this embodiment, the operations related to abnormal conditions in the self-position estimation based on the three-dimensional map will be described.

[0371] The use of autonomous movement of moving bodies such as the autonomous driving of motor vehicles, robots, or flying objects such as drones will expand in the future. As an example of a method for realizing such autonomous movement, there is a method in which the moving body estimates its own position in the three-dimensional map (self-position estimation) and travels according to the map.

[0372] Self-position estimation is achieved by matching the three-dimensional map with the three-dimensional information around the own vehicle (hereinafter referred to as the own vehicle detection three-dimensional data) obtained by sensors such as a rangefinder (LIDAR, etc.) or a stereo camera mounted on the own vehicle, and estimating the position of the own vehicle in the three-dimensional map.

[0373] A 3D map, such as the HD map proposed by HERE Technologies, etc., is not only a 3D point cloud, but may also include 2D map data such as road and intersection shape information, or information that changes in real time such as traffic jams and accidents. The 3D map is composed of multiple layers such as 3D data, 2D data, and metadata that changes in real time. The device can obtain only the required data, or can also refer to the required data.

[0374] The data of the point cloud can be the above-mentioned SWLD, or can also include point group data that is not feature points. And the transmission and reception of the data of the point cloud are basically performed in one or more random access units.

[0375] As a method for matching the 3D map with the 3D data detected by the vehicle itself, the following method can be adopted. For example, the device compares the shapes of the point groups in the point clouds of each other, and determines the part with a high similarity between the feature points as the same position. And when the 3D map is composed of SWLD, the device compares the feature points that make up the SWLD with the 3D feature points extracted from the 3D data detected by the vehicle itself and performs matching.

[0376] Here, in order to estimate the vehicle's own position with high accuracy, the following (A) and (B) need to be satisfied. (A) The 3D map and the 3D data detected by the vehicle itself can already be obtained. (B) Their accuracy meets a pre-determined standard. However, in the following abnormal situations, (A) or (B) cannot be satisfied.

[0377] (1) The 3D map cannot be obtained through the communication path.

[0378] (2) There is no 3D map, or the obtained 3D map is damaged.

[0379] (3) The sensors of the vehicle itself malfunction, or due to bad weather, the generation accuracy of the 3D data detected by the vehicle itself is insufficient.

[0380] The following describes the operations for coping with these abnormal situations. Although the following describes the operations by taking a vehicle as an example, the following methods can also be applied to all moving objects that perform autonomous movement such as robots and drones.

[0381] What will be described below is the configuration and operation of the 3D information processing device according to the present embodiment for coping with abnormal situations in the 3D map or the 3D data detected by the vehicle itself. Fig.26 It is a block diagram showing a configuration example of the 3D information processing device 700 according to the present embodiment.

[0382] The 3D information processing device 700 is mounted on a moving object such as a motor vehicle, for example. As Fig.26As shown in the figure, the three-dimensional information processing device 700 includes: a three-dimensional map acquisition unit 701, an own vehicle detection data acquisition unit 702, an abnormal situation determination unit 703, a response operation determination unit 704, and an operation control unit 705.

[0383] In addition, the three-dimensional information processing device 700 may also include a camera for acquiring a two-dimensional image, or may include a two-dimensional or one-dimensional sensor (not shown) such as a sensor for one-dimensional data using ultrasonic waves or lasers for detecting structural objects or moving objects around the own vehicle. Further, the three-dimensional information processing device 700 may also include a communication unit (not shown) for acquiring a three-dimensional map through a mobile communication network such as 4G or 5G, vehicle-to-vehicle communication, or road-to-vehicle communication.

[0384] The three-dimensional map acquisition unit 701 acquires a three-dimensional map 711 near the driving route. For example, the three-dimensional map acquisition unit 701 acquires the three-dimensional map 711 through a mobile communication network, vehicle-to-vehicle communication, or road-to-vehicle communication.

[0385] Next, the own vehicle detection data acquisition unit 702 acquires own vehicle detection three-dimensional data 712 based on sensor information. For example, the own vehicle detection data acquisition unit 702 generates own vehicle detection three-dimensional data 712 based on the sensor information obtained by the sensors equipped on the own vehicle.

[0386] Next, the abnormal situation determination unit 703 detects an abnormal situation by performing a pre-determined check on at least one of the acquired three-dimensional map 711 and own vehicle detection three-dimensional data 712. That is, the abnormal situation determination unit 703 determines whether at least one of the acquired three-dimensional map 711 and own vehicle detection three-dimensional data 712 is abnormal.

[0387] When an abnormal situation is detected, the response operation determination unit 704 determines a response operation for the abnormal situation. Next, the operation control unit 705 controls the operations of the various processing units such as the three-dimensional map acquisition unit 701 required for the implementation of the response operation.

[0388] In addition, when no abnormal situation is detected, the three-dimensional information processing device 700 ends the processing.

[0389] Furthermore, the three-dimensional information processing device 700 estimates the own position of the vehicle equipped with the three-dimensional information processing device 700 by using the three-dimensional map 711 and the own vehicle detection three-dimensional data 712. Next, the three-dimensional information processing device 700 uses the result of the own position estimation to perform autonomous driving of the vehicle.

[0390] Accordingly, the three-dimensional information processing device 700 obtains map data (three-dimensional map 711) including first three-dimensional position information via a channel. For example, the first three-dimensional position information is encoded in units of partial spaces having three-dimensional coordinate information, and the first three-dimensional position information includes a plurality of random access units. Each of the plurality of random access units is an aggregate of one or more partial spaces and can be independently decoded. For example, the first three-dimensional position information is data (SWLD) in which feature points where three-dimensional feature amounts exceed a specified threshold are encoded.

[0391] Furthermore, the three-dimensional information processing device 700 generates second three-dimensional position information (own vehicle detection three-dimensional data 712) based on information detected by a sensor. Next, the three-dimensional information processing device 700 determines whether the first three-dimensional position information or the second three-dimensional position information is abnormal by performing an abnormality determination process on the first three-dimensional position information or the second three-dimensional position information.

[0392] When the three-dimensional information processing device 700 determines that the first three-dimensional position information or the second three-dimensional position information is abnormal, it determines a response operation for the abnormality. Next, the three-dimensional information processing device 700 executes control required for the implementation of the response operation.

[0393] Accordingly, the three-dimensional information processing device 700 can detect an abnormality in the first three-dimensional position information or the second three-dimensional position information and can perform a response operation.

[0394] (Embodiment 5)

[0395] In the present embodiment, a method for transmitting three-dimensional data to a following vehicle and the like will be described.

[0396] Fig. 27 FIG. is a block diagram showing a configuration example of a three-dimensional data production device 810 according to the present embodiment. The three-dimensional data production device 810 is mounted on a vehicle, for example. The three-dimensional data production device 810 transmits and receives three-dimensional data to and from external traffic cloud monitoring, a preceding vehicle, or a following vehicle, and simultaneously produces and stores the three-dimensional data.

[0397] The three-dimensional data production device 810 includes: a data receiving unit 811, a communication unit 812, a reception control unit 813, a format conversion unit 814, a plurality of sensors 815, a three-dimensional data production unit 816, a three-dimensional data synthesis unit 817, a three-dimensional data storage unit 818, a communication unit 819, a transmission control unit 820, a format conversion unit 821, and a data transmission unit 822.

[0398] The data receiving unit 811 receives three-dimensional data 831 from the traffic cloud monitoring or the vehicle ahead. The three-dimensional data 831 includes, for example, point clouds, visible light images, depth information, sensor position information, or speed information that contain information on areas that cannot be detected by the sensors 815 of the host vehicle itself.

[0399] The communication unit 812 communicates with the traffic cloud monitoring or the vehicle ahead, and sends data transmission requests and the like to the traffic cloud monitoring or the vehicle ahead.

[0400] The reception control unit 813 exchanges information such as corresponding formats with the communication partner via the communication unit 812 to establish communication with the communication partner.

[0401] The format conversion unit 814 generates three-dimensional data 832 by performing format conversion and the like on the three-dimensional data 831 received by the data receiving unit 811. Moreover, when the three-dimensional data 831 is compressed or encoded, the format conversion unit 814 performs decompression or decoding processing.

[0402] The multiple sensors 815 are a group of sensors such as LiDAR, visible light cameras, or infrared cameras that obtain information on the outside of the vehicle, and generate sensor information 833. For example, when the sensor 815 is a laser sensor such as LiDAR, the sensor information 833 is three-dimensional data such as a point cloud (point group data). In addition, the sensor 815 may not be multiple.

[0403] The three-dimensional data production unit 816 generates three-dimensional data 834 based on the sensor information 833. The three-dimensional data 834 includes, for example, information such as point clouds, visible light images, depth information, sensor position information, or speed information.

[0404] The three-dimensional data synthesis unit 817 synthesizes the three-dimensional data 832 produced by the traffic cloud monitoring or the vehicle ahead, etc., into the three-dimensional data 834 produced based on the sensor information 833 of the host vehicle itself, so as to be able to construct three-dimensional data 835 that also includes the space in front of the vehicle ahead that cannot be detected by the sensors 815 of the host vehicle itself.

[0405] The three-dimensional data storage unit 818 stores the generated three-dimensional data 835, etc.

[0406] The communication unit 819 communicates with the traffic cloud monitoring or the vehicle behind, and sends data transmission requests and the like to the traffic cloud monitoring or the vehicle behind.

[0407] The transmission control unit 820 exchanges information such as the corresponding format with the communication partner via the communication unit 819 to establish communication with the communication partner. Further, the transmission control unit 820 determines the transmission area of the space of the three-dimensional data to be transmitted based on the three-dimensional data construction information of the three-dimensional data 832 generated by the three-dimensional data synthesis unit 817 and the data transmission request from the communication partner.

[0408] Specifically, the transmission control unit 820 determines the transmission area of the space in front of the host vehicle that cannot be detected by the sensors of the following vehicle in accordance with the data transmission request from the traffic cloud monitoring or the following vehicle. Further, the transmission control unit 820 determines the transmission area by judging, based on the three-dimensional data construction information, whether there is an update to the space that can be transmitted or the transmitted space. For example, the transmission control unit 820 determines as the transmission area the area that is specified by the data transmission request and where the corresponding three-dimensional data 835 exists. Then, the transmission control unit 820 notifies the format conversion unit 821 of the format corresponding to the communication partner and the transmission area.

[0409] The format conversion unit 821 generates the three-dimensional data 837 by converting the three-dimensional data 836 of the transmission area in the three-dimensional data 835 stored in the three-dimensional data storage unit 818 into a format corresponding to the receiving side. Additionally, the format conversion unit 821 may compress or encode the three-dimensional data 837 to reduce the data volume.

[0410] The data transmission unit 822 transmits the three-dimensional data 837 to the traffic cloud monitoring or the following vehicle. The three-dimensional data 837 includes, for example, the point cloud, visible light image, depth information, or sensor position information in front of the host vehicle that includes information on the area that is a blind spot for the following vehicle.

[0411] In addition, although the format conversion units 814 and 821 are taken as examples for format conversion and the like, format conversion may not be performed.

[0412] With this configuration, the three-dimensional data production device 810 obtains the three-dimensional data 831 of the area that cannot be detected by the sensors 815 of the host vehicle from the outside, and generates the three-dimensional data 835 by synthesizing the three-dimensional data 831 and the three-dimensional data 834 based on the sensor information 833 detected by the sensors 815 of the host vehicle. Accordingly, the three-dimensional data production device 810 can generate the three-dimensional data of the range that cannot be detected by the sensors 815 of the host vehicle.

[0413] Further, the three-dimensional data production device 810 can transmit the three-dimensional data of the space in front of the host vehicle that cannot be detected by the sensors of the following vehicle to the traffic cloud monitoring or the following vehicle in accordance with the data transmission request from the traffic cloud monitoring or the following vehicle.

[0414] (Embodiment 6)

[0415] In the example to be described in Embodiment 5, a client device such as a vehicle sends three-dimensional data to another vehicle or a server such as a traffic cloud monitor. In the present embodiment, the client device sends sensor information obtained by a sensor to the server or another client device.

[0416] First, the configuration of the system according to the present embodiment will be described. Fig.28 The configuration of the three-dimensional map and the sensor information transceiver system according to the present embodiment is shown. The system includes a server 901, client devices 902A and 902B. In addition, when the client devices 902A and 902B are not particularly distinguished, they are also denoted as the client device 902.

[0417] The client device 902 is, for example, an in-vehicle device mounted on a moving body such as a vehicle. The server 901 is, for example, a traffic cloud monitor or the like and can communicate with a plurality of client devices 902.

[0418] The server 901 sends a three-dimensional map composed of point clouds to the client device 902. In addition, the composition of the three-dimensional map is not limited to point clouds and can also be represented by other three-dimensional data such as a grid structure.

[0419] The client device 902 sends sensor information obtained by the client device 902 to the server 901. The sensor information includes, for example, at least one of information obtained by LiDAR, visible light images, infrared images, depth images, sensor position information, and speed information.

[0420] Regarding the data transmitted and received between the server 901 and the client device 902, it can be compressed when reducing data is desired, and can be not compressed when maintaining the accuracy of the data is desired. When compressing the data, for example, a three-dimensional compression method based on an octree can be adopted in the point cloud. And, a two-dimensional image compression method can be adopted in visible light images, infrared images, and depth images. The two-dimensional image compression method is, for example, MPEG-4 AVC or HEVC standardized by MPEG.

[0421] Furthermore, the server 901 sends the 3D map managed by the server 901 to the client device 902 in accordance with a transmission request for the 3D map from the client device 902. Additionally, the server 901 may also send the 3D map without waiting for a transmission request for the 3D map from the client device 902. For example, the server 901 may broadcast the 3D map to one or more client devices 902 in a pre-specified space. Also, the server 901 may send a 3D map suitable for the position of the client device 902 to the client device 902 that has received a transmission request once at regular intervals. Moreover, the server 901 may send the 3D map to the client device 902 whenever the 3D map managed by the server 901 is updated.

[0422] The client device 902 issues a transmission request for the 3D map to the server 901. For example, when the client device 902 wants to estimate its own position while in motion, the client device 902 sends a transmission request for the 3D map to the server 901.

[0423] In addition, in the following cases, the client device 902 may also issue a transmission request for the 3D map to the server 901. When the 3D map held by the client device 902 is relatively old, the client device 902 may issue a transmission request for the 3D map to the server 901. For example, when a certain period has elapsed since the client device 902 obtained the 3D map, the client device 902 may issue a transmission request for the 3D map to the server 901.

[0424] It may also be that, a certain time before the client device 902 is about to leave the space shown in the 3D map held by the client device 902, the client device 902 issues a transmission request for the 3D map to the server 901. For example, it may also be that when the client device 902 is within a pre-specified distance from the boundary of the space shown in the 3D map held by the client device 902, the client device 902 issues a transmission request for the 3D map to the server 901. Also, when the movement path and movement speed of the client device 902 are known, the time when the client device 902 leaves the space shown in the 3D map held by the client device 902 can be predicted based on the known movement path and movement speed.

[0425] When the error in the position comparison between the 3D data created by the client device 902 based on sensor information and the 3D map is above a certain range, the client device 902 may issue a transmission request for the 3D map to the server 901.

[0426] The client device 902 sends the sensor information to the server 901 in accordance with the transmission request of the sensor information sent from the server 901. In addition, the client device 902 may also send the sensor information to the server 901 without waiting for the transmission request of the sensor information from the server 901. For example, in the case where the client device 902 has received a transmission request for sensor information from the server 901 once, it may periodically send the sensor information to the server 901 within a certain period. And it may also be that when the error in the position comparison between the three-dimensional data generated based on the sensor information by the client device 902 and the three-dimensional map obtained from the server 901 is above a certain range, the client device 902 determines that there is a possibility that the three-dimensional map around the client device 902 has changed, and sends this judgment result together with the sensor information to the server 901.

[0427] The server 901 sends a transmission request for sensor information to the client device 902. For example, the server 901 receives the position information of the client device 902 such as GPS from the client device 902. Based on the position information of the client device 902, when it is determined that the client device 902 is approaching a space with less information in the three-dimensional map managed by the server 901, in order to regenerate the three-dimensional map, a transmission request for sensor information is sent to the client device 902. And it may also be that the server 901 may send a transmission request for sensor information when it wants to update the three-dimensional map, when it wants to confirm the road conditions such as during snow accumulation or disasters, or when it wants to confirm the congestion conditions or accident conditions, etc.

[0428] And it may also be that the client device 902 sets the data volume of the sensor information sent to the server 901 according to the communication state or frequency band at the time of receiving the transmission request of the sensor information received from the server 901. Setting the data volume of the sensor information sent to the server 901, for example, means increasing or decreasing the data itself, or selecting an appropriate compression method.

[0429] Fig.29 It is a block diagram showing a configuration example of the client device 902. The client device 902 receives a three-dimensional map composed of point clouds, etc. from the server 901, and estimates its own position based on the three-dimensional data generated based on the sensor information of the client device 902. And the client device 902 sends the obtained sensor information to the server 901.

[0430] The client device 902 includes: a data receiving unit 1011, a communication unit 1012, a reception control unit 1013, a format conversion unit 1014, a plurality of sensors 1015, a three-dimensional data creation unit 1016, a three-dimensional image processing unit 1017, a three-dimensional data storage unit 1018, a format conversion unit 1019, a communication unit 1020, a transmission control unit 1021, and a data transmission unit 1022.

[0431] The data receiving unit 1011 receives a three-dimensional map 1031 from the server 901. The three-dimensional map 1031 is data including point clouds such as WLD or SWLD. The three-dimensional map 1031 may include either compressed data or uncompressed data.

[0432] The communication unit 1012 communicates with the server 901 and sends a data transmission request (e.g., a request to send a three-dimensional map) to the server 901.

[0433] The reception control unit 1013 exchanges information such as the corresponding format with the communication partner via the communication unit 1012 and establishes communication with the communication partner.

[0434] The format conversion unit 1014 generates a three-dimensional map 1032 by performing format conversion etc. on the three-dimensional map 1031 received by the data receiving unit 1011. Further, when the three-dimensional map 1031 is compressed or encoded, the format conversion unit 1014 performs decompression or decoding processing. In addition, when the three-dimensional map 1031 is uncompressed data, the format conversion unit 1014 does not perform decompression or decoding processing.

[0435] The plurality of sensors 1015 are a group of sensors mounted on the client device 902 such as LiDAR, visible light cameras, infrared cameras, or depth sensors, which are used to obtain information about the outside of the vehicle, and generate sensor information 1033. For example, when the sensor 1015 is a laser sensor such as LiDAR, the sensor information 1033 is three-dimensional data such as point clouds (point group data). In addition, the sensor 1015 may not be plural.

[0436] The three-dimensional data creation unit 1016 creates three-dimensional data 1034 around its own vehicle based on the sensor information 1033. For example, the three-dimensional data creation unit 1016 uses the information obtained by LiDAR and the visible light images obtained by visible light cameras to create point cloud data with color information around its own vehicle.

[0437] The 3D image processing unit 1017 performs self-position estimation processing of the host vehicle and the like using the received 3D map 1032 such as point cloud and the 3D data 1034 around the host vehicle generated based on the sensor information 1033. Alternatively, the 3D image processing unit 1017 may synthesize the 3D map 1032 and the 3D data 1034 to produce the 3D data 1035 around the host vehicle, and perform self-position estimation processing using the produced 3D data 1035.

[0438] The 3D data storage unit 1018 stores the 3D map 1032, the 3D data 1034, the 3D data 1035, and the like.

[0439] The format conversion unit 1019 generates the sensor information 1037 by converting the sensor information 1033 into the format corresponding to the receiving side. In addition, the format conversion unit 1019 may reduce the data volume by compressing or encoding the sensor information 1037. And when format conversion is not required, the format conversion unit 1019 may omit the processing. Also, the format conversion unit 1019 can control the data volume to be transmitted according to the specified transmission range.

[0440] The communication unit 1020 communicates with the server 901 and receives a data transmission request (a transmission request for sensor information) and the like from the server 901.

[0441] The transmission control unit 1021 exchanges information such as the corresponding format with the communication partner via the communication unit 1020 to establish communication.

[0442] The data transmission unit 1022 transmits the sensor information 1037 to the server 901. The sensor information 1037 includes, for example, information obtained by LiDAR, a luminance image (visible light image) obtained by a visible light camera, an infrared image obtained by an infrared camera, a depth image obtained by a depth sensor, sensor position information, and speed information, etc., which are obtained by the plurality of sensors 1015.

[0443] Next, the configuration of the server 901 will be described. Fig.30 It is a block diagram showing a configuration example of the server 901. The server 901 receives the sensor information sent from the client device 902 and produces 3D data based on the received sensor information. The server 901 updates the 3D map managed by the server 901 using the produced 3D data. And the server 901 sends the updated 3D map to the client device 902 according to the transmission request of the 3D map from the client device 902.

[0444] The server 901 includes: a data receiving unit 1111, a communication unit 1112, a reception control unit 1113, a format conversion unit 1114, a three-dimensional data creation unit 1116, a three-dimensional data synthesis unit 1117, a three-dimensional data storage unit 1118, a format conversion unit 1119, a communication unit 1120, a transmission control unit 1121, and a data transmission unit 1122.

[0445] The data receiving unit 1111 receives sensor information 1037 from the client device 902. The sensor information 1037 includes, for example, information obtained by LiDAR, a luminance image (visible light image) obtained by a visible light camera, an infrared image obtained by an infrared camera, a depth image obtained by a depth sensor, sensor position information, and speed information.

[0446] The communication unit 1112 communicates with the client device 902 and sends a data transmission request (e.g., a transmission request for sensor information) to the client device 902.

[0447] The reception control unit 1113 exchanges information such as the corresponding format with the communication partner via the communication unit 1112 to establish communication.

[0448] When the received sensor information 1037 is compressed or encoded, the format conversion unit 1114 generates sensor information 1132 by performing decompression or decoding processing. Additionally, when the sensor information 1037 is uncompressed data, the format conversion unit 1114 does not perform decompression or decoding processing.

[0449] The three-dimensional data creation unit 1116 creates three-dimensional data 1134 around the client device 902 based on the sensor information 1132. For example, the three-dimensional data creation unit 1116 uses the information obtained by LiDAR and the visible light image obtained by the visible light camera to create point cloud data with color information around the client device 902.

[0450] The three-dimensional data synthesis unit 1117 synthesizes the three-dimensional data 1134 created based on the sensor information 1132 with the three-dimensional map 1135 managed by the server 901, and thereby updates the three-dimensional map 1135.

[0451] The three-dimensional data storage unit 1118 stores the three-dimensional map 1135 and the like.

[0452] The format conversion unit 1119 generates the three-dimensional map 1031 by converting the three-dimensional map 1135 into a format corresponding to the receiving side. Additionally, the format conversion unit 1119 can also reduce the data volume by compressing or encoding the three-dimensional map 1135. Moreover, when format conversion is not required, the format conversion unit 1119 can also omit the processing. And the format conversion unit 1119 can control the data volume to be sent according to the specified transmission range.

[0453] The communication unit 1120 communicates with the client device 902 and receives a data transmission request (a transmission request for a three-dimensional map), etc. from the client device 902.

[0454] The transmission control unit 1121 exchanges information such as the corresponding format with the communication partner via the communication unit 1120 to establish communication.

[0455] The data transmission unit 1122 transmits the three-dimensional map 1031 to the client device 902. The three-dimensional map 1031 is data including point clouds such as WLD or SWLD. Either compressed data or uncompressed data may be included in the three-dimensional map 1031.

[0456] Next, the operation process of the client device 902 will be described. Fig.31 It is a flowchart showing the operations when the client device 902 obtains a three-dimensional map.

[0457] First, the client device 902 requests the transmission of a three-dimensional map (point cloud, etc.) from the server 901 (S1001). At this time, the client device 902 also transmits the position information of the client device 902 obtained through GPS, etc. Accordingly, it is possible to request the server 901 to transmit a three-dimensional map related to this position information.

[0458] Next, the client device 902 receives the three-dimensional map from the server 901 (S1002). If the received three-dimensional map is compressed data, the client device 902 decodes the received three-dimensional map to generate an uncompressed three-dimensional map (S1003).

[0459] Next, the client device 902 creates three-dimensional data 1034 around the client device 902 based on the sensor information 1033 obtained from the plurality of sensors 1015 (S1004). Next, the client device 902 estimates its own position using the three-dimensional map 1032 received from the server 901 and the three-dimensional data 1034 created based on the sensor information 1033 (S1005).

[0460] Fig.32It is a flowchart showing the operation when the client device 902 transmits sensor information. First, the client device 902 receives a sensor information transmission request from the server 901 (S1011). The client device 902 that has received the transmission request transmits the sensor information 1037 to the server 901 (S1012). Additionally, when the sensor information 1033 includes multiple pieces of information obtained by multiple sensors 1015, the client device 902 compresses each piece of information in a compression method suitable for each piece of information to generate the sensor information 1037.

[0461] Next, the operation process of the server 901 will be described. Fig.33 It is a flowchart showing the operation when the server 901 obtains sensor information. First, the server 901 requests the client device 902 to transmit sensor information (S1021). Next, the server 901 receives the sensor information 1037 transmitted from the client device 902 in accordance with this request (S1022). Next, the server 901 uses the received sensor information 1037 to create three-dimensional data 1134 (S1023). Next, the server 901 reflects the created three-dimensional data 1134 onto the three-dimensional map 1135 (S1024).

[0462] Fig.34 It is a flowchart showing the operation when the server 901 transmits the three-dimensional map. First, the server 901 receives a three-dimensional map transmission request from the client device 902 (S1031). The server 901 that has received the three-dimensional map transmission request transmits the three-dimensional map 1031 to the client device 902 (S1032). At this time, the server 901 can extract the three-dimensional map in its vicinity corresponding to the location information of the client device 902 and transmit the extracted three-dimensional map. And it can be that the server 901 compresses the three-dimensional map composed of point clouds, for example, using a compression method such as an octree, and transmits the compressed three-dimensional map.

[0463] Hereinafter, a modification example of the present embodiment will be described.

[0464] The server 901 uses the sensor information 1037 received from the client device 902 to create three-dimensional data 1134 near the location of the client device 902. Next, the server 901 matches the created three-dimensional data 1134 with the three-dimensional map 1135 of the same area managed by the server 901, and calculates the difference between the three-dimensional data 1134 and the three-dimensional map 1135. When the difference is equal to or greater than a predetermined threshold, the server 901 determines that some abnormality has occurred around the client device 902. For example, when the ground surface sinks due to natural disasters such as earthquakes, a large difference may occur between the three-dimensional map 1135 managed by the server 901 and the three-dimensional data 1134 created based on the sensor information 1037.

[0465] The sensor information 1037 may also include at least one of the type of sensor, the performance of the sensor, and the model of the sensor. Also, it may be that a category ID corresponding to the performance of the sensor is attached to the sensor information 1037. For example, when the sensor information 1037 is information obtained by LiDAR, it can be considered to assign an identifier according to the performance of the sensor. For example, category 1 is assigned to a sensor that can obtain information with an accuracy of several millimeters, category 2 is assigned to a sensor that can obtain information with an accuracy of several centimeters, and category 3 is assigned to a sensor that can obtain information with an accuracy of several meters. Also, the server 901 can estimate the performance information of the sensor from the model of the client device 902. For example, when the client device 902 is mounted on a vehicle, the server 901 can determine the specification information of the sensor according to the model of the vehicle. In this case, the server 901 can obtain the information of the vehicle model in advance, or include this information in the sensor information. Also, it may be that the server 901 uses the obtained sensor information 1037 to switch the degree of correction for the three-dimensional data 1134 created using the sensor information 1037. For example, when the sensor performance is high accuracy (category 1), the server 901 does not perform correction on the three-dimensional data 1134. When the sensor performance is low accuracy (category 3), the server 901 applies correction suitable for the accuracy of the sensor to the three-dimensional data 1134. For example, the lower the accuracy of the sensor, the greater the degree (intensity) of correction by the server 901.

[0466] The server 901 can also send a request to multiple client devices 902 existing in a certain space to send sensor information simultaneously. When the server 901 receives multiple sensor information from the multiple client devices 902, it is not necessary to utilize all the sensor information for the production of the three-dimensional data 1134. For example, the sensor information to be utilized can be selected according to the performance of the sensors. For example, when the server 901 updates the three-dimensional map 1135, it can select high-precision sensor information (category 1) from the received multiple sensor information and utilize the selected sensor information to produce the three-dimensional data 1134.

[0467] The server 901 is not limited to servers such as traffic cloud monitoring, and can also be other client devices (in-vehicle). Fig.35 The system configuration in this case is shown.

[0468] For example, the client device 902C sends a request to the client device 902A existing nearby to send sensor information, and obtains the sensor information from the client device 902A. Then, the client device 902C utilizes the obtained sensor information of the client device 902A to produce three-dimensional data and updates the three-dimensional map of the client device 902C. In this way, the client device 902C can utilize the performance of the client device 902C to generate a three-dimensional map of the space that can be obtained from the client device 902A. For example, this situation can be considered when the performance of the client device 902C is high.

[0469] Moreover, in this case, the client device 902A that provided the sensor information is given the right to obtain the high-precision three-dimensional map generated by the client device 902C. The client device 902A receives the high-precision three-dimensional map from the client device 902C according to this right.

[0470] It can also be that the client device 902C sends a request to multiple client devices 902 (client device 902A and client device 902B) existing nearby to send sensor information. When the sensor of the client device 902A or the client device 902B is of high performance, the client device 902C can utilize the sensor information obtained through this high-performance sensor to produce three-dimensional data.

[0471] Fig.36 It is a block diagram showing the functional configurations of the server 901 and the client device 902. The server 901 includes, for example: a three-dimensional map compression / decoding processing unit 1201 for compressing and decoding the three-dimensional map, and a sensor information compression / decoding processing unit 1202 for compressing and decoding the sensor information.

[0472] The client device 902 includes: a 3D map decoding processing unit 1211 and a sensor information compression processing unit 1212. The 3D map decoding processing unit 1211 receives the encoded data of the compressed 3D map, decodes the encoded data, and obtains the 3D map. The sensor information compression processing unit 1212 does not compress the 3D data created from the acquired sensor information, but compresses the sensor information itself, and sends the encoded data of the compressed sensor information to the server 901. With this configuration, the client device 902 can keep the processing unit (device or LSI) for decoding the 3D map (point cloud, etc.) inside, without having to keep the processing unit for compressing the 3D data of the 3D map (point cloud, etc.) inside. In this way, the cost and power consumption of the client device 902 can be suppressed.

[0473] As described above, the client device 902 according to this embodiment is mounted on a moving body, and creates 3D data 1034 of the surroundings of the moving body based on the sensor information 1033 showing the surroundings of the moving body obtained by the sensor 1015 mounted on the moving body. The client device 902 estimates its own position using the created 3D data 1034. The client device 902 sends the acquired sensor information 1033 to the server 901 or another moving body 902.

[0474] Accordingly, the client device 902 sends the sensor information 1033 to the server 901 or the like. In this way, there is a possibility that the amount of data to be transmitted can be reduced compared with the case of transmitting 3D data. And since there is no need to perform processing such as compression or encoding of 3D data in the client device 902, the amount of processing of the client device 902 can be reduced. Therefore, the client device 902 can achieve a reduction in the amount of transmitted data or a simplification of the device configuration.

[0475] Furthermore, the client device 902 further sends a transmission request for the 3D map to the server 901, and receives the 3D map 1031 from the server 901. The client device 902 estimates its own position using the 3D data 1034 and the 3D map 1032 in the estimation of its own position.

[0476] Moreover, the sensor information 1033 includes at least one of the information obtained by a laser sensor, a luminance image (visible light image), an infrared image, a depth image, the position information of the sensor, and the speed information of the sensor.

[0477] Moreover, the sensor information 1033 includes information showing the performance of the sensor.

[0478] Further, the client device 902 encodes or compresses the sensor information 1033, and in the transmission of the sensor information, transmits the encoded or compressed sensor information 1037 to the server 901 or another moving body 902. Accordingly, the client device 902 can reduce the amount of data transmitted.

[0479] For example, the client device 902 includes a processor and a memory, and the processor uses the memory to perform the above processing.

[0480] Further, the server 901 according to the present embodiment can communicate with the client device 902 mounted on the moving body, and receive the sensor information 1037 showing the surrounding conditions of the moving body obtained by the sensor 1015 mounted on the moving body from the client device 902. The server 901 creates three-dimensional data 1134 of the periphery of the moving body according to the received sensor information 1037.

[0481] Accordingly, the server 901 uses the sensor information 1037 transmitted from the client device 902 to create the three-dimensional data 1134. In this way, compared with the case where the client device 902 transmits the three-dimensional data, there is a possibility of reducing the amount of data to be transmitted. And since it is not necessary to perform processing such as compression or encoding of the three-dimensional data on the client device 902, the processing amount of the client device 902 can be reduced. In this way, the server 901 can achieve a reduction in the amount of data transmitted or a simplification of the device configuration.

[0482] Further, the server 901 further sends a transmission request for the sensor information to the client device 902.

[0483] Further, the server 901 further uses the created three-dimensional data 1134 to update the three-dimensional map 1135, and according to the transmission request of the three-dimensional map 1135 from the client device 902, transmits the three-dimensional map 1135 to the client device 902.

[0484] And the sensor information 1037 includes at least one of the information obtained by the laser sensor, the luminance image (visible light image), the infrared image, the depth image, the position information of the sensor, and the speed information of the sensor.

[0485] And the sensor information 1037 includes the information showing the performance of the sensor.

[0486] Further, the server 901 further corrects the three-dimensional data according to the performance of the sensor. Accordingly, the quality of the three-dimensional data can be improved by this three-dimensional data creation method.

[0487] Further, in receiving the sensor information, the server 901 receives a plurality of pieces of sensor information 1037 from a plurality of client devices 902, and selects the sensor information 1037 to be used in the production of the three-dimensional data 1134 based on a plurality of pieces of information included in the plurality of pieces of sensor information 1037 that indicate the performance of the sensors. Accordingly, the server 901 can improve the quality of the three-dimensional data 1134.

[0488] Further, the server 901 decodes or decompresses the received sensor information 1037, and produces three-dimensional data 1134 based on the decoded or decompressed sensor information 1132. Accordingly, the server 901 can reduce the amount of data to be transmitted.

[0489] For example, the server 901 includes a processor and a memory, and the processor uses the memory to perform the above processing.

[0490] (Embodiment 7)

[0491] In the present embodiment, a method for encoding and decoding three-dimensional data using inter-frame prediction processing will be described.

[0492] Fig.37 is a block diagram of a three-dimensional data encoding apparatus 1300 according to the present embodiment. The three-dimensional data encoding apparatus 1300 generates an encoded bitstream (hereinafter also simply referred to as a bitstream) as an encoded signal by encoding three-dimensional data. As Fig.37 shown, the three-dimensional data encoding apparatus 1300 includes: a segmentation unit 1301, a subtraction unit 1302, a transformation unit 1303, a quantization unit 1304, an inverse quantization unit 1305, an inverse transformation unit 1306, an addition unit 1307, a reference volume memory 1308, an intra-frame prediction unit 1309, a reference space memory 1310, an inter-frame prediction unit 1311, a prediction control unit 1312, and an entropy encoding unit 1313.

[0493] The segmentation unit 1301 divides each space (SPC) included in the three-dimensional data into a plurality of volumes (VLM) as encoding units. Further, the segmentation unit 1301 performs octree representation (Octreeization) on the voxels within each volume. In addition, the segmentation unit 1301 may make the space and the volume the same size and perform octree representation on the space. Further, the segmentation unit 1301 may attach information (depth information, etc.) required for octreeization to the head of the bitstream or the like.

[0494] The subtraction unit 1302 calculates the difference between the volume (encoding target volume) output from the segmentation unit 1301 and the predicted volume generated by intra-frame prediction or inter-frame prediction described later, and outputs the calculated difference as a prediction residual to the transformation unit 1303. Fig.38An example of calculating the prediction residual is shown. In addition, the bit strings of the volume to be encoded and the predicted volume shown here are, for example, position information indicating the positions of three-dimensional points (e.g., point cloud) included in the volume.

[0495] Hereinafter, the octree representation and the scanning order of voxels will be described. After the volume is transformed into an octree structure (octree conversion), it is encoded. The octree structure consists of nodes and leaf nodes. Each node has eight nodes or leaf nodes, and each leaf node has voxel (VXL) information. Fig.39 A configuration example of a volume including a plurality of voxels is shown. Fig.40 An example is shown of Fig.39 transforming the volume shown into an octree structure. Here, Fig.40 Among the leaf nodes shown, leaf nodes 1, 2, and 3 respectively represent Fig.39 the voxels VXL1, VXL2, and VXL3 shown, and represent a VXL including a point group (hereinafter referred to as an effective VXL).

[0496] The octree is represented, for example, by a binary sequence of 0 and 1. For example, when a node or an effective VXL is set to a value of 1 and the others are set to a value of 0, the binary sequence shown is assigned to each node and leaf node. Then, in accordance with a breadth-first or depth-first scanning order, this binary sequence is scanned. For example, when scanning is performed in a breadth-first manner, the binary sequence shown in A of Fig.40 is obtained. When scanning is performed in a depth-first manner, the binary sequence shown in B of Fig.41 is obtained. The binary sequence obtained by this scanning is encoded by entropy coding, thereby reducing the amount of information. Fig.41

[0497] Next, the depth information in the octree representation will be described. The depth in the octree representation is used for controlling up to which granularity the point cloud information included in the volume is retained. If the depth is set large, the point cloud information can be reproduced at a finer level, but the amount of data for representing the nodes and leaf nodes will increase. On the contrary, if the depth is set small, although the amount of data can be reduced, point cloud information at multiple different positions and with different colors will be regarded as being at the same position and having the same color, and thus the information originally possessed by the point cloud information will be lost.

[0498] For example, Fig.42 an example is shown of Fig.40 representing an octree with a depth of 2 shown as an octree with a depth of 1. Fig.42 The octree shown has less data volume than Fig.40 the octree shown. That is, Fig.42 the octree shown and Fig.42Compared with the octree shown, the number of bits after binary serialization is less. Fig.40 The leaf nodes 1 and 2 shown in the figure become Fig.41 The leaf node 1 shown is shown. That is, Fig.40 The leaf node 1 and the leaf node 2 shown are information of different locations.

[0499] Fig.43 Shown with Fig.42 The volume corresponding to the octree shown. Fig.39 The VXL1 and VXL2 shown are Fig.43 In this case, the three-dimensional data encoding device 1300 corresponds to VXL12 shown in FIG. Fig.39 The color information of VXL1 and VXL2 shown in the figure generates Fig.43 For example, the three-dimensional data encoding device 1300 calculates the color information of VXL1 and VXL2 as the color information of VXL12 using the average value, median value, or weighted average value. In this way, the three-dimensional data encoding device 1300 can control the reduction of the data amount by changing the depth of the octree.

[0500] The three-dimensional data encoding device 1300 may also use any one of the world space units, space units, and volume units to set the depth information of the octree. In addition, at this time, the three-dimensional data encoding device 1300 may also attach the depth information to the header information of the world space, the header information of the space, or the header information of the volume. In addition, the same value may be used as the depth information in all world spaces, spaces, and volumes at different times. In this case, the three-dimensional data encoding device 1300 may also attach the depth information to the header information that manages the world space at all times.

[0501] When the voxel contains color information, the transformation unit 1303 applies a frequency transformation such as an orthogonal transformation to the prediction residual of the color information of the voxels in the volume. For example, the transformation unit 1303 scans the prediction residual in a certain scanning order to produce a one-dimensional arrangement. Thereafter, the transformation unit 1303 transforms the one-dimensional arrangement into the frequency domain by applying a one-dimensional orthogonal transformation to the produced one-dimensional arrangement. Accordingly, when the value of the prediction residual in the volume is close, the value of the frequency component of the low frequency band becomes larger, and the value of the frequency component of the high frequency band becomes smaller. Therefore, the quantization unit 1304 can more effectively reduce the amount of coding.

[0502] Further, the transformation unit 1303 may use an orthogonal transformation of two or more dimensions instead of a one-dimensional orthogonal transformation. For example, the transformation unit 1303 maps the prediction residual into a two-dimensional arrangement in a certain scanning order, and applies a two-dimensional orthogonal transformation to the obtained two-dimensional arrangement. Further, the transformation unit 1303 may select the orthogonal transformation method to be used from among a plurality of orthogonal transformation methods. In this case, the three-dimensional data encoding apparatus 1300 attaches information indicating which orthogonal transformation method is used to the bitstream. Further, it may be that the transformation unit 1303 selects the orthogonal transformation method to be used from among a plurality of orthogonal transformation methods with different dimensions. In this case, the three-dimensional data encoding apparatus 1300 attaches information indicating which dimensional orthogonal transformation method is used to the bitstream.

[0503] For example, the transformation unit 1303 matches the scanning order of the prediction residual with the scanning order (such as breadth-first or depth-first) in the octree within the volume. Accordingly, since it is not necessary to attach information indicating the scanning order of the prediction residual to the bitstream, the overhead can be reduced. Further, the transformation unit 1303 may apply a scanning order different from the scanning order of the octree. In this case, the three-dimensional data encoding apparatus 1300 attaches information indicating the scanning order of the prediction residual to the bitstream. Accordingly, the three-dimensional data encoding apparatus 1300 can efficiently encode the prediction residual. Further, it may be that the three-dimensional data encoding apparatus 1300 attaches information (such as a flag) indicating whether the scanning order of the octree is applied to the bitstream, and in the case where the scanning order of the octree is not applied, attaches information indicating the scanning order of the prediction residual to the bitstream.

[0504] The transformation unit 1303 may transform not only the prediction residual of the color information but also other attribute information possessed by the voxels. For example, it may be that the transformation unit 1303 transforms and encodes information such as reflectance obtained when acquiring a point cloud by LiDAR or the like.

[0505] When the space does not have attribute information such as color information, the transformation unit 1303 may skip the processing. Further, the three-dimensional data encoding apparatus 1300 may attach information (flag) indicating whether to skip the processing of the transformation unit 1303 to the bitstream.

[0506] The quantization unit 1304 quantizes the frequency components of the prediction residuals generated by the transformation unit 1303 using quantization control parameters, thereby generating quantization coefficients. This reduces the amount of information. The generated quantization coefficients are output to the entropy coding unit 1313. The quantization unit 1304 can control the quantization control parameters in world space units, spatial units, or volume units. At this time, the three-dimensional data encoding device 1300 attaches the quantization control parameters to respective header information and the like. Also, the quantization unit 1304 can change weights for quantization control according to the frequency components of each prediction residual. For example, the quantization unit 1304 can perform detailed quantization on low-frequency components and rough quantization on high-frequency components. In this case, the three-dimensional data encoding device 1300 can attach a parameter indicating the weights of the respective frequency components to the header.

[0507] When the quantization unit 1304 does not have attribute information such as color information in the space, the processing can be skipped. Also, the three-dimensional data encoding device 1300 can attach information (flag) indicating whether the processing of the quantization unit 1304 is skipped to the bitstream.

[0508] The inverse quantization unit 1305 performs inverse quantization on the quantization coefficients generated by the quantization unit 1304 using the quantization control parameters, thereby generating inverse quantization coefficients of the prediction residuals, and outputs the generated inverse quantization coefficients to the inverse transformation unit 1306.

[0509] The inverse transformation unit 1306 applies an inverse transformation to the inverse quantization coefficients generated by the inverse quantization unit 1305, thereby generating a prediction residual after the inverse transformation is applied. Since the prediction residual after the inverse transformation is applied is the prediction residual generated after quantization, it may not be exactly the same as the prediction residual output by the transformation unit 1303.

[0510] The addition unit 1307 adds the prediction volume generated by the inverse transformation unit 1306 after the inverse transformation is applied and the prediction volume generated by intra-frame prediction or inter-frame prediction, which is used in the generation of the prediction residuals before quantization, to generate a reconstructed volume. This reconstructed volume is stored in the reference volume memory 1308 or the reference space memory 1310.

[0511] The intra-frame prediction unit 1309 generates a prediction volume of the volume to be encoded using the attribute information of adjacent volumes stored in the reference volume memory 1308. The attribute information includes voxel color information or reflectivity. The intra-frame prediction unit 1309 generates a predicted value of the color information or reflectivity of the volume to be encoded.

[0512] Fig.44 is a diagram for explaining the operation of the intra-frame prediction unit 1309. For example, Fig.44As shown, the intra prediction unit 1309 generates a predicted volume of the encoding target volume (volume idx = 3) based on the adjacent volume (volume idx = 0). Here, the volume idx is identifier information attached to the volumes in the space, and different values are assigned to each volume. The order of assignment of the volume idx may be the same as the encoding order or different from the encoding order. For example, as Fig.44 the predicted value of the color information of the encoding target volume shown, the intra prediction unit 1309 uses the average value of the color information of the voxels included in the volume with volume idx = 0 which is the adjacent volume. In this case, by subtracting the predicted value of the color information from the color information of each voxel included in the encoding target volume, a prediction residual is generated. The processes after the transform unit 1303 are performed on this prediction residual. And, in this case, the three-dimensional data encoding device 1300 attaches the adjacent volume information and the prediction mode information to the bitstream. Here, the adjacent volume information is information showing the adjacent volume used in the prediction, for example, showing the volume idx of the adjacent volume used in the prediction. And, the prediction mode information shows the mode used in the generation of the predicted volume. The mode is, for example, the average value mode that generates a predicted value based on the average value of the voxels in the adjacent volume, or the median value mode that generates a predicted value based on the median value of the voxels in the adjacent volume, etc.

[0513] The intra prediction unit 1309 may also generate a predicted volume based on multiple adjacent volumes. For example, in the Fig.44 configuration shown, the intra prediction unit 1309 generates a predicted volume 0 based on the volume with volume idx = 0, and generates a predicted volume 1 based on the volume with volume idx = 1. Then, the intra prediction unit 1309 generates the average of the predicted volume 0 and the predicted volume 1 as the final predicted volume. In this case, the three-dimensional data encoding device 1300 may also attach the multiple volume idxs of the multiple volumes used in the generation of the predicted volume to the bitstream.

[0514] Fig.45 The inter prediction process according to this embodiment is shown in terms of the mode. The inter prediction unit 1311 encodes (inter prediction) the space (SPC) at a certain time T_Cur using the encoded spaces at different times T_LX. In this case, the inter prediction unit 1311 applies rotation and translation processes to the encoded spaces at different times T_LX for encoding processing.

[0515] Furthermore, the 3D data encoding device 1300 attaches RT information related to the rotation and translation processing of the space applicable to different times T_LX to the bitstream. The different time T_LX is, for example, a time T_L0 before a certain time T_Cur. At this time, the 3D data encoding device 1300 may also attach RT information RT_L0 related to the rotation and translation processing of the space applicable to time T_L0 to the bitstream.

[0516] Alternatively, the different time T_LX is, for example, a time T_L1 after the certain time T_Cur. At this time, the 3D data encoding device 1300 may attach RT information RT_L1 related to the rotation and translation processing of the space applicable to time T_L1 to the bitstream.

[0517] Alternatively, the inter-frame prediction unit 1311 performs encoding (bi-prediction) by referring to the spaces at both different times T_L0 and time T_L1. In this case, the 3D data encoding device 1300 may attach both RT information RT_L0 and RT_L1 related to the rotation and translation applicable to the respective spaces to the bitstream.

[0518] In addition, although T_L0 is set as a time before T_Cur and T_L1 is set as a time after T_Cur above, it is not limited thereto. For example, both T_L0 and T_L1 may be times before T_Cur. Or, both T_L0 and T_L1 may be times after T_Cur.

[0519] And it can also be that, when the 3D data encoding device 1300 performs encoding by referring to the spaces at multiple different times, it attaches RT information related to the rotation and translation applicable to each space to the bitstream. For example, the 3D data encoding device 1300 manages the multiple encoded spaces it refers to through two reference lists (L0 list and L1 list). When the first reference space in the L0 list is set as L0R0, the second reference space in the L0 list is set as L0R1, the first reference space in the L1 list is set as L1R0, and the second reference space in the L1 list is set as L1R1, the 3D data encoding device 1300 attaches the RT information RT_L0R0 of L0R0, the RT information RT_L0R1 of L0R1, the RT information RT_L1R0 of L1R0, and the RT information RT_L1R1 of L1R1 to the bitstream. For example, the 3D data encoding device 1300 attaches these RT information to the head of the bitstream, etc.

[0520] Alternatively, when the three-dimensional data encoding device 1300 performs encoding with reference to reference spaces at multiple different times, it determines whether rotation and translation are applied for each reference space. At this time, the three-dimensional data encoding device 1300 may attach information (such as an RT application flag) indicating whether rotation and translation are applied for each reference space to the header information of the bitstream. For example, the three-dimensional data encoding device 1300 calculates RT information and an ICP error value for each reference space to be referred to according to the encoding target space, using the ICP (Interactive Closest Point) algorithm. When the ICP error value is equal to or less than a predetermined fixed value, the three-dimensional data encoding device 1300 determines that rotation and translation are not required and sets the RT application flag to OFF (invalid). In addition, when the ICP error value is greater than the above-mentioned fixed value, the three-dimensional data encoding device 1300 sets the RT application flag to ON (valid) and attaches the RT information to the bitstream.

[0521] Fig.46 An example of the syntax for attaching RT information and the RT application flag to the header is shown. In addition, the number of bits allocated to each syntax can be determined according to the range that the syntax can take. For example, when the number of reference spaces included in the reference list L0 is 8, 3 bits can be allocated to MaxRefSpc_l0. The number of allocated bits can be changed according to the values that each syntax can take, or the number of allocated bits can be fixed regardless of the values that can be taken. When the number of allocated bits is fixed, the three-dimensional data encoding device 1300 can attach the fixed number of bits to other header information.

[0522] Here, Fig.46 the shown MaxRefSpc_l0 indicates the number of reference spaces included in the reference list L0. RT_flag_l0[i] is the RT application flag for the reference space i in the reference list L0. When RT_flag_l0[i] is 1, rotation and translation are applied to the reference space i. When RT_flag_l0[i] is 0, rotation and translation are not applied to the reference space i.

[0523] R_l0[i] and T_l0[i] are the RT information for the reference space i in the reference list L0. R_l0[i] is the rotation information for the reference space i in the reference list L0. The rotation information shows the content of the applied rotation process, such as a rotation matrix or a quaternion, etc. T_l0[i] is the translation information for the reference space i in the reference list L0. The translation information shows the content of the applied translation process, such as a translation vector, etc.

[0524] MaxRefSpc_l1 shows the number of reference spaces included in reference list L1. RT_flag_l1[i] is the RT application flag for reference space i in reference list L1. When RT_flag_l1[i] is 1, rotation and translation are applied to reference space i. When RT_flag_l1[i] is 0, rotation and translation are not applied to reference space i.

[0525] R_l1[i] and T_l1[i] are the RT information for reference space i in reference list L1. R_l1[i] is the rotation information for reference space i in reference list L1. The rotation information shows the content of the applied rotation process, such as a rotation matrix or quaternion, etc. T_l1[i] is the translation information for reference space i in reference list L1. The translation information shows the content of the applied translation process, such as a translation vector, etc.

[0526] The inter-frame prediction unit 1311 generates a predicted volume of the volume to be encoded by using the information of the encoded reference spaces stored in the reference space memory 1310. As described above, before generating the predicted volume of the volume to be encoded, the inter-frame prediction unit 1311 uses the ICP (Interactive Closest Point) algorithm in the volume to be encoded space and the reference space to find the RT information in order to make the positional relationship between the volume to be encoded space and the entire reference space closer. Then, the inter-frame prediction unit 1311 applies rotation and translation processing to the reference space by using the obtained RT information, thereby obtaining reference space B. After that, the inter-frame prediction unit 1311 generates a predicted volume of the volume to be encoded in the volume to be encoded space by using the information in reference space B. Here, the three-dimensional data encoding device 1300 attaches the RT information used to obtain reference space B to the header information, etc., of the volume to be encoded space.

[0527] In this way, the inter-frame prediction unit 1311 applies rotation and translation processing to the reference space, so that after making the positional relationship between the volume to be encoded space and the entire reference space closer, it uses the information of the reference space to generate the predicted volume. In this way, the accuracy of the predicted volume can be improved. And since the prediction residual can be suppressed, the encoding amount can be reduced. In addition, although an example of using the volume to be encoded space and the reference space to perform ICP is shown here, it is not limited thereto. For example, in order to reduce the processing amount, the inter-frame prediction unit 1311 can also use at least one of the volume to be encoded space with voxel or point cloud number extracted and the reference space with voxel or point cloud number extracted to perform ICP, thereby obtaining the RT information.

[0528] Furthermore, when the ICP error value obtained from the ICP result is smaller than a predetermined first threshold value, that is, when the positional relationship between the encoding object space and the reference space is close, the inter-frame prediction unit 1311 may determine that rotation and translation processing are not required, and does not perform rotation and translation. In this case, the three-dimensional data encoding device 1300 may not add RT information to the bitstream, thereby suppressing additional overhead.

[0529] Furthermore, when the ICP error value is greater than a predetermined second threshold, the inter-frame prediction unit 1311 determines that the shape change in space is large, and intra-frame prediction can be applied to all volumes of the encoding object space. Hereinafter, the space to which intra-frame prediction is applied is referred to as intra-frame space. Furthermore, the second threshold is a value greater than the above-mentioned first threshold. Furthermore, it is not limited to ICP, and any method can be applied as long as the method of obtaining RT information from two voxel sets or two point cloud sets.

[0530] Furthermore, when the three-dimensional data contains attribute information such as shape or color, the inter-frame prediction unit 1311 searches, for example, a volume in the reference space that is closest to the shape or color attribute information of the encoding target volume as a prediction volume of the encoding target volume in the encoding target space. Furthermore, the reference space is, for example, a reference space after the above-mentioned rotation and translation processing. The inter-frame prediction unit 1311 generates a prediction volume based on the volume (reference volume) obtained by the search. Fig.47 is a diagram for explaining the generation of the prediction volume. Fig.47 When the encoding target volume (volume idx=0) shown in the figure is encoded by using inter-frame prediction, the reference volumes in the reference space are scanned in sequence while searching for a volume in which the difference between the encoding target volume and the reference volume, i.e., the prediction residual, is the smallest. The inter-frame prediction unit 1311 selects the volume with the smallest prediction residual as the prediction volume. The prediction residual between the encoding target volume and the prediction volume is encoded by the processing after the transformation unit 1303. Here, the prediction residual refers to the difference between the attribute information of the encoding target volume and the attribute information of the prediction volume. In addition, the three-dimensional data encoding device 1300 adds the volume idx of the reference volume in the reference space referred to as the prediction volume to the header of the bit stream, etc.

[0531] exist Fig.47 In the example shown, the reference volume idx=4 of the reference space L0R0 is selected as the prediction volume of the encoding target volume. Then, the prediction residual between the encoding target volume and the reference volume and the reference volume idx=4 are encoded and added to the bit stream.

[0532] In addition, although the prediction volume of the attribute information has been described as an example here, the same process can be performed for the prediction volume of the position information.

[0533] The prediction control unit 1312 controls which of intra prediction and inter prediction is used to encode the encoding target volume. Here, the mode including intra prediction and inter prediction is called the prediction mode. For example, the prediction control unit 1312 calculates the prediction residual when the encoding target volume is predicted by intra prediction and the prediction residual when it is predicted by inter prediction as evaluation values, and selects the prediction mode with the smaller evaluation value. Alternatively, the prediction control unit 1312 may perform orthogonal transformation, quantization, and entropy coding on the prediction residual of intra prediction and the prediction residual of inter prediction respectively to calculate the actual coding amount, and use the calculated coding amount as the evaluation value to select the prediction mode. Also, overhead information other than the prediction residual (such as the reference volume idx information) may be added to the evaluation value. Further, when the encoding target space is predetermined to be encoded in the intra space, the prediction control unit 1312 may generally select intra prediction.

[0534] The entropy coding unit 1313 generates an encoded signal (encoded bitstream) by performing variable-length coding on the quantization coefficients that are the input from the quantization unit 1304. Specifically, the entropy coding unit 1313 binarizes the quantization coefficients, for example, and performs arithmetic coding on the obtained binary signal.

[0535] Next, a three-dimensional data decoding device that decodes the encoded signal generated by the three-dimensional data encoding device 1300 will be described. Fig.48 It is a block diagram of the three-dimensional data decoding device 1400 according to the present embodiment. The three-dimensional data decoding device 1400 includes: an entropy decoding unit 1401, an inverse quantization unit 1402, an inverse transformation unit 1403, an addition unit 1404, a reference volume memory 1405, an intra prediction unit 1406, a reference space memory 1407, an inter prediction unit 1408, and a prediction control unit 1409.

[0536] The entropy decoding unit 1401 performs variable-length decoding on the encoded signal (encoded bitstream). For example, the entropy decoding unit 1401 performs arithmetic decoding on the encoded signal to generate a binary signal, and generates quantization coefficients based on the generated binary signal.

[0537] The inverse quantization unit 1402 performs inverse quantization on the quantization coefficients input from the entropy decoding unit 1401 using the quantization parameters attached to the bitstream or the like, thereby generating inverse quantization coefficients.

[0538] The inverse transform unit 1403 performs an inverse transform on the inverse quantized coefficients input from the inverse quantization unit 1402 to generate a prediction residual. For example, the inverse transform unit 1403 performs an inverse orthogonal transform on the inverse quantized coefficients according to the information appended to the bitstream to generate a prediction residual.

[0539] The addition unit 1404 adds the prediction residual generated by the inverse transform unit 1403 and the prediction volume generated by intra prediction or inter prediction to generate a reconstructed volume. This reconstructed volume is output as decoded three-dimensional data and stored in the reference volume memory 1405 or the reference space memory 1407.

[0540] The intra prediction unit 1406 generates a prediction volume by intra prediction using the reference volume in the reference volume memory 1405 and the information appended to the bitstream. Specifically, the intra prediction unit 1406 obtains prediction mode information and adjacent volume information (such as volume idx) appended to the bitstream, and uses the adjacent volume indicated by the adjacent volume information to generate a prediction volume in the mode indicated by the prediction mode information. In addition, the details of these processes are the same as those of the above intra prediction unit 1309 except for using the information appended to the bitstream.

[0541] The inter prediction unit 1408 generates a prediction volume by inter prediction using the reference space in the reference space memory 1407 and the information appended to the bitstream. Specifically, the inter prediction unit 1408 uses the RT information of each reference space appended to the bitstream, applies rotation and translation processing to the reference space, and uses the processed reference space to generate a prediction volume. In addition, when the RT application flag for each reference space exists in the bitstream, the inter prediction unit 1408 applies rotation and translation processing to the reference space according to the RT application flag. In addition, the details of the above processes are the same as those of the above inter prediction unit 1311 except for using the information appended to the bitstream.

[0542] Whether to decode the volume to be decoded by intra prediction or inter prediction will be controlled by the prediction control unit 1409. For example, the prediction control unit 1409 selects intra prediction or inter prediction according to the information appended to the bitstream and indicating the prediction mode to be used. In addition, the prediction control unit 1409 may usually select intra prediction when it is predetermined that the object space to be decoded is decoded in the intra space.

[0543] The following describes a modification example of the present embodiment. In the present embodiment, although rotation and translation are applied in units of space as an example, rotation and translation can also be applied in smaller units. For example, the three-dimensional data encoding device 1300 can divide the space into sub-spaces and apply rotation and translation in units of sub-spaces. In this case, the three-dimensional data encoding device 1300 generates RT information for each sub-space and attaches the generated RT information to the head of the bitstream or the like. Further, the three-dimensional data encoding device 1300 can apply rotation and translation in units of volume as the encoding unit. In this case, the three-dimensional data encoding device 1300 generates RT information in units of encoding volume and attaches the generated RT information to the head of the bitstream or the like. Moreover, the above can be combined. That is, the three-dimensional data encoding device 1300 can apply rotation and translation in a large unit first, and then apply rotation and translation in a smaller unit. For example, the three-dimensional data encoding device 1300 can apply rotation and translation in units of space, and apply different rotations and translations to each of a plurality of volumes included in the obtained space.

[0544] Moreover, in the present embodiment, although rotation and translation are applied to the reference space as an example, it is not limited thereto. For example, the three-dimensional data encoding device 1300 can apply a scaling process to change the size of the three-dimensional data. Further, the three-dimensional data encoding device 1300 can also apply any one or two of rotation, translation, and scaling. Moreover, as described above, when processing is applied in different units in multiple stages, the types of processing applied in each unit can be different. For example, rotation and translation can be applied in units of space, and translation can be applied in units of volume.

[0545] In addition, regarding these modification examples, the same can be applied to the three-dimensional data decoding device 1400.

[0546] As described above, the three-dimensional data encoding device 1300 according to the present embodiment performs the following processing. Fig.48 It is a flowchart of the inter-frame prediction process performed by the three-dimensional data encoding device 1300.

[0547] First, the three-dimensional data encoding device 1300 generates prediction position information (e.g., predicted volume) using the position information of three-dimensional points included in the object three-dimensional data (e.g., the encoding object space) and the reference three-dimensional data (e.g., the reference space) at different times (S1301). Specifically, the three-dimensional data encoding device 1300 generates prediction position information by applying rotation and translation processes to the position information of the three-dimensional points included in the reference three-dimensional data.

[0548] In addition, the three-dimensional data encoding device 1300 performs rotation and translation processing in a first unit (e.g., a space), and generates predicted position information in a second unit (e.g., a volume) that is finer than the first unit. For example, the three-dimensional data encoding device 1300 may search for a volume in the plurality of volumes included in the reference space after the rotation and translation processing, where the difference between the encoded object volume included in the encoded object space and the position information is minimized, and use the obtained volume as the predicted volume. In addition, the three-dimensional data encoding device 1300 may perform the rotation and translation processing and the generation of the predicted position information in the same unit.

[0549] Moreover, it can be that the three-dimensional data encoding device 1300 applies the first rotation and translation processing to the position information of the three-dimensional points included in the reference three-dimensional data in a first unit (e.g., a space), and applies the second rotation and translation processing to the position information of the three-dimensional points obtained through the first rotation and translation processing in a second unit (e.g., a volume) that is finer than the first unit, thereby generating predicted position information.

[0550] Here, the position information of the three-dimensional points and the predicted position information are Fig.41 represented in an octree structure as shown. For example, the position information of the three-dimensional points and the predicted position information are represented in a width-first scan order of the depth and width in the octree structure. Alternatively, the position information of the three-dimensional points and the predicted position information are represented in a depth-first scan order of the depth and width in the octree structure.

[0551] Furthermore, as Fig.46 shown, the three-dimensional data encoding device 1300 encodes an RT application flag indicating whether rotation and translation processing is applied to the position information of the three-dimensional points included in the reference three-dimensional data. That is, the three-dimensional data encoding device 1300 generates an encoded signal (encoded bitstream) including the RT application flag. And the three-dimensional data encoding device 1300 encodes RT information indicating the content of the rotation and translation processing. That is, the three-dimensional data encoding device 1300 generates an encoded signal (encoded bitstream) including the RT information. Additionally, it can be that the three-dimensional data encoding device 1300 encodes the RT information when the RT application flag indicates that rotation and translation processing is applied, and does not encode the RT information when the RT application flag indicates that rotation and translation processing is not applied.

[0552] Moreover, the three-dimensional data includes, for example, the position information of the three-dimensional points and the attribute information (such as color information) of each three-dimensional point. The three-dimensional data encoding device 1300 generates predicted attribute information (S1302) using the attribute information of the three-dimensional points included in the reference three-dimensional data.

[0553] Next, the three-dimensional data encoding device 1300 uses the predicted position information to encode the position information of the three-dimensional points included in the target three-dimensional data. For example, as shown in Fig.38 , the three-dimensional data encoding device 1300 calculates the difference, i.e., the differential position information, between the position information of the three-dimensional points included in the target three-dimensional data and the predicted position information (S1303).

[0554] In addition, the three-dimensional data encoding device 1300 uses the predicted attribute information to encode the attribute information of the three-dimensional points included in the target three-dimensional data. For example, the three-dimensional data encoding device 1300 calculates the difference, i.e., the differential attribute information, between the attribute information of the three-dimensional points included in the target three-dimensional data and the predicted attribute information (S1304). Next, the three-dimensional data encoding device 1300 performs transformation and quantization on the calculated differential attribute information (S1305).

[0555] Finally, the three-dimensional data encoding device 1300 encodes the differential position information and the quantized differential attribute information (e.g., entropy encoding) (S1306). That is, the three-dimensional data encoding device 1300 generates an encoded signal (encoded bitstream) including the differential position information and the differential attribute information.

[0556] In addition, when the attribute information is not included in the three-dimensional data, the three-dimensional data encoding device 1300 may not perform steps S1302, S1304, and S1305. Moreover, the three-dimensional data encoding device 1300 may only perform one of the encoding of the position information of the three-dimensional points and the encoding of the attribute information of the three-dimensional points.

[0557] Moreover, Fig.49 the order of the processes shown is only an example and is not limited thereto. For example, since the processing for the position information (S1301, S1303) and the processing for the attribute information (S1302, S1304, S1305) are independent of each other, they can be executed in any order, or a part of them can be processed in parallel.

[0558] As described above, in the present embodiment, the three-dimensional data encoding device 1300 uses the position information of the three-dimensional points included in the target three-dimensional data and the reference three-dimensional data at different times to generate the predicted position information, and encodes the difference, i.e., the differential position information, between the position information of the three-dimensional points included in the target three-dimensional data and the predicted position information. Accordingly, since the data amount of the encoded signal can be reduced, the encoding efficiency can be improved.

[0559] Further, in the present embodiment, the three-dimensional data encoding device 1300 generates predicted attribute information by using the attribute information of the three-dimensional points included in the reference three-dimensional data, and encodes the difference, i.e., the differential attribute information, between the attribute information of the three-dimensional points included in the object three-dimensional data and the predicted attribute information. Accordingly, since the data amount of the encoded signal can be reduced, the encoding efficiency can be improved.

[0560] For example, the three-dimensional data encoding device 1300 includes a processor and a memory, and the processor uses the memory to perform the above-described processing.

[0561] Fig.48 It is a flowchart of the inter-frame prediction process performed by the three-dimensional data decoding device 1400.

[0562] First, the three-dimensional data decoding device 1400 decodes (e.g., entropy decodes) the differential position information and the differential attribute information according to the encoded signal (encoded bitstream) (S1401).

[0563] Further, the three-dimensional data decoding device 1400 decodes the RT applicability flag indicating whether rotation and translation processing are applicable to the position information of the three-dimensional points included in the reference three-dimensional data according to the encoded signal. And the three-dimensional data decoding device 1400 decodes the RT information indicating the content of the rotation and translation processing. In addition, the three-dimensional data decoding device 1400 decodes the RT information when the RT applicability flag indicates that rotation and translation processing are applicable, and does not decode the RT information when the RT applicability flag indicates that rotation and translation processing are not applicable.

[0564] Next, the three-dimensional data decoding device 1400 performs inverse quantization and inverse transformation on the decoded differential attribute information (S1402).

[0565] Next, the three-dimensional data decoding device 1400 generates predicted position information (e.g., predicted volume) by using the position information of the three-dimensional points included in the object three-dimensional data (e.g., decoding target space) and the reference three-dimensional data (e.g., reference space) at different times (S1403). Specifically, the three-dimensional data decoding device 1400 generates the predicted position information by applying rotation and translation processing to the position information of the three-dimensional points included in the reference three-dimensional data.

[0566] More specifically, when the RT applicability flag indicates that rotation and translation processing are applicable, the three-dimensional data decoding device 1400 applies rotation and translation processing to the position information of the three-dimensional points included in the reference three-dimensional data indicated by the RT information. And when the RT applicability flag indicates that rotation and translation processing are not applicable, the three-dimensional data decoding device 1400 does not apply rotation and translation processing to the position information of the three-dimensional points included in the reference three-dimensional data.

[0567] In addition, the three-dimensional data decoding device 1400 can perform rotation and translation processing in a first unit (e.g., a space), and can generate predicted position information in a second unit (e.g., a volume) that is finer than the first unit. In addition, the three-dimensional data decoding device 1400 can also perform rotation and translation processing and generation of predicted position information in the same unit.

[0568] It can be that the three-dimensional data decoding device 1400 applies first rotation and translation processing to the position information of the three-dimensional points included in the reference three-dimensional data in a first unit (e.g., a space), and applies second rotation and translation processing to the position information of the three-dimensional points obtained by the first rotation and translation processing in a second unit (e.g., a volume) that is finer than the first unit, thereby generating predicted position information.

[0569] Here, the position information of the three-dimensional points and the predicted position information are, for example Fig.41 As shown, they are represented in an octree structure. For example, the position information of the three-dimensional points and the predicted position information are represented in a scan order that gives priority to the width among the depth and width in the octree structure. Alternatively, the position information of the three-dimensional points and the predicted position information are represented in a scan order that gives priority to the depth among the depth and width in the octree structure.

[0570] The three-dimensional data decoding device 1400 generates predicted attribute information by using the attribute information of the three-dimensional points included in the reference three-dimensional data (S1404).

[0571] Next, the three-dimensional data decoding device 1400 decodes the encoded position information included in the encoded signal by using the predicted position information, thereby restoring the position information of the three-dimensional points included in the target three-dimensional data. Here, the encoded position information is, for example, differential position information, and the three-dimensional data decoding device 1400 restores the position information of the three-dimensional points included in the target three-dimensional data by adding the differential position information and the predicted position information (S1405).

[0572] In addition, the three-dimensional data decoding device 1400 decodes the encoded attribute information included in the encoded signal by using the predicted attribute information, thereby restoring the attribute information of the three-dimensional points included in the target three-dimensional data. Here, the encoded attribute information is, for example, differential attribute information, and the three-dimensional data decoding device 1400 restores the attribute information of the three-dimensional points included in the target three-dimensional data by adding the differential attribute information and the predicted attribute information (S1406).

[0573] Alternatively, when the three-dimensional data does not include attribute information, the three-dimensional data decoding apparatus 1400 may also not execute steps S1402, S1404, and S1406. Further, the three-dimensional data decoding apparatus 1400 may also perform only one of the decoding of the position information of the three-dimensional points and the decoding of the attribute information of the three-dimensional points.

[0574] Also, Fig.50 The order of the processes shown is an example and is not limited thereto. For example, since the processes for the position information (S1403, S1405) and the processes for the attribute information (S1402, S1404, S1406) are independent of each other, they can be performed in any order, and a part of them can also be processed in parallel.

[0575] (Embodiment 8)

[0576] The information of the three-dimensional point group includes position information (geometry) and attribute information (attribute). The position information includes coordinates (x coordinate, y coordinate, z coordinate) with respect to a certain point. When encoding the position information, instead of directly encoding the coordinates of each three-dimensional point, a method is used in which the positions of each three-dimensional point are represented by an octree representation and the information of the octree is encoded to reduce the amount of encoding.

[0577] On the other hand, the attribute information includes information such as color information (RGB, YUV, etc.), reflectance, and normal vector indicating each three-dimensional point. For example, the three-dimensional data encoding apparatus can encode the attribute information using an encoding method different from the position information.

[0578] In the present embodiment, an encoding method for the attribute information will be described. Further, in the present embodiment, integer values are used as the values of the attribute information for description. For example, when each color component of the color information RGB or YUV has an 8-bit precision, each color component takes an integer value from 0 to 255. When the value of the reflectance has a 10-bit precision, the value of the reflectance takes an integer value from 0 to 1023. Further, when the bit precision of the attribute information is a fractional precision, the three-dimensional data encoding apparatus may also multiply the value by a scaling value and then round it to an integer value so that the value of the attribute information becomes an integer value. Additionally, the three-dimensional data encoding apparatus may also attach the scaling value to the head of the bit stream or the like.

[0579] As a method for encoding attribute information of three-dimensional points, consider calculating a predicted value of the attribute information of the three-dimensional points, and encoding the difference (prediction residual) between the value of the original attribute information and the predicted value. For example, when the value of the attribute information of the three-dimensional point p is Ap and the predicted value is Pp, the three-dimensional data encoding device encodes the absolute value of the difference Diffp = |Ap - Pp|. In this case, if the predicted value Pp can be generated with high accuracy, the value of the absolute difference Diffp becomes smaller. Therefore, for example, by using an encoding table in which the smaller the value, the smaller the number of bits generated, to perform entropy encoding on the absolute difference Diffp, the amount of encoding can be reduced.

[0580] As a method for generating a predicted value of attribute information, consider using the attribute information of other three-dimensional points, that is, reference three-dimensional points, located around the object three-dimensional point to be encoded. Here, the reference three-dimensional points refer to three-dimensional points within a pre-specified distance range from the object three-dimensional point. For example, when there are an object three-dimensional point p = (x1, y1, z1) and a three-dimensional point q = (x2, y2, z2), the three-dimensional data encoding device calculates the Euclidean distance d(p, q) between the three-dimensional point p and the three-dimensional point q as shown in (Equation A1).

[0581]

Equation 1

[0582]

[0583] When the Euclidean distance d(p, q) is less than a predetermined threshold THd, the three-dimensional data encoding device determines that the position of the three-dimensional point q is close to the position of the object three-dimensional point p, and determines that the value of the attribute information of the three-dimensional point q is used in the generation of the predicted value of the attribute information of the object three-dimensional point p. In addition, the distance calculation method can also be other methods, for example, the Mahalanobis distance can also be used. In addition, the three-dimensional data encoding device can also determine not to use three-dimensional points outside the pre-specified distance range from the object three-dimensional point for prediction processing. For example, when there is a three-dimensional point r and the distance d(p, r) between the object three-dimensional point p and the three-dimensional point r is greater than or equal to the threshold THd, the three-dimensional data encoding device can also determine not to use the three-dimensional point r for prediction. In addition, the three-dimensional data encoding device can also attach information indicating the threshold THd to the head of the bitstream, etc.

[0584] Fig.51 It is a diagram showing an example of three-dimensional points. In this example, the distance d(p, q) between the object three-dimensional point p and the three-dimensional point q is less than the threshold THd. Therefore, the three-dimensional data encoding device determines the three-dimensional point q as a reference three-dimensional point of the object three-dimensional point p, and determines that the value of the attribute information Aq of the three-dimensional point q is used in the generation of the predicted value Pp of the attribute information Ap of the object three-dimensional point p.

[0585] On the other hand, the distance d(p, r) between the object three-dimensional point p and the three-dimensional point r is equal to or greater than the threshold THd. Therefore, the three-dimensional data encoding device determines that the three-dimensional point r is not the reference three-dimensional point of the object three-dimensional point p, and determines that the value of the attribute information Ar of the three-dimensional point r is not used in generating the predicted value Pp of the attribute information Ap of the object three-dimensional point p.

[0586] In addition, in the case of encoding the attribute information of the object three-dimensional point using the predicted value, the three-dimensional data encoding device uses the three-dimensional point for which the encoding and decoding of the attribute information have been completed as the reference three-dimensional point. Similarly, in the case of decoding the attribute information of the object three-dimensional point of the decoding target using the predicted value, the three-dimensional data decoding device uses the three-dimensional point for which the attribute information has been decoded as the reference three-dimensional point. Thereby, the same predicted value can be generated during encoding and decoding, and thus the bitstream of the three-dimensional point generated by encoding can be correctly decoded on the decoding side.

[0587] In addition, in the case of encoding the attribute information of the three-dimensional point, it is considered to classify each three-dimensional point into a plurality of levels using the position information of the three-dimensional point and then perform encoding. Here, each classified level is referred to as LoD (Level of Detail). Fig.52 A method for generating LoD will be described.

[0588] First, the three-dimensional data encoding device selects the initial point a0 and assigns it to LoD0. Next, the three-dimensional data encoding device extracts the point a1 whose distance from the point a0 is greater than the threshold Thres_LoD[0] of LoD0 and assigns it to LoD0. Next, the three-dimensional data encoding device extracts the point a2 whose distance from the point a1 is greater than the threshold Thres_LoD[0] of LoD0 and assigns it to LoD0. In this way, the three-dimensional data encoding device constructs LoD0 such that the distance between each point within LoD0 is greater than the threshold Thres_LoD[0].

[0589] Next, the three-dimensional data encoding device selects the point b0 that has not been assigned to LoD and assigns it to LoD1. Next, the three-dimensional data encoding device extracts the point b1 whose distance from the point b0 is greater than the threshold Thres_LoD[1] of LoD1 and that has not been assigned to LoD and assigns it to LoD1. Next, the three-dimensional data encoding device extracts the point b2 whose distance from the point b1 is greater than the threshold Thres_LoD[1] of LoD1 and that has not been assigned to LoD and assigns it to LoD1. In this way, the three-dimensional data encoding device constructs LoD1 such that the distance between each point within LoD1 is greater than the threshold Thres_LoD[1].

[0590] Next, the three-dimensional data encoding device selects a point c0 for which LoD has not been assigned and assigns it to LoD2. Next, the three-dimensional data encoding device extracts a point c1 that is at a distance greater than the threshold Thres_LoD[2] of LoD2 from the point c0 and for which LoD has not been assigned, and assigns it to LoD2. Next, the three-dimensional data encoding device extracts a point c2 that is at a distance greater than the threshold Thres_LoD[2] of LoD2 from the point c1 and for which LoD has not been assigned, and assigns it to LoD2. In this way, the three-dimensional data encoding device constructs LoD2 such that the distance between each point within LoD2 is greater than the threshold Thres_LoD[2]. For example, as Fig.53 shown, the thresholds Thres_LoD[0], Thres_LoD[1], and Thres_LoD[2] for each LoD are set.

[0591] In addition, the three-dimensional data encoding device may also attach information indicating the thresholds of each LoD to the header of the bitstream or the like. For example, in the Fig.53 case of the example shown, the three-dimensional data encoding device may also attach the thresholds Thres_LoD[0], Thres_LoD[1], and Thres_LoD[2] to the header.

[0592] In addition, the three-dimensional data encoding device may also assign all three-dimensional points for which LoD has not been assigned to the lowest layer of LoD. In this case, the three-dimensional data encoding device can reduce the encoding amount of the header by not attaching the threshold of the lowest layer of LoD to the header. For example, in the Fig.53 case of the example shown, the three-dimensional data encoding device attaches the thresholds Thres_LoD[0] and Thres_LoD[1] to the header and does not attach Thres_LoD[2] to the header. In this case, the three-dimensional data decoding device may also estimate the value of Thres_LoD[2] to be 0. In addition, the three-dimensional data encoding device may also attach the number of levels of LoD to the header. Thereby, the three-dimensional data decoding device can use the number of levels of LoD to determine the lowest layer of LoD.

[0593] In addition, as Fig.53 shown, the values of the thresholds for each layer of LoD are set to be larger for the upper layers, so that the upper layers (the layers closer to LoD0) are sparser groups of points with a greater distance between three-dimensional points, and the lower layers are denser groups of points with a closer distance between three-dimensional points. In addition, in the Fig.53 example shown, LoD0 is the uppermost layer.

[0594] In addition, the method for selecting the initial three-dimensional points when setting each LoD can also depend on the encoding order during the encoding of the position information. For example, the three-dimensional data encoding device selects the three-dimensional point that was first encoded during the encoding of the position information as the initial point a0 of LoD0, and selects points a1 and a2 with the initial point a0 as the base point to form LoD0. Moreover, the three-dimensional data encoding device can also select the three-dimensional point with the earliest encoded position information among the three-dimensional points that do not belong to LoD0 as the initial point b0 of LoD1. That is, the three-dimensional data encoding device can also select the three-dimensional point with the earliest encoded position information among the three-dimensional points in the upper layer (LoD0 to LoDn-1) that do not belong to LoDn as the initial point n0 of LoDn. Thus, the three-dimensional data decoding device can form the same LoD as during encoding by using the same initial point selection method during decoding, and thus can appropriately decode the bitstream. Specifically, the three-dimensional data decoding device selects the three-dimensional point with the earliest decoded position information among the three-dimensional points in the upper layer that do not belong to LoDn as the initial point n0 of LoDn.

[0595] Hereinafter, a method for generating a predicted value of the attribute information of the three-dimensional points using the information of LoD will be described. For example, in the case of encoding the three-dimensional points included in LoD0 in sequence, the three-dimensional data encoding device uses the encoded and decoded (hereinafter, also simply referred to as "encoded") attribute information included in LoD0 and LoD1 to generate the target three-dimensional points included in LoD1. In this way, the three-dimensional data encoding device uses the encoded attribute information included in LoDn' (n' <= n) to generate a predicted value of the attribute information of the three-dimensional points included in LoDn. That is, the three-dimensional data encoding device does not use the attribute information of the three-dimensional points included in the lower layer of LoDn in the calculation of the predicted value of the attribute information of the three-dimensional points included in LoDn.

[0596] For example, the three-dimensional data encoding device generates a predicted value of the attribute information of the three-dimensional points by calculating the average value of the attribute values of N or fewer three-dimensional points among the encoded three-dimensional points around the target three-dimensional points to be encoded. In addition, the three-dimensional data encoding device can attach the value of N to the head of the bitstream or the like. In addition, the three-dimensional data encoding device can also change the value of N for each three-dimensional point and attach the value of N to each three-dimensional point. Thus, an appropriate N can be selected for each three-dimensional point, and therefore the accuracy of the predicted value can be improved. Therefore, the prediction residual can be reduced. In addition, the three-dimensional data encoding device can also attach the value of N to the head of the bitstream and fix the value of N within the bitstream. Thus, it is not necessary to encode or decode the value of N for each three-dimensional point, and therefore the processing amount can be reduced. In addition, the three-dimensional data encoding device can also encode the value of N separately for each LoD. Thus, by selecting an appropriate N for each LoD, the encoding efficiency can be improved.

[0597] Alternatively, the 3D data encoding device may also calculate the predicted value of the attribute information of the 3D point by the weighted average of the attribute information of the N encoded 3D points around it. For example, the 3D data encoding device calculates the weight using the distance information between the target 3D point and each of the N surrounding 3D points.

[0598] When the 3D data encoding device encodes the value of N for each LoD respectively, for example, the larger the LoD upper layer, the larger the value of N is set, and the smaller the LoD lower layer, the smaller the value of N is set. In the upper layer of LoD, the distance between the 3D points belonging to this layer is far. Therefore, it is possible to improve the prediction accuracy by setting a large value of N and selecting multiple surrounding 3D points for averaging. In addition, since in the lower layer of LoD, the distance between the 3D points belonging to this layer is close, it is possible to perform efficient prediction while suppressing the processing amount of averaging by setting a small value of N.

[0599] Fig.54 It is a diagram showing an example of the attribute information used in the predicted value. As described above, the predicted value of the point P included in LoDn is generated using the encoded surrounding points P' included in LoD N' (N' <= N). Here, the surrounding points P' are selected based on the distance from the point P. For example, the attribute information of points a0, a1, a2, b0, b1 is used to generate Fig.54 the predicted value of the attribute information of the point b2 shown.

[0600] According to the above value of N, the selected surrounding points change. For example, when N = 5, as the surrounding points of the point b2, a0, a1, a2, b0, b1 are selected. When N = 4, based on the distance information, points a0, a1, a2, b1 are selected.

[0601] The predicted value is calculated by weighted average depending on the distance. For example, in Fig.54 the example shown, the predicted value a2p of the point a2 is calculated by the weighted average of the attribute information of the points a0 and a1 as shown in (Equation A2) and (Equation A3). In addition, A i is the value of the attribute information of the point ai.

[0602] [Equation 2]

[0603]

[0604]

[0605] In addition, the predicted value b2p of the point b2 is calculated by the weighted average of the attribute information of the points a0, a1, a2, b0, b1 as shown in (Equation A4) to (Equation A6). In addition, B i is the value of the attribute information of the point bi.

[0606]

Mathematical formula 3

[0607]

[0608]

[0609]

[0610] In addition, the three-dimensional data encoding device can also calculate the difference value (prediction residual) between the value of the attribute information of the three-dimensional point and the predicted value generated from the surrounding points, and quantize the calculated prediction residual. For example, the three-dimensional data encoding device quantizes by dividing the prediction residual by the quantization scale (also referred to as the quantization step). In this case, the smaller the quantization scale, the smaller the error (quantization error) generated due to quantization. On the contrary, the larger the quantization scale, the larger the quantization error.

[0611] Furthermore, the three-dimensional data encoding device can also change the quantization scale used for each LoD. For example, the higher the layer of the three-dimensional data encoding device, the smaller the quantization scale, and the lower the layer, the larger the quantization scale. The value of the attribute information of the three-dimensional points belonging to the upper layer may be used as the predicted value of the attribute information of the three-dimensional points belonging to the lower layer. Therefore, the quantization scale of the upper layer can be reduced to suppress the quantization error generated in the upper layer, and by improving the accuracy of the predicted value, the encoding efficiency can be improved. In addition, the three-dimensional data encoding device can also attach the quantization scale used for each LoD to the header or the like. As a result, the three-dimensional data decoding device can correctly decode the quantization scale, and thus can appropriately decode the bitstream.

[0612] In addition, the three-dimensional data encoding device can also transform the signed integer value (signed quantization value) that is the quantized prediction residual into an unsigned integer value (unsigned quantization value). As a result, when entropy encoding the prediction residual, there is no need to consider the generation of negative integers. In addition, the three-dimensional data encoding device does not necessarily need to transform the signed integer value into an unsigned integer value. For example, it can also perform entropy encoding on the sign bit separately.

[0613] The prediction residual is calculated by subtracting the predicted value from the original value. For example, as shown in (Equation A7), the prediction residual a2r of point a2 is calculated by subtracting the predicted value a2p of point a2 from the value A2 of the attribute information of point a2. As shown in (Equation A8), the prediction residual b2r of point b2 is calculated by subtracting the predicted value b2p of point b2 from the value B2 of the attribute information of point b2.

[0614] a2r = A2 - a2p...(Equation A7)

[0615] b2r = B2 - b2p...(Equation A8)

[0616] In addition, the prediction residual is quantized by dividing it by QS (Quantization Step). For example, the quantized value a2q of point a2 is calculated by (Equation A9). The quantized value b2q of point b2 is calculated by (Equation A10). Here, QS_LoD0 is the QS for LoD0, and QS_LoD1 is the QS for LoD1. That is, the QS can be changed according to LoD.

[0617] a2q = a2r / QS_LoD0 … (Equation A9)

[0618] b2q = b2r / QS_LoD1 … (Equation A10)

[0619] In addition, as described below, the three-dimensional data encoding device converts the signed integer value, which is the above-mentioned quantized value, into an unsigned integer value. When the signed integer value a2q is less than 0, the three-dimensional data encoding device sets the unsigned integer value a2u to -1 - (2 × a2q). When the signed integer value a2q is 0 or more, the three-dimensional data encoding device sets the unsigned integer value a2u to 2 × a2q.

[0620] Similarly, when the signed integer value b2q is less than 0, the three-dimensional data encoding device sets the unsigned integer value b2u to -1 - (2 × b2q). When the signed integer value b2q is 0 or more, the three-dimensional data encoding device sets the unsigned integer value b2u to 2 × b2q.

[0621] In addition, the three-dimensional data encoding device can also encode the quantized prediction residual (unsigned integer value) by entropy encoding. For example, binary arithmetic coding can also be applied after binarizing the unsigned integer value.

[0622] In addition, in this case, the three-dimensional data encoding device can also switch the binarization method according to the value of the prediction residual. For example, when the prediction residual pu is less than the threshold R_TH, the three-dimensional data encoding device binarizes the prediction residual pu with a fixed number of bits required to represent the threshold R_TH. In addition, when the prediction residual pu is equal to or greater than the threshold R_TH, the three-dimensional data encoding device binarizes the binarized data of the threshold R_TH and the value of (pu - R_TH) using, for example, Exponential-Golomb.

[0623] For example, when the threshold R_TH is 63 and the prediction residual pu is less than 63, the three-dimensional data encoding device binarizes the prediction residual pu with 6 bits. Additionally, when the prediction residual pu is 63 or more, the three-dimensional data encoding device uses Golomb code to binarize the binary data (111111) of the threshold R_TH and (pu - 63), and then performs arithmetic coding accordingly.

[0624] In a more specific example, when the prediction residual pu is 32, the three-dimensional data encoding device generates 6-bit binary data (100000) and performs arithmetic coding on this bit string. Additionally, when the prediction residual pu is 66, the three-dimensional data encoding device generates a bit string (00100) representing the binary data (111111) of the threshold R_TH using Golomb code and the value 3 (66 - 63), and performs arithmetic coding on this bit string (111111 + 00100).

[0625] In this way, the three-dimensional data encoding device switches the binarization method according to the magnitude of the prediction residual, thereby enabling encoding while suppressing a sharp increase in the number of binarization bits when the prediction residual becomes large. In addition, the three-dimensional data encoding device can also attach the threshold R_TH to the head of the bit stream or the like.

[0626] For example, when encoding at a high bit rate, that is, when the quantization scale is small, the quantization error becomes small and the prediction accuracy becomes high. As a result, the prediction residual may not become large. Therefore, in this case, the three-dimensional data encoding device sets the threshold R_TH large. Thereby, the possibility of encoding the binary data of the threshold R_TH becomes low, and the encoding efficiency is improved. On the contrary, when encoding at a low bit rate, that is, when the quantization scale is large, the quantization error becomes large and the prediction accuracy deteriorates. As a result, it is possible that the prediction residual becomes large. Therefore, in this case, the three-dimensional data encoding device sets the threshold R_TH small. Thereby, a sharp increase in the bit length of the binarized data can be prevented.

[0627] In addition, the three-dimensional data encoding device can also switch the threshold R_TH for each LoD and attach the threshold R_TH of each LoD to the head or the like. That is, the three-dimensional data encoding device can also switch the binarization method for each LoD. For example, in the upper layer, since the distance between three-dimensional points is far, the prediction accuracy deteriorates, and as a result, the prediction residual may become large. Therefore, the three-dimensional data encoding device sets the threshold R_TH small for the upper layer to prevent a sharp increase in the bit length of the binarized data. Additionally, in the lower layer, since the distance between three-dimensional points is close, the prediction accuracy becomes high, and as a result, it is possible that the prediction residual becomes small. Therefore, the three-dimensional data encoding device sets the threshold R_TH large for the layer to improve the encoding efficiency.

[0628] Fig.55 This is a diagram showing an example of an exponential Golomb code, which represents the relationship between the values before binarization (multi-valued) and the bits (codes) after binarization. Additionally, it is also possible to Fig.55 invert the 0 and 1 shown.

[0629] Furthermore, the three-dimensional data encoding device applies arithmetic coding to the binarized data of the prediction residual. Thereby, the coding efficiency can be improved. In addition, when applying arithmetic coding, in the binarized data, in the part binarized with n bits, i.e., the n-bit code, and the part binarized using exponential Golomb, i.e., the remaining code, the tendency of the occurrence probabilities of 0 and 1 for each bit may be different. Therefore, the three-dimensional data encoding device can also switch the application method of arithmetic coding based on the n-bit code and the remaining code.

[0630] For example, for the n-bit code, the three-dimensional data encoding device performs arithmetic coding on each bit using a different coding table (probability table). At this time, the three-dimensional data encoding device can also change the number of coding tables used for each bit. For example, the three-dimensional data encoding device uses 1 coding table to perform arithmetic coding on the leading bit b0 of the n-bit code. Additionally, the three-dimensional data encoding device uses 2 coding tables for the next bit b1. Furthermore, the three-dimensional data encoding device switches the coding table used in the arithmetic coding of bit b1 according to the value of b0 (0 or 1). Similarly, the three-dimensional data encoding device uses 4 coding tables for the next bit b2. Additionally, the three-dimensional data encoding device switches the coding table used in the arithmetic coding of bit b2 according to the values of b0 and b1 (0 to 3).

[0631] In this way, when the three-dimensional data encoding device performs arithmetic coding on each bit bn-1 of the n-bit code, it uses 2 n-1 coding tables. Additionally, the three-dimensional data encoding device switches the coding table used according to the values (occurrence patterns) of the bits before bn-1. Thereby, the three-dimensional data encoding device can use an appropriate coding table for each bit, and thus can improve the coding efficiency.

[0632] Furthermore, the three-dimensional data encoding device can also reduce the number of coding tables used for each bit. For example, when performing arithmetic coding on each bit bn-1, the three-dimensional data encoding device can also switch 2 mA coding table. Thus, the number of coding tables used in each bit can be suppressed, and the coding efficiency can be improved. In addition, the three-dimensional data coding device can also update the occurrence probabilities of 0 and 1 in each coding table according to the value of the actually generated binarized data. Additionally, the three-dimensional data coding device can also fix the occurrence probabilities of 0 and 1 in the coding tables of a part of the bits. Thus, the number of times of updating the occurrence probability can be suppressed, and therefore the processing amount can be reduced.

[0633] For example, when n bits are encoded as b0b1b2…bn-1, the coding table used for b0 is 1 (CTb0). The coding tables used for b1 are 2 (CTb10, CTb11). Additionally, the coding table to be used is switched according to the value of b0 (0 to 1). The coding tables used for b2 are 4 (CTb20, CTb21, CTb22, CTb23). Furthermore, the coding table to be used is switched according to the values of b0 and b1 (0 to 3). The coding tables used for bn-1 are 2 n-1 (CTbn0, CTbn1, …, CTbn(2 n-1 -1)). Additionally, the coding table to be used is switched according to the value of b0b1…bn-2 (0 to 2 n-1 -1).

[0634] In addition, the three-dimensional data coding device can also apply m-ary arithmetic coding (m = 2 n ) that sets values from 0 to 2 n -1 without binarizing the n-bit encoding. In addition, when the three-dimensional data coding device performs arithmetic coding on the n-bit encoding using m-ary, the three-dimensional data decoding device can also restore the n-bit encoding through m-ary arithmetic decoding.

[0635] Fig.56 is a diagram for explaining the processing when, for example, the residual coding is an exponential Golomb code. As Fig.56 shown, the part where binarization is performed using exponential Golomb, that is, the residual coding, includes a prefix part and a suffix part. For example, the three-dimensional data coding device switches the coding table in the prefix part and the suffix part. That is, the three-dimensional data coding device performs arithmetic coding on each bit included in the prefix part using the coding table for the prefix, and performs arithmetic coding on each bit included in the suffix part using the coding table for the suffix.

[0636] In addition, the three-dimensional data encoding device can also update the occurrence probabilities of 0 and 1 in each encoding table according to the values of the actually generated binarized data. Alternatively, the three-dimensional data encoding device can also fix the occurrence probabilities of 0 and 1 in a certain encoding table. Thereby, the number of times of updating the occurrence probabilities can be suppressed, and thus the processing amount can be reduced. For example, the three-dimensional data encoding device can update the occurrence probabilities for the prefix part and fix the occurrence probabilities for the suffix part.

[0637] In addition, the three-dimensional data encoding device decodes the quantized prediction residual through inverse quantization and reconstruction, and uses the decoded prediction residual, that is, the decoded value, for prediction after the three-dimensional point to be encoded. Specifically, the three-dimensional data encoding device calculates the inverse quantization value by multiplying the quantized prediction residual (quantized value) by the quantization scale, and obtains the decoded value (reconstructed value) by adding the inverse quantization value and the predicted value.

[0638] For example, the inverse quantization value a2iq of point a2 is calculated using the quantized value a2q of point a2 through (Equation A11). The inverse quantization value b2iq of point b2 is calculated using the quantized value b2q of point b2 through (Equation A12). Here, QS_LoD0 is the QS for LoD0, and QS_LoD1 is the QS for LoD1. That is, the QS can be changed according to LoD.

[0639] a2iq = a2q × QS_LoD0…(Equation A11)

[0640] b2iq = b2q × QS_LoD1…(Equation A12)

[0641] For example, as shown in (Equation A13), the decoded value a2rec of point a2 is calculated by adding the inverse quantization value a2iq of point a2 to the predicted value a2p of point a2. As shown in (Equation A14), the decoded value b2rec of point b2 is calculated by adding the inverse quantization value b2iq of point b2 to the predicted value b2p of point b2.

[0642] a2rec = a2iq + a2p…(Equation A13)

[0643] b2rec = b2iq + b2p…(Equation A14)

[0644] Hereinafter, a syntax example of the bitstream of the present embodiment will be described. Fig.57 It is a diagram showing a syntax example of the attribute header of the present embodiment. The attribute header is the header information of the attribute information. As Fig.57As shown, the attribute header includes the number of levels information (NumLoD), the three-dimensional point number information (NumOfPoint[i]), the level threshold (Thres_Lod[i]), the surrounding point number information (NumNeighborPoint[i]), the prediction threshold (THd[i]), the quantization scale (QS[i]), and the binarization threshold (R_TH[i]).

[0645] The number of levels information (NumLoD) represents the number of levels of LoD used.

[0646] The three-dimensional point number information (NumOfPoint[i]) represents the number of three-dimensional points belonging to level i. In addition, the three-dimensional data encoding device may also attach the three-dimensional point total number information (AllNumOfPoint) representing the total number of three-dimensional points to other headers. In this case, the three-dimensional data encoding device may not attach NumOfPoint[NumLoD - 1] representing the number of three-dimensional points belonging to the bottommost level to the header. In this case, the three-dimensional data decoding device can calculate NumOfPoint[NumLoD - 1] through (Equation A15). Thereby, the encoding amount of the header can be reduced.

[0647]

Equation 4

[0648]

[0649] The level threshold (Thres_Lod[i]) is a threshold for setting level i. The three-dimensional data encoding device and the three-dimensional data decoding device construct LoDi such that the distance between each point within LoDi is greater than the threshold Thres_LoD[i]. In addition, the three-dimensional data encoding device may not attach the value of Thres_Lod[NumLoD - 1] (the bottommost level) to the header. In this case, the three-dimensional data decoding device estimates the value of Thres_Lod[NumLoD - 1] as 0. Thereby, the encoding amount of the header can be reduced.

[0650] The surrounding point number information (NumNeighborPoint[i]) represents the upper limit value of the number of surrounding points used in the generation of the predicted value of the three-dimensional points belonging to level i. When the number of surrounding points M is less than NumNeighborPoint[i] (M < NumNeighborPoint[i]), the three-dimensional data encoding device may also use M surrounding points to calculate the predicted value. In addition, when it is not necessary to separate the value of NumNeighborPoint[i] in each LoD, the three-dimensional data encoding device may also attach 1 surrounding point number information (NumNeighborPoint) used in all LoDs to the header.

[0651] The prediction threshold (THd[i]) represents the upper limit value of the distance between the surrounding three-dimensional points used in the prediction of the object three-dimensional points for encoding or decoding the object at level i. The three-dimensional data encoding device and the three-dimensional data decoding device do not use three-dimensional points whose distance from the object three-dimensional point is farther than THd[i] for prediction. Additionally, when it is not necessary to separate the values of THd[i] for each LoD, the three-dimensional data encoding device may also attach one prediction threshold (THd) used for all LoDs to the header.

[0652] The quantization scale (QS[i]) represents the quantization scale used in quantization and inverse quantization at level i.

[0653] The binarization threshold (R_TH[i]) is a threshold for switching the binarization method of the prediction residual of the three-dimensional points belonging to level i. For example, when the prediction residual is less than the threshold R_TH, the three-dimensional data encoding device binarizes the prediction residual pu with a fixed number of bits, and when the prediction residual is equal to or greater than the threshold R_TH, it binarizes the binarized data of the threshold R_TH and the value of (pu - R_TH) using Exponential Golomb. Additionally, when it is not necessary to switch the values of R_TH[i] for each LoD, the three-dimensional data encoding device may also attach one binarization threshold (R_TH) used for all LoDs to the header.

[0654] Furthermore, R_TH[i] may also be the maximum value represented by nbit. For example, in 6 bits, R_TH is 63, and in 8 bits, R_TH is 255. Additionally, instead of encoding the maximum value represented by nbit as the binarization threshold, the three-dimensional data encoding device may encode the number of bits. For example, the three-dimensional data encoding device may attach the value 6 to the header when R_TH[i] = 63, and attach the value 8 to the header when R_TH[i] = 255. Additionally, the three-dimensional data encoding device may also define the minimum value (minimum number of bits) representing the number of bits of R_TH[i] and attach the relative number of bits based on the minimum value to the header. For example, the three-dimensional data encoding device may attach the value 0 to the header when R_TH[i] = 63 and the minimum number of bits is 6, and attach the value 2 to the header when R_TH[i] = 255 and the minimum number of bits is 6.

[0655] Additionally, the three-dimensional data encoding device may also perform entropy encoding on at least one of NumLoD, Thres_Lod[i], NumNeighborPoint[i], THd[i], QS[i], and R_TH[i] and attach it to the header. For example, the three-dimensional data encoding device may also perform arithmetic encoding by binarizing each value. Additionally, in order to suppress the processing amount, the three-dimensional data encoding device may also encode each value with a fixed length.

[0656] In addition, the three-dimensional data encoding device may not attach at least one of NumLoD, Thres_Lod[i], NumNeighborPoint[i], THd[i], QS[i], and R_TH[i] to the header. For example, the value of at least one of them may also be specified by a profile or level of a standard or the like. Thereby, the bit amount of the header can be reduced.

[0657] Fig.58 FIG. is a syntax example of attribute data of the present embodiment. The attribute data includes encoded data of attribute information of a plurality of three-dimensional points. As Fig.58 shown, the attribute data includes an n-bit code and a remaining code.

[0658] The n-bit code is encoded data of a prediction residual of the value of the attribute information or a part thereof. The bit length of the n-bit code depends on the value of R_TH[i]. For example, when the value shown by R_TH[i] is 63, the n-bit code is 6 bits, and when the value shown by R_TH[i] is 255, the n-bit code is 8 bits.

[0659] The remaining code is the encoded data after exponential Golomb coding in the encoded data of the prediction residual of the value of the attribute information. When the n-bit code is the same as R_TH[i], the remaining code is encoded or decoded. In addition, the three-dimensional data decoding device adds the value of the n-bit code and the value of the remaining code to decode the prediction residual. In addition, when the n-bit code is not the same value as R_TH[i], the remaining code may not be encoded or decoded.

[0660] Hereinafter, the flow of processing in the three-dimensional data encoding device will be described. Fig.59 FIG. is a flowchart of three-dimensional data encoding processing performed by the three-dimensional data encoding device.

[0661] First, the three-dimensional data encoding device encodes the geometry (S3001). For example, the three-dimensional data encoding device performs encoding using an octree representation.

[0662] After encoding the position information, when the positions of the three-dimensional points change due to quantization or the like, the three-dimensional data encoding device reassigns the attribute information of the original three-dimensional points to the changed three-dimensional points (S3002). For example, the three-dimensional data encoding device performs reassignment by interpolating the values of the attribute information according to the amount of change in position. For example, the three-dimensional data encoding device detects N three-dimensional points before the change that are close to the changed three-dimensional position, and performs weighted averaging on the values of the attribute information of the N three-dimensional points. For example, in the weighted averaging, the three-dimensional data encoding device determines the weights based on the distances from the changed three-dimensional position to each of the N three-dimensional points. Then, the three-dimensional data encoding device determines the value obtained by weighted averaging as the value of the attribute information of the changed three-dimensional point. In addition, when two or more three-dimensional points change to the same three-dimensional position due to quantization or the like, the three-dimensional data encoding device may also assign the average value of the attribute information of the two or more three-dimensional points before the change as the value of the attribute information of the changed three-dimensional point.

[0663] Next, the three-dimensional data encoding device encodes the reassigned attribute information (S3003). For example, when encoding multiple types of attribute information, the three-dimensional data encoding device may also encode the multiple types of attribute information sequentially. For example, when encoding color and reflectance as the attribute information, the three-dimensional data encoding device may also generate a bitstream in which the encoding result of the reflectance is appended after the encoding result of the color. In addition, the order of the multiple encoding results of the attribute information appended to the bitstream is not limited to this order and can be any order.

[0664] In addition, the three-dimensional data encoding device may also attach information indicating the start position of the encoded data of each attribute information in the bitstream to the header or the like. As a result, the three-dimensional data decoding device can selectively decode the attribute information that needs to be decoded, and thus can omit the decoding process of the attribute information that does not need to be decoded. Therefore, the processing amount of the three-dimensional data decoding device can be reduced. In addition, the three-dimensional data encoding device may also encode multiple types of attribute information in parallel and merge the encoding results into one bitstream. As a result, the three-dimensional data encoding device can encode multiple types of attribute information at high speed.

[0665] Fig.60 It is a flowchart of the attribute information encoding process (S3003). First, the three-dimensional data encoding device sets the LoD (S3011). That is, the three-dimensional data encoding device assigns each three-dimensional point to any one of the multiple LoDs.

[0666] Next, the three-dimensional data encoding device starts a loop for each LoD (S3012). That is, the three-dimensional data encoding device repeatedly performs the processing of steps S3013 to S3021 for each LoD.

[0667] Next, the three-dimensional data encoding device starts a loop for each three-dimensional point (S3013). That is, the three-dimensional data encoding device repeatedly performs the processes of steps S3014 to S3020 for each three-dimensional point.

[0668] First, the three-dimensional data encoding device searches for a plurality of surrounding points, which are three-dimensional points existing around the object three-dimensional point used in the calculation of the predicted value of the object three-dimensional point to be processed (S3014). Next, the three-dimensional data encoding device calculates the weighted average of the values of the attribute information of the plurality of surrounding points, and sets the obtained value as the predicted value P (S3015). Next, the three-dimensional data encoding device calculates the difference between the attribute information of the object three-dimensional point and the predicted value, that is, the prediction residual (S3016). Next, the three-dimensional data encoding device calculates the quantization value by quantizing the prediction residual (S3017). Next, the three-dimensional data encoding device performs arithmetic coding on the quantization value (S3018).

[0669] In addition, the three-dimensional data encoding device calculates the inverse quantization value by inverse quantizing the quantization value (S3019). Next, the three-dimensional data encoding device generates a decoded value by adding the inverse quantization value to the predicted value (S3020). Next, the three-dimensional data encoding device ends the loop for each three-dimensional point (S3021). In addition, the three-dimensional data encoding device ends the loop for each LoD unit (S3022).

[0670] Hereinafter, the three-dimensional data decoding process in the three-dimensional data decoding device that decodes the bitstream generated by the above three-dimensional data encoding device will be described.

[0671] The three-dimensional data decoding device generates decoded binary data by performing arithmetic decoding on the binary data of the attribute information in the bitstream generated by the three-dimensional data encoding device in the same manner as the three-dimensional data encoding device. In addition, in the three-dimensional data encoding device, when the application method of arithmetic coding is switched between the part binarized with n bits (n-bit coding) and the part binarized with exponential Golomb (residual coding), the three-dimensional data decoding device performs decoding accordingly when applying arithmetic decoding.

[0672] For example, in an arithmetic decoding method of n-bit encoding, a three-dimensional data decoding device performs arithmetic decoding on each bit using a different encoding table (decoding table). At this time, the three-dimensional data decoding device can also change the number of encoding tables used for each bit. For example, for the first bit b0 of n-bit encoding, 1 encoding table is used for arithmetic decoding. In addition, the three-dimensional data decoding device uses 2 encoding tables for the next bit b1. Furthermore, the three-dimensional data decoding device switches the encoding table used in the arithmetic decoding of bit b1 according to the value (0 or 1) of b0. Similarly, the three-dimensional data decoding device further uses 4 encoding tables for the next bit b2. In addition, the three-dimensional data decoding device switches the encoding table used in the arithmetic decoding of bit b2 according to the values (0 to 3) of b0 and b1.

[0673] In this way, when the three-dimensional data decoding device performs arithmetic decoding on each bit bn-1 of n-bit encoding, it uses 2 n-1 encoding tables. In addition, the three-dimensional data decoding device switches the encoding table used according to the values (occurrence patterns) of the bits before bn-1. Thereby, the three-dimensional data decoding device can use an appropriate encoding table for each bit to appropriately decode a bit stream with improved encoding efficiency.

[0674] In addition, the three-dimensional data decoding device can also reduce the number of encoding tables used in each bit. For example, the three-dimensional data decoding device can also switch 2 m encoding tables according to the values (occurrence patterns) of the m bits (m < n-1) before bn-1 when performing arithmetic decoding on each bit bn-1. Thereby, the three-dimensional data decoding device can appropriately decode a bit stream with improved encoding efficiency while suppressing the number of encoding tables used in each bit. In addition, the three-dimensional data decoding device can also update the occurrence probabilities of 0 and 1 in each encoding table according to the values of the actually generated binarized data. In addition, the three-dimensional data decoding device can also fix the occurrence probabilities of 0 and 1 in the encoding tables of a part of the bits. Thereby, the number of updates of the occurrence probability can be suppressed, and thus the processing amount can be reduced.

[0675] For example, when n-bit encoding is b0b1b2…bn-1, the encoding table used for b0 is 1 (CTb0). The encoding tables used for b1 are 2 (CTb10, CTb11). In addition, the encoding table is switched according to the value (0 to 1) of b0. The encoding tables used for b2 are 4 (CTb20, CTb21, CTb22, CTb23). In addition, the encoding table is switched according to the values (0 to 3) of b0 and b1. The encoding tables used for bn-1 are 2 n-1 (CTbn0, CTbn1, …, CTbn(2 n-1 -1)). In addition, according to the values (0 to 2 n-1-1) to switch the code table.

[0676] For example, Fig.61 is a diagram for explaining the processing in the case where the remaining code is an exponential Golomb code. As Fig.61 shown, the part (remaining code) encoded by binarization using exponential Golomb in the three-dimensional data encoding device includes a prefix part and a suffix part. For example, the three-dimensional data decoding device switches the code table between the prefix part and the suffix part. That is, the three-dimensional data decoding device performs arithmetic decoding on each bit included in the prefix part using the code table for the prefix, and performs arithmetic decoding on each bit included in the suffix part using the code table for the suffix.

[0677] In addition, the three-dimensional data decoding device may update the occurrence probabilities of 0 and 1 in each code table according to the value of the binarized data generated during decoding. Alternatively, the three-dimensional data decoding device may fix the occurrence probabilities of 0 and 1 in a certain code table. Thereby, the number of times of updating the occurrence probability can be suppressed, and thus the processing amount can be reduced. For example, the three-dimensional data decoding device may update the occurrence probability for the prefix part and fix the occurrence probability for the suffix part.

[0678] In addition, the three-dimensional data decoding device multi-valued the binarized data of the predicted residual obtained by arithmetic decoding in accordance with the coding method used in the three-dimensional data encoding device, thereby decoding the quantized predicted residual (unsigned integer value). The three-dimensional data decoding device first calculates the value of the n-bit code decoded by performing arithmetic decoding on the binarized data encoded by n bits. Next, the three-dimensional data decoding device compares the value of the n-bit code with the value of R_TH.

[0679] When the value of the n-bit code is consistent with the value of R_TH, the three-dimensional data decoding device determines that there are bits encoded by exponential Golomb next, and performs arithmetic decoding on the binarized data encoded by exponential Golomb, that is, the remaining code. Then, the three-dimensional data decoding device calculates the value of the remaining code using the inverse table showing the relationship between the remaining code and the value. Fig.62 is a diagram showing an example of the inverse table representing the relationship between the remaining code and its value. Next, the three-dimensional data decoding device obtains the multi-valued quantized predicted residual by adding the value of the obtained remaining code to R_TH.

[0680] On the other hand, when the value of the n-bit encoding does not match the value of R_TH (the value is less than R_TH), the three-dimensional data decoding device directly determines the value of the n-bit encoding as the quantized prediction residual after multi-valued quantization. Thereby, the three-dimensional data decoding device can appropriately decode the bitstream generated by switching the binarization method according to the value of the prediction residual in the three-dimensional data encoding device.

[0681] In addition, when the threshold R_TH is attached to the head of the bitstream or the like, the three-dimensional data decoding device can also decode the value of the threshold R_TH from the head and use the decoded value of the threshold R_TH to switch the decoding method. Further, when the three-dimensional data decoding device attaches the threshold R_TH to the head or the like for each LoD, it switches the decoding method for each LoD using the decoded threshold R_TH.

[0682] For example, when the threshold R_TH is 63 and the decoded value of the n-bit encoding is 63, the three-dimensional data decoding device decodes the remaining encoding using exponential Golomb to obtain the value of the remaining encoding. For example, in Fig.62 the example shown, the remaining encoding is 00100, and 3 is obtained as the value of the remaining encoding. Then, the three-dimensional data decoding device adds the value 63 of the threshold R_TH and the value 3 of the remaining encoding to obtain the value 66 of the prediction residual.

[0683] In addition, when the decoded value of the n-bit encoding is 32, the three-dimensional data decoding device sets the value 32 of the n-bit encoding as the value of the prediction residual.

[0684] In addition, the three-dimensional data decoding device transforms the decoded quantized prediction residual from an unsigned integer value to a signed integer value through, for example, a process opposite to that in the three-dimensional data encoding device. Thereby, when performing entropy coding on the prediction residual, the three-dimensional data decoding device can appropriately decode the bitstream generated without considering the generation of negative integers. In addition, the three-dimensional data decoding device does not necessarily need to transform the unsigned integer value to a signed integer value. For example, when decoding a bitstream generated by separately performing entropy coding on the sign bit, it can also decode the sign bit.

[0685] The three-dimensional data decoding device decodes the quantized prediction residual transformed into a signed integer value through inverse quantization and reconstruction, thereby generating a decoded value. In addition, the three-dimensional data decoding device uses the generated decoded value for prediction after the three-dimensional point to be decoded. Specifically, the three-dimensional data decoding device calculates the inverse quantization value by multiplying the quantized prediction residual by the decoded quantization scale, and adds the inverse quantization value and the prediction value to obtain the decoded value.

[0686] The decoded unsigned integer value (unsigned quantization value) is transformed into a signed integer value through the following process. When the least significant bit (LSB) of the decoded unsigned integer value a2u is 1, the 3D data decoding device sets the signed integer value a2q to -((a2u + 1) >> 1). When the LSB of the unsigned integer value a2u is not 1, the 3D data decoding device sets the signed integer value a2q to (a2u >> 1).

[0687] Similarly, when the LSB of the decoded unsigned integer value b2u is 1, the 3D data decoding device sets the signed integer value b2q to -((b2u + 1) >> 1). When the LSB of the unsigned integer value n2u is not 1, the 3D data decoding device sets the signed integer value b2q to (b2u >> 1).

[0688] In addition, the details of the inverse quantization and reconstruction processing performed by the 3D data decoding device are the same as those of the inverse quantization and reconstruction processing in the 3D data encoding device.

[0689] Hereinafter, the processing flow in the 3D data decoding device will be described. Fig.63 It is a flowchart of the 3D data decoding process performed by the 3D data decoding device. First, the 3D data decoding device decodes the position information (geometry) from the bitstream (S3031). For example, the 3D data decoding device performs decoding using an octree representation.

[0690] Next, the 3D data decoding device decodes the attribute information (Attribute) from the bitstream (S3032). For example, when decoding multiple pieces of attribute information, the 3D data decoding device can also decode the multiple pieces of attribute information sequentially. For example, when decoding color and reflectance as attribute information, the 3D data decoding device decodes the encoding result of the color and the encoding result of the reflectance in the order attached to the bitstream. For example, when the encoding result of the reflectance is attached after the encoding result of the color in the bitstream, the 3D data decoding device decodes the encoding result of the color and then decodes the encoding result of the reflectance. In addition, the 3D data decoding device can decode the encoding results of the attribute information attached to the bitstream in any order.

[0691] In addition, the three-dimensional data decoding device can also obtain information indicating the start position of the encoded data of each attribute information in the bitstream by decoding the header and the like. Thus, the three-dimensional data decoding device can selectively decode the attribute information that needs to be decoded, and therefore can omit the decoding process of the attribute information that does not need to be decoded. Therefore, the processing amount of the three-dimensional data decoding device can be reduced. In addition, the three-dimensional data decoding device can also decode multiple types of attribute information in parallel and merge the decoding results into one three-dimensional point cloud. Thus, the three-dimensional data decoding device can decode multiple types of attribute information at high speed.

[0692] Fig.64 is a flowchart of the attribute information decoding process (S3032). First, the three-dimensional data decoding device sets the LoD (S3041). That is, the three-dimensional data decoding device assigns each of the multiple three-dimensional points with decoded position information to any one of the multiple LoDs. For example, this assignment method is the same as the assignment method used in the three-dimensional data encoding device.

[0693] Next, the three-dimensional data decoding device starts a loop for each LoD (S3042). That is, the three-dimensional data decoding device repeatedly performs the processes of steps S3043 to S3049 for each LoD.

[0694] Next, the three-dimensional data decoding device starts a loop for each three-dimensional point (S3043). That is, the three-dimensional data decoding device repeatedly performs the processes of steps S3044 to S3048 for each three-dimensional point.

[0695] First, the three-dimensional data decoding device searches for a plurality of surrounding points, which are three-dimensional points existing around the object three-dimensional point used in the calculation of the predicted value of the object three-dimensional point to be processed (S3044). Next, the three-dimensional data decoding device calculates the weighted average of the values of the attribute information of the plurality of surrounding points and sets the obtained value as the predicted value P (S3045). In addition, these processes are the same as the processes in the three-dimensional data encoding device.

[0696] Next, the three-dimensional data decoding device performs arithmetic decoding on the quantization value from the bitstream (S3046). In addition, the three-dimensional data decoding device calculates the inverse quantization value by inverse quantizing the decoded quantization value (S3047). Next, the three-dimensional data decoding device generates a decoded value by adding the inverse quantization value to the predicted value (S3048). Next, the three-dimensional data decoding device ends the loop for each three-dimensional point (S3049). In addition, the three-dimensional data decoding device ends the loop for each LoD (S3050).

[0697] Next, the structures of the three-dimensional data encoding device and the three-dimensional data decoding device of the present embodiment will be described. Fig.65It is a block diagram showing the structure of the three-dimensional data encoding device 3000 according to the present embodiment. The three-dimensional data encoding device 3000 includes a position information encoding unit 3001, an attribute information redistribution unit 3002, and an attribute information encoding unit 3003.

[0698] The attribute information encoding unit 3003 encodes the position information (geometry) of a plurality of three-dimensional points included in the input point cloud. The attribute information redistribution unit 3002 redistributes the values of the attribute information of the plurality of three-dimensional points included in the input point cloud by using the encoding and decoding results of the position information. The attribute information encoding unit 3003 encodes the redistributed attribute information (attribute). In addition, the three-dimensional data encoding device 3000 generates a bitstream including the encoded position information and the encoded attribute information.

[0699] Fig.66 It is a block diagram showing the structure of the three-dimensional data decoding device 3010 according to the present embodiment. The three-dimensional data decoding device 3010 includes a position information decoding unit 3011 and an attribute information decoding unit 3012.

[0700] The position information decoding unit 3011 decodes the position information (geometry) of a plurality of three-dimensional points from the bitstream. The attribute information decoding unit 3012 decodes the attribute information (attribute) of a plurality of three-dimensional points from the bitstream. In addition, the three-dimensional data decoding device 3010 generates an output point cloud by combining the decoded position information and the decoded attribute information.

[0701] As described above, the three-dimensional data encoding device according to the present embodiment performs Fig.67 the processing shown in. The three-dimensional data encoding device encodes three-dimensional points having attribute information. First, the three-dimensional data encoding device calculates a predicted value of the attribute information of the three-dimensional points (S3061). Next, the three-dimensional data encoding device calculates the difference between the attribute information of the three-dimensional points and the predicted value, that is, the prediction residual (S3062). Then, the three-dimensional data encoding device generates binary data by binarizing the prediction residual (S3063). Then, the three-dimensional data encoding device performs arithmetic coding on the binary data (S3064).

[0702] Thus, the three-dimensional data encoding device can reduce the amount of encoded data of the attribute information by calculating the prediction residual of the attribute information, and then binarizing and arithmetically coding the prediction residual.

[0703] For example, in the arithmetic coding (S3064), the three-dimensional data encoding device uses a different coding table for each bit of the binary data. Thus, the three-dimensional data encoding device can improve the coding efficiency.

[0704] For example, in arithmetic coding (S3064), the lower the bit position of the binary data, the larger the number of coding tables used.

[0705] For example, in arithmetic coding (S3064), the three-dimensional data coding device selects a coding table to be used in the arithmetic coding of the target bit according to the value of the upper bit of the target bit included in the binary data. Thus, the three-dimensional data coding device can select a coding table according to the value of the upper bit, and therefore can improve the coding efficiency.

[0706] For example, in binarization (S3063), when the prediction residual is less than the threshold (R_TH), the three-dimensional data coding device generates binary data by binarizing the prediction residual with a fixed number of bits. When the prediction residual is greater than or equal to the threshold (R_TH), the three-dimensional data coding device generates binary data including a first code (n-bit code) representing the threshold (R_TH) with a fixed number of bits and a second code (residual code) obtained by binarizing the value obtained by subtracting the threshold (R_TH) from the prediction residual using exponential Golomb. The three-dimensional data coding device uses different arithmetic coding methods for the first code and the second code in arithmetic coding (S3064).

[0707] Thus, the three-dimensional data coding device can, for example, perform arithmetic coding on the first code and the second code by arithmetic coding methods respectively suitable for the first code and the second code, and therefore can improve the coding efficiency.

[0708] For example, the three-dimensional data coding device quantizes the prediction residual and binarizes the quantized prediction residual in binarization (S3063). The threshold (R_TH) is changed according to the quantization scale in quantization. Thus, the three-dimensional data coding device can use an appropriate threshold corresponding to the quantization scale, and therefore can improve the coding efficiency.

[0709] For example, the second code includes a prefix part and a suffix part. The three-dimensional data coding device uses different coding tables for the prefix part and the suffix part in arithmetic coding (S3064). Thus, the three-dimensional data coding device can improve the coding efficiency.

[0710] For example, the three-dimensional data coding device includes a processor and a memory, and the processor uses the memory to perform the above processing.

[0711] In addition, the three-dimensional data decoding device of the present embodiment performs Fig.68The processing shown. The three-dimensional data decoding device decodes three-dimensional points with attribute information. First, the three-dimensional data decoding device calculates a predicted value of the attribute information of the three-dimensional points (S3071). Next, the three-dimensional data decoding device generates binary data by performing arithmetic decoding on the encoded data included in the bitstream (S3072). Next, the three-dimensional data decoding device generates a prediction residual by multi-valuing the binary data (S3073). Next, the three-dimensional data decoding device calculates a decoded value of the attribute information of the three-dimensional points by adding the predicted value and the prediction residual (S3074).

[0712] Thus, the three-dimensional data decoding device can calculate the prediction residual of the attribute information, and thus appropriately decode the bitstream of the attribute information generated by binarizing and arithmetic encoding the prediction residual.

[0713] For example, in arithmetic decoding (S3072), the three-dimensional data decoding device uses different coding tables for each bit of the binary data. Thus, the three-dimensional data decoding device can appropriately decode the bitstream with improved coding efficiency.

[0714] For example, in arithmetic decoding (S3072), the lower the bit position of the binary data, the larger the number of coding tables used.

[0715] For example, in arithmetic decoding (S3072), the three-dimensional data decoding device selects the coding table used in the arithmetic decoding of the target bit according to the value of the upper bit of the target bit included in the binary data. Thus, the three-dimensional data decoding device can appropriately decode the bitstream with improved coding efficiency.

[0716] For example, in multi-valuing (S3073), the three-dimensional data decoding device generates a first value by multi-valuing the first encoding (n-bit encoding) with a fixed number of bits included in the binary data. When the first value is less than the threshold (R_TH), the three-dimensional data decoding device determines the first value as the prediction residual. When the first value is equal to or greater than the threshold (R_TH), the three-dimensional data decoding device generates a second value by multi-valuing the exponential Golomb code, i.e., the second encoding (remaining encoding) included in the binary data, and adds the first value and the second value to generate the prediction residual. The three-dimensional data decoding device uses different arithmetic decoding methods for the first encoding and the second encoding in arithmetic decoding (S3072).

[0717] Thus, the three-dimensional data decoding device can appropriately decode the bitstream with improved coding efficiency.

[0718] For example, the three-dimensional data decoding device performs inverse quantization on the prediction residual, and in the addition operation (S3074), adds the predicted value and the inverse quantized prediction residual. The threshold value (R_TH) is changed according to the quantization scale in the inverse quantization. Thereby, the three-dimensional data decoding device can appropriately decode the bitstream with improved coding efficiency.

[0719] For example, the second encoding includes a prefix part and a suffix part. The three-dimensional data decoding device uses different coding tables for the prefix part and the suffix part in the arithmetic decoding (S3072). Thereby, the three-dimensional data decoding device can appropriately decode the bitstream with improved coding efficiency.

[0720] For example, the three-dimensional data decoding device includes a processor and a memory, and the processor uses the memory to perform the above processing.

[0721] (Embodiment 9)

[0722] The predicted value may also be generated by a method different from that of Embodiment 8. Hereinafter, the three-dimensional points to be encoded may sometimes be referred to as the first three-dimensional points, and the surrounding three-dimensional points may be referred to as the second three-dimensional points.

[0723] For example, in the generation of the predicted value of the attribute information of the three-dimensional points, the attribute value of the three-dimensional point with the shortest distance among the encoded and decoded surrounding three-dimensional points of the three-dimensional point to be encoded may be directly used as the predicted value. In addition, in the generation of the predicted value, prediction mode information (PredMode) may be attached to each three-dimensional point, and the predicted value may be generated by selecting one predicted value from a plurality of predicted values. That is, for example, it may be considered that in a total of M prediction modes, the average value is assigned to prediction mode 0, the attribute value of three-dimensional point A is assigned to prediction mode 1,..., the attribute value of three-dimensional point Z is assigned to prediction mode M - 1, and the prediction mode used in the prediction is attached to the bitstream for each three-dimensional point. In this way, it may also be that the first prediction mode value indicating the first prediction mode in which the average of the attribute information of the surrounding three-dimensional points is calculated as the predicted value is smaller than the second prediction mode value indicating the second prediction mode in which the attribute information of the surrounding three-dimensional points itself is calculated as the predicted value. Here, the predicted value calculated in prediction mode 0, that is, the "average value", is the average of the attribute values of the surrounding three-dimensional points of the three-dimensional point to be encoded.

[0724] Fig.69 It is a first example of a diagram showing a table of predicted values calculated in each prediction mode of Embodiment 9. Fig.70 It is a diagram showing an example of the attribute information used in the predicted value of Embodiment 9. Fig.71 It is a second example of a diagram showing a table of predicted values calculated in each prediction mode of Embodiment 9.

[0725] The number of prediction modes M can also be appended to the bitstream. Additionally, the number of prediction modes M may not be appended to the bitstream, and the value may be specified by the standard profile, level, etc. Further, the number of prediction modes M may also use a value calculated based on the number of three-dimensional points N used in the prediction. For example, the number of prediction modes M may be calculated by M = N + 1.

[0726] In addition, Fig.69 The table shown is an example in the case where the number of three-dimensional points N used in the prediction is 4 and the number of prediction modes M is 5. The attribute information of point a0, a1, a2, b1 can be used to generate the predicted value of the attribute information of point b2. In the case of selecting one prediction mode from multiple prediction modes, the prediction mode that generates the attribute values of points a0, a1, a2, b1 as predicted values can also be selected based on the distance information from point b2 to each of points a0, a1, a2, b1. A prediction mode is appended to each three-dimensional point of the coding object. The predicted value is calculated according to the value corresponding to the appended prediction mode.

[0727] Fig.71 The table shown is the same as Fig.69 Similarly, it is an example in the case where the number of three-dimensional points N used in the prediction is 4 and the number of prediction modes M is 5. The predicted value of the attribute information of point a2 can be generated using the attribute information of points a0, a1. In the case of selecting one prediction mode from multiple prediction modes, the prediction mode that generates the attribute values of points a0, a1 as predicted values can also be selected based on the distance information from point a2 to each of points a0, a1. A prediction mode is appended to each three-dimensional point of the coding object. The predicted value is calculated according to the value corresponding to the appended prediction mode.

[0728] In addition, in the case where the number of adjacent points, that is, the number of surrounding three-dimensional points N, is less than 4 as in point a2 above, the prediction mode for which no predicted value is assigned in the table may be set to not available.

[0729] In addition, the assignment of the values of the prediction modes can also be determined in the order of the distance from the three-dimensional points of the coding object. For example, the closer the distance from the three-dimensional points of the coding object to the surrounding three-dimensional points having the attribute information used as the predicted value, the smaller the prediction mode value representing the multiple prediction modes. In Fig.69In the example, it indicates the case where the distances to point b2 of the three-dimensional points to be encoded are close in the order of points b1, a2, a1, and a0. For example, in the calculation of the predicted value, in the prediction mode where the prediction mode value in two or more prediction modes is as shown in "1", the attribute information of point b1 is calculated as the predicted value, and in the prediction mode where the prediction mode value is as shown in "2", the attribute information of point a2 is calculated as the predicted value. In this way, it indicates that the prediction mode value representing the prediction mode in which the attribute information of point b1 is calculated as the predicted value is smaller than the prediction mode value representing the prediction mode in which the attribute information of point a2 is calculated as the predicted value, and point a2 is located at a position farther from point b2 than point b1.

[0730] Thus, a small prediction mode value can be assigned to a certain point that is likely to be easily selected and is easy to predict due to its close distance, and the number of bits used to encode the prediction mode value can be reduced. In addition, small prediction mode values can also be preferentially assigned to three-dimensional points belonging to the same LoD as the three-dimensional points to be encoded.

[0731] Fig.72 It is a diagram showing a third example of a table indicating the predicted values calculated in each prediction mode of Embodiment 9. Specifically, the third example is an example where the attribute information used in the predicted value is a value based on the color information (YUV) of surrounding three-dimensional points. In this way, the attribute information for the predicted value can also be color information representing the color of the three-dimensional points.

[0732] As Fig.72 shown, the predicted value calculated in the prediction mode where the prediction mode value is as shown in "0" is the average of the respective components of YUV that define the YUV color space. Specifically, this predicted value includes: the weighted average Yave of the values of the Y components Yb1, Ya2, Ya1, and Ya0 corresponding to points b1, a2, a1, and a0 respectively; the weighted average Uave of the values of the U components Ub1, Ua2, Ua1, and Ua0 corresponding to points b1, a2, a1, and a0 respectively; and the weighted average Vave of the values of the V components Vb1, Va2, Va1, and Va0 corresponding to points b1, a2, a1, and a0 respectively. In addition, the predicted values calculated in the prediction modes where the prediction mode values are from "1" to "4" respectively include the color information of the surrounding three-dimensional points b1, a2, a1, and a0. The color information is represented by a combination of the values of the Y component, U component, and V component.

[0733] In addition, in Fig.72 , the color information is represented by the values defined by the YUV color space, but it is not limited to the YUV color space. It can also be represented by the values defined by the RGB color space or the values defined by other color spaces.

[0734] Thus, it is also possible that, in the calculation of the predicted value, two or more averages or attribute information are calculated as the predicted value of the prediction mode. Additionally, the two or more averages or attribute information may also respectively represent the values of two or more components defining a color space.

[0735] In addition, for example, in Fig.72 when the prediction mode shown by the prediction mode value "2" is selected in the table, the Y component, U component, and V component of the attribute value of the three-dimensional point of the coding target may also be used as the predicted values Ya2, Ua2, Va2, respectively, for coding. In this case, "2" as the prediction mode value is appended to the bit stream.

[0736] Fig.73 FIG. is a diagram showing a fourth example of a table showing the predicted values calculated in each prediction mode of Embodiment 9. Specifically, the fourth example is an example where the attribute information used in the predicted value is a value based on the reflectance information of surrounding three-dimensional points. The reflectance information is, for example, information indicating the reflectance R.

[0737] As Fig.73 shown, the predicted value calculated in the prediction mode shown by the prediction mode value "0" is the weighted average Rave of the reflectances Rb1, Ra2, Ra1, Ra0 corresponding to the points b1, a2, a1, a0, respectively. Additionally, the predicted values calculated in the prediction modes shown by the prediction mode values "1" to "4" are the reflectances Rb1, Ra2, Ra1, Ra0 of the surrounding three-dimensional points b1, a2, a1, a0, respectively.

[0738] In addition, for example, in Fig.73 when the prediction mode shown by the prediction mode value "3" is selected in the table, the reflectance of the attribute value of the three-dimensional point of the coding target may also be used as the predicted value Ra1 for coding. In this case, "3" as the prediction mode value is appended to the bit stream.

[0739] As Fig.72 and Fig.73 shown, the attribute information may include first attribute information and second attribute information of a different type from the first attribute information. The first attribute information is, for example, color information. The second attribute information is, for example, reflectance information. In the calculation of the predicted value, the first attribute information may also be used to calculate the first predicted value, and the second attribute information may be used to calculate the second predicted value.

[0740] (Embodiment 10)

[0741] As another example of encoding the attribute information of three-dimensional points using the information of LoD, a method of encoding multiple three-dimensional points sequentially starting from the three-dimensional points included in the upper layer of LoD will be described. For example, when the three-dimensional data encoding device calculates the predicted value (attribute information) of the attribute value of the three-dimensional points included in LoDn, it can also use a flag or the like to switch which LoD's included three-dimensional point attribute value can be referred to. For example, the three-dimensional data encoding device generates information indicating whether to permit reference to other three-dimensional points within the same LoD as the object three-dimensional point to be encoded, i.e., EnableReferringSameLoD (same layer reference permission flag). For example, when EnableReferringSameLoD has a value of 1, reference within the same LoD is permitted, and when EnableReferringSameLoD has a value of 0, reference within the same LoD is prohibited.

[0742] For example, the three-dimensional data encoding device selects the three-dimensional points around the object three-dimensional point based on EnableReferringSameLoD, calculates the average of the attribute values of a predetermined number of three-dimensional points or less among the selected surrounding three-dimensional points, and thereby generates a predicted value of the attribute information of the object three-dimensional point. In addition, the three-dimensional data encoding device attaches the value of N to the head of the bitstream or the like. Alternatively, the three-dimensional data encoding device can also attach the value of N to each three-dimensional point for which the predicted value is generated. Thereby, an appropriate N can be selected for each three-dimensional point for which the predicted value is generated, so that the accuracy of the predicted value can be improved and the prediction residual can be reduced.

[0743] Or, the three-dimensional data encoding device can also attach the value of N to the head of the bitstream and fix the value of N within the bitstream. Thereby, it is not necessary to encode or decode the value of N for each three-dimensional point, and the processing amount can be reduced.

[0744] Or, the three-dimensional data encoding device can also encode the information representing the value of N separately for each LoD. Thereby, by selecting an appropriate value of N for each LoD, the encoding efficiency can be improved. In addition, the three-dimensional data encoding device can also calculate the predicted value of the attribute information of the three-dimensional point based on the weighted average of the attribute information of the surrounding N three-dimensional points. For example, the three-dimensional data encoding device uses the distance information between the object three-dimensional point and each of the N three-dimensional points to calculate the weight.

[0745] In this way, EnableReferringSameLoD is information indicating whether to permit reference to the three-dimensional points within the same LoD. For example, a value of 1 indicates that reference is possible, and a value of 0 indicates that reference is not possible. Alternatively, it can be that when the value is 1, reference to the three-dimensional points within the same LoD that have already been encoded or decoded is possible.

[0746] Fig.74This is a diagram showing an example of the reference relationship when EnableReferringSameLoD = 0. The predicted value of the point P included in LoD N is generated using the reconstructed value P' included in a LoD N' (N' < N) that is one level above LoD N. Here, the reconstructed value P' is an attribute value (attribute information) that has been encoded and decoded. For example, the reconstructed value P' of adjacent points based on distance is used.

[0747] In addition, in Fig.74 In the example shown, for example, the predicted value of b2 is generated using any one of the attribute values of a0, a1, and a2. Even when b0 and b1 have been encoded and decoded, reference to b0 and b1 is prohibited.

[0748] Thus, the three-dimensional point data encoding device and the three-dimensional data decoding device can generate the predicted value of b2 without waiting for the encoding or decoding process of b0 and b1 to complete. That is, the three-dimensional point data encoding device and the three-dimensional data decoding device can calculate multiple predicted values of the attribute values of multiple three-dimensional points within the same LoD in parallel, and thus can reduce the processing time.

[0749] Fig.75 This is a diagram showing an example of the reference relationship when EnableReferringSameLoD = 1. The predicted value of the point P included in LoD N is generated using the reconstructed value P' included in a LoD N' (N' ≤ N) that is at the same level or one level above LoD N. Here, the reconstructed value P' is an attribute value (attribute information) that has been encoded and decoded. For example, the reconstructed value P' of adjacent points based on distance is used.

[0750] In addition, in Fig.75 In the example shown, for example, the predicted value of b2 is generated using any one of the attribute values of a0, a1, a2, b0, and b1. That is, it can be referred to when b0 and b1 have already been encoded and decoded.

[0751] Thus, the three-dimensional data encoding device can generate the predicted value of b2 using the attribute information of a large number of adjacent three-dimensional points. Therefore, the prediction accuracy is improved and the encoding efficiency is improved.

[0752] Hereinafter, a method for restricting the number of search times when selecting N three-dimensional points for generating the predicted value of the attribute information of the three-dimensional points will be described. Thereby, the processing amount can be reduced.

[0753] For example, define SearchNumPoint (search point number information). SearchNumPoint represents the number of searches when selecting N three-dimensional points for prediction from a three-dimensional point group within the LoD. For example, the three-dimensional data encoding device may also select the same number of three-dimensional points as represented by SearchNumPoint from among a total of T three-dimensional points included in the LoD, and select N three-dimensional points for prediction from the selected three-dimensional points. Thereby, the three-dimensional data encoding device does not need to search all T three-dimensional points included in the LoD, and thus can reduce the processing amount.

[0754] In addition, the three-dimensional data encoding device may also switch the selection method of the value of SearchNumPoint according to the position of the LoD being referred to. The following shows examples.

[0755] For example, when the reference LoD of the three-dimensional data encoding device is a higher-level LoD than the LoD to which the target three-dimensional point belongs, the three-dimensional data encoding device searches for the three-dimensional point A in the reference LoD that is closest in distance to the target three-dimensional point. Next, the three-dimensional data encoding device selects the number of three-dimensional points represented by SearchNumPoint that are adjacent before and after the three-dimensional point A. Thereby, the three-dimensional data encoding device can efficiently search for three-dimensional points in the upper layer that are close in distance to the target three-dimensional point, and thus can improve the prediction efficiency.

[0756] For example, when the reference LoD and the LoD to which the target three-dimensional point belongs are in the same layer, the three-dimensional data encoding device selects the number of three-dimensional points represented by SearchNumPoint that were encoded and decoded earlier than the target three-dimensional point. For example, the three-dimensional data encoding device selects the number of three-dimensional points represented by SearchNumPoint that were encoded and decoded immediately before the target three-dimensional point.

[0757] Thereby, the three-dimensional data encoding device can select the number of three-dimensional points represented by SearchNumPoint with a low processing amount. In addition, the three-dimensional data encoding device may also select a three-dimensional point B that is close in distance to the target three-dimensional point from among the three-dimensional points that were encoded and decoded earlier than the target three-dimensional point, and select the number of three-dimensional points represented by SearchNumPoint that are adjacent before and after the three-dimensional point B. Thereby, the three-dimensional data encoding device can efficiently search for three-dimensional points in the same layer that are close in distance to the target three-dimensional point, and thus can improve the prediction efficiency. ...

Claims

1. A coding method, which is executed by a three-dimensional point coding device for generating a predicted value by selecting attribute information of a three-dimensional point for predicting a prediction object, wherein, evaluate at least one of a plurality of three-dimensional point candidates, based on the evaluation result, determine whether to include the evaluated three-dimensional point candidate in a set of three-dimensional points for generating a predicted value composed of N three-dimensional points, select at least one three-dimensional point from the set of three-dimensional points for generating a predicted value by using the Morton code assigned to the three-dimensional point candidate.

2. The coding method according to claim 1, wherein, in the determination, when the evaluation value of the evaluated three-dimensional point candidate is equal to or less than the minimum evaluation value, the evaluated three-dimensional point candidate is not included in the set of three-dimensional points for generating a predicted value, and the minimum evaluation value is the minimum evaluation value among the evaluation values of the three-dimensional points included in the set of three-dimensional points for generating a predicted value.

3. The coding method according to claim 1 or 2, wherein, in the determination, when the evaluation value of the evaluated three-dimensional point candidate is greater than the minimum evaluation value, the evaluated three-dimensional point candidate is included in the set of three-dimensional points for generating a predicted value, and the minimum evaluation value is the minimum evaluation value among the evaluation values of the three-dimensional points included in the set of three-dimensional points for generating a predicted value.

4. The coding method according to claim 3, wherein, when the evaluated three-dimensional point candidate is included in the set of three-dimensional points for generating a predicted value, correspondingly, the three-dimensional point with the minimum evaluation value that has already been included in the set of three-dimensional points for generating a predicted value is removed from the set of three-dimensional points for generating a predicted value.

5. The coding method according to claim 1, wherein, calculate the evaluation value of the evaluated three-dimensional point candidate based on the distance between the three-dimensional point of the prediction object and the three-dimensional point candidate of the evaluation object.

6. The coding method according to claim 1, wherein, the plurality of three-dimensional point candidates are three-dimensional points belonging to a hierarchy higher than the hierarchy to which the three-dimensional point of the prediction object belongs.

7. A decoding method, which is executed by a three-dimensional point decoding device for generating a predicted value by selecting attribute information of a three-dimensional point for predicting a prediction object, wherein, evaluate at least one of a plurality of three-dimensional point candidates, based on the evaluation result, determine whether to include the evaluated three-dimensional point candidate in a set of three-dimensional points for generating a predicted value composed of N three-dimensional points, select at least one three-dimensional point from the set of three-dimensional points for generating a predicted value by using the Morton code assigned to the three-dimensional point candidate.

8. The decoding method according to claim 7, wherein, in the determination, when the evaluation value of the evaluated three-dimensional point candidate is equal to or less than the minimum evaluation value, the evaluated three-dimensional point candidate is not included in the set of three-dimensional points for generating a predicted value, and the minimum evaluation value is the minimum evaluation value among the evaluation values of the three-dimensional points included in the set of three-dimensional points for generating a predicted value.

9. The decoding method according to claim 7 or 8, wherein, In the decision, when the evaluation value of the three-dimensional point candidate to be evaluated is greater than the minimum evaluation value, the three-dimensional point candidate to be evaluated is included in the set of three-dimensional points for generating a predicted value, where the minimum evaluation value is the minimum evaluation value among the evaluation values of the three-dimensional points included in the set of three-dimensional points for generating a predicted value.

10. The decoding method according to claim 9, wherein when the three-dimensional point candidate to be evaluated is included in the set of three-dimensional points for generating a predicted value, correspondingly, the three-dimensional point having the minimum evaluation value that has already been included in the set of three-dimensional points for generating a predicted value is removed from the set of three-dimensional points for generating a predicted value.

11. The decoding method according to claim 7, wherein the evaluation value of the three-dimensional point candidate to be evaluated is calculated based on the distance between the three-dimensional point of the prediction object and the three-dimensional point candidate of the evaluation object.

12. The decoding method according to claim 7, wherein the plurality of three-dimensional point candidates are three-dimensional points belonging to a hierarchy higher than the hierarchy to which the three-dimensional point of the prediction object belongs.

13. An encoding device that selects a three-dimensional point for generating a predicted value of attribute information of a three-dimensional point to be predicted, wherein, Comprising: a processor; and a memory, wherein the processor uses the memory to evaluate at least one of a plurality of three-dimensional point candidates, and based on the evaluation result, determine whether to include the three-dimensional point candidate to be evaluated in the set of three-dimensional points for generating a predicted value composed of N three-dimensional points, and use the Morton code assigned to the three-dimensional point candidate to select at least one three-dimensional point from the set of three-dimensional points for generating a predicted value.

14. A computer program product for causing a computer to execute an encoding method, the encoding method selecting three-dimensional points for generating a predicted value for predicting attribute information of a three-dimensional point of a prediction object, wherein for causing the computer to execute: evaluating at least one of a plurality of three-dimensional point candidates, and based on the evaluation result, determine whether to include the three-dimensional point candidate to be evaluated in the set of three-dimensional points for generating a predicted value composed of N three-dimensional points, and use the Morton code assigned to the three-dimensional point candidate to select at least one three-dimensional point from the set of three-dimensional points for generating a predicted value.

15. A decoding device that selects a three-dimensional point for generating a predicted value of attribute information of a three-dimensional point for predicting a prediction target, wherein, Comprising: a processor; and a memory, wherein the processor uses the memory to evaluate at least one of a plurality of three-dimensional point candidates, and based on the evaluation result, determine whether to include the three-dimensional point candidate to be evaluated in the set of three-dimensional points for generating a predicted value composed of N three-dimensional points, and use the Morton code assigned to the three-dimensional point candidate to select at least one three-dimensional point from the set of three-dimensional points for generating a predicted value.

16. A computer program product for causing a computer to execute a decoding method, the decoding method selecting three-dimensional points for generating a predicted value for predicting attribute information of a three-dimensional point of a prediction object, wherein for causing the computer to execute: evaluating at least one of a plurality of three-dimensional point candidates, and based on the evaluation result, determine whether to include the three-dimensional point candidate to be evaluated in the set of three-dimensional points for generating a predicted value composed of N three-dimensional points, Select at least one three-dimensional point from the set of three-dimensional points for generating the predicted value using the Morton code assigned to the three-dimensional point candidate.

Citation Information

Patent Citations

  • Map display device

    WO2014020663A1