Encoding method, decoding method, encoding device, decoding device, and program
The encoding method improves efficiency by selecting three-dimensional points for predicted values and evaluating candidates using a Morton code, addressing the inefficiencies in current three-dimensional data encoding techniques.
Patent Information
- Application Number
- JP2025040561
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2019-02-28
- Filing Date
- 2025-03-13
- Publication Date
- 2025-06-05
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Current methods for encoding three-dimensional data are inefficient, leading to high data storage and transmission requirements for point clouds, which are expected to become a mainstream method for expressing three-dimensional data.
An encoding method that selects three-dimensional points for generating predicted values, evaluates candidates using a Morton code, and determines whether to include them in a set of points for predicting attribute information, thereby improving encoding efficiency.
The proposed method enhances encoding efficiency by effectively selecting and predicting attribute information for three-dimensional points, reducing the amount of data required for storage and transmission.
Smart Images

Figure 2025085721000001_ABST
Abstract
Description
[Technical field]
[0001] The present disclosure relates to an encoding method, a decoding method, an encoding device, a decoding device, and a program. [Background technology]
[0002] In the future, devices and services that utilize 3D data are expected to become widespread in a wide range of fields, such as computer vision for autonomous operation of automobiles or robots, map information, monitoring, infrastructure inspection, video distribution, etc. 3D data is acquired in various ways, such as distance sensors such as range finders, stereo cameras, or a combination of multiple monocular cameras.
[0003] One method of expressing three-dimensional data is a method called a point cloud, which represents the shape of a three-dimensional structure using a group of points in a three-dimensional space. In a point cloud, the positions and colors of the points are stored. Point clouds are expected to become the mainstream method of expressing three-dimensional data, but point clouds have a very large amount of data. Therefore, when storing or transmitting three-dimensional data, it is essential to compress the amount of data by encoding, just like two-dimensional video images (examples include MPEG-4 AVC or HEVC standardized by MPEG).
[0004] In addition, compression of point clouds is partially supported by public libraries (Point Cloud Library) that perform point cloud-related processing.
[0005] Furthermore, a technique is known that uses three-dimensional map data to search for and display facilities located around a vehicle (see, for example, Patent Document 1). [Prior art documents] [Patent documents]
[0006] [Patent Document 1] International Publication No. 2014 / 020663 Summary of the Invention [Problem to be solved by the invention]
[0007] It is desirable to be able to improve the coding efficiency in coding three-dimensional data.
[0008] An object of the present disclosure is to provide an encoding method, a decoding method, an encoding device, a decoding device, and a program that can improve encoding efficiency. [Means for solving the problem]
[0009] An encoding method according to one embodiment of the present disclosure is an encoding method executed by an encoding device that selects a three-dimensional point for generating a predicted value to be used for predicting attribute information of a three-dimensional point to be predicted, evaluates at least one of a plurality of three-dimensional point candidates, and based on a result of the evaluation, determines whether or not to include the evaluated three-dimensional point candidate in a set of three-dimensional points for generating a predicted value consisting of N three-dimensional points, and at least one three-dimensional point from the set of three-dimensional points for generating a predicted value is selected using a Morton code assigned to the three-dimensional point candidate.
[0010] A decoding method according to one embodiment of the present disclosure is a decoding method executed by a decoding device that selects a three-dimensional point for generating a predicted value to be used for predicting attribute information of a three-dimensional point to be predicted, and evaluates at least one of a plurality of three-dimensional point candidates, and based on a result of the evaluation, determines whether or not to include the evaluated three-dimensional point candidate in a set of three-dimensional points for generating a predicted value consisting of N three-dimensional points, and at least one three-dimensional point from the set of three-dimensional points for generating a predicted value is selected using a Morton code assigned to the three-dimensional point candidate.
[0011] An encoding method according to one embodiment of the present disclosure is an encoding method that encodes three-dimensional points using attribute information of N three-dimensional points for generating predicted values, and evaluates each of a plurality of referenceable three-dimensional points in order of priority based on a Morton code. If the evaluation value of the three-dimensional point to be evaluated is the same as the smallest evaluation value among the points included in the N three-dimensional points for generating predicted values that are set as initial values, the three-dimensional point evaluated first in the order of priority is set as the N three-dimensional point for generating predicted values.
[0012] A three-dimensional data encoding method according to one embodiment of the present disclosure is a three-dimensional data encoding method for encoding a plurality of three-dimensional points, comprising the steps of: selecting, from among a plurality of second three-dimensional points surrounding a first three-dimensional point, N second three-dimensional points in order of distance to the first three-dimensional point as candidates for calculating a predicted value of attribute information of the first three-dimensional point; calculating a predicted value using attribute information of the N second three-dimensional points selected as candidates; calculating a prediction residual, which is the difference between the attribute information of the first three-dimensional point and the calculated predicted value; and generating a bitstream including the prediction residual; and in selecting the candidates, if there are a plurality of third three-dimensional points among the plurality of second three-dimensional points that are equidistant to the first three-dimensional point, selecting the candidate from the plurality of third three-dimensional points in order of priority based on a first Morton code of the first three-dimensional point.
[0013] Furthermore, a three-dimensional data decoding method according to one embodiment of the present disclosure is a three-dimensional data decoding method for decoding a plurality of three-dimensional points, comprising the steps of: acquiring a bit stream to acquire a prediction residual of a first three-dimensional point among the plurality of three-dimensional points; selecting, from among a plurality of second three-dimensional points surrounding the first three-dimensional point among the plurality of three-dimensional points, N second three-dimensional points in order of proximity to the first three-dimensional point as candidates for calculating a predicted value of attribute information of the first three-dimensional point; calculating a predicted value using attribute information of the N second three-dimensional points selected as the candidates; and calculating the attribute information of the first three-dimensional point by adding together the predicted value and the prediction residual; and in selecting the candidates, if there are a plurality of third three-dimensional points among the plurality of second three-dimensional points that are equidistant to the first three-dimensional point, selecting the candidate from the plurality of third three-dimensional points in a priority order based on a first Morton code of the first three-dimensional point.
[0014] Furthermore, these general or specific aspects may be realized by a system, an apparatus, an integrated circuit, a computer program, or a recording medium such as a computer-readable CD-ROM, or may be realized by any combination of a system, an apparatus, an integrated circuit, a computer program, and a recording medium. Effect of the Invention
[0015] The present disclosure can provide an encoding method that can improve encoding efficiency. [Brief description of the drawings]
[0016] [Figure 1] FIG. 1 is a diagram showing a structure of encoded three-dimensional data according to the first embodiment. [Diagram 2] FIG. 2 is a diagram showing an example of a prediction structure between SPCs belonging to the lowest layer of a GOS according to the first embodiment. [Diagram 3] FIG. 3 is a diagram showing an example of an inter-layer prediction structure according to the first embodiment. [Figure 4] FIG. 4 is a diagram showing an example of the coding order of GOS according to the first embodiment. [Diagram 5]FIG. 5 is a diagram showing an example of the coding order of GOS according to the first embodiment. [Figure 6] FIG. 6 is a block diagram of the three-dimensional data encoding device according to the first embodiment. [Figure 7] FIG. 7 is a flowchart of the encoding process according to the first embodiment. [Figure 8] FIG. 8 is a block diagram of the three-dimensional data decoding device according to the first embodiment. [Figure 9] FIG. 9 is a flowchart of the decoding process according to the first embodiment. [Figure 10] FIG. 10 is a diagram illustrating an example of meta information according to the first embodiment. [Figure 11] FIG. 11 is a diagram illustrating an example of a configuration of an SWLD according to the second embodiment. [Figure 12] FIG. 12 is a diagram illustrating an example of the operation of the server and the client according to the second embodiment. [Figure 13] FIG. 13 is a diagram illustrating an example of the operation of the server and the client according to the second embodiment. [Figure 14] FIG. 14 is a diagram illustrating an example of the operation of the server and the client according to the second embodiment. [Figure 15] FIG. 15 is a diagram illustrating an example of the operation of the server and the client according to the second embodiment. [Figure 16] FIG. 16 is a block diagram of a three-dimensional data encoding device according to the second embodiment. [Figure 17] FIG. 17 is a flowchart of the encoding process according to the second embodiment. [Figure 18] FIG. 18 is a block diagram of a three-dimensional data decoding device according to the second embodiment. [Figure 19] FIG. 19 is a flowchart of the decoding process according to the second embodiment. [Figure 20] FIG. 20 is a diagram illustrating an example of a configuration of a WLD according to the second embodiment. [Figure 21] FIG. 21 is a diagram illustrating an example of an octree structure of a WLD according to the second embodiment. [Figure 22]FIG. 22 is a diagram illustrating an example of a configuration of an SWLD according to the second embodiment. [Diagram 23] FIG. 23 is a diagram illustrating an example of an octree structure of the SWLD according to the second embodiment. [Figure 24] FIG. 24 is a block diagram of a three-dimensional data creation device according to the third embodiment. [Diagram 25] FIG. 25 is a block diagram of a three-dimensional data transmission device according to the third embodiment. [Figure 26] FIG. 26 is a block diagram of a three-dimensional information processing device according to the fourth embodiment. [Figure 27] FIG. 27 is a block diagram of a three-dimensional data creation device according to the fifth embodiment. [Figure 28] FIG. 28 is a diagram showing a configuration of a system according to the sixth embodiment. [Figure 29] FIG. 29 is a block diagram of a client device according to the sixth embodiment. [Diagram 30] FIG. 30 is a block diagram of a server according to the sixth embodiment. [Diagram 31] FIG. 31 is a flowchart of three-dimensional data creation processing by a client device according to the sixth embodiment. [Diagram 32] FIG. 32 is a flowchart of a sensor information transmission process by a client device according to the sixth embodiment. [Diagram 33] FIG. 33 is a flowchart of three-dimensional data creation processing by the server according to the sixth embodiment. [Diagram 34] FIG. 34 is a flowchart of a 3D map transmission process performed by the server according to the sixth embodiment. [Diagram 35] FIG. 35 is a diagram showing a configuration of a modified example of the system according to the sixth embodiment. In FIG. [Diagram 36] FIG. 36 is a diagram showing configurations of a server and a client device according to the sixth embodiment. In FIG. [Figure 37] FIG. 37 is a block diagram of a three-dimensional data encoding device according to the seventh embodiment. [Figure 38]FIG. 38 is a diagram showing an example of a prediction residual according to the seventh embodiment. In FIG. [Figure 39] FIG. 39 is a diagram illustrating an example of a volume according to the seventh embodiment. [Diagram 40] FIG. 40 is a diagram showing an example of an octree representation of a volume according to the seventh embodiment. [Diagram 41] FIG. 41 is a diagram showing an example of a bit string of a volume according to the seventh embodiment. [Diagram 42] FIG. 42 is a diagram showing an example of an octree representation of a volume according to the seventh embodiment. [Diagram 43] FIG. 43 is a diagram illustrating an example of a volume according to the seventh embodiment. [Diagram 44] FIG. 44 is a diagram for explaining the intra prediction process according to the seventh embodiment. [Diagram 45] FIG. 45 is a diagram for explaining the rotation and translation processing according to the seventh embodiment. In FIG. [Figure 46] FIG. 46 is a diagram showing an example of the syntax of the RT application flag and the RT information according to the seventh embodiment. [Figure 47] FIG. 47 is a diagram for explaining the inter prediction process according to the seventh embodiment. [Figure 48] FIG. 48 is a block diagram of a three-dimensional data decoding device according to the seventh embodiment. [Figure 49] FIG. 49 is a flowchart of three-dimensional data encoding processing by the three-dimensional data encoding device according to the seventh embodiment. [Figure 50] FIG. 50 is a flowchart of three-dimensional data decoding processing by the three-dimensional data decoding device according to the seventh embodiment. [Figure 51] FIG. 51 is a diagram showing an example of three-dimensional points according to the eighth embodiment. [Figure 52] FIG. 52 is a diagram showing an example of setting the LoD according to the eighth embodiment. In FIG. [Diagram 53] FIG. 53 is a diagram showing an example of threshold values used for setting the LoD according to the eighth embodiment. [Figure 54]FIG. 54 is a diagram showing an example of attribute information used for a predicted value according to the eighth embodiment. In FIG. [Figure 55] FIG. 55 is a diagram illustrating an example of an exponential-Golomb code according to the eighth embodiment. [Figure 56] FIG. 56 is a diagram illustrating a process for an exponential-Golomb code according to the eighth embodiment. [Figure 57] FIG. 57 is a diagram illustrating an example of the syntax of the attribute header according to the eighth embodiment. [Figure 58] FIG. 58 is a diagram illustrating an example of the syntax of attribute data according to the eighth embodiment. [Figure 59] FIG. 59 is a flowchart of three-dimensional data encoding processing according to the eighth embodiment. [Figure 60] FIG. 60 is a flowchart of the attribute information encoding process according to the eighth embodiment. [Figure 61] FIG. 61 is a diagram illustrating a process for an exponential-Golomb code according to the eighth embodiment. [Figure 62] FIG. 62 is a diagram showing an example of a reverse lookup table showing the relationship between the remaining codes and their values according to the eighth embodiment. [Figure 63] FIG. 63 is a flowchart of three-dimensional data decoding processing according to the eighth embodiment. [Figure 64] FIG. 64 is a flowchart of the attribute information decoding process according to the eighth embodiment. [Figure 65] FIG. 65 is a block diagram of a three-dimensional data encoding device according to the eighth embodiment. [Figure 66] FIG. 66 is a block diagram of a three-dimensional data decoding device according to the eighth embodiment. [Figure 67] FIG. 67 is a flowchart of three-dimensional data encoding processing according to the eighth embodiment. [Figure 68] FIG. 68 is a flowchart of three-dimensional data decoding processing according to the eighth embodiment. [Figure 69] FIG. 69 is a diagram showing a first example of a table indicating predicted values calculated in each prediction mode according to the ninth embodiment. [Figure 70] FIG. 70 is a diagram showing an example of attribute information used for a predicted value according to the ninth embodiment. In FIG. [Figure 71] FIG. 71 is a diagram showing a second example of a table indicating predicted values calculated in each prediction mode according to the ninth embodiment. [Figure 72] FIG. 72 is a diagram showing a third example of a table indicating predicted values calculated in each prediction mode according to the ninth embodiment. [Figure 73] FIG. 73 is a diagram showing a fourth example of a table indicating predicted values calculated in each prediction mode according to the ninth embodiment. [Figure 74] FIG. 74 is a diagram showing an example of a reference relationship according to the tenth embodiment. In FIG. [Figure 75] FIG. 75 is a diagram showing an example of a reference relationship according to the tenth embodiment. In FIG. [Figure 76] FIG. 76 is a diagram showing an example of setting the number of searches for each LoD according to the tenth embodiment. In FIG. [Figure 77] FIG. 77 is a diagram showing an example of a reference relationship according to the tenth embodiment. In FIG. [Figure 78] FIG. 78 is a diagram showing an example of a reference relationship according to the tenth embodiment. In FIG. [Figure 79] FIG. 79 is a diagram showing an example of a reference relationship according to the tenth embodiment. In FIG. [Figure 80] FIG. 80 is a diagram illustrating an example of the syntax of an attribute information header according to the tenth embodiment. [Figure 81] FIG. 81 is a diagram illustrating an example of the syntax of an attribute information header according to the tenth embodiment. [Figure 82] FIG. 82 is a flowchart of three-dimensional data encoding processing according to the tenth embodiment. [Figure 83] FIG. 83 is a flowchart of the attribute information encoding process according to the tenth embodiment. [Figure 84] FIG. 84 is a flowchart of three-dimensional data decoding processing according to the tenth embodiment. [Figure 85] FIG. 85 is a flowchart of the attribute information decoding process according to the tenth embodiment. [Figure 86] FIG. 86 is a flowchart of the surrounding point search process according to the tenth embodiment. [Figure 87] FIG. 87 is a flowchart of the surrounding point search process according to the tenth embodiment. [Figure 88] FIG. 88 is a flowchart of the surrounding point search process according to the tenth embodiment. [Figure 89] FIG. 89 is a flowchart of three-dimensional data encoding processing according to the tenth embodiment. [Figure 90] FIG. 90 is a flowchart of three-dimensional data decoding processing according to the tenth embodiment. [Figure 91] FIG. 91 is a diagram for explaining a method of selecting N three-dimensional points according to the eleventh embodiment. In FIG. [Figure 92] FIG. 92 is a diagram showing an example of a bounding box of a group Gk according to the eleventh embodiment. [Figure 93] FIG. 93 is a diagram for explaining a process of selecting N candidates for three-dimensional points when a first three-dimensional point and a plurality of second three-dimensional points belong to the same group according to the eleventh embodiment. [Figure 94] FIG. 94 is a diagram for explaining a process of selecting N candidates for three-dimensional points when a first three-dimensional point and a plurality of second three-dimensional points belong to different groups according to the eleventh embodiment. [Figure 95] FIG. 95 is a diagram for explaining the process of selecting a 3D point candidate from a plurality of second 3D points belonging to different layers according to the eleventh embodiment. [Figure 96] FIG. 96 is a diagram for explaining the process of selecting or updating 3D point candidates from groups before and after the initial group in embodiment 11. [Figure 97] FIG. 97 is a diagram for explaining an example in which a group having a smaller bounding box is given priority according to the eleventh embodiment. [Figure 98] FIG. 98 is a flowchart of three-dimensional data encoding processing by the three-dimensional data encoding device according to the eleventh embodiment. [Figure 99] FIG. 99 is a flowchart of the attribute information encoding process according to the eleventh embodiment. [Figure 100] FIG. 100 is a flowchart of three-dimensional data decoding processing by the three-dimensional data decoding device according to the eleventh embodiment. [Figure 101] FIG. 101 is a flowchart of the attribute information decoding process according to the eleventh embodiment. [Figure 102] FIG. 102 is a flowchart of a surrounding point search process according to the eleventh embodiment. [Figure 103] FIG. 103 is a flowchart of a surrounding point search process according to the eleventh embodiment. [Figure 104] FIG. 104 is a block diagram showing a configuration of an attribute information encoding unit included in the three-dimensional data encoding device according to the eleventh embodiment. [Figure 105] FIG. 105 is a block diagram showing a configuration of an attribute information decoding unit included in the three-dimensional data decoding device according to the eleventh embodiment. As shown in FIG. [Fig. 106] FIG. 106 is a flowchart of three-dimensional data encoding processing according to the eleventh embodiment. [Figure 107] FIG. 107 is a flowchart of three-dimensional data decoding processing according to the eleventh embodiment. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS
[0017] A three-dimensional data encoding method according to one embodiment of the present disclosure is a three-dimensional data encoding method for encoding a plurality of three-dimensional points, comprising the steps of: selecting, from among a plurality of second three-dimensional points surrounding a first three-dimensional point, N second three-dimensional points in order of distance to the first three-dimensional point as candidates for calculating a predicted value of attribute information of the first three-dimensional point; calculating a predicted value using attribute information of the N second three-dimensional points selected as candidates; calculating a prediction residual, which is the difference between the attribute information of the first three-dimensional point and the calculated predicted value; and generating a bitstream including the prediction residual; and in selecting the candidates, if there are a plurality of third three-dimensional points among the plurality of second three-dimensional points that are equidistant to the first three-dimensional point, selecting the candidate from the plurality of third three-dimensional points in order of priority based on a first Morton code of the first three-dimensional point.
[0018] This makes it possible to select the second 3D point close to the first 3D point to be coded as a candidate for use in calculating a predicted value, thereby improving coding efficiency.
[0019] For example, the priority order may be an order determined by the Morton codes of the third three-dimensional points, and may be an order of third three-dimensional points having Morton codes closer to the first Morton code.
[0020] For example, the priority order may be the order of third three-dimensional points closest to the second Morton code when the candidates are selected alternately one by one from a first group including a plurality of third three-dimensional points having Morton codes smaller than the second Morton code of a fourth three-dimensional point having a Morton code close to the first Morton code among the plurality of third three-dimensional points, and a second group including a plurality of third three-dimensional points having a Morton code larger than the second Morton code.
[0021] For example, the first 3D point and the plurality of third 3D points may belong to different layers.
[0022] For example, the priority order may be an order determined by the Morton codes of a plurality of third three-dimensional points belonging to either a first group including a plurality of third three-dimensional points having Morton codes smaller than the first Morton code, or a second group including a plurality of third three-dimensional points having Morton codes larger than the first Morton code, and may be an order of third three-dimensional points closest to the first Morton code.
[0023] For example, the first 3D point and the plurality of third 3D points may belong to the same layer.
[0024] A three-dimensional data decoding method according to one embodiment of the present disclosure is a three-dimensional data decoding method for decoding a plurality of three-dimensional points, comprising the steps of: acquiring a bit stream to acquire a prediction residual of a first three-dimensional point among the plurality of three-dimensional points; selecting, from among a plurality of second three-dimensional points surrounding the first three-dimensional point among the plurality of three-dimensional points, N second three-dimensional points in order of proximity to the first three-dimensional point as candidates for calculating a predicted value of attribute information of the first three-dimensional point; calculating a predicted value using attribute information of the N second three-dimensional points selected as candidates; and calculating the attribute information of the first three-dimensional point by adding together the predicted value and the prediction residual; and in selecting the candidates, if there are a plurality of third three-dimensional points among the plurality of second three-dimensional points that are equidistant to the first three-dimensional point, selecting the candidate from the plurality of third three-dimensional points in a priority order based on a first Morton code of the first three-dimensional point.
[0025] This makes it possible to appropriately decode the attribute information of the first 3D point to be processed.
[0026] For example, the priority order may be an order determined by the Morton codes of the third three-dimensional points, and may be an order of third three-dimensional points having Morton codes closer to the first Morton code.
[0027] For example, the priority order may be the order of third three-dimensional points closest to the second Morton code when the candidates are selected alternately one by one from a first group including a plurality of third three-dimensional points having Morton codes smaller than the second Morton code of a fourth three-dimensional point having a Morton code close to the first Morton code among the plurality of third three-dimensional points, and a second group including a plurality of third three-dimensional points having a Morton code larger than the second Morton code.
[0028] For example, the first 3D point and the plurality of third 3D points may belong to different layers.
[0029] For example, the priority order may be an order determined by the Morton codes of a plurality of third three-dimensional points belonging to either a first group including a plurality of third three-dimensional points having Morton codes smaller than the first Morton code, or a second group including a plurality of third three-dimensional points having Morton codes larger than the first Morton code, and may be an order of third three-dimensional points closest to the first Morton code.
[0030] For example, the first 3D point and the plurality of third 3D points may belong to the same layer.
[0031] A three-dimensional data encoding device according to one embodiment of the present disclosure is a three-dimensional data encoding device that encodes a plurality of three-dimensional points, and includes a processor and a memory. The processor uses the memory to select, from among a plurality of second three-dimensional points surrounding a first three-dimensional point, N second three-dimensional points in order of closest distance to the first three-dimensional point as candidates for calculating a predicted value of attribute information of the first three-dimensional point, calculate a predicted value using attribute information of the N second three-dimensional points selected as candidates, calculate a prediction residual which is the difference between the attribute information of the first three-dimensional point and the calculated predicted value, and generate a bit stream including the prediction residual. In selecting the candidates, if there are a plurality of third three-dimensional points among the plurality of second three-dimensional points which are equidistant to the first three-dimensional point, the processor selects the candidate from the plurality of third three-dimensional points in a priority order based on a first Morton code of the first three-dimensional point.
[0032] This makes it possible to select the second 3D point close to the first 3D point to be coded as a candidate for use in calculating a predicted value, thereby improving coding efficiency.
[0033] A three-dimensional data decoding device according to one embodiment of the present disclosure is a three-dimensional data decoding device that decodes a plurality of three-dimensional points, and includes a processor and a memory. The three-dimensional data decoding device acquires a prediction residual of a first three-dimensional point among the plurality of three-dimensional points by acquiring a bit stream, selects N second three-dimensional points from among a plurality of second three-dimensional points surrounding the first three-dimensional point among the plurality of three-dimensional points in order of distance to the first three-dimensional point as candidates for calculating a predicted value of attribute information of the first three-dimensional point, calculates a predicted value using attribute information of the N second three-dimensional points selected as candidates, and calculates the attribute information of the first three-dimensional point by adding the predicted value and the prediction residual. In selecting the candidates, if there are a plurality of third three-dimensional points among the plurality of second three-dimensional points that are equidistant to the first three-dimensional point, the candidate is selected from the plurality of third three-dimensional points in a priority order based on a first Morton code of the first three-dimensional point.
[0034] This makes it possible to appropriately decode the attribute information of the first 3D point to be processed.
[0035] These comprehensive or specific aspects may be realized as a system, a method, an integrated circuit, a computer program, or a recording medium such as a computer-readable CD-ROM, or may be realized as any combination of a system, a method, an integrated circuit, a computer program, and a recording medium.
[0036] Hereinafter, the embodiments will be described in detail with reference to the drawings. Note that each of the embodiments described below shows a specific example of the present disclosure. The numerical values, shapes, materials, components, arrangement and connection forms of the components, steps, and order of steps shown in the following embodiments are merely examples and are not intended to limit the present disclosure. In addition, among the components in the following embodiments, components that are not described in an independent claim showing a top concept are described as optional components.
[0037] (Embodiment 1) First, the data structure of encoded three-dimensional data (hereinafter, also referred to as encoded data) according to this embodiment will be described. Fig. 1 is a diagram showing the structure of encoded three-dimensional data according to this embodiment.
[0038] In this embodiment, the three-dimensional space is divided into spaces (SPC) corresponding to pictures in video coding, and the three-dimensional data is coded using the spaces as units. The space is further divided into volumes (VLM) corresponding to macroblocks in video coding, and prediction and conversion are performed using the VLM as units. The volume includes a plurality of voxels (VXL), which are the smallest units to which position coordinates are associated. Note that prediction refers to, as with prediction performed on two-dimensional images, generating predicted three-dimensional data similar to the processing unit to be processed by referring to other processing units, and coding the difference between the predicted three-dimensional data and the processing unit to be processed. In addition, this prediction includes not only spatial prediction that refers to other prediction units at the same time, but also temporal prediction that refers to prediction units at different times.
[0039] For example, when a three-dimensional data encoding device (hereinafter also referred to as an encoding device) encodes a three-dimensional space represented by point cloud data such as a point cloud, it encodes each point of the point cloud or a group of points contained in a voxel according to the size of the voxel. By subdividing the voxel, the three-dimensional shape of the point cloud can be expressed with high accuracy, and by increasing the size of the voxel, the three-dimensional shape of the point cloud can be expressed roughly.
[0040] In the following, an example will be described in which the three-dimensional data is a point cloud; however, the three-dimensional data is not limited to a point cloud, and may be three-dimensional data in any format.
[0041] Also, voxels with a hierarchical structure may be used. In this case, in the nth layer, it may be indicated in order whether a sample point exists in the n-1th or lower layer (a layer below the nth layer). For example, when decoding only the nth layer, if a sample point exists in the n-1th or lower layer, the sample point can be decoded by considering it to exist at the center of the voxel in the nth layer.
[0042] In addition, the encoding device acquires the point cloud data using a distance sensor, a stereo camera, a monocular camera, a gyro, an inertial sensor, or the like.
[0043] As in video coding, the space is classified into at least three prediction structures, including an independently decodable intra space (I-SPC), a unidirectionally referable predictive space (P-SPC), and a bidirectionally referable bidirectional space (B-SPC). The space also has two types of time information: the decoding time and the display time.
[0044] As shown in Fig. 1, there is a random access unit called a Group Of Space (GOS) which is a processing unit including multiple spaces. Furthermore, there is a World (WLD) which is a processing unit including multiple GOS.
[0045] The spatial region that the world occupies is associated with an absolute position on the earth by GPS or latitude and longitude information, etc. This position information is stored as meta information. Note that the meta information may be included in the encoded data or may be transmitted separately from the encoded data.
[0046] Furthermore, within a GOS, all SPCs may be three-dimensionally adjacent, or there may be SPCs that are not three-dimensionally adjacent to other SPCs.
[0047] In the following, the process of encoding, decoding, referencing, etc. of three-dimensional data included in a processing unit such as a GOS, SPC, or VLM will also be simply referred to as encoding, decoding, referencing, etc. of a processing unit. Also, the three-dimensional data included in a processing unit includes at least one pair of a spatial position such as three-dimensional coordinates and a characteristic value such as color information.
[0048] Next, the prediction structure of SPCs in a GOS will be explained. Multiple SPCs in the same GOS, or multiple VLMs in the same SPC, occupy different spaces from each other, but have the same time information (decoding time and display time).
[0049] In addition, the first SPC in a GOS in decoding order is the I-SPC. There are two types of GOS: closed GOS and open GOS. A closed GOS is a GOS in which all SPCs in the GOS can be decoded when decoding starts from the first I-SPC. In an open GOS, some SPCs in the GOS whose display time precedes the first I-SPC refer to a different GOS, and cannot be decoded using only that GOS.
[0050] In addition, in coded data such as map information, the WLD may be decoded in the reverse order from the coding order, and if there is a dependency between the GOS, reverse playback is difficult. Therefore, in such cases, closed GOS is basically used.
[0051] Furthermore, the GOS has a layer structure in the height direction, and encoding or decoding is performed in order starting from the SPC in the lower layer.
[0052] Fig. 2 is a diagram showing an example of a prediction structure between SPCs belonging to the lowest layer of a GOS, and Fig. 3 is a diagram showing an example of a prediction structure between layers.
[0053] There are one or more I-SPCs in a GOS. Objects such as humans, animals, cars, bicycles, traffic lights, and landmark buildings exist in a three-dimensional space, and it is effective to encode small objects as I-SPCs. For example, a three-dimensional data decoding device (hereinafter also referred to as a decoding device) decodes only the I-SPCs in the GOS when decoding the GOS with low processing load or at high speed.
[0054] Furthermore, the encoding device may switch the encoding interval or occurrence frequency of I-SPC depending on the density of objects in the WLD.
[0055] In addition, in the configuration shown in Fig. 3, the encoding device or decoding device encodes or decodes multiple layers in order from the lower layer (layer 1). This allows, for example, an autonomous vehicle to increase the priority of data near the ground, which contains more information.
[0056] In addition, in the case of encoded data used by drones, etc., encoding or decoded may be performed in order from the SPC of the upper layer in the height direction within the GOS.
[0057] Alternatively, the encoding device or decoding device may encode or decode multiple layers so that the decoding device can roughly grasp the GOS and gradually increase the resolution. For example, the encoding device or decoding device may encode or decode layers 3, 8, 1, 9, etc. in that order.
[0058] Next, how to handle static and dynamic objects will be described.
[0059] In a three-dimensional space, there exist static objects or scenes such as buildings or roads (hereinafter collectively referred to as static objects), and dynamic objects such as cars or people (hereinafter referred to as dynamic objects). Object detection is performed separately by extracting feature points from point cloud data or camera images such as a stereo camera. Here, an example of a method for encoding dynamic objects will be described.
[0060] The first method is to encode the static objects without distinguishing between the static objects and the dynamic objects, and the second method is to distinguish between the static objects and the dynamic objects by using identification information.
[0061] For example, a GOS is used as an identification unit. In this case, a GOS including SPCs constituting static objects and a GOS including SPCs constituting dynamic objects are distinguished by identification information stored in the coded data or separately from the coded data.
[0062] Alternatively, the SPC may be used as the identification unit. In this case, the SPC including the VLM constituting the static object and the SPC including the VLM constituting the dynamic object are distinguished by the above-mentioned identification information.
[0063] Alternatively, the VLM or VXL may be used as the identification unit. In this case, the VLM or VXL including a static object is distinguished from the VLM or VXL including a dynamic object by the above-mentioned identification information.
[0064] The encoding device may also encode a dynamic object as one or more VLMs or SPCs, and encode a VLM or SPC including a static object and an SPC including a dynamic object as different GOSs. In addition, when the size of the GOS varies depending on the size of the dynamic object, the encoding device stores the size of the GOS separately as meta information.
[0065] The encoding device may also encode the static object and the dynamic object independently of each other, and overlay the dynamic object on the world composed of the static objects. In this case, the dynamic object is composed of one or more SPCs, and each SPC is associated with one or more SPCs constituting the static object on which the SPC is overlaid. Note that the dynamic object may be represented by one or more VLMs or VXLs instead of SPCs.
[0066] The encoding device may also encode static objects and dynamic objects as different streams.
[0067] The encoding device may also generate a GOS including one or more SPCs that constitute a dynamic object. Furthermore, the encoding device may set the GOS including the dynamic object (GOS_M) and the GOS of the static object corresponding to the spatial area of GOS_M to the same size (occupy the same spatial area). This allows the superimposition process to be performed on a GOS-by-GOS basis.
[0068] A P-SPC or B-SPC constituting a dynamic object may refer to an SPC included in a different encoded GOS. In cases where the position of a dynamic object changes over time and the same dynamic object is encoded as a GOS at a different time, referencing across GOS is effective in terms of compression ratio.
[0069] Also, the first method and the second method may be switched depending on the use of the encoded data. For example, when the encoded three-dimensional data is used as a map, it is desirable to be able to separate dynamic objects, so the encoding device uses the second method. On the other hand, when encoding three-dimensional data of an event such as a concert or sporting event, if there is no need to separate dynamic objects, the encoding device uses the first method.
[0070] Furthermore, the decode time and display time of the GOS or SPC can be stored in the encoded data or as meta information. Furthermore, all time information of static objects may be the same. In this case, the actual decode time and display time may be determined by the decoding device. Alternatively, a different value may be assigned as the decode time for each GOS or SPC, and the same value may be assigned as the display time for all. Furthermore, a model may be introduced in which the decoder has a buffer of a predetermined size, and guarantees that decoding can be performed without failure if a bitstream is read at a predetermined bit rate according to the decode time, as in a decoder model in video coding such as the HRD (Hypothetical Reference Decoder) of HEVC.
[0071] Next, the arrangement of GOS in the world will be described. The coordinates of the three-dimensional space in the world are expressed by three mutually orthogonal coordinate axes (x-axis, y-axis, z-axis). By setting a predetermined rule for the coding order of GOS, coding can be performed so that spatially adjacent GOS are continuous in the coded data. For example, in the example shown in FIG. 4, GOS in the xz plane are coded continuously. After coding of all GOS in a certain xz plane is completed, the value of the y axis is updated. In other words, as coding progresses, the world extends in the y axis direction. Also, the index numbers of GOS are set in the coding order.
[0072] Here, the three-dimensional space of the world is associated one-to-one with absolute geographical coordinates such as GPS or latitude and longitude. Alternatively, the three-dimensional space may be expressed by a relative position from a preset reference position. The directions of the x-axis, y-axis, and z-axis of the three-dimensional space are expressed as direction vectors determined based on the latitude and longitude, and the direction vectors are stored as meta information together with the encoded data.
[0073] Moreover, the size of the GOS is fixed, and the encoding device stores the size as meta information. The size of the GOS may be switched depending on, for example, whether it is an urban area or not, or whether it is indoors or outdoors. That is, the size of the GOS may be switched depending on the amount or nature of objects that have information value. Alternatively, the encoding device may adaptively switch the size of the GOS or the interval of I-SPCs in the GOS depending on the density of objects within the same world. For example, the higher the density of objects, the smaller the size of the GOS and the shorter the interval of I-SPCs in the GOS.
[0074] In the example of Figure 5, the third to tenth GOS areas have a high density of objects, so the GOS are subdivided to achieve fine-grained random access. Note that the seventh to tenth GOS are located behind the third to sixth GOS, respectively.
[0075] Next, the configuration and operation flow of the three-dimensional data encoding device according to this embodiment will be described. Fig. 6 is a block diagram of the three-dimensional data encoding device 100 according to this embodiment. Fig. 7 is a flowchart showing an example of the operation of the three-dimensional data encoding device 100.
[0076] 6 generates encoded three-dimensional data 112 by encoding three-dimensional data 111. The three-dimensional data encoding device 100 includes an acquisition unit 101, an encoding region determination unit 102, a division unit 103, and an encoding unit 104.
[0077] As shown in FIG. 7, first, the acquisition unit 101 acquires three-dimensional data 111, which is point cloud data (S101).
[0078] Next, the coding region determination unit 102 determines a region to be coded from among the spatial regions corresponding to the acquired point cloud data (S102). For example, the coding region determination unit 102 determines a spatial region around a position of a user or a vehicle as a region to be coded, depending on the position of the user or the vehicle.
[0079] Next, the division unit 103 divides the point cloud data included in the region to be coded into each processing unit. Here, the processing units are the above-mentioned GOS and SPC, etc. Furthermore, this region to be coded corresponds to, for example, the above-mentioned world. Specifically, the division unit 103 divides the point cloud data into processing units based on a preset size of the GOS, or the presence or absence or size of a dynamic object (S103). Furthermore, the division unit 103 determines the start position of the SPC that is the first in the coding order in each GOS.
[0080] Next, the encoding unit 104 generates encoded three-dimensional data 112 by sequentially encoding a plurality of SPCs in each GOS (S104).
[0081] In this example, the area to be coded is divided into GOS and SPC, and then each GOS is coded, but the processing procedure is not limited to the above. For example, a procedure may be used in which the configuration of one GOS is determined, the GOS is coded, and then the configuration of the next GOS is determined.
[0082] In this way, the three-dimensional data encoding device 100 generates encoded three-dimensional data 112 by encoding the three-dimensional data 111. Specifically, the three-dimensional data encoding device 100 divides the three-dimensional data into first processing units (GOS), which are random access units and each of which corresponds to a three-dimensional coordinate, divides the first processing units (GOS) into a plurality of second processing units (SPC), and divides the second processing units (SPC) into a plurality of third processing units (VLM). Furthermore, the third processing units (VLM) include one or more voxels (VXL), which are the smallest units to which position information corresponds.
[0083] Next, the three-dimensional data encoding device 100 generates encoded three-dimensional data 112 by encoding each of the multiple first processing units (GOS). Specifically, the three-dimensional data encoding device 100 encodes each of the multiple second processing units (SPC) in each first processing unit (GOS). Also, the three-dimensional data encoding device 100 encodes each of the multiple third processing units (VLM) in each second processing unit (SPC).
[0084] For example, when the first processing unit (GOS) to be processed is a closed GOS, the three-dimensional data encoding device 100 encodes the second processing unit (SPC) to be processed included in the first processing unit (GOS) to be processed by referring to other second processing units (SPC) included in the first processing unit (GOS) to be processed. In other words, the three-dimensional data encoding device 100 does not refer to the second processing unit (SPC) included in a first processing unit (GOS) different from the first processing unit (GOS) to be processed.
[0085] On the other hand, when the first processing unit (GOS) to be processed is an open GOS, the second processing unit (SPC) to be processed included in the first processing unit (GOS) to be processed is encoded by referring to other second processing units (SPC) included in the first processing unit (GOS) to be processed, or a second processing unit (SPC) included in a first processing unit (GOS) different from the first processing unit (GOS) to be processed.
[0086] In addition, the three-dimensional data encoding device 100 selects, as the type of the second processing unit (SPC) to be processed, one of a first type (I-SPC) that does not reference other second processing units (SPCs), a second type (P-SPC) that references one other second processing unit (SPC), and a third type that references two other second processing units (SPCs), and encodes the second processing unit (SPC) to be processed according to the selected type.
[0087] Next, the configuration and operation flow of the three-dimensional data decoding device according to this embodiment will be described. Fig. 8 is a block diagram of the blocks of the three-dimensional data decoding device 200 according to this embodiment. Fig. 9 is a flowchart showing an example of the operation of the three-dimensional data decoding device 200.
[0088] 8 generates decoded three-dimensional data 212 by decoding encoded three-dimensional data 211. Here, the encoded three-dimensional data 211 is, for example, the encoded three-dimensional data 112 generated by the three-dimensional data encoding device 100. This three-dimensional data decoding device 200 includes an acquisition unit 201, a decoding start GOS determination unit 202, a decoding SPC determination unit 203, and a decoding unit 204.
[0089] First, the acquisition unit 201 acquires the encoded 3D data 211 (S201). Next, the decoding start GOS determination unit 202 determines a GOS to be decoded (S202). Specifically, the decoding start GOS determination unit 202 refers to meta information stored in the encoded 3D data 211 or separately from the encoded 3D data, and determines a GOS including an SPC corresponding to a spatial position, object, or time at which decoding starts as a GOS to be decoded.
[0090] Next, the decoding SPC determination unit 203 determines the type (I, P, B) of the SPC to be decoded in the GOS (S203). For example, the decoding SPC determination unit 203 determines whether to (1) decode only the I-SPC, (2) decode the I-SPC and P-SPC, or (3) decode all types. Note that if the type of SPC to be decoded has been determined in advance, such as when all SPCs are to be decoded, this step does not need to be performed.
[0091] Next, the decoding unit 204 obtains the address position at which the first SPC in the GOS in decoding order (the same as the encoding order) starts in the encoded 3D data 211, obtains the encoded data of the first SPC from the address position, and sequentially decodes each SPC in order starting from the first SPC (S204). Note that the address position is stored in meta information or the like.
[0092] In this manner, the three-dimensional data decoding device 200 decodes the decoded three-dimensional data 212. Specifically, the three-dimensional data decoding device 200 generates the decoded three-dimensional data 212 of the first processing unit (GOS) by decoding each of the encoded three-dimensional data 211 of the first processing unit (GOS), which is a random access unit and corresponds to a three-dimensional coordinate. More specifically, the three-dimensional data decoding device 200 decodes each of the multiple second processing units (SPC) in each first processing unit (GOS). Also, the three-dimensional data decoding device 200 decodes each of the multiple third processing units (VLM) in each second processing unit (SPC).
[0093] The meta information for random access will be described below. This meta information is generated by the three-dimensional data encoding device 100 and is included in the encoded three-dimensional data 112 (211).
[0094] In conventional random access for two-dimensional video, decoding is started from the first frame of a random access unit that is close to a specified time. On the other hand, in the world, random access to space (coordinates, objects, etc.) is assumed in addition to time.
[0095] Therefore, in order to realize random access to at least three elements, coordinates, object, and time, a table is prepared that associates each element with a GOS index number. Furthermore, the GOS index number is associated with the address of the I-SPC at the beginning of the GOS. FIG. 10 is a diagram showing an example of a table included in the meta information. Note that it is not necessary to use all of the tables shown in FIG. 10, and it is sufficient to use at least one table.
[0096] Hereinafter, as an example, random access starting from coordinates will be described. When accessing coordinates (x2, y2, z2), first, the coordinate-GOS table is referenced and it is found that the point with coordinates (x2, y2, z2) is included in the second GOS. Next, the GOS address table is referenced and it is found that the address of the first I-SPC in the second GOS is addr(2), so the decoding unit 204 obtains data from this address and starts decoding.
[0097] The address may be an address in a logical format or a physical address of a HDD or memory. Information that identifies a file segment may be used instead of the address. For example, a file segment is a unit obtained by segmenting one or more GOS.
[0098] In addition, if an object spans multiple GOSs, the object-GOS table may indicate multiple GOSs to which the object belongs. If the multiple GOSs are closed GOSs, the encoding device and the decoding device can perform encoding or decoding in parallel. On the other hand, if the multiple GOSs are open GOSs, the multiple GOSs can refer to each other to improve compression efficiency.
[0099] Examples of objects include humans, animals, cars, bicycles, traffic lights, landmark buildings, etc. For example, the three-dimensional data encoding device 100 can extract feature points specific to objects from a three-dimensional point cloud or the like when encoding a world, detect objects based on the feature points, and set the detected objects as random access points.
[0100] In this way, the three-dimensional data encoding device 100 generates first information indicating a plurality of first processing units (GOS) and three-dimensional coordinates associated with each of the plurality of first processing units (GOS). The encoded three-dimensional data 112 (211) includes this first information. The first information further indicates at least one of an object, a time, and a data storage destination associated with each of the plurality of first processing units (GOS).
[0101] The three-dimensional data decoding device 200 acquires first information from the encoded three-dimensional data 211, and uses the first information to identify the encoded three-dimensional data 211 of a first processing unit corresponding to a specified three-dimensional coordinate, object or time, and decodes the encoded three-dimensional data 211.
[0102] Other examples of meta information will be described below. In addition to the meta information for random access, the three-dimensional data encoding device 100 may generate and store the following meta information. Furthermore, the three-dimensional data decoding device 200 may use this meta information during decoding.
[0103] When using three-dimensional data as map information, a profile may be defined according to the application, and information indicating the profile may be included in the meta information. For example, a profile for urban areas or suburban areas, or for flying objects may be defined, and the maximum or minimum size of the world, SPC, or VLM may be defined for each. For example, the minimum size of the VLM is set to be smaller for urban areas, since more detailed information is required than for suburban areas.
[0104] The meta information may include a tag value indicating the type of object. This tag value is associated with the VLM, SPC, or GOS that constitutes the object. For example, a tag value may be set for each type of object, such as a tag value of "0" indicating a "person," a tag value of "1" indicating a "car," and a tag value of "2" indicating a "traffic light." Alternatively, when it is difficult to determine the type of object or there is no need to determine it, a tag value indicating a property such as size or whether the object is a dynamic or static object may be used.
[0105] The meta information may also include information indicating the extent of the spatial region that the world occupies.
[0106] The meta information may also store the size of an SPC or VXL as header information common to a plurality of SPCs, such as an SPC in an entire stream of encoded data or a GOS.
[0107] The meta information may also include identification information for the range sensor or camera used to generate the point cloud, or information indicating the positional accuracy of the points in the point cloud.
[0108] The meta information may also include information indicating whether the world is made up of only static objects or whether it also includes dynamic objects.
[0109] A modification of this embodiment will now be described.
[0110] The encoding device or the decoding device may encode or decode two or more different SPCs or GOSs in parallel. The GOSs to be encoded or decoded in parallel can be determined based on meta-information indicating the spatial positions of the GOSs.
[0111] In cases where three-dimensional data is used as a spatial map for the movement of a vehicle or flying object, or where such a spatial map is to be generated, the encoding device or decoding device may encode or decode a GOS or SPC contained in a space identified based on a GPS, route information, zoom magnification, etc.
[0112] The decoding device may also decode spaces in order from the closest to the self-location or the driving route. The encoding device or decoding device may code or decode spaces farther from the self-location or the driving route by lowering the priority compared to closer spaces. Here, lowering the priority means lowering the processing order, lowering the resolution (thinning out the data before processing), or lowering the image quality (increasing the encoding efficiency, for example, by increasing the quantization step), etc.
[0113] Furthermore, when decoding encoded data that has been hierarchically encoded in space, the decoding device may decode only the lower hierarchical layers.
[0114] The decoding device may also give priority to decoding from the lower layers depending on the zoom factor or purpose of the map.
[0115] In addition, for applications such as self-position estimation or object recognition performed when a car or robot is driving autonomously, the encoding device or decoding device may perform encoding or decoding with reduced resolution except for areas within a specific height from the road surface (area where recognition is performed).
[0116] The encoding device may also encode point clouds that represent the spatial shapes of the interior and exterior of a room separately. For example, by separating the GOS that represent the interior of a room (indoor GOS) from the GOS that represent the exterior of a room (outdoor GOS), the decoding device can select the GOS to be decoded according to the viewpoint position when using the encoded data.
[0117] The encoding device may also encode the indoor GOS and the outdoor GOS that are close to each other in the encoded stream so that they are adjacent to each other. For example, the encoding device may associate identifiers of the two and store information indicating the associated identifiers in the encoded stream or in meta information stored separately. This allows the decoding device to identify the indoor GOS and the outdoor GOS that are close to each other by referring to the information in the meta information.
[0118] The encoding device may also switch the size of the GOS or SPC between the indoor GOS and the outdoor GOS. For example, the encoding device may set the size of the GOS smaller indoors than outdoors. The encoding device may also change the accuracy of extracting feature points from the point cloud or the accuracy of object detection between the indoor GOS and the outdoor GOS.
[0119] The encoding device may also add information to the encoded data that allows the decoding device to distinguish dynamic objects from static objects. This allows the decoding device to display dynamic objects together with red frames or explanatory text. The decoding device may display only red frames or explanatory text instead of dynamic objects. The decoding device may also display more specific object types. For example, a red frame may be used for cars and a yellow frame for people.
[0120] Furthermore, the encoding device or decoding device may determine whether to encode or decode dynamic objects and static objects as different SPCs or GOSs, depending on the frequency of appearance of the dynamic objects, or the ratio of static objects to dynamic objects, etc. For example, if the frequency of appearance or the ratio of dynamic objects exceeds a threshold, an SPC or GOS in which dynamic objects and static objects are mixed is permitted, and if the frequency of appearance or the ratio of dynamic objects does not exceed the threshold, an SPC or GOS in which dynamic objects and static objects are mixed is not permitted.
[0121] When detecting a dynamic object from two-dimensional image information from a camera, rather than from a point cloud, the encoding device may separately obtain information for identifying the detection result (such as a frame or text) and the object position, and encode this information as part of the three-dimensional encoded data. In this case, the decoding device displays auxiliary information (a frame or text) indicating the dynamic object by superimposing it on the decoding result of the static object.
[0122] The encoding device may also change the density of the VXL or VLM in the SPC depending on the complexity of the shape of the static object, etc. For example, the encoding device sets the VXL or VLM to be denser as the shape of the static object becomes more complex. Furthermore, the encoding device may determine the quantization step, etc., when quantizing spatial position or color information, depending on the density of the VXL or VLM. For example, the encoding device sets the quantization step to be smaller as the VXL or VLM becomes denser.
[0123] As described above, the encoding device or decoding device according to this embodiment performs spatial encoding or decoding in units of spaces having coordinate information.
[0124] Furthermore, the encoding device and the decoding device perform encoding and decoding in units of volumes within a space. A volume includes voxels, which are the smallest units to which position information can be associated.
[0125] The encoding device and the decoding device perform encoding or decoding by associating any elements with each other using a table that associates each element of spatial information including coordinates, objects, and time with a GOP, or a table that associates each element with each other. The decoding device determines coordinates using the value of a selected element, identifies a volume, voxel, or space from the coordinates, and decodes the space including the volume or voxel, or the identified space.
[0126] The encoding device also determines volumes, voxels, or spaces that can be selected by elements by feature point extraction or object recognition, and encodes them as randomly accessible volumes, voxels, or spaces.
[0127] Spaces are classified into three types: I-SPC, which can be encoded or decoded by itself; P-SPC, which is encoded or decoded by reference to any one processed space; and B-SPC, which is encoded or decoded by reference to any two processed spaces.
[0128] One or more volumes correspond to static objects or dynamic objects. The space containing the static objects and the space containing the dynamic objects are encoded or decoded as different GOS. That is, the SPC containing the static objects and the SPC containing the dynamic objects are assigned to different GOS.
[0129] Dynamic objects are encoded or decoded on an object-by-object basis and associated with one or more spaces containing static objects, i.e., multiple dynamic objects are encoded individually and the resulting encoded data for the multiple dynamic objects are associated with an SPC containing the static objects.
[0130] The encoding device and the decoding device increase the priority of the I-SPC in the GOS and perform encoding or decoding. For example, the encoding device performs encoding so as to reduce degradation of the I-SPC (so that the original three-dimensional data is reproduced more faithfully after decoding). Also, the decoding device decodes only the I-SPC, for example.
[0131] The encoding device may perform encoding by changing the frequency of using I-SPC depending on the density or number (amount) of objects in the world. In other words, the encoding device changes the frequency of selecting I-SPC depending on the number or density of objects included in the three-dimensional data. For example, the encoding device increases the frequency of using I-space as the density of objects in the world increases.
[0132] Furthermore, the encoding device sets random access points in units of GOS, and stores information indicating a spatial region corresponding to the GOS in the header information.
[0133] The encoding device uses, for example, a default value as the spatial size of the GOS. The encoding device may change the size of the GOS depending on the number (quantity) or density of objects or dynamic objects. For example, the encoding device reduces the spatial size of the GOS as the objects or dynamic objects become denser or the number of objects or dynamic objects becomes greater.
[0134] The space or volume may also include a set of feature points derived using information obtained by sensors such as a depth sensor, a gyro, or a camera. The coordinates of the feature points are set at the center positions of the voxels. Furthermore, the position information can be highly accurate by subdividing the voxels.
[0135] The feature points are derived using a plurality of pictures having at least two types of time information: actual time information and time information that is the same for a plurality of pictures associated with the space (e.g., encoding time used for rate control, etc.).
[0136] Moreover, encoding or decoding is performed in units of GOS including one or more spaces.
[0137] The encoding device and the decoding device refer to spaces in a GOS that have already been processed to predict a P space or a B space in a GOS to be processed.
[0138] Alternatively, the encoding device and the decoding device predict the P space or the B space in the GOS to be processed by using a processed space in the GOS to be processed, without referring to a different GOS.
[0139] Furthermore, the encoding device and the decoding device transmit or receive an encoded stream in units of worlds, each world including one or more GOSs.
[0140] Furthermore, the GOS has a layer structure in at least one direction within a world, and the encoding device and decoding device encode or decode from the lower layer. For example, a randomly accessible GOS belongs to the lowest layer. A GOS belonging to a higher layer refers to a GOS belonging to the same layer or lower. In other words, the GOS is spatially divided in a predetermined direction and includes multiple layers, each of which includes one or more SPCs. The encoding device and decoding device encode or decode each SPC by referring to an SPC included in the same layer as the SPC or in a layer lower than the SPC.
[0141] In addition, the encoding device and the decoding device encode or decode the GOSs continuously within a world unit including multiple GOSs. The encoding device and the decoding device write or read information indicating the order (direction) of encoding or decoding as metadata. In other words, the encoded data includes information indicating the encoding order of multiple GOSs.
[0142] Furthermore, the encoding device and the decoding device encode or decode two or more different spaces or GOS in parallel.
[0143] In addition, the encoding device and the decoding device encode and decode spatial information (coordinates, size, etc.) of the space or GOS.
[0144] In addition, the encoding device and the decoding device encode or decode a space or GOS included in a specific space that is specified based on external information related to its own position and / or area size, such as GPS, route information, or magnification.
[0145] The encoding device or decoding device encodes or decodes a space far from its own position with a lower priority than a space close to its own position.
[0146] The encoding device sets one direction in a world according to a magnification or a use, and encodes a GOS having a layer structure in the direction. The decoding device decodes the GOS having a layer structure in one direction in a world set according to a magnification or a use, preferentially from a lower layer.
[0147] The encoding device changes the feature point extraction, object recognition accuracy, or spatial region size included in the indoor and outdoor spaces. However, the encoding device and the decoding device encode or decode the indoor GOS and the outdoor GOS that are close in coordinates as adjacent in the world, and also encode or decode their identifiers in association with each other.
[0148] (Embodiment 2) When using the encoded data of a point cloud in an actual device or service, it is desirable to transmit and receive the necessary information according to the purpose in order to reduce the network bandwidth. However, until now, such a function has not existed in the encoding structure of three-dimensional data, and there has been no encoding method for this purpose.
[0149] In this embodiment, we describe a three-dimensional data encoding method and a three-dimensional data encoding device that provide the function of transmitting and receiving only the information necessary depending on the application in encoded data of a three-dimensional point cloud, as well as a three-dimensional data decoding method and a three-dimensional data decoding device that decodes the encoded data.
[0150] A voxel (VXL) having a certain amount of features or more is defined as a feature voxel (FVXL), and a world (WLD) composed of FVXL is defined as a sparse world (SWLD). FIG. 11 is a diagram showing an example of the configuration of a sparse world and a world. SWLD includes FGOS, which is a GOS composed of FVXL, FSPC, which is an SPC composed of FVXL, and FVLM, which is a VLM composed of FVXL. The data structures and prediction structures of FGOS, FSPC, and FVLM may be the same as those of GOS, SPC, and VLM.
[0151] The feature amount is a feature amount that expresses three-dimensional position information of the VXL or visible light information of the VXL position, and is a feature amount that is often detected especially at corners and edges of a three-dimensional object. Specifically, this feature amount is a three-dimensional feature amount or visible light feature amount as described below, but any feature amount that expresses the position, brightness, or color information of the VXL, etc., may be used.
[0152] As 3D features, SHOT features (Signature of Histograms of Orientations) and PFH features (Point Feature Heuristics) are used. Histograms, or PPF features (Point Pair Feature) are used.
[0153] The SHOT feature is obtained by dividing the VXL area, calculating the inner product of the reference point and the normal vector of the divided area, and forming a histogram. This SHOT feature has a high dimensionality and is characterized by its high expressiveness.
[0154] The PFH feature is obtained by selecting many pairs of two points near the VXL, calculating normal vectors from those two points, and creating a histogram. Since the PFH feature is a histogram feature, it is robust against some disturbances and has high feature expression power.
[0155] The PPF feature is calculated using normal vectors for each of two VXL points. Since the entire VXL is used for the PPF feature, it is robust against occlusion.
[0156] In addition, SIFT (Scale-Invariant Feature Transform), SURF (Speeded Up Robust Features), or HOG (Histogram Array) are used to estimate the visible light feature quantity, using information such as the luminance gradient of the image. of Oriented Gradients, etc. can be used.
[0157] The SWLD is generated by calculating the above feature amount from each VXL of the WLD and extracting the FVXL. Here, the SWLD may be updated every time the WLD is updated, or may be updated periodically after a certain period of time has elapsed, regardless of the timing of updating the WLD.
[0158] A SWLD may be generated for each feature. For example, a separate SWLD may be generated for each feature, such as SWLD1 based on SHOT features and SWLD2 based on SIFT features, and the SWLDs may be used according to the purpose. Also, the calculated feature of each FVXL may be stored in each FVXL as feature information.
[0159] Next, we explain how to use sparse world learning (SWLD). SWLDs contain only feature voxels (FVXL), so their data size is generally smaller than WLDs that contain all VXL.
[0160] In applications that use features to achieve some purpose, the time required to read from a hard disk, as well as the bandwidth and transfer time required for network transfer, can be reduced by using SWLD information instead of WLD information. For example, by storing WLD and SWLD as map information on a server and switching the map information to be sent to WLD or SWLD in response to a request from a client, the network bandwidth and transfer time can be reduced. A specific example is shown below.
[0161] 12 and 13 are diagrams showing examples of using SWLD and WLD. As shown in FIG. 12, when a client 1, which is an in-vehicle device, needs map information for self-location determination, the client 1 sends a request for map data for self-location estimation to the server (S301). The server transmits an SWLD to the client 1 in response to the request (S302). The client 1 performs self-location determination using the received SWLD (S303). At this time, the client 1 obtains VXL information around the client 1 using various methods such as a distance sensor such as a range finder, a stereo camera, or a combination of multiple monocular cameras, and estimates self-location information from the obtained VXL information and SWLD. Here, the self-location information includes three-dimensional position information and orientation of the client 1.
[0162] As shown in Fig. 13, when the client 2, which is an in-vehicle device, needs map information for the purpose of drawing a map such as a three-dimensional map, the client 2 sends a request to the server to obtain map data for map drawing (S311). The server transmits a WLD to the client 2 in response to the request (S312). The client 2 uses the received WLD to draw the map (S313). At this time, the client 2 creates a rendering image using, for example, an image captured by itself using a visible light camera or the like and the WLD obtained from the server, and draws the created image on the screen of a car navigation system or the like.
[0163] As described above, the server transmits SWLD to the client for applications that mainly require the feature values of each VXL, such as self-location estimation, and transmits WLD to the client for applications that require detailed VXL information, such as map drawing. This makes it possible to transmit and receive map data efficiently.
[0164] The client may determine whether it needs a SWLD or a WLD and request the server to send either a SWLD or a WLD. The server may determine whether it should send a SWLD or a WLD depending on the client or network conditions.
[0165] Next, we will explain how to switch between sending and receiving data in the sparse world (SWLD) and the world (WLD).
[0166] Whether to receive WLD or SWLD may be switched depending on the network band. FIG. 14 is a diagram showing an example of operation in this case. For example, in LTE (Long Term When a low-speed network with limited available network bandwidth, such as in a 3G (Internet of Things) Evolution environment, is used, the client accesses the server via the low-speed network (S321) and acquires SWLD as map information from the server (S322). On the other hand, when a high-speed network with ample network bandwidth, such as in a Wi-Fi (registered trademark) environment, is used, the client accesses the server via the high-speed network (S323) and acquires WLD from the server (S324). This allows the client to acquire map information appropriate to the client's network bandwidth.
[0167] Specifically, the client receives SWLD via LTE when outdoors, and obtains WLD via Wi-Fi (registered trademark) when inside a facility, etc. This enables the client to obtain more detailed map information about the indoor area.
[0168] In this way, the client may request a WLD or SWLD from the server depending on the bandwidth of the network that the client uses. Alternatively, the client may send information indicating the bandwidth of the network that the client uses to the server, and the server may send data (WLD or SWLD) suitable for the client depending on the information. Alternatively, the server may determine the network bandwidth of the client and send data (WLD or SWLD) suitable for the client.
[0169] Also, whether to receive the WLD or SWLD may be switched depending on the moving speed. FIG. 15 is a diagram showing an example of the operation in this case. For example, when the client is moving at high speed (S331), the client receives the SWLD from the server (S332). On the other hand, when the client is moving at low speed (S333), the client receives the WLD from the server (S334). This allows the client to obtain map information suited to the speed while suppressing the network bandwidth. Specifically, the client can update rough map information at an appropriate speed by receiving SWLD with a small amount of data while traveling on a highway. On the other hand, the client can obtain more detailed map information by receiving WLD while traveling on an ordinary road.
[0170] In this way, the client may request the server for the WLD or SWLD depending on its own moving speed. Alternatively, the client may transmit information indicating its own moving speed to the server, and the server may transmit data (WLD or SWLD) suitable for the client depending on the information. Alternatively, the server may determine the moving speed of the client and transmit data (WLD or SWLD) suitable for the client.
[0171] Alternatively, the client may first obtain SWLD from the server, and then obtain WLD for important areas from that. For example, when obtaining map data, the client may first obtain rough map information in SWLD, narrow down the map information to areas where buildings, signs, people, or other features appear frequently, and obtain WLD for the narrowed down area later. This allows the client to obtain detailed information for the required area while suppressing the amount of data received from the server.
[0172] Also, the server may create separate SWLDs for each object from the WLD, and the client may receive each according to the purpose. This can reduce network bandwidth. For example, the server may recognize people or cars from the WLD in advance, and create SWLDs for people and SWLDs for cars. The client receives SWLDs for people when it wants to obtain information about people in the vicinity, and SWLDs for cars when it wants to obtain information about cars. Also, the types of SWLDs may be distinguished by information (flags, types, etc.) added to the header, etc.
[0173] Next, the configuration and operation flow of a three-dimensional data encoding device (e.g., a server) according to this embodiment will be described. Fig. 16 is a block diagram of a three-dimensional data encoding device 400 according to this embodiment. Fig. 17 is a flowchart of a three-dimensional data encoding process by the three-dimensional data encoding device 400.
[0174] 16 generates encoded three-dimensional data 413 and 414, which are encoded streams, by encoding input three-dimensional data 411. Here, the encoded three-dimensional data 413 is encoded three-dimensional data corresponding to a WLD, and the encoded three-dimensional data 414 is encoded three-dimensional data corresponding to a SWLD. This three-dimensional data encoding device 400 includes an acquisition unit 401, an encoding region determination unit 402, an SWLD extraction unit 403, a WLD encoding unit 404, and an SWLD encoding unit 405.
[0175] As shown in FIG. 17, first, the acquisition unit 401 acquires input three-dimensional data 411, which is point cloud data in a three-dimensional space (S401).
[0176] Next, the coding region determination unit 402 determines a spatial region to be coded based on the spatial region in which the point cloud data exists (S402).
[0177] Next, the SWLD extraction unit 403 defines the spatial region to be encoded as a WLD, and calculates a feature amount from each VXL included in the WLD.The SWLD extraction unit 403 then extracts VXLs whose feature amount is equal to or greater than a predetermined threshold, defines the extracted VXLs as FVXLs, and adds the FVXLs to the SWLD to generate extracted three-dimensional data 412 (S403).That is, extracted three-dimensional data 412 whose feature amount is equal to or greater than a threshold is extracted from the input three-dimensional data 411.
[0178] Next, the WLD encoding unit 404 generates encoded 3D data 413 corresponding to the WLD by encoding the input 3D data 411 corresponding to the WLD (S404). At this time, the WLD encoding unit 404 adds information to the header of the encoded 3D data 413 for distinguishing that the encoded 3D data 413 is a stream including a WLD.
[0179] Furthermore, the SWLD encoding unit 405 generates encoded three-dimensional data 414 corresponding to the SWLD by encoding the extracted three-dimensional data 412 corresponding to the SWLD (S405). At this time, the SWLD encoding unit 405 adds information to the header of the encoded three-dimensional data 414 for distinguishing that the encoded three-dimensional data 414 is a stream including a SWLD.
[0180] The order of the process for generating the encoded three-dimensional data 413 and the process for generating the encoded three-dimensional data 414 may be reversed. Also, some or all of these processes may be performed in parallel.
[0181] For example, a parameter "world_type" is defined as information to be added to the headers of the encoded 3D data 413 and 414. When world_type=0, it indicates that the stream includes a WLD, and when world_type=1, it indicates that the stream includes a SWLD. When defining many other types, the assigned numerical value may be increased, such as world_type=2. Also, a specific flag may be included in one of the encoded 3D data 413 and 414. For example, a flag indicating that the stream includes a SWLD may be added to the encoded 3D data 414. In this case, the decoding device can determine whether the stream includes a WLD or a SWLD depending on the presence or absence of the flag.
[0182] Furthermore, the encoding method used by the WLD encoding unit 404 when encoding the WLD may be different from the encoding method used by the SWLD encoding unit 405 when encoding the SWLD.
[0183] For example, in SWLD, data is thinned, so that the correlation with surrounding data may be lower than in WLD. Therefore, in the encoding method used in SWLD, among intra prediction and inter prediction, inter prediction may be prioritized over the encoding method used in WLD.
[0184] In addition, the coding method used for SWLD and the coding method used for WLD may differ in the way of expressing the three-dimensional position. For example, in SWLD, the three-dimensional position of FVXL may be expressed by three-dimensional coordinates, and in WLD, the three-dimensional position may be expressed by an octree, which will be described later, or vice versa.
[0185] Furthermore, the SWLD encoding unit 405 performs encoding so that the data size of the encoded three-dimensional data 414 of SWLD is smaller than the data size of the encoded three-dimensional data 413 of WLD. For example, as described above, the correlation between data of SWLD may be lower than that of WLD. This may result in a decrease in encoding efficiency, and the data size of the encoded three-dimensional data 414 may be larger than the data size of the encoded three-dimensional data 413 of WLD. Therefore, when the data size of the obtained encoded three-dimensional data 414 is larger than the data size of the encoded three-dimensional data 413 of WLD, the SWLD encoding unit 405 regenerates the encoded three-dimensional data 414 with a reduced data size by performing re-encoding.
[0186] For example, the SWLD extraction unit 403 regenerates the extracted three-dimensional data 412 with a reduced number of extracted feature points, and the SWLD encoding unit 405 encodes the extracted three-dimensional data 412. Alternatively, the degree of quantization in the SWLD encoding unit 405 may be made coarser. For example, in an octree structure described later, the degree of quantization can be made coarser by rounding the data in the lowest layer.
[0187] Furthermore, if the data size of the encoded three-dimensional data 414 of SWLD cannot be made smaller than the data size of the encoded three-dimensional data 413 of WLD, the SWLD encoding unit 405 may not generate the encoded three-dimensional data 414 of SWLD. Alternatively, the encoded three-dimensional data 413 of WLD may be copied to the encoded three-dimensional data 414 of SWLD. In other words, the encoded three-dimensional data 413 of WLD may be used as it is as the encoded three-dimensional data 414 of SWLD.
[0188] Next, the configuration and operation flow of a three-dimensional data decoding device (e.g., a client) according to this embodiment will be described. Fig. 18 is a block diagram of a three-dimensional data decoding device 500 according to this embodiment. Fig. 19 is a flowchart of three-dimensional data decoding processing by the three-dimensional data decoding device 500.
[0189] 18 generates decoded three-dimensional data 512 or 513 by decoding encoded three-dimensional data 511. Here, the encoded three-dimensional data 511 is, for example, the encoded three-dimensional data 413 or 414 generated by the three-dimensional data encoding device 400.
[0190] This three-dimensional data decoding device 500 includes an acquisition unit 501 , a header analysis unit 502 , a WLD decoding unit 503 , and a SWLD decoding unit 504 .
[0191] 19, first, an acquisition unit 501 acquires encoded 3D data 511 (S501). Next, a header analysis unit 502 analyzes the header of the encoded 3D data 511 and determines whether the encoded 3D data 511 is a stream including a WLD or a stream including a SWLD (S502). For example, the above-mentioned world_type parameter is referenced to perform the determination.
[0192] If the encoded three-dimensional data 511 is a stream including a WLD (Yes in S503), the WLD decoding unit 503 generates decoded three-dimensional data 512 of the WLD by decoding the encoded three-dimensional data 511 (S504). On the other hand, if the encoded three-dimensional data 511 is a stream including a SWLD (No in S503), the SWLD decoding unit 504 generates decoded three-dimensional data 513 of the SWLD by decoding the encoded three-dimensional data 511 (S505).
[0193] Also, similarly to the encoding device, the decoding method used by the WLD decoding unit 503 when decoding the WLD may be different from the decoding method used by the SWLD decoding unit 504 when decoding the SWLD. For example, in the decoding method used for the SWLD, inter prediction may be prioritized over the decoding method used for the WLD, out of intra prediction and inter prediction.
[0194] In addition, the decoding method used for SWLD and the decoding method used for WLD may use different methods for expressing three-dimensional positions. For example, in SWLD, the three-dimensional position of FVXL may be expressed by three-dimensional coordinates, and in WLD, the three-dimensional position may be expressed by an octet tree, which will be described later, or vice versa.
[0195] Next, an octree representation, which is a method of representing a three-dimensional position, will be described. The VXL data included in the three-dimensional data is converted into an octree structure and then encoded. FIG. 20 is a diagram showing an example of a VXL in a WLD. FIG. 21 is a diagram showing the octree structure of the WLD shown in FIG. 20. In the example shown in FIG. 20, there are three VXLs VXL1 to 3 that are VXLs (hereinafter, valid VXLs) that include point groups. As shown in FIG. 21, the octree structure is composed of nodes and leaves. Each node has a maximum of eight nodes or leaves. Each leaf has VXL information. Here, among the leaves shown in FIG. 21, leaves 1, 2, and 3 respectively represent VXL1, VXL2, and VXL3 shown in FIG. 20.
[0196] Specifically, each node and leaf corresponds to a three-dimensional position. Node 1 corresponds to the entire block shown in FIG. 20. The block corresponding to node 1 is divided into eight blocks, and among the eight blocks, the block containing a valid VXL is set as a node, and the other blocks are set as leaves. The block corresponding to the node is further divided into eight nodes or leaves, and this process is repeated for the number of layers of the tree structure. Also, all the blocks in the lowest layer are set as leaves.
[0197] FIG. 22 is a diagram showing an example of a SWLD generated from the WLD shown in FIG. 20. As a result of feature extraction, VXL1 and VXL2 shown in FIG. 20 are determined to be FVXL1 and FVXL2, and are added to the SWLD. On the other hand, VXL3 is not determined to be FVXL, and is not included in the SWLD. FIG. 23 is a diagram showing an octree structure of the SWLD shown in FIG. 22. In the octree structure shown in FIG. 23, leaf 3 corresponding to VXL3 shown in FIG. 21 is deleted. As a result, node 3 shown in FIG. 21 no longer has a valid VXL, and is changed to a leaf. In this way, the number of leaves in the SWLD is generally smaller than the number of leaves in the WLD, and the encoded three-dimensional data of the SWLD is also smaller than the encoded three-dimensional data of the WLD.
[0198] A modification of this embodiment will now be described.
[0199] For example, when a client such as an in-vehicle device estimates its own location, it receives an SWLD from a server and estimates its own location using the SWLD. When performing obstacle detection, the client may perform obstacle detection based on three-dimensional information of the surrounding area that it has acquired using various methods such as a distance sensor such as a range finder, a stereo camera, or a combination of multiple monocular cameras.
[0200] In addition, SWLDs generally do not include VXL data for flat areas. Therefore, the server may hold a subsampled world (subWLD) that is a subsample of WLD for static obstacle detection, and transmit SWLD and subWLD to the client. This allows the client to perform self-location estimation and obstacle detection while suppressing network bandwidth.
[0201] Also, when a client wants to quickly draw 3D map data, it may be more convenient for the map information to have a mesh structure. Therefore, the server may generate a mesh from the WLD and store it in advance as a mesh world (MWLD). For example, a client receives an MWLD when it needs a rough 3D drawing, and receives a WLD when it needs a detailed 3D drawing. This can reduce the network bandwidth.
[0202] In addition, the server sets the VXL whose feature amount is equal to or greater than the threshold value as the FVXL among the VXLs, but the FVXL may be calculated by a different method. For example, the server may determine that the VXL, VLM, SPC, or GOS constituting a signal or an intersection is necessary for self-location estimation, driving assistance, or automatic driving, and may include the VXL, VLM, FSPC, or GOS in the SWLD. The above determination may be performed manually. The FVXL, etc. obtained by the above method may be added to the FVXL, etc. set based on the feature amount. That is, the SWLD extraction unit 403 may further extract data corresponding to an object having a predetermined attribute from the input three-dimensional data 411 as the extracted three-dimensional data 412.
[0203] In addition, the fact that it is necessary for those uses may be labeled separately from the feature. In addition, the server may separately hold FVXL necessary for self-location estimation at signals or intersections, driving assistance, autonomous driving, etc., as a higher layer (e.g., lane world) of SWLD.
[0204] The server may also add attributes to the VXL in the WLD for each random access unit or for each predetermined unit. The attributes include, for example, information indicating whether it is necessary or unnecessary for self-location estimation, or information indicating whether it is important as traffic information such as a signal or an intersection. The attributes may also include a correspondence relationship with a feature (such as an intersection or a road) in lane information (such as GDF: Geographic Data Files).
[0205] In addition, the following method may be used as a method for updating the WLD or SWLD.
[0206] Updates, such as changes in people, construction, or tree-lined streets (for trucks), are uploaded to the server as point clouds or metadata. The server updates the WLD based on the upload, and then updates the SWLD with the updated WLD.
[0207] In addition, when the client detects an inconsistency between the 3D information generated by itself during self-location estimation and the 3D information received from the server, the client may transmit the 3D information generated by itself together with an update notification to the server. In this case, the server updates the SWLD using the WLD. If the SWLD is not updated, the server determines that the WLD itself is out of date.
[0208] In addition, although information for distinguishing between WLD and SWLD is added as header information of the encoded stream, when there are many kinds of worlds, such as mesh world and lane world, information for distinguishing them may be added to the header information. In addition, when there are many SWLDs with different features, information for distinguishing them may be added to the header information.
[0209] Although the SWLD is described as being composed of FVXL, it may include VXL that is not determined to be FVXL. For example, the SWLD may include adjacent VXL used when calculating the feature amount of FVXL. This allows the client to calculate the feature amount of FVXL when receiving the SWLD, even if feature amount information is not added to each FVXL in the SWLD. In this case, the SWLD may include information for distinguishing whether each VXL is FVXL or VXL.
[0210] As described above, the three-dimensional data encoding device 400 extracts extracted three-dimensional data 412 (second three-dimensional data) whose features are greater than or equal to a threshold value from the input three-dimensional data 411 (first three-dimensional data), and generates encoded three-dimensional data 414 (first encoded three-dimensional data) by encoding the extracted three-dimensional data 412.
[0211] According to this, the three-dimensional data encoding device 400 generates encoded three-dimensional data 414 by encoding data whose feature amount is equal to or greater than a threshold value. This makes it possible to reduce the amount of data compared to when the input three-dimensional data 411 is encoded as is. Therefore, the three-dimensional data encoding device 400 can reduce the amount of data to be transmitted.
[0212] Moreover, the three-dimensional data encoding device 400 further encodes the input three-dimensional data 411 to generate encoded three-dimensional data 413 (second encoded three-dimensional data).
[0213] According to this, the three-dimensional data encoding device 400 can selectively transmit the encoded three-dimensional data 413 and the encoded three-dimensional data 414 depending on, for example, the intended use.
[0214] Moreover, the extracted three-dimensional data 412 is encoded by a first encoding method, and the input three-dimensional data 411 is encoded by a second encoding method different from the first encoding method.
[0215] This allows the three-dimensional data encoding device 400 to use an encoding method suitable for each of the input three-dimensional data 411 and the extracted three-dimensional data 412.
[0216] Furthermore, in the first encoding method, of intra prediction and inter prediction, inter prediction is given priority over the second encoding method.
[0217] This allows the 3D data encoding device 400 to increase the priority of inter prediction for the extracted 3D data 412 in which the correlation between adjacent data is likely to be low.
[0218] In addition, the first and second encoding methods differ in the way they represent three-dimensional positions. For example, the second encoding method represents three-dimensional positions using an octree, whereas the first encoding method represents three-dimensional positions using three-dimensional coordinates.
[0219] This enables the three-dimensional data encoding device 400 to use a more suitable three-dimensional position expression method for three-dimensional data having a different number of pieces of data (the number of VXLs or FVXLs).
[0220] Furthermore, at least one of the encoded three-dimensional data 413 and 414 includes an identifier indicating whether the encoded three-dimensional data is encoded three-dimensional data obtained by encoding the input three-dimensional data 411, or encoded three-dimensional data obtained by encoding a portion of the input three-dimensional data 411. In other words, the identifier indicates whether the encoded three-dimensional data is encoded three-dimensional data 413 of WLD or encoded three-dimensional data 414 of SWLD.
[0221] This allows the decoding device to easily determine whether the acquired encoded three-dimensional data is encoded three-dimensional data 413 or encoded three-dimensional data 414.
[0222] Furthermore, the three-dimensional data encoding device 400 encodes the extracted three-dimensional data 412 so that the amount of encoded three-dimensional data 414 is smaller than the amount of encoded three-dimensional data 413 .
[0223] According to this, the three-dimensional data encoding device 400 can make the amount of encoded three-dimensional data 414 smaller than the amount of encoded three-dimensional data 413 .
[0224] Moreover, the three-dimensional data encoding device 400 further extracts data corresponding to an object having a predetermined attribute from the input three-dimensional data 411 as extracted three-dimensional data 412. For example, the object having the predetermined attribute is an object necessary for self-location estimation, driving assistance, automatic driving, or the like, such as a traffic light or an intersection.
[0225] This allows the three-dimensional data encoding device 400 to generate encoded three-dimensional data 414 that includes data required by the decoding device.
[0226] Moreover, the three-dimensional data encoding device 400 (server) further transmits one of the encoded three-dimensional data 413 and 414 to the client depending on the state of the client.
[0227] This allows the three-dimensional data encoding device 400 to transmit appropriate data depending on the state of the client.
[0228] The state of the client also includes the communication status of the client (eg, network bandwidth) or the movement speed of the client.
[0229] Moreover, the three-dimensional data encoding device 400 further transmits one of the encoded three-dimensional data 413 and 414 to the client in response to a request from the client.
[0230] This allows the three-dimensional data encoding device 400 to transmit appropriate data in response to a request from a client.
[0231] Moreover, the three-dimensional data decoding device 500 according to this embodiment decodes the encoded three-dimensional data 413 or 414 generated by the three-dimensional data encoding device 400 described above.
[0232] That is, the three-dimensional data decoding device 500 decodes, by a first decoding method, the encoded three-dimensional data 414 obtained by encoding the extracted three-dimensional data 412, the feature amount of which is equal to or greater than a threshold value, extracted from the input three-dimensional data 411. Also, the three-dimensional data decoding device 500 decodes, by a second decoding method different from the first decoding method, the encoded three-dimensional data 413 obtained by encoding the input three-dimensional data 411.
[0233] According to this, the three-dimensional data decoding device 500 can selectively receive the encoded three-dimensional data 414, which is obtained by encoding data whose feature amount is equal to or greater than a threshold value, and the encoded three-dimensional data 413, depending on, for example, the intended use. This allows the three-dimensional data decoding device 500 to reduce the amount of data to be transmitted. Furthermore, the three-dimensional data decoding device 500 can use a decoding method suitable for each of the input three-dimensional data 411 and the extracted three-dimensional data 412.
[0234] Furthermore, in the first decoding method, of intra prediction and inter prediction, inter prediction is given priority over the second decoding method.
[0235] This allows the 3D data decoding device 500 to increase the priority of inter prediction for extracted 3D data in which the correlation between adjacent data is likely to be low.
[0236] In addition, the first and second decoding methods have different methods of expressing three-dimensional positions. For example, in the second decoding method, a three-dimensional position is expressed by an octree, whereas in the first decoding method, a three-dimensional position is expressed by three-dimensional coordinates.
[0237] This allows the three-dimensional data decoding device 500 to use a more suitable three-dimensional position expression method for three-dimensional data with different numbers of data (numbers of VXLs or FVXLs).
[0238] Furthermore, at least one of the encoded three-dimensional data 413 and 414 includes an identifier indicating whether the encoded three-dimensional data is encoded three-dimensional data obtained by encoding the input three-dimensional data 411, or encoded three-dimensional data obtained by encoding a portion of the input three-dimensional data 411. The three-dimensional data decoding device 500 identifies the encoded three-dimensional data 413 and 414 by referring to the identifier.
[0239] This allows the three-dimensional data decoding device 500 to easily determine whether the acquired encoded three-dimensional data is the encoded three-dimensional data 413 or the encoded three-dimensional data 414.
[0240] Furthermore, the three-dimensional data decoding device 500 further notifies the server of the state of the client (the three-dimensional data decoding device 500). The three-dimensional data decoding device 500 receives one of the encoded three-dimensional data 413 and 414 transmitted from the server depending on the state of the client.
[0241] This allows the three-dimensional data decoding device 500 to receive appropriate data depending on the state of the client.
[0242] The state of the client also includes the communication status of the client (eg, network bandwidth) or the movement speed of the client.
[0243] Moreover, the three-dimensional data decoding device 500 further requests one of the encoded three-dimensional data 413 and 414 from the server, and receives one of the encoded three-dimensional data 413 and 414 transmitted from the server in response to the request.
[0244] This allows the three-dimensional data decoding device 500 to receive appropriate data according to the application.
[0245] (Embodiment 3) In this embodiment, a method for transmitting and receiving three-dimensional data between vehicles will be described. For example, three-dimensional data is transmitted and received between a vehicle and a surrounding vehicle.
[0246] 24 is a block diagram of a three-dimensional data creation device 620 according to this embodiment. The three-dimensional data creation device 620 is included in, for example, a vehicle, and creates a more detailed third three-dimensional data 636 by synthesizing a first three-dimensional data 632 created by the three-dimensional data creation device 620 with a second three-dimensional data 635 received.
[0247] This three-dimensional data creation device 620 includes a three-dimensional data creation unit 621 , a requested range determination unit 622 , a search unit 623 , a reception unit 624 , a decoding unit 625 , and a synthesis unit 626 .
[0248] First, a three-dimensional data creation unit 621 creates first three-dimensional data 632 using sensor information 631 detected by a sensor equipped in the vehicle. Next, a required range determination unit 622 determines a required range, which is a three-dimensional spatial range in which data is insufficient in the created first three-dimensional data 632.
[0249] Next, the search unit 623 searches for nearby vehicles that have three-dimensional data in the requested range, and transmits requested range information 633 indicating the requested range to the nearby vehicles identified by the search. Next, the receiving unit 624 receives encoded three-dimensional data 634, which is an encoded stream of the requested range, from the nearby vehicles (S624). Note that the search unit 623 may indiscriminately issue a request to all vehicles present in a specific range, and receive the encoded three-dimensional data 634 from those that respond. Also, the search unit 623 may issue a request to an object, such as a traffic light or a sign, in addition to a vehicle, and receive the encoded three-dimensional data 634 from the object.
[0250] Next, the decoding unit 625 obtains second three-dimensional data 635 by decoding the received encoded three-dimensional data 634. Next, the synthesis unit 626 creates third three-dimensional data 636, which is denser, by synthesizing the first three-dimensional data 632 and the second three-dimensional data 635.
[0251] Next, a configuration and operation of three-dimensional data transmission device 640 according to this embodiment will be described.
[0252] The three-dimensional data transmission device 640 is, for example, included in the surrounding vehicle described above, processes the fifth three-dimensional data 652 created by the surrounding vehicle into sixth three-dimensional data 654 requested by the host vehicle, encodes the sixth three-dimensional data 654 to generate encoded three-dimensional data 634, and transmits the encoded three-dimensional data 634 to the host vehicle.
[0253] The three-dimensional data transmission device 640 includes a three-dimensional data creation unit 641 , a receiving unit 642 , an extracting unit 643 , an encoding unit 644 , and a transmitting unit 645 .
[0254] First, the three-dimensional data creation unit 641 creates fifth three-dimensional data 652 using sensor information 651 detected by a sensor equipped in the surrounding vehicle. Next, the receiving unit 642 receives the requested range information 633 transmitted from the host vehicle.
[0255] Next, the extraction unit 643 extracts three-dimensional data of the requested range indicated by the requested range information 633 from the fifth three-dimensional data 652, thereby processing the fifth three-dimensional data 652 into sixth three-dimensional data 654. Next, the encoding unit 644 encodes the sixth three-dimensional data 654 to generate encoded three-dimensional data 634, which is an encoded stream. Then, the transmission unit 645 transmits the encoded three-dimensional data 634 to the host vehicle.
[0256] Here, an example is described in which the vehicle itself is equipped with a three-dimensional data creation device 620 and the surrounding vehicles are equipped with a three-dimensional data transmission device 640, but each vehicle may have the functions of both the three-dimensional data creation device 620 and the three-dimensional data transmission device 640.
[0257] (Embodiment 4) In this embodiment, the operation of an abnormality system in self-location estimation based on a three-dimensional map will be described.
[0258] It is expected that applications such as self-driving cars, or autonomous movement of mobile objects such as robots or flying objects such as drones will expand in the future. One example of a means for realizing such autonomous movement is a method in which a mobile object estimates its own position within a three-dimensional map (self-location estimation) and travels according to the map.
[0259] Self-location estimation can be achieved by matching a 3D map with 3D information about the surroundings of the vehicle (hereinafter referred to as vehicle-detected 3D data) obtained by sensors such as a range finder (such as LiDAR) or a stereo camera mounted on the vehicle, and estimating the vehicle's position within the 3D map.
[0260] A 3D map, such as the HD map proposed by HERE, may include not only a 3D point cloud, but also 2D map data such as shape information of roads and intersections, or information that changes in real time such as traffic jams and accidents. A 3D map is made up of multiple layers such as 3D data, 2D data, and metadata that changes in real time, and a device can obtain or refer to only the necessary data.
[0261] The point cloud data may be the above-mentioned SWLD, or may include point group data other than feature points. Furthermore, the transmission and reception of the point cloud data is performed on a basis of one or multiple random access units.
[0262] The following methods can be used as a method for matching the 3D map with the vehicle-detected 3D data. For example, the device compares the shapes of the point groups in each point cloud, and determines that parts with high similarity between feature points are in the same position. In addition, when the 3D map is composed of SWLDs, the device performs matching by comparing the feature points that compose the SWLDs with the 3D feature points extracted from the vehicle-detected 3D data.
[0263] Here, to estimate the vehicle's position with high accuracy, (A) it is necessary to acquire a 3D map and 3D vehicle detection data, and (B) the accuracy of these must satisfy a predetermined standard. However, in the following abnormal cases, (A) or (B) cannot be satisfied.
[0264] (1) 3D maps cannot be obtained via communication.
[0265] (2) The 3D map does not exist, or the 3D map has been acquired but is corrupted.
[0266] (3) The vehicle's sensor is out of order or the weather is bad, so the accuracy of generating the vehicle-detected 3D data is insufficient.
[0267] The operation for dealing with these abnormal cases will be described below. Although the operation will be described below using a car as an example, the following method can be applied to any autonomously moving animal, such as a robot or a drone.
[0268] The configuration and operation of a three-dimensional information processing device according to this embodiment for dealing with abnormal cases in a three-dimensional map or vehicle-detected three-dimensional data will be described below. Fig. 26 is a block diagram showing an example of the configuration of a three-dimensional information processing device 700 according to this embodiment.
[0269] The three-dimensional information processing device 700 is mounted on a moving object such as an automobile. As shown in FIG. 26 , the three-dimensional information processing device 700 includes a three-dimensional map acquisition unit 701, a host vehicle detection data acquisition unit 702, an abnormality case determination unit 703, a countermeasure operation decision unit 704, and an operation control unit 705.
[0270] The three-dimensional information processing device 700 may include a two-dimensional or one-dimensional sensor (not shown) for detecting structures or animals around the vehicle, such as a camera for acquiring a two-dimensional image, or a sensor for one-dimensional data using ultrasonic waves or a laser. The three-dimensional information processing device 700 may also include a communication unit (not shown) for acquiring a three-dimensional map through a mobile communication network such as 4G or 5G, or through vehicle-to-vehicle communication or road-to-vehicle communication.
[0271] The three-dimensional map acquisition unit 701 acquires a three-dimensional map 711 of the vicinity of the travel route. For example, the three-dimensional map acquisition unit 701 acquires the three-dimensional map 711 through a mobile communication network, or through vehicle-to-vehicle communication or road-to-vehicle communication.
[0272] Next, the host vehicle detection data acquisition unit 702 acquires host vehicle detection three-dimensional data 712 based on the sensor information. For example, the host vehicle detection data acquisition unit 702 generates the host vehicle detection three-dimensional data 712 based on sensor information acquired by a sensor equipped in the host vehicle.
[0273] Next, the abnormality case determination unit 703 detects an abnormality case by performing a predetermined check on at least one of the acquired three-dimensional map 711 and the host vehicle detected three-dimensional data 712. In other words, the abnormality case determination unit 703 determines whether at least one of the acquired three-dimensional map 711 and the host vehicle detected three-dimensional data 712 is abnormal.
[0274] When an abnormal case is detected, a countermeasure action decision unit 704 decides a countermeasure action for the abnormal case. Next, an action control unit 705 controls the operation of each processing unit required for carrying out the countermeasure action, such as the three-dimensional map acquisition unit 701.
[0275] On the other hand, if no abnormal case is detected, the three-dimensional information processing apparatus 700 ends the process.
[0276] Furthermore, the three-dimensional information processing device 700 uses the three-dimensional map 711 and the vehicle-detected three-dimensional data 712 to estimate the self-position of the vehicle having the three-dimensional information processing device 700. Next, the three-dimensional information processing device 700 automatically drives the vehicle using the result of the self-position estimation.
[0277] In this way, the three-dimensional information processing device 700 acquires map data (three-dimensional map 711) including the first three-dimensional position information via a communication path. For example, the first three-dimensional position information is encoded in units of subspaces having three-dimensional coordinate information, and each is a collection of one or more subspaces, and includes a plurality of random access units that can be independently decoded. For example, the first three-dimensional position information is data (SWLD) in which feature points whose three-dimensional feature amount is equal to or greater than a predetermined threshold value are encoded.
[0278] Furthermore, the three-dimensional information processing device 700 generates second three-dimensional position information (vehicle-detected three-dimensional data 712) from information detected by the sensor. Next, the three-dimensional information processing device 700 performs an abnormality determination process on the first three-dimensional position information or the second three-dimensional position information to determine whether the first three-dimensional position information or the second three-dimensional position information is abnormal.
[0279] When the first three-dimensional position information or the second three-dimensional position information is determined to be abnormal, the three-dimensional information processing apparatus 700 determines a countermeasure action for the abnormality. Next, the three-dimensional information processing apparatus 700 performs control required for carrying out the countermeasure action.
[0280] This allows the three-dimensional information processing apparatus 700 to detect an abnormality in the first three-dimensional position information or the second three-dimensional position information, and to take appropriate action.
[0281] (Embodiment 5) In this embodiment, a method of transmitting three-dimensional data to a following vehicle will be described.
[0282] 27 is a block diagram showing an example of the configuration of a three-dimensional data creation device 810 according to this embodiment. This three-dimensional data creation device 810 is mounted on, for example, a vehicle. The three-dimensional data creation device 810 transmits and receives three-dimensional data to and from an external traffic monitoring cloud, a leading vehicle, or a following vehicle, and creates and stores three-dimensional data.
[0283] The three-dimensional data creation device 810 includes a data receiving unit 811, a communication unit 812, a receiving control unit 813, a format conversion unit 814, multiple sensors 815, a three-dimensional data creation unit 816, a three-dimensional data synthesis unit 817, a three-dimensional data storage unit 818, a communication unit 819, a transmission control unit 820, a format conversion unit 821, and a data transmission unit 822.
[0284] The data receiving unit 811 receives three-dimensional data 831 from a traffic monitoring cloud or a preceding vehicle. The three-dimensional data 831 includes information such as a point cloud, a visible light image, depth information, sensor position information, or speed information, including an area that cannot be detected by the sensor 815 of the vehicle itself.
[0285] The communication unit 812 communicates with the traffic monitoring cloud or the vehicle ahead, and transmits data transmission requests, etc. to the traffic monitoring cloud or the vehicle ahead.
[0286] The reception control unit 813 exchanges information such as compatible formats with the communication destination via the communication unit 812, and establishes communication with the communication destination.
[0287] The format conversion unit 814 generates three-dimensional data 832 by performing format conversion or the like on the three-dimensional data 831 received by the data receiving unit 811. Furthermore, when the three-dimensional data 831 is compressed or encoded, the format conversion unit 814 performs decompression or decoding processing.
[0288] The multiple sensors 815 are a group of sensors such as LiDAR, a visible light camera, or an infrared camera that acquires information about the outside of the vehicle, and generate sensor information 833. For example, when the sensor 815 is a laser sensor such as LiDAR, the sensor information 833 is three-dimensional data such as a point cloud (point cloud data). Note that the number of sensors 815 does not need to be multiple.
[0289] The three-dimensional data creation unit 816 generates three-dimensional data 834 from the sensor information 833. The three-dimensional data 834 includes information such as a point cloud, a visible light image, depth information, sensor position information, or velocity information.
[0290] The three-dimensional data synthesis unit 817 synthesizes three-dimensional data 834 created based on the host vehicle's sensor information 833 with three-dimensional data 832 created by a traffic monitoring cloud or a preceding vehicle, etc., to construct three-dimensional data 835 that includes the space in front of the preceding vehicle that cannot be detected by the host vehicle's sensor 815.
[0291] The three-dimensional data storage unit 818 stores the generated three-dimensional data 835 and the like.
[0292] The communication unit 819 communicates with the traffic monitoring cloud or the following vehicle, and transmits data transmission requests, etc. to the traffic monitoring cloud or the following vehicle.
[0293] The transmission control unit 820 exchanges information such as compatible formats with the communication destination and establishes communication with the communication destination via the communication unit 819. In addition, the transmission control unit 820 determines a transmission region, which is the space of the three-dimensional data to be transmitted, based on three-dimensional data construction information of the three-dimensional data 832 generated by the three-dimensional data synthesis unit 817 and a data transmission request from the communication destination.
[0294] Specifically, the transmission control unit 820 determines a transmission region including a space in front of the vehicle that cannot be detected by the sensor of the following vehicle in response to a data transmission request from the traffic monitoring cloud or the following vehicle. The transmission control unit 820 also determines the transmission region by judging whether or not a transmittable space or a transmitted space has been updated based on the three-dimensional data construction information. For example, the transmission control unit 820 determines the region specified in the data transmission request and in which the corresponding three-dimensional data 835 exists as the transmission region. Then, the transmission control unit 820 notifies the format conversion unit 821 of the format supported by the communication destination and the transmission region.
[0295] The format conversion unit 821 converts three-dimensional data 836 of the transmission region, among the three-dimensional data 835 stored in the three-dimensional data storage unit 818, into a format supported by the receiving side, thereby generating three-dimensional data 837. Note that the format conversion unit 821 may reduce the amount of data by compressing or encoding the three-dimensional data 837.
[0296] The data transmission unit 822 transmits the three-dimensional data 837 to the traffic monitoring cloud or the following vehicle. The three-dimensional data 837 includes information such as a point cloud, visible light image, depth information, or sensor position information in front of the vehicle, including blind spots of the following vehicle.
[0297] In the above description, the format conversion units 814 and 821 perform format conversion, but the format conversion does not necessarily have to be performed.
[0298] With this configuration, the three-dimensional data creation device 810 externally acquires three-dimensional data 831 of an area that cannot be detected by the sensor 815 of the host vehicle, and generates three-dimensional data 835 by synthesizing the three-dimensional data 831 with three-dimensional data 834 based on sensor information 833 detected by the sensor 815 of the host vehicle. In this way, the three-dimensional data creation device 810 can generate three-dimensional data of an area that cannot be detected by the sensor 815 of the host vehicle.
[0299] In addition, in response to a data transmission request from a traffic monitoring cloud or a following vehicle, the three-dimensional data creation device 810 can transmit three-dimensional data including the space in front of the vehicle that cannot be detected by the sensors of the following vehicle to the traffic monitoring cloud or the following vehicle, etc.
[0300] (Embodiment 6) In the fifth embodiment, an example has been described in which a client device such as a vehicle transmits three-dimensional data to another vehicle or a server such as a traffic monitoring cloud. In the present embodiment, the client device transmits sensor information obtained by a sensor to the server or another client device.
[0301] First, the configuration of the system according to this embodiment will be described. Fig. 28 is a diagram showing the configuration of a transmission / reception system for a 3D map and sensor information according to this embodiment. This system includes a server 901 and client devices 902A and 902B. When there is no particular need to distinguish between the client devices 902A and 902B, they are also referred to as client devices 902.
[0302] The client device 902 is, for example, an in-vehicle device mounted on a moving object such as a vehicle. The server 901 is, for example, a traffic monitoring cloud or the like, and is capable of communicating with a plurality of client devices 902.
[0303] The server 901 transmits a three-dimensional map composed of a point cloud to the client device 902. Note that the composition of the three-dimensional map is not limited to a point cloud, and may represent other three-dimensional data such as a mesh structure.
[0304] The client device 902 transmits sensor information acquired by the client device 902 to the server 901. The sensor information includes, for example, at least one of LiDAR acquisition information, a visible light image, an infrared image, a depth image, sensor position information, and speed information.
[0305] Data transmitted between the server 901 and the client device 902 may be compressed to reduce data, or may be left uncompressed to maintain data accuracy. When compressing data, a three-dimensional compression method based on an octree structure, for example, can be used for the point cloud. Also, a two-dimensional image compression method can be used for the visible light image, the infrared image, and the depth image. The two-dimensional image compression method is, for example, MPEG-4 AVC or HEVC standardized by MPEG.
[0306] Furthermore, the server 901 transmits the three-dimensional map managed by the server 901 to the client device 902 in response to a transmission request for the three-dimensional map from the client device 902. Note that the server 901 may transmit the three-dimensional map without waiting for a transmission request for the three-dimensional map from the client device 902. For example, the server 901 may broadcast the three-dimensional map to one or more client devices 902 in a predetermined space. Furthermore, the server 901 may transmit a three-dimensional map suitable for the position of the client device 902 at regular intervals to the client device 902 that has received a transmission request once. Furthermore, the server 901 may transmit the three-dimensional map to the client device 902 every time the three-dimensional map managed by the server 901 is updated.
[0307] The client device 902 issues a request for transmitting a three-dimensional map to the server 901. For example, when the client device 902 wants to estimate its own position while driving, the client device 902 transmits a request for transmitting a three-dimensional map to the server 901.
[0308] In the following cases, the client device 902 may issue a transmission request for a three-dimensional map to the server 901. If the three-dimensional map held by the client device 902 is old, the client device 902 may issue a transmission request for a three-dimensional map to the server 901. For example, if a certain period of time has passed since the client device 902 acquired the three-dimensional map, the client device 902 may issue a transmission request for a three-dimensional map to the server 901.
[0309] The client device 902 may issue a request to send a three-dimensional map to the server 901 a certain time before the client device 902 leaves the space shown in the three-dimensional map held by the client device 902. For example, when the client device 902 is within a predetermined distance from the boundary of the space shown in the three-dimensional map held by the client device 902, the client device 902 may issue a request to send a three-dimensional map to the server 901. In addition, when the movement route and movement speed of the client device 902 are known, the time when the client device 902 will leave the space shown in the three-dimensional map held by the client device 902 may be predicted based on these.
[0310] If the error in aligning the three-dimensional data created by the client device 902 from sensor information with the three-dimensional map is equal to or greater than a certain level, the client device 902 may issue a request to the server 901 to transmit the three-dimensional map.
[0311] The client device 902 transmits the sensor information to the server 901 in response to a request to transmit the sensor information transmitted from the server 901. The client device 902 may transmit the sensor information to the server 901 without waiting for a request to transmit the sensor information from the server 901. For example, once the client device 902 receives a request to transmit the sensor information from the server 901, the client device 902 may transmit the sensor information to the server 901 periodically for a certain period of time. In addition, when an error occurs in aligning three-dimensional data created by the client device 902 based on the sensor information with the three-dimensional map obtained from the server 901, the client device 902 may determine that a change may have occurred in the three-dimensional map around the client device 902 and transmit a message to that effect along with the sensor information to the server 901.
[0312] The server 901 issues a transmission request for sensor information to the client device 902. For example, the server 901 receives location information of the client device 902, such as a GPS, from the client device 902. When the server 901 determines, based on the location information of the client device 902, that the client device 902 is approaching a space with little information in the three-dimensional map managed by the server 901, it issues a transmission request for sensor information to the client device 902 in order to generate a new three-dimensional map. The server 901 may also issue a transmission request for sensor information when it is desired to update the three-dimensional map, when it is desired to check road conditions such as during snowfall or a disaster, when it is desired to check traffic congestion conditions, or when it is desired to check incident and accident conditions, etc.
[0313] Furthermore, the client device 902 may set the amount of data of the sensor information to be transmitted to the server 901 depending on the communication state or bandwidth at the time of receiving a request to transmit the sensor information from the server 901. Setting the amount of data of the sensor information to be transmitted to the server 901 means, for example, increasing or decreasing the amount of the data itself or appropriately selecting a compression method.
[0314] 29 is a block diagram showing an example of the configuration of a client device 902. The client device 902 receives a three-dimensional map composed of a point cloud or the like from the server 901, and estimates the self-position of the client device 902 from three-dimensional data created based on sensor information of the client device 902. In addition, the client device 902 transmits the acquired sensor information to the server 901.
[0315] The client device 902 includes a data receiving unit 1011, a communication unit 1012, a reception control unit 1013, a format conversion unit 1014, multiple sensors 1015, a three-dimensional data creation unit 1016, a three-dimensional image processing unit 1017, a three-dimensional data storage unit 1018, a format conversion unit 1019, a communication unit 1020, a transmission control unit 1021, and a data transmission unit 1022.
[0316] The data receiving unit 1011 receives a three-dimensional map 1031 from the server 901. The three-dimensional map 1031 is data including a point cloud such as a WLD or a SWLD. The three-dimensional map 1031 may include either compressed data or uncompressed data.
[0317] The communication unit 1012 communicates with the server 901, and transmits data transmission requests (for example, requests to transmit a three-dimensional map) to the server 901.
[0318] The reception control unit 1013 exchanges information such as compatible formats with the communication destination via the communication unit 1012, and establishes communication with the communication destination.
[0319] The format conversion unit 1014 generates a three-dimensional map 1032 by performing format conversion and the like on the three-dimensional map 1031 received by the data receiving unit 1011. Furthermore, if the three-dimensional map 1031 is compressed or encoded, the format conversion unit 1014 performs decompression or decoding processing. Note that if the three-dimensional map 1031 is uncompressed data, the format conversion unit 1014 does not perform decompression or decoding processing.
[0320] The multiple sensors 1015 are a group of sensors, such as LiDAR, a visible light camera, an infrared camera, or a depth sensor, that acquire information about the outside of the vehicle in which the client device 902 is mounted, and generate sensor information 1033. For example, when the sensor 1015 is a laser sensor such as LiDAR, the sensor information 1033 is three-dimensional data such as a point cloud (point cloud data). Note that the number of sensors 1015 does not need to be multiple.
[0321] The three-dimensional data creation unit 1016 creates three-dimensional data 1034 of the surroundings of the vehicle based on the sensor information 1033. For example, the three-dimensional data creation unit 1016 creates point cloud data with color information of the surroundings of the vehicle using information acquired by the LiDAR and a visible light image acquired by a visible light camera.
[0322] The three-dimensional image processing unit 1017 performs self-position estimation processing of the vehicle using a three-dimensional map 1032 such as a received point cloud and three-dimensional data 1034 of the surroundings of the vehicle generated from the sensor information 1033. The three-dimensional image processing unit 1017 may generate three-dimensional data 1035 of the surroundings of the vehicle by combining the three-dimensional map 1032 and the three-dimensional data 1034, and perform self-position estimation processing using the generated three-dimensional data 1035.
[0323] The three-dimensional data storage unit 1018 stores a three-dimensional map 1032, three-dimensional data 1034, and three-dimensional data 1035, and the like.
[0324] The format conversion unit 1019 generates sensor information 1037 by converting the sensor information 1033 into a format supported by the receiving side. The format conversion unit 1019 may reduce the amount of data by compressing or encoding the sensor information 1037. Furthermore, the format conversion unit 1019 may omit processing when format conversion is not necessary. Furthermore, the format conversion unit 1019 may control the amount of data to be transmitted in accordance with a designated transmission range.
[0325] The communication unit 1020 communicates with the server 901, and receives data transmission requests (sensor information transmission requests) and the like from the server 901.
[0326] The transmission control unit 1021 exchanges information such as compatible formats with the communication destination via the communication unit 1020, and establishes communication.
[0327] The data transmission unit 1022 transmits the sensor information 1037 to the server 901. The sensor information 1037 includes information acquired by a plurality of sensors 1015, such as information acquired by a LiDAR, a luminance image acquired by a visible light camera, an infrared image acquired by an infrared camera, a depth image acquired by a depth sensor, sensor position information, and speed information.
[0328] Next, the configuration of the server 901 will be described. Fig. 30 is a block diagram showing an example of the configuration of the server 901. The server 901 receives sensor information transmitted from the client device 902, and creates three-dimensional data based on the received sensor information. The server 901 updates a three-dimensional map managed by the server 901 using the created three-dimensional data. In addition, the server 901 transmits the updated three-dimensional map to the client device 902 in response to a transmission request for the three-dimensional map from the client device 902.
[0329] The server 901 includes a data receiving unit 1111, a communication unit 1112, a receiving control unit 1113, a format conversion unit 1114, a three-dimensional data creation unit 1116, a three-dimensional data synthesis unit 1117, a three-dimensional data storage unit 1118, a format conversion unit 1119, a communication unit 1120, a transmission control unit 1121, and a data transmission unit 1122.
[0330] The data receiving unit 1111 receives sensor information 1037 from the client device 902. The sensor information 1037 includes, for example, information acquired by a LiDAR, a luminance image acquired by a visible light camera, an infrared image acquired by an infrared camera, a depth image acquired by a depth sensor, sensor position information, and speed information.
[0331] The communication unit 1112 communicates with the client device 902 and transmits data transmission requests (for example, requests to transmit sensor information) to the client device 902 .
[0332] The reception control unit 1113 exchanges information such as compatible formats with the communication destination via the communication unit 1112, and establishes communication.
[0333] When the received sensor information 1037 is compressed or encoded, the format conversion unit 1114 performs decompression or decoding processing to generate the sensor information 1132. Note that, when the sensor information 1037 is non-compressed data, the format conversion unit 1114 does not perform decompression or decoding processing.
[0334] The three-dimensional data creation unit 1116 creates three-dimensional data 1134 of the periphery of the client device 902 based on the sensor information 1132. For example, the three-dimensional data creation unit 1116 creates point cloud data with color information of the periphery of the client device 902 using information acquired by the LiDAR and visible light images acquired by a visible light camera.
[0335] A three-dimensional data synthesis unit 1117 synthesizes three-dimensional data 1134 created based on the sensor information 1132 with a three-dimensional map 1135 managed by the server 901, thereby updating the three-dimensional map 1135.
[0336] The three-dimensional data storage unit 1118 stores a three-dimensional map 1135 and the like.
[0337] The format conversion unit 1119 generates the three-dimensional map 1031 by converting the three-dimensional map 1135 into a format supported by the receiving side. The format conversion unit 1119 may reduce the amount of data by compressing or encoding the three-dimensional map 1135. Furthermore, the format conversion unit 1119 may omit processing when format conversion is not necessary. Furthermore, the format conversion unit 1119 may control the amount of data to be transmitted in accordance with a designated transmission range.
[0338] The communication unit 1120 communicates with the client device 902 and receives a data transmission request (a request to transmit a three-dimensional map) or the like from the client device 902 .
[0339] A transmission control unit 1121 exchanges information such as compatible formats with a communication destination via a communication unit 1120, and establishes communication.
[0340] The data transmission unit 1122 transmits the three-dimensional map 1031 to the client device 902. The three-dimensional map 1031 is data including a point cloud such as a WLD or a SWLD. The three-dimensional map 1031 may include either compressed data or uncompressed data.
[0341] Next, a description will be given of the operation flow of the client device 902. Fig. 31 is a flowchart showing the operation of the client device 902 when acquiring a three-dimensional map.
[0342] First, the client device 902 requests the server 901 to transmit a three-dimensional map (such as a point cloud) (S1001). At this time, the client device 902 may also transmit position information of the client device 902 obtained by GPS or the like, thereby requesting the server 901 to transmit a three-dimensional map related to the position information.
[0343] Next, the client device 902 receives the three-dimensional map from the server 901 (S1002). If the received three-dimensional map is compressed data, the client device 902 decodes the received three-dimensional map to generate an uncompressed three-dimensional map (S1003).
[0344] Next, the client device 902 creates three-dimensional data 1034 of the periphery of the client device 902 from sensor information 1033 obtained by the multiple sensors 1015 (S1004). Next, the client device 902 estimates its own position using the three-dimensional map 1032 received from the server 901 and the three-dimensional data 1034 created from the sensor information 1033 (S1005).
[0345] 32 is a flowchart showing an operation of the client device 902 when transmitting sensor information. First, the client device 902 receives a request to transmit sensor information from the server 901 (S1011). The client device 902 that has received the transmission request transmits sensor information 1037 to the server 901 (S1012). Note that, when the sensor information 1033 includes a plurality of pieces of information obtained by a plurality of sensors 1015, the client device 902 may generate the sensor information 1037 by compressing each piece of information using a compression method suitable for each piece of information.
[0346] Next, the operation flow of the server 901 will be described. Fig. 33 is a flowchart showing the operation when the server 901 acquires sensor information. First, the server 901 requests the client device 902 to transmit sensor information (S1021). Next, the server 901 receives the sensor information 1037 transmitted from the client device 902 in response to the request (S1022). Next, the server 901 creates three-dimensional data 1134 using the received sensor information 1037 (S1023). Next, the server 901 reflects the created three-dimensional data 1134 in the three-dimensional map 1135 (S1024).
[0347] 34 is a flowchart showing the operation of the server 901 when transmitting a three-dimensional map. First, the server 901 receives a request to transmit a three-dimensional map from the client device 902 (S1031). The server 901 that has received the request to transmit the three-dimensional map transmits the three-dimensional map 1031 to the client device 902 (S1032). At this time, the server 901 may extract a three-dimensional map of the vicinity according to the position information of the client device 902, and transmit the extracted three-dimensional map. The server 901 may also compress the three-dimensional map composed of a point cloud using, for example, a compression method based on an octet tree structure, and transmit the compressed three-dimensional map.
[0348] A modification of this embodiment will now be described.
[0349] The server 901 creates three-dimensional data 1134 of the vicinity of the position of the client device 902 using the sensor information 1037 received from the client device 902. Next, the server 901 calculates the difference between the three-dimensional data 1134 and the three-dimensional map 1135 by matching the created three-dimensional data 1134 with a three-dimensional map 1135 of the same area managed by the server 901. If the difference is equal to or greater than a predetermined threshold, the server 901 determines that some abnormality has occurred in the vicinity of the client device 902. For example, when ground subsidence or the like occurs due to a natural disaster such as an earthquake, it is considered that a large difference occurs between the three-dimensional map 1135 managed by the server 901 and the three-dimensional data 1134 created based on the sensor information 1037.
[0350] The sensor information 1037 may include information indicating at least one of the type of sensor, the performance of the sensor, and the model number of the sensor. In addition, a class ID or the like according to the performance of the sensor may be added to the sensor information 1037. For example, when the sensor information 1037 is information acquired by LiDAR, it is possible to assign an identifier to the performance of the sensor, such as class 1 for a sensor that can acquire information with an accuracy of several millimeters, class 2 for a sensor that can acquire information with an accuracy of several centimeters, and class 3 for a sensor that can acquire information with an accuracy of several meters. In addition, the server 901 may estimate the performance information of the sensor from the model number of the client device 902. For example, when the client device 902 is mounted on a vehicle, the server 901 may determine the spec information of the sensor from the model of the vehicle. In this case, the server 901 may acquire information on the model of the vehicle in advance, or the information may be included in the sensor information. In addition, the server 901 may use the acquired sensor information 1037 to switch the degree of correction for the three-dimensional data 1134 created using the sensor information 1037. For example, when the sensor performance is high accuracy (class 1), the server 901 does not perform correction on the three-dimensional data 1134. When the sensor performance is low accuracy (class 3), the server 901 applies correction according to the accuracy of the sensor to the three-dimensional data 1134. For example, the server 901 increases the degree (strength) of correction as the accuracy of the sensor decreases.
[0351] The server 901 may simultaneously issue requests to send sensor information to multiple client devices 902 in a certain space. When the server 901 receives multiple pieces of sensor information from multiple client devices 902, the server 901 does not need to use all of the sensor information to create the three-dimensional data 1134, and may select the sensor information to be used according to the performance of the sensor, for example. For example, when updating the three-dimensional map 1135, the server 901 may select high-precision sensor information (class 1) from the multiple pieces of sensor information received, and create the three-dimensional data 1134 using the selected sensor information.
[0352] The server 901 is not limited to a server such as a traffic monitoring cloud, but may be another client device (mounted in a vehicle). Figure 35 is a diagram showing the system configuration in this case.
[0353] For example, client device 902C issues a request to send sensor information to nearby client device 902A and acquires the sensor information from client device 902A. Client device 902C then creates three-dimensional data using the acquired sensor information of client device 902A and updates the three-dimensional map of client device 902C. This allows client device 902C to generate a three-dimensional map of the space that can be acquired from client device 902A by taking advantage of the performance of client device 902C. For example, it is considered that such a case occurs when client device 902C has high performance.
[0354] In this case, the client device 902A that provided the sensor information is given the right to obtain the high-precision 3D map generated by the client device 902C. The client device 902A receives the high-precision 3D map from the client device 902C in accordance with the right.
[0355] In addition, the client device 902C may issue a request to transmit sensor information to multiple nearby client devices 902 (client device 902A and client device 902B). If the sensor of the client device 902A or the client device 902B is high performance, the client device 902C can create three-dimensional data using the sensor information obtained by the high performance sensor.
[0356] 36 is a block diagram showing the functional configuration of a server 901 and a client device 902. The server 901 includes, for example, a 3D map compression / decoding processing unit 1201 that compresses and decodes a 3D map, and a sensor information compression / decoding processing unit 1202 that compresses and decodes sensor information.
[0357] The client device 902 includes a three-dimensional map decoding processing unit 1211 and a sensor information compression processing unit 1212. The three-dimensional map decoding processing unit 1211 receives encoded data of the compressed three-dimensional map and decodes the encoded data to obtain the three-dimensional map. The sensor information compression processing unit 1212 compresses the sensor information itself instead of the three-dimensional data created from the acquired sensor information, and transmits the encoded data of the compressed sensor information to the server 901. With this configuration, the client device 902 only needs to internally hold a processing unit (device or LSI) that performs processing to decode the three-dimensional map (point cloud, etc.), and does not need to internally hold a processing unit that performs processing to compress the three-dimensional data of the three-dimensional map (point cloud, etc.). This makes it possible to reduce the cost and power consumption of the client device 902.
[0358] As described above, the client device 902 according to this embodiment is mounted on a moving body, and creates three-dimensional data 1034 of the surroundings of the moving body from sensor information 1033 indicating the surrounding conditions of the moving body, obtained by a sensor 1015 mounted on the moving body. The client device 902 estimates the self-position of the moving body using the created three-dimensional data 1034. The client device 902 transmits the acquired sensor information 1033 to the server 901 or another moving body 902.
[0359] According to this, the client device 902 transmits the sensor information 1033 to the server 901 or the like. This may reduce the amount of data to be transmitted compared to the case of transmitting three-dimensional data. Also, since the client device 902 does not need to perform processing such as compression or encoding of the three-dimensional data, the amount of processing by the client device 902 can be reduced. Therefore, the client device 902 can reduce the amount of data to be transmitted or simplify the device configuration.
[0360] Moreover, the client device 902 further transmits a request to the server 901 to transmit a three-dimensional map, and receives a three-dimensional map 1031 from the server 901. The client device 902 estimates its own location using the three-dimensional data 1034 and the three-dimensional map 1032.
[0361] The sensor information 1033 includes at least one of information obtained by a laser sensor, a luminance image, an infrared image, a depth image, sensor position information, and sensor speed information.
[0362] Furthermore, the sensor information 1033 includes information indicating the performance of the sensor.
[0363] Furthermore, the client device 902 encodes or compresses the sensor information 1033, and transmits the encoded or compressed sensor information 1037 to the server 901 or another mobile body 902. This allows the client device 902 to reduce the amount of data to be transmitted.
[0364] For example, the client device 902 includes a processor and a memory, and the processor uses the memory to perform the above-mentioned processing.
[0365] Furthermore, the server 901 according to this embodiment is capable of communicating with a client device 902 mounted on the mobile object, and receives sensor information 1037 indicating the surrounding conditions of the mobile object, obtained by a sensor 1015 mounted on the mobile object, from the client device 902. The server 901 creates three-dimensional data 1134 of the surroundings of the mobile object from the received sensor information 1037.
[0366] According to this, the server 901 creates three-dimensional data 1134 using the sensor information 1037 transmitted from the client device 902. This may reduce the amount of data transmitted compared to when the client device 902 transmits three-dimensional data. In addition, since the client device 902 does not need to perform processing such as compression or encoding of the three-dimensional data, the amount of processing by the client device 902 can be reduced. Therefore, the server 901 can reduce the amount of data transmitted or simplify the device configuration.
[0367] Moreover, the server 901 further transmits a request to the client device 902 to transmit the sensor information.
[0368] Furthermore, the server 901 further updates a three-dimensional map 1135 using the created three-dimensional data 1134 , and transmits the three-dimensional map 1135 to the client device 902 in response to a transmission request for the three-dimensional map 1135 from the client device 902 .
[0369] The sensor information 1037 includes at least one of information obtained by a laser sensor, a luminance image, an infrared image, a depth image, sensor position information, and sensor speed information.
[0370] Additionally, the sensor information 1037 includes information indicating the performance of the sensor.
[0371] Moreover, the server 901 further corrects the three-dimensional data in accordance with the performance of the sensor, thereby enabling the three-dimensional data creation method to improve the quality of the three-dimensional data.
[0372] Furthermore, in receiving the sensor information, the server 901 receives a plurality of pieces of sensor information 1037 from a plurality of client devices 902, and selects the sensor information 1037 to be used for creating the three-dimensional data 1134 based on a plurality of pieces of information indicating the performance of the sensors included in the plurality of pieces of sensor information 1037. This allows the server 901 to improve the quality of the three-dimensional data 1134.
[0373] Furthermore, the server 901 decodes or expands the received sensor information 1037, and creates three-dimensional data 1134 from the decoded or expanded sensor information 1132. This allows the server 901 to reduce the amount of data to be transmitted.
[0374] For example, the server 901 includes a processor and a memory, and the processor uses the memory to perform the above-mentioned processing.
[0375] (Embodiment 7) In this embodiment, a method for encoding and decoding three-dimensional data using inter prediction processing will be described.
[0376] Fig. 37 is a block diagram of a three-dimensional data encoding device 1300 according to this embodiment. The three-dimensional data encoding device 1300 generates an encoded bit stream (hereinafter, simply referred to as a bit stream) which is an encoded signal by encoding three-dimensional data. As shown in Fig. 37, the three-dimensional data encoding device 1300 includes a division unit 1301, a subtraction unit 1302, a transformation unit 1303, a quantization unit 1304, an inverse quantization unit 1305, an inverse transformation unit 1306, an addition unit 1307, a reference volume memory 1308, an intra prediction unit 1309, a reference space memory 1310, an inter prediction unit 1311, a prediction control unit 1312, and an entropy encoding unit 1313.
[0377] The dividing unit 1301 divides each space (SPC) included in the three-dimensional data into a plurality of volumes (VLM), which are coding units. The dividing unit 1301 also converts the voxels in each volume into an octree representation. The dividing unit 1301 may make the space and the volume the same size and convert the space into an octree representation. The dividing unit 1301 may also add information required for the octree representation (depth information, etc.) to the header of the bitstream, etc.
[0378] The subtraction unit 1302 calculates a difference between the volume (volume to be coded) output from the division unit 1301 and a prediction volume generated by intra prediction or inter prediction, which will be described later, and outputs the calculated difference as a prediction residual to the conversion unit 1303. Fig. 38 is a diagram showing an example of calculation of a prediction residual. Note that the bit strings of the volume to be coded and the prediction volume shown here are, for example, position information indicating the positions of three-dimensional points (for example, a point cloud) included in the volume.
[0379] The octree representation and the voxel scan order will be described below. The volume is converted (octreeized) into an octree structure and then encoded. The octree structure is composed of nodes and leaves. Each node has eight nodes or leaves, and each leaf has voxel (VXL) information. Fig. 39 is a diagram showing an example of the structure of a volume including multiple voxels. Fig. 40 is a diagram showing an example of the volume shown in Fig. 39 converted into an octree structure. Here, among the leaves shown in Fig. 40, leaves 1, 2, and 3 respectively represent voxels VXL1, VXL2, and VXL3 shown in Fig. 39, and represent a VXL including a point cloud (hereinafter, effective VXL).
[0380] An octree is represented by a binary sequence of, for example, 0 and 1. For example, if a node or a valid VXL is set to value 1 and the rest to value 0, then the binary sequence shown in FIG. 40 is assigned to each node and leaf. Then, this binary sequence is scanned according to the scan order of breadth-first or depth-first. For example, when breadth-first scanning is performed, the binary sequence shown in A of FIG. 41 is obtained. When depth-first scanning is performed, the binary sequence shown in B of FIG. 41 is obtained. The binary sequence obtained by this scan is coded by entropy coding to reduce the amount of information.
[0381] Next, the depth information in the octree representation will be explained. The depth in the octree representation is used to control the granularity of the point cloud information contained in the volume to be retained. If the depth is set large, the point cloud information can be reproduced to a finer level, but the amount of data required to represent the nodes and leaves increases. Conversely, if the depth is set small, the amount of data decreases, but since multiple point cloud information with different positions and colors is considered to be in the same position and with the same color, the information contained in the original point cloud information will be lost.
[0382] For example, FIG. 42 is a diagram showing an example in which the octree with depth=2 shown in FIG. 40 is expressed as an octree with depth=1. The octree shown in FIG. 42 has a smaller amount of data than the octree shown in FIG. 40. In other words, the octree shown in FIG. 42 has a smaller number of bits after binary string conversion than the octree shown in FIG. 42. Here, leaf 1 and leaf 2 shown in FIG. 40 are expressed as leaf 1 shown in FIG. 41. In other words, the information that leaf 1 and leaf 2 shown in FIG. 40 were in different positions is lost.
[0383] FIG. 43 is a diagram showing a volume corresponding to the octree shown in FIG. 42. VXL1 and VXL2 shown in FIG. 39 correspond to VXL12 shown in FIG. 43. In this case, the three-dimensional data encoding device 1300 generates color information of VXL12 shown in FIG. 43 from color information of VXL1 and VXL2 shown in FIG. 39. For example, the three-dimensional data encoding device 1300 calculates an average value, a median value, or a weighted average value of the color information of VXL1 and VXL2 as the color information of VXL12. In this way, the three-dimensional data encoding device 1300 may control the reduction of the data amount by changing the depth of the octree.
[0384] The three-dimensional data encoding device 1300 may set the depth information of the octree in any unit of world, space, or volume. In addition, in this case, the three-dimensional data encoding device 1300 may add the depth information to the header information of the world, the header information of the space, or the header information of the volume. Also, the same value may be used as the depth information for all worlds, spaces, and volumes of different times. In this case, the three-dimensional data encoding device 1300 may add the depth information to the header information that manages the worlds of all times.
[0385] When the voxel includes color information, the transform unit 1303 applies a frequency transform such as an orthogonal transform to the prediction residual of the color information of the voxel in the volume. For example, the transform unit 1303 creates a one-dimensional array by scanning the prediction residual in a certain scan order. After that, the transform unit 1303 applies a one-dimensional orthogonal transform to the created one-dimensional array to transform the one-dimensional array into the frequency domain. As a result, when the values of the prediction residuals in the volume are close to each other, the values of the low-frequency components become large and the values of the high-frequency components become small. Therefore, the amount of code can be reduced more efficiently in the quantization unit 1304.
[0386] Furthermore, the transform unit 1303 may use a two-dimensional or higher-dimensional orthogonal transform instead of a one-dimensional one. For example, the transform unit 1303 maps the prediction residuals to a two-dimensional array in a certain scan order, and applies a two-dimensional orthogonal transform to the obtained two-dimensional array. Furthermore, the transform unit 1303 may select an orthogonal transform method to be used from a plurality of orthogonal transform methods. In this case, the three-dimensional data encoding device 1300 adds information indicating which orthogonal transform method has been used to the bit stream. Furthermore, the transform unit 1303 may select an orthogonal transform method to be used from a plurality of orthogonal transform methods of different dimensions. In this case, the three-dimensional data encoding device 1300 adds information indicating which orthogonal transform method has been used to the bit stream.
[0387] For example, the conversion unit 1303 matches the scan order of the prediction residuals to the scan order (such as breadth-first or depth-first) in the octree in the volume. This eliminates the need to add information indicating the scan order of the prediction residuals to the bitstream, thereby reducing overhead. The conversion unit 1303 may also apply a scan order different from the scan order of the octree. In this case, the three-dimensional data encoding device 1300 adds information indicating the scan order of the prediction residuals to the bitstream. This allows the three-dimensional data encoding device 1300 to efficiently encode the prediction residuals. The three-dimensional data encoding device 1300 may also add information (such as a flag) indicating whether or not to apply the scan order of the octree to the bitstream, and add information indicating the scan order of the prediction residuals to the bitstream when the scan order of the octree is not applied.
[0388] The conversion unit 1303 may convert not only the prediction residual of the color information but also other attribute information of the voxels. For example, the conversion unit 1303 may convert and encode information such as reflectance obtained when the point cloud is acquired by LiDAR or the like.
[0389] If the space does not have attribute information such as color information, the conversion unit 1303 may skip the process. Furthermore, the three-dimensional data encoding device 1300 may add information (a flag) indicating whether or not the process of the conversion unit 1303 is to be skipped to the bitstream.
[0390] The quantization unit 1304 generates quantization coefficients by quantizing the frequency components of the prediction residuals generated by the conversion unit 1303 using the quantization control parameters. This reduces the amount of information. The generated quantization coefficients are output to the entropy coding unit 1313. The quantization unit 1304 may control the quantization control parameters on a world-by-world, space-by-space, or volume-by-volume basis. In this case, the three-dimensional data coding device 1300 adds the quantization control parameters to the respective header information or the like. The quantization unit 1304 may also control quantization by changing the weight for each frequency component of the prediction residual. For example, the quantization unit 1304 may finely quantize low-frequency components and coarsely quantize high-frequency components. In this case, the three-dimensional data coding device 1300 may add a parameter representing the weight of each frequency component to the header.
[0391] If the space does not have attribute information such as color information, the quantization unit 1304 may skip the process. Furthermore, the three-dimensional data encoding device 1300 may add information (a flag) indicating whether or not the process of the quantization unit 1304 is to be skipped to the bitstream.
[0392] The inverse quantization unit 1305 uses the quantization control parameter to perform inverse quantization on the quantized coefficients generated by the quantization unit 1304 to generate inverse quantized coefficients of the prediction residuals, and outputs the generated inverse quantized coefficients to the inverse transform unit 1306.
[0393] The inverse transform unit 1306 generates a prediction residual after inverse transform application by applying inverse transform to the inverse quantized coefficients generated by the inverse quantization unit 1305. This prediction residual after inverse transform application is a prediction residual generated after quantization, and therefore does not need to completely match the prediction residual output by the transform unit 1303.
[0394] The adder 1307 generates a reconstructed volume by adding the prediction residual after inverse transformation applied generated by the inverse transformer 1306 and a prediction volume generated by intra prediction or inter prediction, which will be described later, and used to generate the prediction residual before quantization. This reconstructed volume is stored in a reference volume memory 1308 or a reference space memory 1310.
[0395] The intra prediction unit 1309 generates a predicted volume of the volume to be encoded using attribute information of the adjacent volume stored in the reference volume memory 1308. The attribute information includes color information or reflectance of a voxel. The intra prediction unit 1309 generates a predicted value of the color information or reflectance of the volume to be encoded.
[0396] FIG. 44 is a diagram for explaining the operation of the intra prediction unit 1309. For example, the intra prediction unit 1309 generates a prediction volume of the encoding target volume (volume idx=3) shown in FIG. 44 from an adjacent volume (volume idx=0). Here, the volume idx is identifier information added to a volume in a space, and a different value is assigned to each volume. The order of allocation of the volumes idx may be the same as the encoding order, or may be different from the encoding order. For example, the intra prediction unit 1309 uses an average value of color information of voxels included in the adjacent volume, volume idx=0, as a predicted value of color information of the encoding target volume shown in FIG. 44. In this case, a prediction residual is generated by subtracting the predicted value of color information from the color information of each voxel included in the encoding target volume. The processing of the conversion unit 1303 and subsequent processes is performed on this prediction residual. In this case, the three-dimensional data encoding device 1300 adds adjacent volume information and prediction mode information to the bit stream. Here, the adjacent volume information is information indicating the adjacent volume used for prediction, for example, the volume idx of the adjacent volume used for prediction. Also, the prediction mode information indicates the mode used for generating the prediction volume. The mode is, for example, an average mode that generates a predicted value from the average value of the voxels in the adjacent volume, or an intermediate value mode that generates a predicted value from the intermediate value of the voxels in the adjacent volume.
[0397] The intra prediction unit 1309 may generate a prediction volume from a plurality of adjacent volumes. For example, in the configuration shown in Fig. 44, the intra prediction unit 1309 generates a prediction volume 0 from the volume of volume idx=0, and generates a prediction volume 1 from the volume of volume idx=1. Then, the intra prediction unit 1309 generates an average of the prediction volume 0 and the prediction volume 1 as a final prediction volume. In this case, the three-dimensional data encoding device 1300 may add a plurality of volumes idx of the plurality of volumes used to generate the prediction volume to the bitstream.
[0398] Fig. 45 is a diagram illustrating an inter prediction process according to the present embodiment. The inter prediction unit 1311 performs coding (inter prediction) on a space (SPC) at a certain time T_Cur by using a coded space at a different time T_LX. In this case, the inter prediction unit 1311 performs coding by applying rotation and translation processing to the coded space at the different time T_LX.
[0399] Furthermore, the three-dimensional data encoding device 1300 adds RT information related to the rotation and translation processing applied to the space at a different time T_LX to the bit stream. The different time T_LX is, for example, a time T_L0 prior to the certain time T_Cur. At this time, the three-dimensional data encoding device 1300 may add RT information RT_L0 related to the rotation and translation processing applied to the space at the time T_L0 to the bit stream.
[0400] Alternatively, the different time T_LX is, for example, a time T_L1 that is later than the certain time T_Cur. At this time, the three-dimensional data encoding device 1300 may add RT information RT_L1 related to the rotation and translation processing applied to the space of the time T_L1 to the bitstream.
[0401] Alternatively, the inter prediction unit 1311 performs encoding (bi-prediction) by referring to both spaces at different times T_L0 and T_L1. In this case, the 3D data encoding device 1300 may add both pieces of RT information RT_L0 and RT_L1 related to the rotation and translation applied to the respective spaces to the bitstream.
[0402] In the above, T_L0 is a time before T_Cur, and T_L1 is a time after T_Cur, but this is not necessarily limited to this. For example, T_L0 and T_L1 may both be times before T_Cur. Or, T_L0 and T_L1 may both be times after T_Cur.
[0403] Furthermore, when the three-dimensional data encoding device 1300 performs encoding by referring to a plurality of spaces at different times, the three-dimensional data encoding device 1300 may add RT information related to the rotation and translation applied to each space to the bit stream. For example, the three-dimensional data encoding device 1300 manages a plurality of encoded spaces to be referenced in two reference lists (L0 list and L1 list). If the first reference space in the L0 list is L0R0, the second reference space in the L0 list is L0R1, the first reference space in the L1 list is L1R0, and the second reference space in the L1 list is L1R1, the three-dimensional data encoding device 1300 adds RT information RT_L0R0 of L0R0, RT information RT_L0R1 of L0R1, RT information RT_L1R0 of L1R0, and RT information RT_L1R1 of L1R1 to the bit stream. For example, the three-dimensional data encoding device 1300 adds these pieces of RT information to the header of the bit stream, etc.
[0404] Furthermore, when the three-dimensional data encoding device 1300 performs encoding with reference to a plurality of reference spaces at different times, it determines whether or not rotation and translation are applied for each reference space. At that time, the three-dimensional data encoding device 1300 may add information (such as an RT application flag) indicating whether or not rotation and translation are applied for each reference space to header information or the like of the bit stream. For example, the three-dimensional data encoding device 1300 calculates RT information and an ICP error value using an ICP (Interactive Closest Point) algorithm for each reference space referenced from the encoding target space. If the ICP error value is equal to or less than a predetermined fixed value, the three-dimensional data encoding device 1300 determines that rotation and translation are not required and sets the RT application flag to OFF. On the other hand, if the ICP error value is greater than the above-mentioned fixed value, the three-dimensional data encoding device 1300 sets the RT application flag to ON and adds RT information to the bit stream.
[0405] Fig. 46 is a diagram showing an example of syntax for adding RT information and an RT application flag to a header. The number of bits to be allocated to each syntax may be determined within the range that the syntax can take. For example, when the number of reference spaces included in the reference list L0 is eight, 3 bits may be allocated to MaxRefSpc_l0. The number of bits to be allocated may be variable according to the value that each syntax can take, or may be fixed regardless of the value that each syntax can take. When the number of bits to be allocated is fixed, the three-dimensional data encoding device 1300 may add the fixed number of bits to another header information.
[0406] Here, MaxRefSpc_l0 shown in Fig. 46 indicates the number of reference spaces included in the reference list L0. RT_flag_l0[i] is the RT application flag of the reference space i in the reference list L0. When RT_flag_l0[i] is 1, rotation and translation are applied to the reference space i. When RT_flag_l0[i] is 0, rotation and translation are not applied to the reference space i.
[0407] R_l0[i] and T_l0[i] are RT information of the reference space i in the reference list L0. R_l0[i] is rotation information of the reference space i in the reference list L0. The rotation information indicates the content of the applied rotation process, for example, a rotation matrix or a quaternion. T_l0[i] is translation information of the reference space i in the reference list L0. The translation information indicates the content of the applied translation process, for example, a translation vector.
[0408] MaxRefSpc_l1 indicates the number of reference spaces included in reference list L1. RT_flag_l1[i] is the RT application flag for reference space i in reference list L1. When RT_flag_l1[i] is 1, rotation and translation are applied to reference space i. When RT_flag_l1[i] is 0, rotation and translation are not applied to reference space i.
[0409] R_l1[i] and T_l1[i] are RT information of the reference space i in the reference list L1. R_l1[i] is rotation information of the reference space i in the reference list L1. The rotation information indicates the content of the applied rotation process, for example, a rotation matrix or a quaternion. T_l1[i] is translation information of the reference space i in the reference list L1. The translation information indicates the content of the applied translation process, for example, a translation vector.
[0410] The inter prediction unit 1311 generates a prediction volume of the volume to be encoded using information of the reference space that has already been encoded and that is stored in the reference space memory 1310. As described above, before generating a prediction volume of the volume to be encoded, the inter prediction unit 1311 obtains RT information using an ICP (Interactive Closest Point) algorithm in the space to be encoded and the reference space in order to approximate the overall positional relationship between the space to be encoded and the reference space. Then, the inter prediction unit 1311 obtains a reference space B by applying a rotation and translation process to the reference space using the obtained RT information. Then, the inter prediction unit 1311 generates a prediction volume of the volume to be encoded in the space to be encoded using information in the reference space B. Here, the three-dimensional data encoding device 1300 adds the RT information used to obtain the reference space B to header information or the like of the space to be encoded.
[0411] In this way, the inter prediction unit 1311 can improve the accuracy of the predicted volume by applying rotation and translation processing to the reference space to bring the overall positional relationship between the encoding target space and the reference space closer, and then generating a predicted volume using information on the reference space. In addition, since the prediction residual can be suppressed, the amount of coding can be reduced. Note that, although an example of performing ICP using the encoding target space and the reference space has been shown here, this is not necessarily limited to this. For example, in order to reduce the amount of processing, the inter prediction unit 1311 may obtain RT information by performing ICP using at least one of the encoding target space in which the number of voxels or point clouds has been thinned out and the reference space in which the number of voxels or point clouds has been thinned out.
[0412] Furthermore, if an ICP error value obtained as a result of the ICP is smaller than a first threshold value determined in advance, that is, for example, if the positional relationship between the encoding target space and the reference space is close, the inter prediction unit 1311 may determine that rotation and translation processing is not necessary, and may not perform rotation and translation. In this case, the three-dimensional data encoding device 1300 may suppress overhead by not adding RT information to the bitstream.
[0413] In addition, when the ICP error value is greater than a second threshold value determined in advance, the inter prediction unit 1311 may determine that the shape change between the spaces is large, and apply intra prediction to all volumes of the encoding target space. Hereinafter, the space to which intra prediction is applied is called an intra space. In addition, the second threshold value is a value greater than the first threshold value. In addition, any method may be applied as long as it is a method of obtaining RT information from two voxel sets or two point cloud sets, without being limited to ICP.
[0414] Furthermore, when the three-dimensional data includes attribute information such as shape or color, the inter prediction unit 1311 searches for a volume in the reference space that has attribute information such as shape or color closest to that of the volume to be encoded as a prediction volume of the volume to be encoded in the space to be encoded. Moreover, this reference space is, for example, a reference space after the above-mentioned rotation and translation processing is performed. The inter prediction unit 1311 generates a prediction volume from a volume (reference volume) obtained by the search. FIG. 47 is a diagram for explaining the operation of generating a prediction volume. When the inter prediction unit 1311 encodes the volume to be encoded (volume idx=0) shown in FIG. 47 using inter prediction, the inter prediction unit 1311 searches for a volume with the smallest prediction residual, which is the difference between the volume to be encoded and the reference volume, while sequentially scanning the reference volumes in the reference space. The inter prediction unit 1311 selects the volume with the smallest prediction residual as the prediction volume. The prediction residual between the volume to be encoded and the prediction volume is encoded by the processing after the conversion unit 1303. Here, the prediction residual is the difference between the attribute information of the volume to be encoded and the attribute information of the prediction volume. Furthermore, the three-dimensional data encoding device 1300 adds the volume idx of the reference volume in the reference space referred to as the prediction volume to the header of the bit stream or the like.
[0415] In the example shown in Fig. 47, the reference volume of volume idx=4 in the reference space L0R0 is selected as a prediction volume of the volume to be encoded. Then, the prediction residual between the volume to be encoded and the reference volume, and the reference volume idx=4 are encoded and added to the bitstream.
[0416] Although an example of generating a predicted volume for attribute information has been described here, a similar process may be performed for a predicted volume for position information.
[0417] The prediction control unit 1312 controls whether to use intra prediction or inter prediction to encode the volume to be encoded. Here, a mode including intra prediction and inter prediction is called a prediction mode. For example, the prediction control unit 1312 calculates a prediction residual when the volume to be encoded is predicted by intra prediction and a prediction residual when the volume to be encoded is predicted by inter prediction as evaluation values, and selects a prediction mode with a smaller evaluation value. Note that the prediction control unit 1312 may calculate an actual code amount by applying orthogonal transformation, quantization, and entropy coding to the prediction residual of intra prediction and the prediction residual of inter prediction, respectively, and select a prediction mode using the calculated code amount as an evaluation value. Also, overhead information other than the prediction residual (such as reference volume idx information) may be added to the evaluation value. Also, the prediction control unit 1312 may always select intra prediction when it is previously determined that the space to be encoded is to be encoded by intra space.
[0418] The entropy coding unit 1313 generates a coded signal (coded bit stream) by variable-length coding the quantized coefficients that are input from the quantization unit 1304. Specifically, the entropy coding unit 1313, for example, binarizes the quantized coefficients and arithmetically codes the obtained binary signal.
[0419] Next, a three-dimensional data decoding device that decodes the encoded signal generated by the three-dimensional data encoding device 1300 will be described. Fig. 48 is a block diagram of a three-dimensional data decoding device 1400 according to this embodiment. This three-dimensional data decoding device 1400 includes an entropy decoding unit 1401, an inverse quantization unit 1402, an inverse transformation unit 1403, an addition unit 1404, a reference volume memory 1405, an intra prediction unit 1406, a reference space memory 1407, an inter prediction unit 1408, and a prediction control unit 1409.
[0420] The entropy decoding unit 1401 performs variable length decoding on the coded signal (coded bit stream). For example, the entropy decoding unit 1401 arithmetically decodes the coded signal to generate a binary signal, and generates a quantization coefficient from the generated binary signal.
[0421] The inverse quantization unit 1402 inversely quantizes the quantized coefficients input from the entropy decoding unit 1401 using a quantization parameter added to the bit stream or the like, thereby generating inverse quantized coefficients.
[0422] The inverse transform unit 1403 generates a prediction residual by inverse transforming the inverse quantized coefficients input from the inverse quantization unit 1402. For example, the inverse transform unit 1403 generates a prediction residual by inverse orthogonal transforming the inverse quantized coefficients based on information added to the bitstream.
[0423] The adder 1404 generates a reconstructed volume by adding the prediction residual generated by the inverse transformer 1403 and a prediction volume generated by intra prediction or inter prediction. This reconstructed volume is output as decoded three-dimensional data, and is also stored in a reference volume memory 1405 or a reference space memory 1407.
[0424] The intra prediction unit 1406 generates a predicted volume by intra prediction using a reference volume in the reference volume memory 1405 and information added to the bit stream. Specifically, the intra prediction unit 1406 acquires adjacent volume information (e.g., volume idx) and prediction mode information added to the bit stream, and generates a predicted volume in a mode indicated by the prediction mode information using adjacent volumes indicated by the adjacent volume information. Details of these processes are similar to the processes by the intra prediction unit 1309 described above, except that information added to the bit stream is used.
[0425] The inter prediction unit 1408 generates a prediction volume by inter prediction using the reference space in the reference space memory 1407 and information added to the bit stream. Specifically, the inter prediction unit 1408 applies rotation and translation processing to the reference space using RT information for each reference space added to the bit stream, and generates a prediction volume using the reference space after application. Note that, if an RT application flag for each reference space exists in the bit stream, the inter prediction unit 1408 applies rotation and translation processing to the reference space according to the RT application flag. Note that details of these processes are similar to the processes by the inter prediction unit 1311 described above, except that information added to the bit stream is used.
[0426] The prediction control unit 1409 controls whether the volume to be decoded is decoded by intra prediction or inter prediction. For example, the prediction control unit 1409 selects intra prediction or inter prediction according to information indicating a prediction mode to be used, which is added to a bit stream. Note that the prediction control unit 1409 may always select intra prediction when it is previously determined that the space to be decoded is decoded by intra space.
[0427] A modified example of this embodiment will be described below. In this embodiment, an example in which rotation and translation are applied in units of spaces has been described, but rotation and translation may be applied in smaller units. For example, the three-dimensional data encoding device 1300 may divide a space into subspaces and apply rotation and translation in units of subspaces. In this case, the three-dimensional data encoding device 1300 generates RT information for each subspace and adds the generated RT information to a header of a bit stream or the like. The three-dimensional data encoding device 1300 may also apply rotation and translation in units of volumes, which are encoding units. In this case, the three-dimensional data encoding device 1300 generates RT information in units of encoding volumes and adds the generated RT information to a header of a bit stream or the like. Furthermore, the above may be combined. That is, the three-dimensional data encoding device 1300 may apply rotation and translation in large units, and then apply rotation and translation in small units. For example, the three-dimensional data encoding device 1300 may apply rotation and translation in units of spaces, and apply different rotations and translations to each of a plurality of volumes included in the obtained space.
[0428] In addition, in the present embodiment, an example of applying rotation and translation to the reference space has been described, but this is not necessarily limited to this. For example, the three-dimensional data encoding device 1300 may change the size of the three-dimensional data by applying a scale process. Also, the three-dimensional data encoding device 1300 may apply any one or two of rotation, translation, and scale. Also, when applying processes in different units in multiple stages as described above, the type of process applied to each unit may be different. For example, rotation and translation may be applied in space units, and translation may be applied in volume units.
[0429] These modifications can also be applied to the three-dimensional data decoding device 1400 in the same manner.
[0430] As described above, the three-dimensional data encoding device 1300 according to this embodiment performs the following processes.
[0431] First, the three-dimensional data encoding device 1300 generates predicted position information (e.g., predicted volume) using position information of three-dimensional points included in reference three-dimensional data (e.g., reference space) at a time different from that of the target three-dimensional data (e.g., encoding target space) (S1301). Specifically, the three-dimensional data encoding device 1300 generates predicted position information by applying rotation and translation processing to the position information of three-dimensional points included in the reference three-dimensional data.
[0432] The three-dimensional data encoding device 1300 may perform the rotation and translation processing in a first unit (e.g., space) and generate the predicted position information in a second unit (e.g., volume) that is smaller than the first unit. For example, the three-dimensional data encoding device 1300 searches for a volume that has the smallest difference in position information from the encoding target volume included in the encoding target space among a plurality of volumes included in the reference space after the rotation and translation processing, and uses the obtained volume as the predicted volume. The three-dimensional data encoding device 1300 may perform the rotation and translation processing and the generation of the predicted position information in the same unit.
[0433] In addition, the three-dimensional data encoding device 1300 may generate predicted position information by applying a first rotation and translation process in a first unit (e.g., space) to position information of three-dimensional points included in the reference three-dimensional data, and applying a second rotation and translation process in a second unit (e.g., volume) finer than the first unit to the position information of the three-dimensional points obtained by the first rotation and translation process.
[0434] Here, the position information and the predicted position information of the three-dimensional point are expressed in an octet tree structure, for example, as shown in Fig. 41. For example, the position information and the predicted position information of the three-dimensional point are expressed in a scan order that prioritizes the width among the depth and width in the octet tree structure. Alternatively, the position information and the predicted position information of the three-dimensional point are expressed in a scan order that prioritizes the depth among the depth and width in the octet tree structure.
[0435] Also, as shown in FIG. 46, the three-dimensional data encoding device 1300 encodes an RT application flag indicating whether or not rotation and translation processing is applied to position information of a three-dimensional point included in the reference three-dimensional data. That is, the three-dimensional data encoding device 1300 generates an encoding signal (encoded bit stream) including an RT application flag. Also, the three-dimensional data encoding device 1300 encodes RT information indicating the contents of the rotation and translation processing. That is, the three-dimensional data encoding device 1300 generates an encoding signal (encoded bit stream) including the RT information. Note that the three-dimensional data encoding device 1300 may encode the RT information when the RT application flag indicates that the rotation and translation processing is applied, and may not encode the RT information when the RT application flag indicates that the rotation and translation processing is not applied.
[0436] The three-dimensional data includes, for example, position information of the three-dimensional points and attribute information (such as color information) of each three-dimensional point. The three-dimensional data encoding device 1300 generates predicted attribute information using the attribute information of the three-dimensional points included in the reference three-dimensional data (S1302).
[0437] Next, the three-dimensional data encoding device 1300 encodes the position information of the three-dimensional point included in the target three-dimensional data using the predicted position information. For example, the three-dimensional data encoding device 1300 calculates differential position information, which is the difference between the position information of the three-dimensional point included in the target three-dimensional data and the predicted position information, as shown in Fig. 38 (S1303).
[0438] Furthermore, the three-dimensional data encoding device 1300 encodes attribute information of a three-dimensional point included in the target three-dimensional data using the predicted attribute information. For example, the three-dimensional data encoding device 1300 calculates differential attribute information that is a difference between the attribute information of a three-dimensional point included in the target three-dimensional data and the predicted attribute information (S1304). Next, the three-dimensional data encoding device 1300 converts and quantizes the calculated differential attribute information (S1305).
[0439] Finally, the three-dimensional data encoding device 1300 encodes (for example, entropy encodes) the difference position information and the quantized difference attribute information (S1306). That is, the three-dimensional data encoding device 1300 generates an encoded signal (encoded bit stream) including the difference position information and the difference attribute information.
[0440] In addition, when the three-dimensional data does not include attribute information, the three-dimensional data encoding device 1300 may not perform steps S1302, S1304, and S1305. In addition, the three-dimensional data encoding device 1300 may only encode the position information of the three-dimensional point and encode the attribute information of the three-dimensional point.
[0441] Also, the order of processing shown in Fig. 49 is an example and is not limited to this. For example, the processing for the position information (S1301, S1303) and the processing for the attribute information (S1302, S1304, S1305) are independent of each other, so they may be performed in any order, or some of them may be processed in parallel.
[0442] As described above, the three-dimensional data encoding device 1300 according to the present embodiment generates predicted position information using position information of three-dimensional points included in reference three-dimensional data at a time different from that of the target three-dimensional data, and encodes differential position information that is the difference between the position information of the three-dimensional points included in the target three-dimensional data and the predicted position information. This makes it possible to reduce the amount of data in the encoded signal, thereby improving encoding efficiency.
[0443] In addition, the three-dimensional data encoding device 1300 in this embodiment generates predicted attribute information using attribute information of three-dimensional points included in the reference three-dimensional data, and encodes difference attribute information that is the difference between the attribute information of three-dimensional points included in the target three-dimensional data and the predicted attribute information. This makes it possible to reduce the data amount of the encoded signal, thereby improving encoding efficiency.
[0444] For example, the three-dimensional data encoding device 1300 includes a processor and a memory, and the processor uses the memory to perform the above-mentioned processing.
[0445] FIG. 48 is a flowchart of inter prediction processing by the three-dimensional data decoding device 1400.
[0446] First, the three-dimensional data decoding device 1400 decodes (for example, entropy decodes) the differential position information and the differential attribute information from the coded signal (coded bit stream) (S1401).
[0447] Furthermore, the three-dimensional data decoding device 1400 decodes an RT application flag indicating whether or not rotation and translation processing is applied to position information of a three-dimensional point included in the reference three-dimensional data from the encoded signal. Furthermore, the three-dimensional data decoding device 1400 decodes RT information indicating the contents of the rotation and translation processing. Note that the three-dimensional data decoding device 1400 may decode the RT information when the RT application flag indicates that rotation and translation processing is applied, and may not need to decode the RT information when the RT application flag indicates that rotation and translation processing is not applied.
[0448] Next, the three-dimensional data decoding device 1400 performs inverse quantization and inverse transformation on the decoded differential attribute information (S1402).
[0449] Next, the three-dimensional data decoding device 1400 generates predicted position information (e.g., predicted volume) using position information of three-dimensional points included in reference three-dimensional data (e.g., reference space) at a time different from that of the target three-dimensional data (e.g., decoding target space) (S1403). Specifically, the three-dimensional data decoding device 1400 generates predicted position information by applying rotation and translation processing to the position information of three-dimensional points included in the reference three-dimensional data.
[0450] More specifically, when the RT application flag indicates that rotation and translation processing is to be applied, the 3D data decoding device 1400 applies rotation and translation processing to position information of 3D points included in the reference 3D data indicated by the RT information. On the other hand, when the RT application flag indicates that rotation and translation processing is not to be applied, the 3D data decoding device 1400 does not apply rotation and translation processing to position information of 3D points included in the reference 3D data.
[0451] The three-dimensional data decoding device 1400 may perform the rotation and translation processing in a first unit (e.g., space) and generate the predicted position information in a second unit (e.g., volume) that is smaller than the first unit. The three-dimensional data decoding device 1400 may perform the rotation and translation processing and the generation of the predicted position information in the same unit.
[0452] In addition, the three-dimensional data decoding device 1400 may generate predicted position information by applying a first rotation and translation process in a first unit (e.g., space) to the position information of the three-dimensional points included in the reference three-dimensional data, and applying a second rotation and translation process in a second unit (e.g., volume) that is finer than the first unit to the position information of the three-dimensional points obtained by the first rotation and translation process.
[0453] Here, the position information and the predicted position information of the three-dimensional point are expressed in an octet tree structure, for example, as shown in Fig. 41. For example, the position information and the predicted position information of the three-dimensional point are expressed in a scan order that prioritizes the width among the depth and width in the octet tree structure. Alternatively, the position information and the predicted position information of the three-dimensional point are expressed in a scan order that prioritizes the depth among the depth and width in the octet tree structure.
[0454] The three-dimensional data decoding device 1400 generates predicted attribute information using attribute information of the three-dimensional points included in the reference three-dimensional data (S1404).
[0455] Next, the three-dimensional data decoding device 1400 restores the position information of the three-dimensional point included in the target three-dimensional data by decoding the encoded position information included in the encoded signal using the predicted position information. Here, the encoded position information is, for example, differential position information, and the three-dimensional data decoding device 1400 restores the position information of the three-dimensional point included in the target three-dimensional data by adding the differential position information and the predicted position information (S1405).
[0456] Furthermore, the three-dimensional data decoding device 1400 restores the attribute information of the three-dimensional point included in the target three-dimensional data by decoding the encoded attribute information included in the encoded signal using the predicted attribute information. Here, the encoded attribute information is, for example, differential attribute information, and the three-dimensional data decoding device 1400 restores the attribute information of the three-dimensional point included in the target three-dimensional data by adding the differential attribute information and the predicted attribute information (S1406).
[0457] In addition, when the three-dimensional data does not include attribute information, the three-dimensional data decoding device 1400 may not perform steps S1402, S1404, and S1406. In addition, the three-dimensional data decoding device 1400 may perform only one of decoding the position information of the three-dimensional point and decoding the attribute information of the three-dimensional point.
[0458] 50 is an example, and is not limited to this example. For example, the processing for the position information (S1403, S1405) and the processing for the attribute information (S1402, S1404, S1406) are independent of each other, and therefore may be performed in any order, or some of them may be processed in parallel.
[0459] (Embodiment 8) The information of the 3D point cloud includes position information (geometry) and attribute information (attribute). The position information includes coordinates (x-coordinate, y-coordinate, z-coordinate) based on a certain point. When encoding the position information, instead of directly encoding the coordinates of each 3D point, a method is used in which the position of each 3D point is expressed in an octree representation and the octree information is encoded to reduce the amount of code.
[0460] On the other hand, the attribute information includes information indicating color information (RGB, YUV, etc.) of each three-dimensional point, reflectance, normal vector, etc. For example, the three-dimensional data encoding device can encode the attribute information using an encoding method different from that for the position information.
[0461] In this embodiment, a method for encoding attribute information will be described. Note that in this embodiment, an integer value is used as the value of attribute information. For example, when each color component of color information RGB or YUV has 8-bit precision, each color component takes an integer value of 0 to 255. When the reflectance value has 10-bit precision, the reflectance value takes an integer value of 0 to 1023. Note that, when the bit precision of attribute information is decimal precision, the three-dimensional data encoding device may multiply the value by a scale value and then round the attribute information to an integer value so that the value becomes an integer value. Note that the three-dimensional data encoding device may add this scale value to the header of the bit stream, etc.
[0462] As a method of encoding attribute information of a three-dimensional point, it is possible to calculate a predicted value of the attribute information of the three-dimensional point, and encode the difference (prediction residual) between the value of the original attribute information and the predicted value. For example, when the value of the attribute information of a three-dimensional point p is Ap and the predicted value is Pp, the three-dimensional data encoding device encodes the absolute difference Diffp=|Ap-Pp|. In this case, if the predicted value Pp can be generated with high accuracy, the value of the absolute difference Diffp becomes small. Therefore, for example, the amount of code can be reduced by entropy encoding the absolute difference Diffp using an encoding table in which the smaller the value, the smaller the number of generated bits.
[0463] A method for generating a predicted value of attribute information may be to use attribute information of a reference 3D point, which is another 3D point around the target 3D point to be encoded. Here, the reference 3D point is a 3D point within a predetermined distance range from the target 3D point. For example, when a target 3D point p=(x1,y1,z1) and a 3D point q=(x2,y2,z2) exist, the 3D data encoding device calculates the Euclidean distance d(p,q) between the 3D points p and q shown in (Equation A1).
[0464]
number
[0465] When the Euclidean distance d(p, q) is smaller than a predetermined threshold THd, the three-dimensional data encoding device determines that the position of the three-dimensional point q is close to the position of the target three-dimensional point p, and determines that the value of the attribute information of the three-dimensional point q is used to generate a predicted value of the attribute information of the target three-dimensional point p. Note that the distance calculation method may be another method, for example, Mahalanobis distance, etc. may be used. The three-dimensional data encoding device may also determine that a three-dimensional point outside a predetermined distance range from the target three-dimensional point is not used in the prediction process. For example, when a three-dimensional point r exists and the distance d(p, r) between the target three-dimensional point p and the three-dimensional point r is equal to or greater than the threshold THd, the three-dimensional data encoding device may determine that the three-dimensional point r is not used for prediction. Note that the three-dimensional data encoding device may add information indicating the threshold THd to a header of a bit stream, etc.
[0466] 51 is a diagram showing an example of a three-dimensional point. In this example, the distance d(p, q) between the target three-dimensional point p and the three-dimensional point q is smaller than the threshold value THd. Therefore, the three-dimensional data encoding device determines that the three-dimensional point q is a reference three-dimensional point of the target three-dimensional point p, and determines that the value of the attribute information Aq of the three-dimensional point q is used to generate the predicted value Pp of the attribute information Ap of the target three-dimensional point p.
[0467] On the other hand, the distance d(p, r) between the target 3D point p and the 3D point r is equal to or greater than the threshold THd. Therefore, the 3D data encoding device determines that the 3D point r is not a reference 3D point of the target 3D point p, and determines not to use the value of the attribute information Ar of the 3D point r to generate the predicted value Pp of the attribute information Ap of the target 3D point p.
[0468] Furthermore, when the three-dimensional data encoding device encodes attribute information of a target three-dimensional point using a predicted value, it uses a three-dimensional point whose attribute information has already been encoded and decoded as a reference three-dimensional point. Similarly, when the three-dimensional data decoding device decodes attribute information of a target three-dimensional point to be decoded using a predicted value, it uses a three-dimensional point whose attribute information has already been decoded as a reference three-dimensional point. This allows the same predicted value to be generated at the time of encoding and decoding, so that the bit stream of the three-dimensional point generated by encoding can be correctly decoded on the decoding side.
[0469] In addition, when encoding attribute information of 3D points, it is possible to classify each 3D point into multiple layers using the position information of the 3D point and then encode it. Here, each classified layer is called LoD (Level of Detail). A method for generating LoD will be described with reference to FIG.
[0470] First, the three-dimensional data encoding device selects an initial point a0 and assigns it to LoD0. Next, the three-dimensional data encoding device extracts a1 whose distance from point a0 is greater than the threshold Thres_LoD[0] of LoD0 and assigns it to LoD0. Next, the three-dimensional data encoding device extracts a2 whose distance from point a1 is greater than the threshold Thres_LoD[0] of LoD0 and assigns it to LoD0. In this way, the three-dimensional data encoding device configures LoD0 so that the distance between each point in LoD0 is greater than the threshold Thres_LoD[0].
[0471] Next, the three-dimensional data encoding device selects point b0, which has not yet been assigned a LoD, and assigns it to LoD1. Next, the three-dimensional data encoding device extracts point b1, whose distance from point b0 is greater than the threshold Thres_LoD[1] of LoD1 and whose LoD has not been assigned, and assigns it to LoD1. Next, the three-dimensional data encoding device extracts point b2, whose distance from point b1 is greater than the threshold Thres_LoD[1] of LoD1 and whose LoD has not been assigned, and assigns it to LoD1. In this way, the three-dimensional data encoding device configures LoD1 so that the distance between each point in LoD1 is greater than the threshold Thres_LoD[1].
[0472] Next, the three-dimensional data encoding device selects point c0 to which LoD is not yet assigned, and assigns it to LoD2. Next, the three-dimensional data encoding device extracts point c1 whose distance from point c0 is greater than the threshold Thres_LoD[2] of LoD2 and whose LoD is not assigned, and assigns it to LoD2. Next, the three-dimensional data encoding device extracts point c2 whose distance from point c1 is greater than the threshold Thres_LoD[2] of LoD2 and whose LoD is not assigned, and assigns it to LoD2. In this way, the three-dimensional data encoding device configures LoD2 so that the distance between each point in LoD2 is greater than the threshold Thres_LoD[2]. For example, as shown in FIG. 53, the thresholds Thres_LoD[0], Thres_LoD[1], and Thres_LoD[2] of each LoD are set.
[0473] Furthermore, the three-dimensional data encoding device may add information indicating the threshold value of each LoD to the header of the bit stream, etc. For example, in the case of the example shown in Fig. 53, the three-dimensional data encoding device may add threshold values Thres_LoD[0], Thres_LoD[1], and Thres_LoD[2] to the header.
[0474] Furthermore, the three-dimensional data encoding device may assign all three-dimensional points to which LoD is not assigned to the lowest layer of the LoD. In this case, the three-dimensional data encoding device can reduce the amount of code of the header by not adding the threshold of the lowest layer of the LoD to the header. For example, in the example shown in FIG. 53, the three-dimensional data encoding device adds thresholds Thres_LoD[0] and Thres_LoD[1] to the header, and does not add Thres_LoD[2] to the header. In this case, the three-dimensional data decoding device may estimate the value of Thres_LoD[2] to be 0. Furthermore, the three-dimensional data encoding device may add the number of layers of the LoD to the header. This allows the three-dimensional data decoding device to determine the LoD of the lowest layer using the number of layers of the LoD.
[0475] In addition, by setting the threshold value of each LoD layer larger for higher layers as shown in Fig. 53, the higher layers (layers closer to LoD0) become sparser with 3D points farther apart, and the lower layers become denser with 3D points closer together. In the example shown in Fig. 53, LoD0 is the top layer.
[0476] In addition, the method of selecting the initial 3D point when setting each LoD may depend on the encoding order when encoding the position information. For example, the 3D data encoding device selects the 3D point that was encoded first when encoding the position information as the initial point a0 of LoD0, and configures LoD0 by selecting points a1 and a2 with the initial point a0 as the base point. Then, the 3D data encoding device may select the 3D point whose position information is encoded earliest among the 3D points that do not belong to LoD0 as the initial point b0 of LoD1. That is, the 3D data encoding device may select the 3D point whose position information is encoded earliest among the 3D points that do not belong to the upper layer (LoD0 to LoDn-1) of LoDn as the initial point n0 of LoDn. In this way, the 3D data decoding device can configure the same LoD as that at the time of encoding by using the same initial point selection method at the time of decoding, and therefore can appropriately decode the bit stream. Specifically, the 3D data decoding device selects the 3D point whose position information is decoded earliest among the 3D points that do not belong to the upper layer of LoDn as the initial point n0 of LoDn.
[0477] Hereinafter, a method of generating a predicted value of attribute information of a three-dimensional point using LoD information will be described. For example, when encoding three-dimensional points included in LoD0 in order, a three-dimensional data encoding device generates a target three-dimensional point included in LoD1 using encoded and decoded (hereinafter, simply referred to as "encoded") attribute information included in LoD0 and LoD1. In this way, the three-dimensional data encoding device generates a predicted value of attribute information of a three-dimensional point included in LoDn using encoded attribute information included in LoDn' (n'<=n). In other words, the three-dimensional data encoding device does not use attribute information of a three-dimensional point included in a lower layer of LoDn to calculate a predicted value of attribute information of a three-dimensional point included in LoDn.
[0478] For example, the three-dimensional data encoding device generates a predicted value of attribute information of a three-dimensional point by calculating an average of attribute values of N or less three-dimensional points among the encoded three-dimensional points around the target three-dimensional point to be encoded. The three-dimensional data encoding device may also add the value of N to the header of the bit stream or the like. The three-dimensional data encoding device may also change the value of N for each three-dimensional point and add the value of N to each three-dimensional point. This allows an appropriate N to be selected for each three-dimensional point, thereby improving the accuracy of the predicted value. Therefore, the prediction residual can be reduced. The three-dimensional data encoding device may also add the value of N to the header of the bit stream and fix the value of N in the bit stream. This eliminates the need to encode or decode the value of N for each three-dimensional point, thereby reducing the amount of processing. The three-dimensional data encoding device may also encode the value of N separately for each LoD. This allows an appropriate N to be selected for each LoD, thereby improving the encoding efficiency.
[0479] Alternatively, the three-dimensional data encoding device may calculate a predicted value of attribute information of a three-dimensional point by a weighted average value of attribute information of N surrounding encoded three-dimensional points. For example, the three-dimensional data encoding device calculates the weight using distance information between the target three-dimensional point and each of the N surrounding three-dimensional points.
[0480] When the three-dimensional data encoding device encodes the value of N separately for each LoD, for example, the higher the LoD layer, the larger the value of N is set, and the lower the layer, the smaller the value of N is set. Since the distance between 3D points belonging to the higher LoD layers is greater, it may be possible to improve prediction accuracy by setting the value of N to a large value and selecting and averaging multiple surrounding 3D points. Also, since the distance between 3D points belonging to the lower LoD layers is closer, it is possible to perform efficient prediction by setting the value of N to a small value and reducing the amount of averaging processing.
[0481] Fig. 54 is a diagram showing an example of attribute information used for a predicted value. As described above, a predicted value of a point P included in LoDN is generated using encoded surrounding points P' included in LoDN' (N'<=N). Here, the surrounding points P' are selected based on the distance from the point P. For example, a predicted value of attribute information of a point b2 shown in Fig. 54 is generated using attribute information of points a0, a1, a2, b0, and b1.
[0482] The surrounding points selected change depending on the value of N mentioned above. For example, when N=5, a0, a1, a2, b0, and b1 are selected as surrounding points of point b2. When N=4, points a0, a1, a2, and b1 are selected based on distance information.
[0483] The predicted value is calculated by a distance-dependent weighted average. For example, in the example shown in FIG. 54, the predicted value a2p of the point a2 is calculated by a weighted average of the attribute information of the points a0 and a1, as shown in (Equation A2) and (Equation A3). i is the value of the attribute information of point ai.
[0484]
number
[0485] The predicted value b2p of the point b2 is calculated by the weighted average of the attribute information of the points a0, a1, a2, b0, and b1, as shown in (Equation A4) to (Equation A6). i is the value of the attribute information of point bi.
[0486]
number
[0487] Furthermore, the three-dimensional data encoding device may calculate a difference value (prediction residual) between the value of the attribute information of the three-dimensional point and a predicted value generated from the surrounding points, and quantize the calculated prediction residual. For example, the three-dimensional data encoding device performs quantization by dividing the prediction residual by a quantization scale (also called a quantization step). In this case, the smaller the quantization scale, the smaller the error (quantization error) that may occur due to quantization. Conversely, the larger the quantization scale, the larger the quantization error.
[0488] The three-dimensional data encoding device may change the quantization scale used for each LoD. For example, the three-dimensional data encoding device may reduce the quantization scale for higher layers and increase the quantization scale for lower layers. Since the value of the attribute information of the three-dimensional points belonging to the higher layers may be used as a predicted value of the attribute information of the three-dimensional points belonging to the lower layers, the quantization scale for the higher layers may be reduced to suppress the quantization error that may occur in the higher layers, and the accuracy of the predicted value may be increased, thereby improving the encoding efficiency. The three-dimensional data encoding device may add the quantization scale used for each LoD to a header or the like. This allows the three-dimensional data decoding device to correctly decode the quantization scale, and therefore to appropriately decode the bit stream.
[0489] Furthermore, the three-dimensional data encoding device may convert a signed integer value (signed quantization value), which is a prediction residual after quantization, into an unsigned integer value (unsigned quantization value). This makes it unnecessary to consider the occurrence of negative integers when entropy encoding the prediction residual. Note that the three-dimensional data encoding device does not necessarily need to convert a signed integer value into an unsigned integer value, and for example, the sign bit may be separately entropy encoded.
[0490] The prediction residual is calculated by subtracting the predicted value from the original value. For example, the prediction residual a2r of the point a2 is calculated by subtracting the value A of the attribute information of the point a2 as shown in (Equation A7). 2 The prediction residual b2r of point b2 is calculated by subtracting the predicted value a2p of point a2 from the attribute information value B of point b2, as shown in (Equation A8). 2It is calculated by subtracting the predicted value b2p of point b2 from
[0491] a2r=A 2 -a2p (formula A7)
[0492] b2r=B 2 -b2p...(Formula A8)
[0493] Furthermore, the prediction residual is quantized by dividing it by QS (Quantization Step). For example, the quantized value a2q of point a2 is calculated by (Equation A9). The quantized value b2q of point b2 is calculated by (Equation A10). Here, QS_LoD0 is the QS for LoD0, and QS_LoD1 is the QS for LoD1. That is, the QS may be changed according to the LoD.
[0494] a2q=a2r / QS_LoD0 (Equation A9)
[0495] b2q=b2r / QS_LoD1 (Formula A10)
[0496] Furthermore, the three-dimensional data encoding device converts the signed integer value, which is the quantized value, into an unsigned integer value as follows: If the signed integer value a2q is smaller than 0, the three-dimensional data encoding device sets the unsigned integer value a2u to -1-(2×a2q). If the signed integer value a2q is 0 or greater, the three-dimensional data encoding device sets the unsigned integer value a2u to 2×a2q.
[0497] Similarly, the three-dimensional data encoding device sets the unsigned integer value b2u to -1-(2×b2q) when the signed integer value b2q is less than 0. The three-dimensional data encoding device sets the unsigned integer value b2u to 2×b2q when the signed integer value b2q is 0 or greater.
[0498] Furthermore, the three-dimensional data encoding device may encode the prediction residuals (unsigned integer values) after quantization by entropy encoding. For example, the unsigned integer values may be binarized and then binary arithmetic encoding may be applied.
[0499] In this case, the three-dimensional data encoding device may switch the binarization method according to the value of the prediction residual. For example, when the prediction residual pu is smaller than the threshold R_TH, the three-dimensional data encoding device binarizes the prediction residual pu with a fixed number of bits required to express the threshold R_TH. When the prediction residual pu is equal to or larger than the threshold R_TH, the three-dimensional data encoding device binarizes the binarized data of the threshold R_TH and the value of (pu-R_TH) using Exponential-Golomb or the like.
[0500] For example, the three-dimensional data encoding device binarizes the prediction residual pu in 6 bits when the threshold value R_TH is 63 and the prediction residual pu is smaller than 63. Furthermore, when the prediction residual pu is 63 or greater, the three-dimensional data encoding device performs arithmetic encoding by binarizing the binary data of the threshold value R_TH (111111) and (pu-63) using the exponential Golomb algorithm.
[0501] In a more specific example, when the prediction residual pu is 32, the three-dimensional data encoding device generates 6-bit binary data (100000) and arithmetically codes this bit string. When the prediction residual pu is 66, the three-dimensional data encoding device generates binary data (111111) of the threshold R_TH and a bit string (00100) that expresses the value 3 (66-63) in Exponential Golomb notation and arithmetically codes this bit string (111111+00100).
[0502] In this way, the three-dimensional data encoding device can perform encoding while suppressing a sudden increase in the number of binarization bits when the prediction residual becomes large, by switching the binarization method according to the size of the prediction residual. Note that the three-dimensional data encoding device may add the threshold value R_TH to the header of the bit stream, etc.
[0503] For example, when encoding is performed at a high bit rate, that is, when the quantization scale is small, the quantization error is small and the prediction accuracy is high, and as a result, the prediction residual may not be large. Therefore, in this case, the three-dimensional data encoding device sets the threshold value R_TH to a large value. This reduces the possibility of encoding the binarized data of the threshold value R_TH, and improves the encoding efficiency. Conversely, when encoding is performed at a low bit rate, that is, when the quantization scale is large, the quantization error is large and the prediction accuracy is poor, and as a result, the prediction residual may be large. Therefore, in this case, the three-dimensional data encoding device sets the threshold value R_TH to a small value. This makes it possible to prevent a sudden increase in the bit length of the binarized data.
[0504] Also, the three-dimensional data encoding device may switch the threshold R_TH for each LoD and add the threshold R_TH for each LoD to a header or the like. That is, the three-dimensional data encoding device may switch the binarization method for each LoD. For example, since the distance between three-dimensional points is far in the upper layer, the prediction accuracy is poor and the prediction residual may be large as a result. Therefore, the three-dimensional data encoding device prevents a sudden increase in the bit length of the binarized data by setting the threshold R_TH small for the upper layer. Also, since the distance between three-dimensional points is close in the lower layer, the prediction accuracy is high and the prediction residual may be small as a result. Therefore, the three-dimensional data encoding device improves the encoding efficiency by setting the threshold R_TH large for the layer.
[0505] Fig. 55 is a diagram showing an example of exponential Golomb code, and shows the relationship between values (multiple values) before binarization and bits (codes) after binarization. Note that 0 and 1 shown in Fig. 55 may be inverted.
[0506] Furthermore, the three-dimensional data encoding device applies arithmetic coding to the binary data of the prediction residual. This can improve the encoding efficiency. When applying arithmetic coding, the tendency of the occurrence probability of 0 and 1 for each bit may be different between an n-bit code, which is a portion of the binary data binarized with n bits, and a remaining code, which is a portion binarized using the exponential Golomb algorithm. Therefore, the three-dimensional data encoding device may switch the application method of arithmetic coding between the n-bit code and the remaining code.
[0507] For example, the three-dimensional data encoding device performs arithmetic coding for an n-bit code using a different coding table (probability table) for each bit. In this case, the three-dimensional data encoding device may change the number of coding tables used for each bit. For example, the three-dimensional data encoding device performs arithmetic coding for the first bit b0 of an n-bit code using one coding table. The three-dimensional data encoding device also uses two coding tables for the next bit b1. The three-dimensional data encoding device also switches the coding table used for the arithmetic coding of bit b1 according to the value of b0 (0 or 1). Similarly, the three-dimensional data encoding device also uses four coding tables for the next bit b2. The three-dimensional data encoding device also switches the coding table used for the arithmetic coding of bit b2 according to the values of b0 and b1 (0 to 3).
[0508] In this way, the three-dimensional data encoding device performs arithmetic encoding on each bit bn-1 of the n-bit code. n-1 The three-dimensional data encoding device uses coding tables. Also, the three-dimensional data encoding device switches the coding table to be used depending on the value (occurrence pattern) of the bit before bn-1. This allows the three-dimensional data encoding device to use an appropriate coding table for each bit, thereby improving the coding efficiency.
[0509] Note that the three-dimensional data encoding device may reduce the number of encoding tables used for each bit. For example, when arithmetic-encoding each bit bn-1, the three-dimensional data encoding device may switch between two encoding tables according to the value (occurrence pattern) of the m bits (m < n - 1) before bn-1. This can improve the encoding efficiency while suppressing the number of encoding tables used for each bit. Note that the three-dimensional data encoding device may update the occurrence probabilities of 0 and 1 in each encoding table according to the value of the actually generated binarized data. Also, the three-dimensional data encoding device may fix the occurrence probabilities of 0 and 1 in the encoding tables for some bits. This can suppress the number of updates of the occurrence probabilities and thus reduce the processing amount. m For example, when the n-bit code is b0b1b2…bn-1, there is one encoding table (CTb0) for b0. There are two encoding tables (CTb10, CTb11) for b1. Also, the encoding table to be used is switched according to the value (0 to 1) of b0. There are four encoding tables (CTb20, CTb21, CTb22, CTb23) for b2. Also, the encoding table to be used is switched according to the values (0 to 3) of b0 and b1. There are two
[0510] (CTbn0, CTbn1, …, CTbn(2 n-1 -1)) encoding tables for bn-1. Also, the encoding table to be used is switched according to the value (0 to 2 n-1 -1) of b0b1…bn-2. n-1
[0511] Note that the three-dimensional data encoding device may apply m-ary arithmetic encoding (m = 2 n ) that sets values from 0 to 2 n -1 without binarization to the n-bit code. Also, when the three-dimensional data encoding device arithmetic-encodes the n-bit code in m-ary, the three-dimensional data decoding device may also restore the n-bit code by m-ary arithmetic decoding.
[0512] Fig. 56 is a diagram for explaining a process in the case where the remaining code is an exponential Golomb code, for example. The remaining code, which is a portion binarized using the exponential Golomb code, includes a prefix part and a suffix part as shown in Fig. 56. For example, the three-dimensional data encoding device switches the encoding table between the prefix part and the suffix part. That is, the three-dimensional data encoding device arithmetically encodes each bit included in the prefix part using the encoding table for the prefix, and arithmetically encodes each bit included in the suffix part using the encoding table for the suffix.
[0513] The three-dimensional data encoding device may update the occurrence probability of 0 and 1 in each encoding table according to the value of the binary data that has actually occurred. Alternatively, the three-dimensional data encoding device may fix the occurrence probability of 0 and 1 in one of the encoding tables. This reduces the number of updates to the occurrence probability, thereby reducing the amount of processing. For example, the three-dimensional data encoding device may update the occurrence probability for the prefix part and fix the occurrence probability for the suffix part.
[0514] Furthermore, the three-dimensional data encoding device decodes the prediction residual after quantization by inverse quantization and reconstruction, and uses the decoded value, which is the decoded prediction residual, for prediction of the three-dimensional point to be encoded and thereafter. Specifically, the three-dimensional data encoding device calculates an inverse quantization value by multiplying the prediction residual after quantization (quantization value) by the quantization scale, and obtains a decoded value (reconstruction value) by adding the inverse quantization value and the prediction value.
[0515] For example, the inverse quantization value a2iq of point a2 is calculated by (Equation A11) using the quantization value a2q of point a2. The inverse quantization value b2iq of point b2 is calculated by (Equation A12) using the quantization value b2q of point b2. Here, QS_LoD0 is the QS for LoD0, and QS_LoD1 is the QS for LoD1. That is, the QS may be changed according to the LoD.
[0516] a2iq=a2q×QS_LoD0 (Formula A11)
[0517] b2iq=b2q×QS_LoD1 (Formula A12)
[0518] For example, the decoded value a2rec of point a2 is calculated by adding the predicted value a2p of point a2 to the inverse quantized value a2iq of point a2 as shown in (Equation A13). The decoded value b2rec of point b2 is calculated by adding the predicted value b2p of point b2 to the inverse quantized value b2iq of point b2 as shown in (Equation A14).
[0519] a2rec=a2iq+a2p (formula A13)
[0520] b2rec=b2iq+b2p (formula A14)
[0521] An example of the syntax of a bitstream according to this embodiment will be described below. Fig. 57 is a diagram showing an example of the syntax of an attribute header (attribute_header) according to this embodiment. The attribute header is header information of attribute information. As shown in Fig. 57, the attribute header includes hierarchical number information (NumLoD), three-dimensional point number information (NumOfPoint[i]), hierarchical threshold (Thres_Lod[i]), surrounding point number information (NumNeighorPoint[i]), prediction threshold (THd[i]), quantization scale (QS[i]), and binarization threshold (R_TH[i]).
[0522] The number of layers information (NumLoD) indicates the number of layers of the LoD to be used.
[0523] The three-dimensional point number information (NumOfPoint[i]) indicates the number of three-dimensional points belonging to layer i. The three-dimensional data encoding device may add three-dimensional point total number information (AllNumOfPoint) indicating the total number of three-dimensional points to another header. In this case, the three-dimensional data encoding device does not need to add NumOfPoint[NumLoD-1] indicating the number of three-dimensional points belonging to the lowest layer to the header. In this case, the three-dimensional data decoding device can calculate NumOfPoint[NumLoD-1] by (Equation A15). This allows the amount of code in the header to be reduced.
[0524]
Number
[0525] The hierarchical threshold (Thres_Lod[i]) is the threshold used for the setting of layer i. The three-dimensional data encoding device and the three-dimensional data decoding device configure LoDi such that the distance between each point in LoDi is greater than the threshold Thres_LoD[i]. Also, the three-dimensional data encoding device may not add the value of Thres_Lod[NumLoD - 1] (the bottom layer) to the header. In this case, the three-dimensional data decoding device estimates the value of Thres_Lod[NumLoD - 1] as 0. Thereby, the amount of code of the header can be reduced.
[0526] The number of surrounding points information (NumNeighorPoint[i]) indicates the upper limit value of the number of surrounding points used for generating the predicted value of the three-dimensional points belonging to layer i. When the number of surrounding points M is less than NumNeighorPoint[i] (M < NumNeighorPoint[i]), the three-dimensional data encoding device may calculate the predicted value using M surrounding points. Also, when the three-dimensional data encoding device does not need to divide the value of NumNeighorPoint[i] for each LoD, it may add one piece of surrounding points information (NumNeighorPoint) used for all LoDs to the header.
[0527] The prediction threshold (THd[i]) indicates the upper limit value of the distance between the surrounding three-dimensional points and the target three-dimensional point used for predicting the target three-dimensional point to be encoded or decoded at layer i. The three-dimensional data encoding device and the three-dimensional data decoding device do not use the three-dimensional points whose distance from the target three-dimensional point is farther than THd[i] for prediction. Note that when the three-dimensional data encoding device does not need to divide the value of THd[i] for each LoD, it may add one prediction threshold (THd) used for all LoDs to the header.
[0528] The quantization scale (QS[i]) indicates the quantization scale used for quantization and inverse quantization of layer i.
[0529] The binarization threshold (R_TH[i]) is a threshold for switching the binarization method of the prediction residual of a 3D point belonging to layer i. For example, when the prediction residual is smaller than the threshold R_TH, the 3D data encoding device binarizes the prediction residual pu with a fixed number of bits, and when the prediction residual is equal to or larger than the threshold R_TH, the binarized data of the threshold R_TH and the value (pu-R_TH) are binarized using the exponential Golomb algorithm. If it is not necessary to switch the value of R_TH[i] for each LoD, the 3D data encoding device may add one binarization threshold (R_TH) used for all LoDs to the header.
[0530] Note that R_TH[i] may be the maximum value that can be expressed in n bits. For example, R_TH is 63 in 6 bits, and R_TH is 255 in 8 bits. The three-dimensional data encoding device may also encode the number of bits instead of encoding the maximum value that can be expressed in n bits as the binarization threshold. For example, the three-dimensional data encoding device may add a value of 6 to the header when R_TH[i]=63, and a value of 8 to the header when R_TH[i]=255. The three-dimensional data encoding device may also define a minimum value (minimum number of bits) of the number of bits that represents R_TH[i], and add a relative number of bits from the minimum value to the header. For example, the three-dimensional data encoding device may add a value of 0 to the header when R_TH[i]=63 and the minimum number of bits is 6, and add a value of 2 to the header when R_TH[i]=255 and the minimum number of bits is 6.
[0531] The three-dimensional data encoding device may also entropy-encode at least one of NumLod, Thres_Lod[i], NumNeighborPoint[i], THd[i], QS[i], and R_TH[i] and add the result to the header. For example, the three-dimensional data encoding device may binarize each value and arithmetically encode it. The three-dimensional data encoding device may also encode each value in a fixed length to reduce the amount of processing.
[0532] Furthermore, the three-dimensional data encoding device may not add at least one of NumLod, Thres_Lod[i], NumNeighborPoint[i], THd[i], QS[i], and R_TH[i] to the header. For example, at least one of these values may be specified by a profile or level of a standard, etc. This allows the amount of bits in the header to be reduced.
[0533] Fig. 58 is a diagram showing an example of the syntax of attribute data (attribute_data) according to this embodiment. This attribute data includes encoded data of attribute information of a plurality of three-dimensional points. As shown in Fig. 58, the attribute data includes an n-bit code and a remaining code.
[0534] An n-bit code is the coded data of the prediction residual of the attribute information value or a part of it. The bit length of the n-bit code depends on the value of R_TH[i]. For example, when the value indicated by R_TH[i] is 63, the n-bit code is 6 bits, and when the value indicated by R_TH[i] is 255, the n-bit code is 8 bits.
[0535] The remaining code is the coded data of the prediction residual of the attribute information value that is coded by exponential Golomb coding. This remaining code is coded or decoded when the n-bit code is the same as R_TH[i]. In addition, the three-dimensional data decoding device adds the value of the n-bit code and the value of the remaining code to decode the prediction residual. Note that if the n-bit code is not the same value as R_TH[i], the remaining code does not need to be coded or decoded.
[0536] The flow of processing in the three-dimensional data encoding device will be explained below. Figure 59 is a flowchart of three-dimensional data encoding processing by the three-dimensional data encoding device.
[0537] First, the three-dimensional data encoding device encodes position information (geometry) (S3001). For example, the three-dimensional data is encoded using an octree representation.
[0538] When the position of the three-dimensional point is changed by quantization or the like after encoding the position information, the three-dimensional data encoding device reallocates the attribute information of the original three-dimensional point to the changed three-dimensional point (S3002). For example, the three-dimensional data encoding device performs reallocation by interpolating the value of the attribute information according to the amount of change in the position. For example, the three-dimensional data encoding device detects N three-dimensional points before the change that are close to the changed three-dimensional position, and calculates a weighted average of the attribute information values of the N three-dimensional points. For example, the three-dimensional data encoding device determines a weight in the weighted average based on the distance from the changed three-dimensional position to each of the N three-dimensional points. Then, the three-dimensional data encoding device determines the value obtained by the weighted average as the value of the attribute information of the changed three-dimensional point. In addition, when two or more three-dimensional points are changed to the same three-dimensional position by quantization or the like, the three-dimensional data encoding device may assign the average value of the attribute information of the two or more three-dimensional points before the change as the value of the attribute information of the changed three-dimensional point.
[0539] Next, the three-dimensional data encoding device encodes the reallocated attribute information (Attribute) (S3003). For example, when encoding multiple types of attribute information, the three-dimensional data encoding device may encode the multiple types of attribute information in order. For example, when encoding color and reflectance as attribute information, the three-dimensional data encoding device may generate a bit stream in which the encoding result of reflectance is added after the encoding result of color. Note that the order of the encoding results of the multiple attribute information added to the bit stream is not limited to this order and may be any order.
[0540] Furthermore, the three-dimensional data encoding device may add information indicating the start location of the encoded data of each piece of attribute information in the bit stream to a header or the like. This allows the three-dimensional data decoding device to selectively decode attribute information that needs to be decoded, and therefore omits the decoding process of attribute information that does not need to be decoded. This reduces the amount of processing by the three-dimensional data decoding device. Furthermore, the three-dimensional data encoding device may encode multiple types of attribute information in parallel and integrate the encoding results into one bit stream. This allows the three-dimensional data encoding device to encode multiple types of attribute information at high speed.
[0541] 60 is a flowchart of the attribute information encoding process (S3003). First, the three-dimensional data encoding device sets the LoD (S3011). That is, the three-dimensional data encoding device assigns each three-dimensional point to one of a plurality of LoDs.
[0542] Next, the three-dimensional data encoding device starts a loop for each LoD (S3012). That is, the three-dimensional data encoding device repeats the process of steps S3013 to S3021 for each LoD.
[0543] Next, the three-dimensional data encoding device starts a loop for each three-dimensional point (S3013). That is, the three-dimensional data encoding device repeatedly performs the processes of steps S3014 to S3020 for each three-dimensional point.
[0544] First, the three-dimensional data encoding device searches for a plurality of surrounding points, which are three-dimensional points existing around the target three-dimensional point to be processed, to be used in calculating a predicted value of the target three-dimensional point (S3014). Next, the three-dimensional data encoding device calculates a weighted average of the values of attribute information of the plurality of surrounding points, and sets the obtained value as the predicted value P (S3015). Next, the three-dimensional data encoding device calculates a prediction residual, which is the difference between the attribute information of the target three-dimensional point and the predicted value (S3016). Next, the three-dimensional data encoding device calculates a quantized value by quantizing the prediction residual (S3017). Next, the three-dimensional data encoding device arithmetically encodes the quantized value (S3018).
[0545] The three-dimensional data encoding device also calculates an inverse quantized value by inverse quantizing the quantized value (S3019). Next, the three-dimensional data encoding device generates a decoded value by adding a predicted value to the inverse quantized value (S3020). Next, the three-dimensional data encoding device ends the loop in three-dimensional point units (S3021). Furthermore, the three-dimensional data encoding device ends the loop in LoD units (S3022).
[0546] Hereinafter, a three-dimensional data decoding process in a three-dimensional data decoding device that decodes a bit stream generated by the above-mentioned three-dimensional data encoding device will be described.
[0547] The three-dimensional data decoding device generates decoded binary data by arithmetically decoding the binary data of the attribute information in the bit stream generated by the three-dimensional data encoding device in the same manner as the three-dimensional data encoding device. Note that, in the three-dimensional data encoding device, when the application method of arithmetic coding is switched between the part binarized with n bits (n-bit code) and the part binarized using the exponential Golomb algorithm (remaining code), the three-dimensional data decoding device performs decoding accordingly when applying arithmetic decoding.
[0548] For example, in an arithmetic decoding method for an n-bit code, the three-dimensional data decoding device performs arithmetic decoding using a different coding table (decoding table) for each bit. In this case, the three-dimensional data decoding device may change the number of coding tables used for each bit. For example, the first bit b0 of the n-bit code is arithmetically decoded using one coding table. The three-dimensional data decoding device uses two coding tables for the next bit b1. The three-dimensional data decoding device switches the coding table used for arithmetic decoding of bit b1 according to the value of b0 (0 or 1). Similarly, the three-dimensional data decoding device uses four coding tables for the next bit b2. The three-dimensional data decoding device switches the coding table used for arithmetic decoding of bit b2 according to the values of b0 and b1 (0 to 3).
[0549] Thus, when the three-dimensional data decoding device arithmetically decodes each bit bn-1 of the n-bit code, it uses n-1 two encoding tables. Also, the three-dimensional data decoding device switches the encoding table to be used according to the values (generated patterns) of the bits before bn-1. Thereby, the three-dimensional data decoding device can appropriately decode a bit stream with improved encoding efficiency by using an appropriate encoding table for each bit.
[0550] Note that the three-dimensional data decoding device may reduce the number of encoding tables used for each bit. For example, when the three-dimensional data decoding device arithmetically decodes each bit bn-1, it may switch two m encoding tables according to the values (generated patterns) of m bits (m < n-1) before bn-1. Thereby, the three-dimensional data decoding device can appropriately decode a bit stream with improved encoding efficiency while suppressing the number of encoding tables used for each bit. Note that the three-dimensional data decoding device may update the occurrence probabilities of 0 and 1 in each encoding table according to the values of the actually generated binarized data. Also, the three-dimensional data decoding device may fix the occurrence probabilities of 0 and 1 in the encoding tables of some bits. Thereby, the number of update times of the occurrence probability can be suppressed, so the processing amount can be reduced.
[0551] For example, when the n-bit code is b0b1b2…bn-1, the encoding table for b0 is one (CTb0). The encoding tables for b1 are two (CTb10, CTb11). Also, the encoding table is switched according to the value (0 to 1) of b0. The encoding tables for b2 are four (CTb20, CTb21, CTb22, CTb23). Also, the encoding table is switched according to the values (0 to 3) of b0 and b1. The encoding tables for bn-1 are n-1 two (CTbn0, CTbn1, …, CTbn(2 n-1 -1)). Also, the encoding table is switched according to the values (0 to 2 n-1 -1) of b0b1…bn-2.
[0552] Fig. 61 is a diagram for explaining a process when the remaining code is an exponential Golomb code, for example. The part (the remaining code) binarized and encoded by the three-dimensional data encoding device using the exponential Golomb includes a prefix part and a suffix part as shown in Fig. 61. For example, the three-dimensional data decoding device switches the encoding table between the prefix part and the suffix part. That is, the three-dimensional data decoding device arithmetically decodes each bit included in the prefix part using the encoding table for the prefix, and arithmetically decodes each bit included in the suffix part using the encoding table for the suffix.
[0553] The three-dimensional data decoding device may update the occurrence probability of 0 and 1 in each encoding table according to the value of the binarized data generated during decoding. Alternatively, the three-dimensional data decoding device may fix the occurrence probability of 0 and 1 in one of the encoding tables. This reduces the number of updates to the occurrence probability, thereby reducing the amount of processing. For example, the three-dimensional data decoding device may update the occurrence probability for the prefix part and fix the occurrence probability for the suffix part.
[0554] Furthermore, the three-dimensional data decoding device decodes the quantized prediction residual (unsigned integer value) by multi-valuing the binary data of the arithmetically decoded prediction residual in accordance with the encoding method used in the three-dimensional data encoding device. The three-dimensional data decoding device first calculates the value of the decoded n-bit code by arithmetically decoding the binary data of the n-bit code. Next, the three-dimensional data decoding device compares the value of the n-bit code with the value of R_TH.
[0555] When the value of the n-bit code and the value of R_TH match, the three-dimensional data decoding device determines that a bit encoded by exponential Golomb exists next, and arithmetically decodes the remaining code, which is binary data encoded by exponential Golomb.Then, the three-dimensional data decoding device calculates the value of the remaining code from the decoded remaining code using a reverse lookup table showing the relationship between the remaining code and its value. FIG. 62 is a diagram showing an example of a reverse lookup table showing the relationship between the remaining code and its value.Then, the three-dimensional data decoding device obtains a multi-valued prediction residual after quantization by adding the obtained value of the remaining code to R_TH.
[0556] On the other hand, when the value of the n-bit code does not match the value of R_TH (the value is smaller than R_TH), the three-dimensional data decoding device determines the value of the n-bit code as the multi-valued post-quantization prediction residual as is. This allows the three-dimensional data decoding device to properly decode the bit stream generated by the three-dimensional data encoding device by switching the binarization method according to the value of the prediction residual.
[0557] When the threshold value R_TH is added to the header of the bit stream or the like, the three-dimensional data decoding device may decode the value of the threshold value R_TH from the header and switch the decoding method using the decoded threshold value R_TH. When the threshold value R_TH is added to the header or the like for each LoD, the three-dimensional data decoding device switches the decoding method using the decoded threshold value R_TH for each LoD.
[0558] For example, when the threshold value R_TH is 63 and the value of the decoded n-bit code is 63, the three-dimensional data decoding device obtains the value of the remaining code by decoding the remaining code using the exponential Golomb method. For example, in the example shown in FIG. 62, the remaining code is 00100, and the value of the remaining code is obtained as 3. Next, the three-dimensional data decoding device obtains the value of the prediction residual, 66, by adding the value of the threshold value R_TH, 63, and the value of the remaining code, 3.
[0559] Furthermore, if the value of the decoded n-bit code is 32, the three-dimensional data decoding device sets the value of the n-bit code, 32, as the value of the prediction residual.
[0560] Furthermore, the three-dimensional data decoding device converts the decoded quantized prediction residual from an unsigned integer value to a signed integer value, for example, by a process that is the reverse of the process in the three-dimensional data encoding device. This allows the three-dimensional data decoding device to appropriately decode the generated bit stream without considering the occurrence of negative integers when entropy encoding the prediction residual. Note that the three-dimensional data decoding device does not necessarily need to convert the unsigned integer value to a signed integer value, and may decode the sign bit, for example, when decoding a bit stream generated by separately entropy encoding the sign bit.
[0561] The three-dimensional data decoding device generates a decoded value by decoding the quantized prediction residual converted into a signed integer value through inverse quantization and reconstruction. The three-dimensional data decoding device also uses the generated decoded value for prediction of the three-dimensional point to be decoded and thereafter. Specifically, the three-dimensional data decoding device calculates an inverse quantization value by multiplying the quantized prediction residual by the decoded quantization scale, and obtains the decoded value by adding the inverse quantization value and the prediction value.
[0562] The decoded unsigned integer value (unsigned quantized value) is converted to a signed integer value by the following process. If the LSB (least significant bit) of the decoded unsigned integer value a2u is 1, the three-dimensional data decoding device sets the signed integer value a2q to -((a2u+1)>>1). If the LSB of the unsigned integer value a2u is not 1, the three-dimensional data decoding device sets the signed integer value a2q to (a2u>>1).
[0563] Similarly, if the LSB of the decoded unsigned integer value b2u is 1, the three-dimensional data decoding device sets the signed integer value b2q to -((b2u+1)>>1). If the LSB of the unsigned integer value n2u is not 1, the three-dimensional data decoding device sets the signed integer value b2q to (b2u>>1).
[0564] Moreover, the details of the inverse quantization and reconstruction processing by the three-dimensional data decoding device are similar to those of the inverse quantization and reconstruction processing by the three-dimensional data encoding device.
[0565] The flow of processing in the three-dimensional data decoding device will be described below. Fig. 63 is a flowchart of three-dimensional data decoding processing by the three-dimensional data decoding device. First, the three-dimensional data decoding device decodes position information (geometry) from the bit stream (S3031). For example, the three-dimensional data decoding device performs decoding using an octree representation.
[0566] Next, the three-dimensional data decoding device decodes attribute information (Attribute) from the bit stream (S3032). For example, when decoding multiple types of attribute information, the three-dimensional data decoding device may decode the multiple types of attribute information in order. For example, when decoding color and reflectance as attribute information, the three-dimensional data decoding device decodes the encoding result of color and the encoding result of reflectance according to the order in which they are added to the bit stream. For example, when the encoding result of reflectance is added after the encoding result of color in the bit stream, the three-dimensional data decoding device decodes the encoding result of color, and then decodes the encoding result of reflectance. Note that the three-dimensional data decoding device may decode the encoding results of attribute information added to the bit stream in any order.
[0567] Furthermore, the three-dimensional data decoding device may obtain information indicating the start location of the encoded data of each piece of attribute information in the bit stream by decoding a header or the like. This allows the three-dimensional data decoding device to selectively decode attribute information that needs to be decoded, and therefore omits the decoding process of attribute information that does not need to be decoded. This reduces the amount of processing by the three-dimensional data decoding device. Furthermore, the three-dimensional data decoding device may decode multiple types of attribute information in parallel and integrate the decoding results into one three-dimensional point cloud. This allows the three-dimensional data decoding device to decode multiple types of attribute information at high speed.
[0568] Fig. 64 is a flowchart of the attribute information decoding process (S3032). First, the three-dimensional data decoding device sets the LoD (S3041). That is, the three-dimensional data decoding device assigns each of the multiple three-dimensional points having the decoded position information to one of the multiple LoDs. For example, this assignment method is the same as the assignment method used in the three-dimensional data encoding device.
[0569] Next, the three-dimensional data decoding device starts a loop for each LoD (S3042). That is, the three-dimensional data decoding device repeats the process of steps S3043 to S3049 for each LoD.
[0570] Next, the three-dimensional data decoding device starts a loop for each three-dimensional point (S3043). That is, the three-dimensional data decoding device repeatedly performs the processes of steps S3044 to S3048 for each three-dimensional point.
[0571] First, the three-dimensional data decoding device searches for a plurality of surrounding points, which are three-dimensional points existing around the target three-dimensional point to be processed, and are used to calculate a predicted value of the target three-dimensional point (S3044). Next, the three-dimensional data decoding device calculates a weighted average of the values of the attribute information of the plurality of surrounding points, and sets the obtained value as the predicted value P (S3045). Note that these processes are similar to those in the three-dimensional data encoding device.
[0572] Next, the three-dimensional data decoding device arithmetically decodes the quantized value from the bit stream (S3046). The three-dimensional data decoding device also calculates an inverse quantized value by inverse quantizing the decoded quantized value (S3047). Next, the three-dimensional data decoding device generates a decoded value by adding a predicted value to the inverse quantized value (S3048). Next, the three-dimensional data decoding device ends the loop in three-dimensional point units (S3049). The three-dimensional data decoding device also ends the loop in LoD units (S3050).
[0573] Next, the configurations of a three-dimensional data encoding device and a three-dimensional data decoding device according to this embodiment will be described. Fig. 65 is a block diagram showing the configuration of a three-dimensional data encoding device 3000 according to this embodiment. This three-dimensional data encoding device 3000 includes a position information encoding unit 3001, an attribute information reallocation unit 3002, and an attribute information encoding unit 3003.
[0574] The attribute information encoding unit 3003 encodes position information (geometry) of a plurality of three-dimensional points included in the input point cloud. The attribute information reallocation unit 3002 reallocates values of attribute information of a plurality of three-dimensional points included in the input point cloud using the encoded and decoded results of the position information. The attribute information encoding unit 3003 encodes the reallocated attribute information. Furthermore, the three-dimensional data encoding device 3000 generates a bit stream including the encoded position information and the encoded attribute information.
[0575] 66 is a block diagram showing the configuration of a three-dimensional data decoding device 3010 according to this embodiment. The three-dimensional data decoding device 3010 includes a position information decoding unit 3011 and an attribute information decoding unit 3012.
[0576] The position information decoding unit 3011 decodes position information (geometry) of multiple 3D points from the bit stream. The attribute information decoding unit 3012 decodes attribute information (attribute) of multiple 3D points from the bit stream. In addition, the 3D data decoding device 3010 generates an output point group by combining the decoded position information and the decoded attribute information.
[0577] As described above, the three-dimensional data encoding device according to this embodiment performs the process shown in FIG. 67. The three-dimensional data encoding device encodes a three-dimensional point having attribute information. First, the three-dimensional data encoding device calculates a predicted value of the attribute information of the three-dimensional point (S3061). Next, the three-dimensional data encoding device calculates a prediction residual which is the difference between the attribute information of the three-dimensional point and the predicted value (S3062). Next, the three-dimensional data encoding device generates binary data by binarizing the prediction residual (S3063). Next, the three-dimensional data encoding device arithmetically encodes the binary data (S3064).
[0578] According to this, the three-dimensional data encoding device calculates a prediction residual of the attribute information, and further binarizes and arithmetically codes the prediction residual, thereby making it possible to reduce the amount of code of encoded data of the attribute information.
[0579] For example, in arithmetic coding (S3064), the three-dimensional data coding device uses a different coding table for each bit of binary data, thereby enabling the three-dimensional data coding device to improve coding efficiency.
[0580] For example, in arithmetic coding (S3064), the lower the bit of the binary data, the greater the number of coding tables that are used.
[0581] For example, in arithmetic coding (S3064), the three-dimensional data coding device selects a coding table to be used for arithmetic coding of a target bit in accordance with the value of the most significant bit of the target bit included in the binary data. This allows the three-dimensional data coding device to select a coding table according to the value of the most significant bit, thereby improving coding efficiency.
[0582] For example, in the binarization (S3063), if the prediction residual is smaller than a threshold (R_TH), the three-dimensional data encoding device generates binary data by binarizing the prediction residual with a fixed number of bits, and if the prediction residual is equal to or larger than the threshold (R_TH), generates binary data including a first code (n-bit code) with a fixed number of bits indicating the threshold (R_TH) and a second code (residual code) obtained by binarizing a value obtained by subtracting the threshold (R_TH) from the prediction residual using the Exponential Golomb method. In the arithmetic coding (S3064), the three-dimensional data encoding device uses different arithmetic coding methods for the first code and the second code.
[0583] According to this, the three-dimensional data encoding device can arithmetically encode the first code and the second code using an arithmetic encoding method suitable for each of the first code and the second code, thereby improving encoding efficiency.
[0584] For example, the three-dimensional data encoding device quantizes the prediction residual, and in the binarization (S3063), binarizes the quantized prediction residual. The threshold (R_TH) is changed according to the quantization scale in the quantization. This allows the three-dimensional data encoding device to use an appropriate threshold according to the quantization scale, thereby improving encoding efficiency.
[0585] For example, the second code includes a prefix portion and a suffix portion. In the arithmetic coding (S3064), the three-dimensional data coding device uses different coding tables for the prefix portion and the suffix portion. This allows the three-dimensional data coding device to improve coding efficiency.
[0586] For example, the three-dimensional data encoding device includes a processor and a memory, and the processor uses the memory to perform the above-mentioned processing.
[0587] Moreover, the three-dimensional data decoding device according to this embodiment performs the process shown in FIG. 68. The three-dimensional data decoding device decodes a three-dimensional point having attribute information. First, the three-dimensional data decoding device calculates a predicted value of the attribute information of the three-dimensional point (S3071). Next, the three-dimensional data decoding device generates binary data by arithmetically decoding the encoded data included in the bit stream (S3072). Next, the three-dimensional data decoding device generates a prediction residual by multi-value encoding the binary data (S3073). Next, the three-dimensional data decoding device calculates a decoded value of the attribute information of the three-dimensional point by adding the predicted value and the prediction residual (S3074).
[0588] With this, the three-dimensional data decoding device can calculate a prediction residual of the attribute information, and further appropriately decode a bit stream of the attribute information generated by binarizing and arithmetically coding the prediction residual.
[0589] For example, in the arithmetic decoding (S3072), the three-dimensional data decoding device uses a different coding table for each bit of binary data. This allows the three-dimensional data decoding device to properly decode a bitstream with improved coding efficiency.
[0590] For example, in arithmetic decoding (S3072), the lower the bit of the binary data, the greater the number of coding tables used.
[0591] For example, in the arithmetic decoding (S3072), the three-dimensional data decoding device selects a coding table to be used for arithmetic decoding of a target bit in accordance with the value of the most significant bit of the target bit included in the binary data. This allows the three-dimensional data decoding device to appropriately decode a bitstream with improved coding efficiency.
[0592] For example, in the multi-value conversion (S3073), the three-dimensional data decoding device generates a first value by multi-value conversion of a first code (n-bit code) having a fixed number of bits included in the binary data. If the first value is smaller than a threshold (R_TH), the three-dimensional data decoding device determines the first value as a prediction residual, and if the first value is equal to or greater than the threshold (R_TH), generates a second value by multi-value conversion of a second code (residual code) which is an exponential Golomb code included in the binary data, and generates a prediction residual by adding the first value and the second value. In the arithmetic decoding (S3072), the three-dimensional data decoding device uses different arithmetic decoding methods for the first code and the second code.
[0593] This allows the three-dimensional data decoding device to properly decode a bitstream with improved coding efficiency.
[0594] For example, the three-dimensional data decoding device dequantizes the prediction residual, and in the addition (S3074), adds the predicted value and the dequantized prediction residual. The threshold value (R_TH) is changed according to the quantization scale in the dequantization. This allows the three-dimensional data decoding device to properly decode a bitstream with improved coding efficiency.
[0595] For example, the second code includes a prefix part and a suffix part. In the arithmetic decoding (S3072), the three-dimensional data decoding device uses different coding tables for the prefix part and the suffix part. This allows the three-dimensional data decoding device to appropriately decode a bit stream with improved coding efficiency.
[0596] For example, the three-dimensional data decoding device includes a processor and a memory, and the processor uses the memory to perform the above-mentioned processing.
[0597] (Embodiment 9) The predicted value may be generated by a method other than that of the embodiment 8. In the following, the 3D point to be coded may be referred to as a first 3D point, and the 3D points surrounding it may be referred to as second 3D points.
[0598] For example, in generating a predicted value of attribute information of a three-dimensional point, among the three-dimensional points around the three-dimensional point to be encoded and decoded, the attribute value of the three-dimensional point that is closest to the three-dimensional point to be encoded may be generated as it is as a predicted value. In addition, in generating a predicted value, prediction mode information (PredMode) may be added to each three-dimensional point, and a predicted value may be generated by selecting one predicted value from a plurality of predicted values. That is, for example, in a total number M of prediction modes, it is possible to assign an average value to prediction mode 0, an attribute value of three-dimensional point A to prediction mode 1, ..., an attribute value of three-dimensional point Z to prediction mode M-1, and add the prediction mode used for prediction to the bit stream for each three-dimensional point. In this way, the first prediction mode value indicating the first prediction mode in which the average of the attribute information of the surrounding three-dimensional points is calculated as a predicted value may be smaller than the second prediction mode value indicating the second prediction mode in which the attribute information of the surrounding three-dimensional points itself is calculated as a predicted value. Here, the "average value" which is the predicted value calculated in prediction mode 0 is the average value of the attribute values of the three-dimensional points around the three-dimensional point to be encoded.
[0599] Fig. 69 is a diagram showing a first example of a table indicating predicted values calculated in each prediction mode according to Embodiment 9. Fig. 70 is a diagram showing an example of attribute information used for predicted values according to Embodiment 9. Fig. 71 is a diagram showing a second example of a table indicating predicted values calculated in each prediction mode according to Embodiment 9.
[0600] The number of prediction modes M may be added to the bit stream. The number of prediction modes M may not be added to the bit stream, and may be specified by a profile, level, or the like of a standard. The number of prediction modes M may be a value calculated from the number of three-dimensional points N used for prediction. For example, the number of prediction modes M may be calculated by M=N+1.
[0601] The table shown in FIG. 69 is an example in which the number of three-dimensional points used for prediction is N=4 and the number of prediction modes is M=5. The predicted value of the attribute information of the point b2 may be generated using the attribute information of the points a0, a1, a2, and b1. When selecting one prediction mode from a plurality of prediction modes, a prediction mode may be selected that generates the attribute values of each of the points a0, a1, a2, and b1 as predicted values based on distance information from the point b2 to each of the points a0, a1, a2, and b1. The prediction mode is added to each three-dimensional point to be coded. The predicted value is calculated according to a value according to the added prediction mode.
[0602] The table shown in FIG. 71 is an example in which the number of three-dimensional points used for prediction N=4 and the number of prediction modes M=5, similarly to FIG. 69. A predicted value of the attribute information of the point a2 may be generated using the attribute information of the points a0 and a1. When selecting one prediction mode from a plurality of prediction modes, a prediction mode may be selected that generates the attribute values of the points a0 and a1 as predicted values based on distance information from the point a2 to the points a0 and a1. A prediction mode is added to each three-dimensional point to be coded. A predicted value is calculated according to a value according to the added prediction mode.
[0603] When the number of adjacent points, that is, the number of surrounding three-dimensional points N, is less than four, as in the case of point a2 above, a prediction mode to which a prediction value is not assigned in the table may be set as "not available."
[0604] The allocation of the prediction mode values may be determined in the order of distance from the three-dimensional point to be encoded. For example, the prediction mode value indicating a plurality of prediction modes is smaller as the distance from the three-dimensional point to be encoded to the surrounding three-dimensional points having attribute information used as a predicted value is closer. In the example of FIG. 69, the order of points b1, a2, a1, and a0 indicates that the distance to point b2, which is the three-dimensional point to be encoded, is closer. For example, in the calculation of the predicted value, the attribute information of point b1 is calculated as a predicted value in a prediction mode in which the prediction mode value of two or more prediction modes is indicated as "1", and the attribute information of point a2 is calculated as a predicted value in a prediction mode in which the prediction mode value is indicated as "2". In this way, the prediction mode value indicating the prediction mode in which the attribute information of point b1 is calculated as a predicted value is smaller than the prediction mode value indicating the prediction mode in which the attribute information of point a2, which is located at a position farther away from point b2 than point b1, is calculated as a predicted value.
[0605] This allows a small prediction mode value to be assigned to points that are likely to be selected because they are close to each other, and the number of bits required to encode the prediction mode value can be reduced. Also, a small prediction mode value may be preferentially assigned to 3D points that belong to the same LoD as the 3D point to be encoded.
[0606] 72 is a diagram showing a third example of a table showing predicted values calculated in each prediction mode according to embodiment 9. Specifically, the third example is an example in which attribute information used for the predicted value is a value based on color information (YUV) of surrounding three-dimensional points. In this way, the attribute information used for the predicted value may be color information indicating the color of the three-dimensional point.
[0607] As shown in FIG. 72, the predicted value calculated in the prediction mode indicated by the prediction mode value "0" is the average of each of the YUV components that define the YUV color space. Specifically, the predicted value includes a weighted average Yave of Yb1, Ya2, Ya1, and Ya0, which are Y component values corresponding to the points b1, a2, a1, and a0, respectively, a weighted average Uave of Ub1, Ua2, Ua1, and Ua0, which are U component values corresponding to the points b1, a2, a1, and a0, respectively, and a weighted average Vave of Vb1, Va2, Va1, and Va0, which are V component values corresponding to the points b1, a2, a1, and a0, respectively. In addition, the predicted value calculated in the prediction mode indicated by the prediction mode values "1" to "4" includes color information of the surrounding three-dimensional points b1, a2, a1, and a0, respectively. The color information is indicated by a combination of the values of the Y component, the U component, and the V component.
[0608] In FIG. 72, the color information is shown using values defined in the YUV color space, but it is not limited to the YUV color space, and may be shown using values defined in the RGB color space or in another color space.
[0609] In this way, in the calculation of the predicted value, two or more averages or attribute information may be calculated as the predicted value of the prediction mode. Furthermore, the two or more averages or attribute information may each indicate values of two or more components that define the color space.
[0610] For example, when a prediction mode indicated by a prediction mode value of "2" in the table of Fig. 72 is selected, the Y component, U component, and V component of the attribute value of the 3D point to be encoded may be used as predicted values Ya2, Ua2, and Va2, respectively, for encoding. In this case, "2" as the prediction mode value is added to the bit stream.
[0611] 73 is a diagram showing a fourth example of a table indicating predicted values calculated in each prediction mode according to embodiment 9. Specifically, the fourth example is an example in which attribute information used for the predicted value is a value based on reflectance information of surrounding three-dimensional points. The reflectance information is, for example, information indicating a reflectance R.
[0612] 73, the predicted value calculated in the prediction mode indicated by the prediction mode value "0" is the weighted average Rave of the reflectances Rb1, Ra2, Ra1, and Ra0 corresponding to the points b1, a2, a1, and a0, respectively. Also, the predicted values calculated in the prediction modes indicated by the prediction mode values "1" to "4" are the reflectances Rb1, Ra2, Ra1, and Ra0 of the surrounding three-dimensional points b1, a2, a1, and a0, respectively.
[0613] For example, when a prediction mode indicated by a prediction mode value of "3" in the table of Fig. 73 is selected, the reflectance of the attribute value of the 3D point to be coded may be used as a predicted value Ra1 for coding. In this case, the prediction mode value "3" is added to the bit stream.
[0614] As shown in Fig. 72 and Fig. 73, the attribute information may include first attribute information and second attribute information of a type different from the first attribute information. The first attribute information is, for example, color information. The second attribute information is, for example, reflectance information. In calculating the predicted value, the first predicted value may be calculated using the first attribute information, and the second predicted value may be calculated using the second attribute information.
[0615] (Embodiment 10) As another example of encoding attribute information of a three-dimensional point using information of LoD, a method of encoding a plurality of three-dimensional points in order from the three-dimensional point included in the upper layer of LoD will be described. For example, when calculating a predicted value of an attribute value (attribute information) of a three-dimensional point included in LoDn, a three-dimensional data encoding device may use a flag or the like to switch which LoD the attribute value of the three-dimensional point may be referred to. For example, the three-dimensional data encoding device generates EnableReferringSameLoD (same layer reference permission flag), which is information indicating whether or not to permit reference to other three-dimensional points in the same LoD as the target three-dimensional point to be encoded. For example, when EnableReferringSameLoD is a value of 1, reference within the same LoD is permitted, and when EnableReferringSameLoD is a value of 0, reference within the same LoD is prohibited.
[0616] For example, the three-dimensional data encoding device selects three-dimensional points around the target three-dimensional point based on EnableReferringSameLoD, and calculates an average of attribute values of a predetermined number of N or less three-dimensional points among the selected surrounding three-dimensional points to generate a predicted value of attribute information of the target three-dimensional point. The three-dimensional data encoding device also adds the value of N to a header of a bit stream or the like. The three-dimensional data encoding device may add the value of N to each three-dimensional point for which a predicted value is generated. This allows an appropriate N to be selected for each three-dimensional point for which a predicted value is generated, thereby improving the accuracy of the predicted value and reducing the prediction residual.
[0617] Alternatively, the three-dimensional data encoding device may add the value of N to the header of the bitstream and fix the value of N in the bitstream. This eliminates the need to encode or decode the value of N for each three-dimensional point, thereby reducing the amount of processing.
[0618] Alternatively, the three-dimensional data encoding device may encode information indicating the value of N separately for each LoD. This allows the encoding efficiency to be improved by selecting an appropriate value of N for each LoD. The three-dimensional data encoding device may calculate a predicted value of attribute information of a three-dimensional point from a weighted average value of attribute information of N surrounding three-dimensional points. For example, the three-dimensional data encoding device calculates the weight using distance information between the target three-dimensional point and each of the N three-dimensional points.
[0619] Thus, EnableReferringSameLoD is information indicating whether or not to allow reference of 3D points in the same LoD. For example, a value of 1 indicates that reference is possible, and a value of 0 indicates that reference is not possible. In addition, in the case of a value of 1, among 3D points in the same LoD, 3D points that have already been encoded or decoded may be referenced.
[0620] FIG. 74 is a diagram showing an example of a reference relationship when EnableReferringSameLoD = 0. The predicted value of point P included in LoDN is generated using the reconstructed value P' included in LoDN' (N' < N) at a higher layer than LoDN. Here, the reconstructed value P' is an attribute value (attribute information) that has been encoded and decoded. For example, the reconstructed value P' of adjacent points based on distance is used.
[0621] Also, in the example shown in FIG. 74, for example, the predicted value of b2 is generated using any of the attribute values of a0, a1, and a2. Even when b0 and b1 have been encoded and decoded, reference to b0 and b1 is prohibited.
[0622] As a result, the three-dimensional point data encoding device and the three-dimensional data decoding device can generate the predicted value of b2 without waiting for the encoding or decoding process of b0 and b1 to be completed. That is, the three-dimensional point data encoding device and the three-dimensional data decoding device can calculate a plurality of predicted values for the attribute values of a plurality of three-dimensional points within the same LoD in parallel, thereby reducing the processing time.
[0623] FIG. 75 is a diagram showing an example of a reference relationship when EnableReferringSameLoD = 1. The predicted value of point P included in LoDN is generated using the reconstructed value P' included in LoDN' (N' ≤ N) at the same layer or a higher layer than LoDN. Here, the reconstructed value P' is an attribute value (attribute information) that has been encoded and decoded. For example, the reconstructed value P' of adjacent points based on distance is used.
[0624] Also, in the example shown in FIG. 75, for example, the predicted value of b2 is generated using any of the attribute values of a0, a1, a2, b0, and b1. That is, when b0 and b1 have already been encoded and decoded, they can be referenced.
[0625] As a result, the three-dimensional data encoding device can generate the predicted value of b2 using the attribute information of many adjacent three-dimensional points. Therefore, the prediction accuracy is improved and the encoding efficiency is improved.
[0626] Hereinafter, a method for limiting the number of searches when selecting N 3D points used to generate predicted values of attribute information of 3D points will be described. This makes it possible to reduce the amount of processing.
[0627] For example, SearchNumPoint (search point number information) is defined. SearchNumPoint indicates the number of searches when selecting N three-dimensional points to be used for prediction from the three-dimensional point group in the LoD. For example, the three-dimensional data encoding device may select the same number of three-dimensional points as the number indicated by SearchNumPoint from a total of T three-dimensional points included in the LoD, and select N three-dimensional points to be used for prediction from the selected three-dimensional points. This eliminates the need for the three-dimensional data encoding device to search all T three-dimensional points included in the LoD, thereby reducing the amount of processing.
[0628] The three-dimensional data encoding device may select the value of SearchNumPoint in accordance with the position of the LoD to be referenced. An example is shown below.
[0629] For example, when the reference LoD, which is the LoD to be referred to, is a higher hierarchical layer than the LoD to which the target 3D point belongs, the three-dimensional data encoding device searches for 3D point A that is closest to the target 3D point among the 3D points included in the reference LoD. Next, the three-dimensional data encoding device selects the number of 3D points indicated by SearchNumPoint that are adjacent before and after 3D point A. This allows the three-dimensional data encoding device to efficiently search for 3D points in higher layers that are close to the target 3D point, thereby improving prediction efficiency.
[0630] For example, when the reference LoD is the same layer as the LoD to which the target 3D point belongs, the 3D data encoding device selects the number of 3D points indicated by SearchNumPoint that have been encoded and decoded before the target 3D point. For example, the 3D data encoding device selects the number of 3D points indicated by SearchNumPoint that have been encoded and decoded immediately before the target 3D point.
[0631] This allows the three-dimensional data encoding device to select the number of three-dimensional points indicated by SearchNumPoint with a low processing load. The three-dimensional data encoding device may select three-dimensional point B that is close to the target three-dimensional point from among the three-dimensional points encoded and decoded before the target three-dimensional point, and select the number of three-dimensional points that are adjacent before and after three-dimensional point B, which is indicated by SearchNumPoint. This allows the three-dimensional data encoding device to efficiently search for three-dimensional points in the same layer that are close to the target three-dimensional point, thereby improving prediction efficiency.
[0632] Furthermore, when selecting N 3D points to be used for prediction from the number of 3D points indicated by SearchNumPoint, the 3D data encoding device may select, for example, the top N 3D points closest to the target 3D point. This can improve prediction accuracy and therefore encoding efficiency.
[0633] Note that SearchNumPoint may be prepared for each LoD, and the number of searches may be changed for each LoD. Fig. 76 is a diagram showing an example of setting the number of searches for each LoD. For example, as shown in Fig. 76, SearchNumPoint[LoD0]=3 for LoD0 and SearchNumPoint[LoD1]=2 for LoD1 are defined. In this way, by switching the number of searches for each LoD, it is possible to balance the amount of processing and the coding efficiency.
[0634] In the example shown in Fig. 76, a0, a1, and a2 are selected from LoD0 as 3D points used for predicting b2, and b0 and b1 are selected from LoD1. N 3D points are selected from the selected a0, a1, a2, b0, and b1, and a predicted value is generated using the selected N 3D points.
[0635] In addition, the predicted value of a point P included in the LoDN is generated using a reconstructed value P' included in the LoDN' (N'≦N) in the same layer as the LoDN or in an upper layer. Here, the reconstructed value P' is an attribute value (attribute information) that has been coded and decoded. For example, the reconstructed value P' of an adjacent point based on the distance is used.
[0636] In addition, SearchNumPoint may indicate the total number of searches in all LoDs. For example, when SearchNumPoint=5, if three searches are performed in LoD0, the remaining two searches can be performed in LoD1. This ensures the worst number of searches, stabilizing the processing time.
[0637] The three-dimensional data encoding device may add SearchNumPoint to a header or the like. This allows the three-dimensional data decoding device to generate the same predicted value as the three-dimensional data encoding device by decoding SearchNumPoint from the header, and to properly decode the bit stream. Furthermore, SearchNumPoint does not necessarily have to be added to the header, and for example, the value of SearchNumPoint may be specified by a profile or level of a standard or the like. This allows the amount of bits in the header to be reduced.
[0638] When calculating a predicted value of an attribute value of a 3D point included in LoDn, the following EnableReferenceLoD may be defined. This allows a 3D data encoding device and a 3D data decoding device to refer to EnableReferenceLoD (reference permission layer information) and determine which LoD contains the attribute value of the 3D point that may be referenced.
[0639] EnableReferenceLoD is inf...
Claims
1. An encoding method executed by an encoding device that selects a 3D point for generating a predicted value used to predict attribute information of a 3D point to be predicted, the method comprising: Evaluating at least one of the plurality of candidate 3D points; determining whether to include the evaluated 3D point candidate in a set of 3D points for generating a predicted value, the set being composed of N 3D points, based on a result of the evaluation; At least one of the set of 3D points for generating the predicted value is selected using a Morton code assigned to the 3D point candidates. Encoding method.
2. In the determination, if the evaluation value of the evaluated 3D point candidate is equal to or smaller than the smallest evaluation value of the evaluation values of the 3D points included in the set of 3D points for generating the predicted value, the evaluated 3D point candidate is not included in the set of 3D points for generating the predicted value. The encoding method according to claim 1 .
3. In the determination, if an evaluation value of the evaluated 3D point candidate is greater than the smallest evaluation value of evaluation values of 3D points included in the set of 3D points for generating the predicted value, the evaluated 3D point candidate is included in the set of 3D points for generating the predicted value.
3. The encoding method according to claim 1 or 2.
4. When the evaluated 3D point candidate is to be included in the set of 3D points for generating the predicted value, the 3D point having the smallest evaluation value that is already included in the set of 3D points for generating the predicted value is instead removed from the set of 3D points for generating the predicted value. The encoding method according to claim 3.
5. The evaluation value of the evaluated 3D point candidate is calculated based on the distance between the prediction target 3D point and the evaluation target 3D point candidate. The encoding method according to claim 1 .
6. The plurality of 3D point candidates are 3D points that belong to a higher hierarchy than the hierarchy to which the 3D point to be predicted belongs. The encoding method according to claim 1 .
7. A decoding method executed by a decoding device that selects a 3D point for generating a predicted value used to predict attribute information of a 3D point to be predicted, the method comprising: Evaluating at least one of the plurality of candidate 3D points; determining whether to include the evaluated 3D point candidate in a set of 3D points for generating a predicted value, the set being composed of N 3D points, based on a result of the evaluation; At least one of the set of 3D points for generating the predicted value is selected using a Morton code assigned to the 3D point candidates. Decryption method.
8. In the determination, if the evaluation value of the evaluated 3D point candidate is equal to or smaller than the smallest evaluation value of the evaluation values of the 3D points included in the set of 3D points for generating the predicted value, the evaluated 3D point candidate is not included in the set of 3D points for generating the predicted value. The decoding method according to claim 7.
9. In the determination, if an evaluation value of the evaluated 3D point candidate is greater than the smallest evaluation value of evaluation values of 3D points included in the set of 3D points for generating the predicted value, the evaluated 3D point candidate is included in the set of 3D points for generating the predicted value. A decoding method according to claim 7 or 8.
10. When the evaluated 3D point candidate is to be included in the set of 3D points for generating the predicted value, the 3D point having the smallest evaluation value that is already included in the set of 3D points for generating the predicted value is instead removed from the set of 3D points for generating the predicted value. The decoding method according to claim 9.
11. The evaluation value of the evaluated 3D point candidate is calculated based on the distance between the prediction target 3D point and the evaluation target 3D point candidate. The decoding method according to claim 7.
12. The plurality of 3D point candidates are 3D points that belong to a higher hierarchy than the hierarchy to which the 3D point to be predicted belongs. The decoding method according to claim 7.
13. An encoding device for selecting a 3D point for generating a predicted value used to predict attribute information of a 3D point to be predicted, comprising: A processor; A memory, The processor, using the memory, Evaluating at least one of the plurality of candidate 3D points; determining whether to include the evaluated 3D point candidate in a set of 3D points for generating a predicted value, the set being composed of N 3D points, based on a result of the evaluation; At least one of the set of 3D points for generating the predicted value is selected using a Morton code assigned to the 3D point candidates. Encoding device.
14. A program for causing a computer to execute an encoding method for selecting a 3D point for generating a predicted value used to predict attribute information of a 3D point to be predicted, comprising: Evaluating at least one of the plurality of candidate 3D points; determining whether to include the evaluated 3D point candidate in a set of 3D points for generating a predicted value, the set being composed of N 3D points, based on a result of the evaluation; At least one 3D point among the set of 3D points for generating the predicted value is selected using a Morton code assigned to the 3D point candidate. A program for a computer to execute.
15. A decoding device for selecting a 3D point for generating a predicted value used to predict attribute information of a 3D point to be predicted, comprising: A processor; A memory, The processor, using the memory, Evaluating at least one of the plurality of candidate 3D points; determining whether to include the evaluated 3D point candidate in a set of 3D points for generating a predicted value, the set being composed of N 3D points, based on a result of the evaluation; At least one of the set of 3D points for generating the predicted value is selected using a Morton code assigned to the 3D point candidates. Decryption device.
16. A program for causing a computer to execute a decoding method for selecting a three-dimensional point for generating a predicted value used to predict attribute information of a three-dimensional point to be predicted, Evaluating at least one of the plurality of candidate 3D points; determining whether to include the evaluated 3D point candidate in a set of 3D points for generating a predicted value, the set being composed of N 3D points, based on a result of the evaluation; At least one 3D point among the set of 3D points for generating the predicted value is selected using a Morton code assigned to the 3D point candidate. A program for a computer to execute.
Citation Information
Patent Citations
Region-adaptive hierarchical transform and entropy coding for point cloud compression, and corresponding decompression
US20170347100A1
Scalable point cloud compression with transform, and corresponding decompression
US20170347122A1
Three-dimensional information processing method and three-dimensional information processing apparatus
WO2018038131A1
Serialising a representation of a three dimensional object
WO2018071011A1
Information processing device and method
WO2019012975A1