Three-dimensional data generation method, three-dimensional data acquisition method, three-dimensional data generation device, and three-dimensional data acquisition device
By encoding three-dimensional data into a bitstream with common control information and identifiers for multiple sub-spaces, the method addresses the inefficiencies in existing encoding and decoding processes, resulting in reduced processing requirements for decoding devices.
Patent Information
- Application Number
- JP2024118741
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2018-07-13
- Filing Date
- 2024-07-24
- Publication Date
- 2025-06-30
- Estimated Expiration
- 2039-07-10
AI Technical Summary
Existing methods for encoding and decoding three-dimensional data are inefficient, leading to high processing requirements for decoding devices.
A method that generates a bitstream containing encoded data for multiple sub-spaces within a target three-dimensional space, using common control information and identifiers to reduce processing complexity.
This approach significantly reduces the processing amount required by three-dimensional data decoding devices, enhancing efficiency and performance.
Smart Images

Figure 0007700334000001 
Figure 0007700334000002 
Figure 0007700334000003
Abstract
Description
Technical Field
[0001] The present disclosure relates to a three-dimensional data encoding method, a three-dimensional data decoding method, a three-dimensional data encoding apparatus, and a three-dimensional data decoding apparatus.
Background Art
[0002] In the future, the spread of devices or services that utilize three-dimensional data is expected in a wide range of fields such as computer vision, map information, monitoring, infrastructure inspection, or video distribution for autonomous operation of automobiles or robots. Three-dimensional data is acquired by various methods such as a distance sensor such as a range finder, a stereo camera, or a combination of a plurality of monocular cameras.
[0003] As one of the methods for expressing three-dimensional data, there is a method called point cloud that represents the shape of a three-dimensional structure by a point group in a three-dimensional space. In a point cloud, the positions and colors of the point group are stored. Although the point cloud is expected to become the mainstream as a method for expressing three-dimensional data, the point group has a very large amount of data. Therefore, in the accumulation or transmission of three-dimensional data, similar to two-dimensional moving images (for example, MPEG-4 AVC or HEVC standardized by MPEG), compression of the data amount by encoding is essential.
[0004] Also, regarding the compression of point clouds, it is partially supported by a publicly available library (Point Cloud Library) that performs point cloud-related processing.
[0005] Also, a technique for searching and displaying facilities located around a vehicle using three-dimensional map data is known (see, for example, Patent Document 1).
Prior Art Documents
Patent Documents
[0006]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0007] In the encoding and decoding of three-dimensional data, it is desired to reduce the processing amount of a three-dimensional data decoding device.
[0008] An object of the present disclosure is to provide a three-dimensional data encoding method, a three-dimensional data decoding method, a three-dimensional data encoding device, or a three-dimensional data decoding device capable of reducing the processing amount of a three-dimensional data decoding device.
Means for Solving the Problems
[0009] A three-dimensional data Generate method according to an aspect of the present disclosure of three-dimensional data generates a bit stream including a plurality of sub-spaces included in a target space common first control information, and and a plurality of encoded data each corresponding to the plurality of sub-spaces, and and includes first control information the and information on the plurality of sub-spaces associated with a plurality of identifiers assigned to the plurality of sub-spaces is and headers of each of the plurality of encoded data each corresponding to includes identification is, of a corresponding size sub-space to and includes information for .
[0010] A three-dimensional data Obtain method according to an aspect of the present disclosure of three-dimensional data includes a bit stream including a plurality of sub-spaces included in a target space common first control information, and and a plurality of encoded data each corresponding to the plurality of sub-spaces, and and includes first control information Obtain and information on the plurality of sub-spaces associated with a plurality of identifiers assigned to the plurality of sub-spaces, and headers of each of the plurality of encoded data the including is the first control information each of and the plurality of identifiers assigned to the plurality of sub-spaces is,Correspond to size B space to Identify includes information for None.
Advantages of the Invention
[0011] The present disclosure can provide a three-dimensional data encoding method, a three-dimensional data decoding method, a three-dimensional data encoding device, or a three-dimensional data decoding device that can reduce the processing amount of the three-dimensional data decoding device.
Brief Description of the Drawings
[0012]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
Figure 9
Figure 10
Figure 11
Figure 12
Figure 13
Figure 14
Figure 15
Figure 16
Figure 17
Figure 18
Figure 19
Figure 20
Figure 21
Figure 22
Figure 23
Figure 24
Figure 25
Figure 26
Figure 27
Figure 28
Figure 29
Figure 30
Figure 31
Figure 32
Figure 33
Figure 34
Figure 35
Figure 36
Figure 37
Figure 38
Figure 39
Figure 40
Figure 41
Figure 42
Figure 43
Figure 44
Figure 45
Figure 46
Figure 47
Figure 48
Figure 49
Figure 50
Figure 51
Figure 52
Figure 53
Figure 54
Figure 55
Figure 56
Figure 57
Figure 58
Figure 59
Figure 60
Figure 61
Figure 62
Figure 63
Figure 64
Figure 65
Figure 66
Figure 67
Figure 68
Figure 69
Figure 70
Figure 71
Figure 72
Figure 73
Figure 74
Figure 75
Figure 76
Figure 77
Figure 78
Figure 79
Figure 80
Figure 81
Figure 82
Figure 83
Figure 84
Figure 85
Figure 86
Figure 87
Figure 88
Figure 89
Figure 90
Figure 91
Figure 92
Figure 93
Figure 94
Figure 95
Figure 96
Figure 97
Figure 98
Figure 99
Figure 100
Figure 101
Figure 102
Figure 103
Figure 104
Figure 105
Figure 106
Figure 107
Figure 108
Figure 109
Figure 110
Figure 111
Figure 112
Figure 113
Figure 114
Figure 115
Figure 116
Figure 117
Figure 118
Figure 119
Figure 120
Figure 121
Figure 122
Figure 123
Figure 124
Figure 125
Figure 126
Figure 127
Figure 128
Figure 129
Figure 130
Figure 131
Figure 132
Figure 133
Figure 134
Figure 135
Figure 136
Figure 137
Figure 138
Figure 139
Figure 140
Figure 141
Figure 142
Figure 143
Figure 144
Figure 145
Figure 146
Figure 147
Figure 148
Figure 149
Figure 150
Figure 151
Figure 152
Figure 153
Figure 154
MODE FOR CARRYING OUT THE INVENTION
[0013] A three-dimensional data encoding method according to an aspect of the present disclosure generates a bitstream including a plurality of encoded data corresponding to a plurality of subspaces by encoding the plurality of subspaces included in a target space including a plurality of three-dimensional points. In generating the bitstream, a list of information of the plurality of subspaces associated with a plurality of identifiers assigned to the plurality of subspaces is stored in first control information common to the plurality of encoded data included in the bitstream, and an identifier assigned to the subspace corresponding to the encoded data is stored in the header of each of the plurality of encoded data.
[0014] According to this, when a three-dimensional data decoding device decodes the bitstream generated by the three-dimensional data encoding method, a list of information of the plurality of subspaces associated with the plurality of identifiers stored in the first control information and the identifiers stored in the headers of the plurality of encoded data are referred to, and desired encoded data can be acquired. Thus, the processing amount of the three-dimensional data decoding device can be reduced.
[0015] For example, in the bitstream, the first control information may be arranged before the plurality of encoded data.
[0016] For example, the list may include position information of the plurality of subspaces.
[0017] For example, the list may include size information of the plurality of subspaces.
[0018] For example, the three-dimensional data encoding method may convert the first control information into second control information in the protocol of the system to which the bitstream is transmitted.
[0019] According to this, the three-dimensional data encoding method can convert control information according to the protocol of the system to which the bitstream is transmitted.
[0020] For example, the second control information may be a table for random access in the protocol.
[0021] For example, the second control information may be an mdat box or a track box in ISOBMFF.
[0022] A three-dimensional data decoding method according to an aspect of the present disclosure decodes a bit stream including a plurality of encoded data corresponding to a plurality of subspaces obtained by encoding a plurality of subspaces included in a target space including a plurality of three-dimensional points. In the decoding of the bit stream, a subspace to be decoded among the plurality of subspaces is determined, and a list of information on the plurality of subspaces associated with a plurality of identifiers assigned to the plurality of subspaces, which is included in first control information common to the plurality of encoded data included in the bit stream, and an identifier assigned to the subspace corresponding to the encoded data included in each header of the plurality of encoded data are used to obtain the encoded data of the subspace to be decoded.
[0023] According to this, the three-dimensional data decoding method can obtain desired encoded data by referring to a list of information on a plurality of subspaces associated with a plurality of identifiers stored in the first control information and identifiers stored in each header of the plurality of encoded data. Therefore, the processing amount of the three-dimensional data decoding apparatus can be reduced.
[0024] For example, in the bit stream, the first control information may be arranged before the plurality of encoded data.
[0025] For example, the list may include position information of the plurality of subspaces.
[0026] For example, the list may include size information of the plurality of subspaces.
[0027] Also, a three-dimensional data encoding device according to an aspect of the present disclosure is a three-dimensional data encoding device that encodes a plurality of three-dimensional points having attribute information, and includes a processor and a memory. The processor uses the memory to generate a bitstream including a plurality of encoded data corresponding to a plurality of subspaces by encoding the plurality of subspaces included in a target space including the plurality of three-dimensional points. In the generation of the bitstream, a list of information of the plurality of subspaces associated with a plurality of identifiers assigned to the plurality of subspaces is stored in first control information common to the plurality of encoded data included in the bitstream, and an identifier assigned to the subspace corresponding to the encoded data is stored in the header of each of the plurality of encoded data.
[0028] According to this, when the three-dimensional data decoding device decodes the bitstream generated by the three-dimensional data encoding device, it can acquire desired encoded data by referring to a list of information of the plurality of subspaces associated with the plurality of identifiers stored in the first control information and the identifiers stored in the headers of each of the plurality of encoded data. Therefore, the processing amount of the three-dimensional data decoding device can be reduced.
[0029] Also, a three-dimensional data decoding device according to an aspect of the present disclosure is a three-dimensional data decoding device that decodes a plurality of three-dimensional points having attribute information, and includes a processor and a memory. The processor uses the memory to decode a bitstream including a plurality of encoded data corresponding to a plurality of subspaces obtained by encoding the plurality of subspaces included in a target space including the plurality of three-dimensional points. In the decoding of the bitstream, a subspace to be decoded among the plurality of subspaces is determined, and the encoded data of the subspace to be decoded is acquired using a list of information of the plurality of subspaces associated with a plurality of identifiers assigned to the plurality of subspaces included in first control information common to the plurality of encoded data included in the bitstream and the identifier assigned to the subspace corresponding to the encoded data included in the header of each of the plurality of encoded data.
[0030] According to this, the three-dimensional data decoding apparatus can acquire desired encoded data by referring to a list of information on a plurality of subspaces associated with a plurality of identifiers stored in the first control information and the identifiers stored in the headers of each of the plurality of encoded data. Therefore, the processing amount of the three-dimensional data decoding apparatus can be reduced.
[0031] These general or specific aspects may be implemented in a system, a method, an integrated circuit, a computer program, or a recording medium such as a computer-readable CD-ROM, or may be implemented in any combination of a system, a method, an integrated circuit, a computer program, and a recording medium.
[0032] Hereinafter, embodiments will be specifically described with reference to the drawings. Note that each of the embodiments described below shows a specific example of the present disclosure. The numerical values, shapes, materials, components, arrangement positions and connection forms of the components, steps, order of steps, etc. shown in the following embodiments are merely examples and are not intended to limit the present disclosure. In addition, among the components in the following embodiments, the components not described in the independent claims indicating the highest-level concept are described as optional components.
[0033] (Embodiment 1) First, the data structure of the encoded three-dimensional data (hereinafter also referred to as encoded data) according to the present embodiment will be described. FIG. 1 is a diagram showing the configuration of the encoded three-dimensional data according to the present embodiment.
[0034] In this embodiment, the three-dimensional space is divided into spaces (SPCs) corresponding to pictures in the encoding of a moving image, and three-dimensional data is encoded in units of spaces. Each space is further divided into volumes (VLMs) corresponding to macroblocks and the like in moving image encoding, and prediction and transformation are performed in units of VLMs. A volume includes a plurality of voxels (VXLs) which are the minimum units to which position coordinates are associated. Note that prediction is, similar to the prediction performed on a two-dimensional image, to generate predicted three-dimensional data similar to the processing unit to be processed by referring to other processing units, and to encode the difference between the predicted three-dimensional data and the processing unit to be processed. Also, this prediction includes not only spatial prediction that refers to other prediction units at the same time but also temporal prediction that refers to prediction units at different times.
[0035] For example, when encoding a three-dimensional space represented by point cloud data such as point cloud, a three-dimensional data encoding apparatus (hereinafter also referred to as an encoding apparatus) encodes each point of the point cloud or a plurality of points included in a voxel according to the size of the voxel. If the voxel is subdivided, the three-dimensional shape of the point cloud can be expressed with high precision, and if the size of the voxel is increased, the three-dimensional shape of the point cloud can be expressed roughly.
[0036] Note that hereinafter, the case where the three-dimensional data is point cloud will be described as an example, but the three-dimensional data is not limited to point cloud and may be three-dimensional data in any format.
[0037] Also, hierarchical voxels may be used. In this case, in the n-th hierarchy, it may be sequentially indicated whether there are sample points in the hierarchies below the (n - 1)-th hierarchy (the lower layers of the n-th hierarchy). For example, when decoding only the n-th hierarchy, if there are sample points in the hierarchies below the (n - 1)-th hierarchy, it can be decoded assuming that there are sample points at the center of the voxels in the n-th hierarchy.
[0038] Also, the encoding apparatus acquires point cloud data using a distance sensor, a stereo camera, a monocular camera, a gyro, an inertial sensor, or the like.
[0039] Space is classified into any of at least three prediction structures including an intra-space (I-SPC) that can be decoded independently, a predictive space (P-SPC) that allows only unidirectional reference, and a bidirectional space (B-SPC) that allows bidirectional reference, similar to the encoding of moving images. Also, space has two types of time information: decoding time and display time.
[0040] Also, as shown in FIG. 1, there is a GOS (Group Of Space) which is a random access unit as a processing unit including a plurality of spaces. Further, there is a world (WLD) as a processing unit including a plurality of GOSs.
[0041] The space area occupied by the world is associated with an absolute position on the earth by GPS or latitude and longitude information, etc. This position information is stored as meta information. Note that the meta information may be included in the encoded data or may be transmitted separately from the encoded data.
[0042] Also, within a GOS, all SPCs may be three-dimensionally adjacent, or there may be an SPC that is not three-dimensionally adjacent to other SPCs.
[0043] Note that hereinafter, processing such as encoding, decoding, or reference for three-dimensional data included in a processing unit such as GOS, SPC, or VLM is also simply described as encoding, decoding, or referencing the processing unit. Also, the three-dimensional data included in the processing unit includes at least one set of a spatial position such as three-dimensional coordinates and a characteristic value such as color information.
[0044] Next, the prediction structure of SPC in GOS will be described. A plurality of SPCs within the same GOS, or a plurality of VLMs within the same SPC, occupy different spaces from each other but have the same time information (decoding time and display time).
[0045] Also, the SPC that is the first in the decoding order within the GOS is the I-SPC. Also, there are two types of GOS: closed GOS and open GOS. The closed GOS is a GOS that can decode all the SPCs within the GOS when starting decoding from the first I-SPC. In the open GOS, some SPCs whose display times are earlier than the first I-SPC within the GOS refer to different GOSs, and decoding cannot be performed only with this GOS.
[0046] Note that in encoded data such as map information, the WLD may be decoded in the direction opposite to the encoding order, and reverse playback is difficult if there is a dependency between GOSs. Therefore, in such a case, basically, a closed GOS is used.
[0047] Also, the GOS has a layer structure in the height direction, and encoding or decoding is performed in order from the SPCs in the lower layer.
[0048] FIG. 2 is a diagram showing an example of a prediction structure between SPCs belonging to the bottom layer of the GOS. FIG. 3 is a diagram showing an example of a prediction structure between layers.
[0049] There is one or more I-SPCs within the GOS. In a three-dimensional space, there are objects such as humans, animals, cars, bicycles, signals, or buildings that are landmarks. In particular, objects with a small size are effectively encoded as I-SPCs. For example, when a three-dimensional data decoding device (hereinafter also referred to as a decoding device) decodes the GOS with a low processing amount or at high speed, it decodes only the I-SPCs within the GOS.
[0050] Also, the encoding device may switch the encoding interval or appearance frequency of the I-SPC according to the coarseness of the objects in the WLD.
[0051] Also, in the configuration shown in FIG. 3, the encoding device or the decoding device encodes or decodes a plurality of layers in order from the lower layer (layer 1). Thereby, for example, the priority of data near the ground with a larger amount of information can be increased for an autonomous vehicle or the like.
[0052] In the case of the encoded data used in a drone or the like, in the GOS, encoding or decoding may be performed in order from the SPC of the upper layer in the height direction.
[0053] Further, the encoding device or the decoding device may encode or decode a plurality of layers so that the decoding device can roughly grasp the GOS and gradually increase the resolution. For example, the encoding device or the decoding device may encode or decode in the order of layer 3, 8, 1, 9, ….
[0054] Next, how to handle static objects and dynamic objects will be described.
[0055] In the three-dimensional space, there are static objects or scenes such as buildings or roads (hereinafter collectively referred to as static objects) and dynamic objects such as cars or humans (hereinafter referred to as dynamic objects). Detection of an object is separately performed by extracting feature points from point cloud data or camera images such as a stereo camera. Here, an example of an encoding method for dynamic objects will be described.
[0056] The first method is a method of encoding without distinguishing between static objects and dynamic objects. The second method is a method of distinguishing between static objects and dynamic objects by identification information.
[0057] For example, the GOS is used as an identification unit. In this case, the GOS including the SPC constituting the static object and the GOS including the SPC constituting the dynamic object are distinguished by identification information stored separately in the encoded data or separately from the encoded data.
[0058] Alternatively, the SPC may be used as an identification unit. In this case, the SPC including the VLM constituting the static object and the SPC including the VLM constituting the dynamic object are distinguished by the above identification information.
[0059] Alternatively, VLM or VXL may be used as an identification unit. In this case, a VLM or VXL containing a static object and a VLM or VXL containing a dynamic object are distinguished by the above identification information.
[0060] Further, the encoding device may encode a dynamic object as one or more VLMs or SPCs, and encode a VLM or SPC containing a static object and an SPC containing a dynamic object as different GOSs. Also, when the size of the GOS is variable according to the size of the dynamic object, the encoding device separately stores the size of the GOS as meta information.
[0061] Further, the encoding device may encode a static object and a dynamic object independently of each other, and may superimpose the dynamic object on a world composed of the static object. At this time, the dynamic object is composed of one or more SPCs, and each SPC is associated with one or more SPCs constituting the static object on which the SPC is superimposed. Note that the dynamic object may be represented by one or more VLMs or VXLs instead of SPCs.
[0062] Further, the encoding device may encode a static object and a dynamic object as different streams.
[0063] Further, the encoding device may generate a GOS including one or more SPCs constituting a dynamic object. Furthermore, the encoding device may set a GOS (GOS_M) including a dynamic object and a GOS of a static object corresponding to the spatial region of GOS_M to the same size (occupying the same spatial region). Thereby, superimposition processing can be performed in units of GOS.
[0064] The P-SPC or B-SPC constituting the dynamic object may refer to SPCs included in different encoded GOSs. In the case where the position of the dynamic object changes over time and the same dynamic object is encoded as GOSs at different times, cross-GOS reference is effective from the viewpoint of compression ratio.
[0065] Also, depending on the use of the encoded data, the above first method and second method may be switched. For example, when using the encoded three-dimensional data as a map, it is desirable to be able to separate dynamic objects, so the encoding device uses the second method. On the other hand, when the encoding device encodes three-dimensional data of an event such as a concert or sports, if there is no need to separate dynamic objects, the first method is used.
[0066] Also, the decoding time and display time of GOS or SPC can be stored in the encoded data or as meta information. Also, all the time information of the static objects may be the same. At this time, the actual decoding time and display time may be determined by the decoding device. Alternatively, different values may be assigned for each GOS or SPC as the decoding time, and the same value may be assigned for all as the display time. Furthermore, a model may be introduced that guarantees that the decoder has a buffer of a predetermined size and can decode without breaking by reading the bit stream at a predetermined bit rate according to the decoding time, such as a decoder model in video coding such as HEVC's HRD (Hypothetical Reference Decoder).
[0067] Next, the arrangement of GOS in the world will be described. The coordinates of the three-dimensional space in the world are represented by three mutually orthogonal coordinate axes (x-axis, y-axis, z-axis). By providing a predetermined rule in the encoding order of GOS, encoding can be performed so that spatially adjacent GOS are continuous in the encoded data. For example, in the example shown in FIG. 4, the GOS in the xz plane is encoded continuously. After encoding all the GOS in a certain xz plane, the value of the y-axis is updated. That is, as the encoding progresses, the world extends in the y-axis direction. Also, the index number of GOS is set in the encoding order.
[0068] Here, the three-dimensional space of the world is associated one-to-one with geographical absolute coordinates such as GPS or latitude and longitude. Alternatively, the three-dimensional space may be represented by the relative position from a preset reference position. The directions of the x-axis, y-axis, and z-axis of the three-dimensional space are represented as direction vectors determined based on latitude and longitude, etc., and the direction vectors are stored together with the encoded data as meta information.
[0069] Also, the size of the GOS is fixed, and the encoding device stores the size as meta information. Also, the size of the GOS may be switched according to, for example, whether it is an urban area or whether it is indoor or outdoor. That is, the size of the GOS may be switched according to the amount or nature of the object that is valuable as information. Alternatively, the encoding device may adaptively switch the size of the GOS or the interval of the I-SPC in the GOS according to, for example, the density of the objects in the same world. For example, the higher the density of the objects, the smaller the size of the GOS and the shorter the interval of the I-SPC in the GOS.
[0070] In the example of FIG. 5, in the regions of the 3rd to 10th GOSs, since the density of the objects is high, the GOS is subdivided in order to realize random access with a fine granularity. Note that the 7th to 10th GOSs are respectively located behind the 3rd to 6th GOSs.
[0071] Next, the configuration and operation flow of the three-dimensional data encoding device according to the present embodiment will be described. FIG. 6 is a block diagram of the three-dimensional data encoding device 100 according to the present embodiment. FIG. 7 is a flowchart showing an operation example of the three-dimensional data encoding device 100.
[0072] The three-dimensional data encoding device 100 shown in FIG. 6 generates encoded three-dimensional data 112 by encoding three-dimensional data 111. This three-dimensional data encoding device 100 includes an acquisition unit 101, an encoding region determination unit 102, a division unit 103, and an encoding unit 104.
[0073] As shown in FIG. 7, first, the acquisition unit 101 acquires three-dimensional data 111 which is point cloud data (S101).
[0074] Next, the encoding region determination unit 102 determines the region to be encoded among the spatial regions corresponding to the acquired point cloud data (S102). For example, the encoding region determination unit 102 determines the spatial region around the position as the region to be encoded according to the position of the user or the vehicle.
[0075] Next, the division unit 103 divides the point cloud data included in the region to be encoded into each processing unit. Here, the processing unit is the above-mentioned GOS, SPC, etc. Also, this region to be encoded corresponds to, for example, the above-mentioned world. Specifically, the division unit 103 divides the point cloud data into processing units based on the preset size of the GOS, or the presence or size of the dynamic object (S103). Also, the division unit 103 determines the start position of the SPC that is the head in the encoding order in each GOS.
[0076] Next, the encoding unit 104 generates encoded three-dimensional data 112 by sequentially encoding a plurality of SPCs within each GOS (S104).
[0077] Here, an example in which the region to be encoded is divided into GOS and SPC and then each GOS is encoded is shown, but the processing procedure is not limited to the above. For example, after determining the configuration of one GOS, encoding that GOS, and then using procedures such as determining the configuration of the next GOS may also be used.
[0078] In this way, the three-dimensional data encoding device 100 generates encoded three-dimensional data 112 by encoding three-dimensional data 111. Specifically, the three-dimensional data encoding device 100 divides the three-dimensional data into first processing units (GOS) that are random access units and each of which is associated with three-dimensional coordinates, divides the first processing units (GOS) into a plurality of second processing units (SPC), and divides the second processing units (SPC) into a plurality of third processing units (VLM). Further, the third processing unit (VLM) includes one or more voxels (VXL) that are the minimum units to which position information is associated.
[0079] Next, the three-dimensional data encoding device 100 generates encoded three-dimensional data 112 by encoding each of the plurality of first processing units (GOS). Specifically, the three-dimensional data encoding device 100 encodes each of the plurality of second processing units (SPC) in each first processing unit (GOS). Further, the three-dimensional data encoding device 100 encodes each of the plurality of third processing units (VLM) in each second processing unit (SPC).
[0080] For example, when the first processing unit (GOS) to be processed is a closed GOS, the three-dimensional data encoding device 100 encodes the second processing unit (SPC) to be processed included in the first processing unit (GOS) to be processed with reference to other second processing units (SPC) included in the first processing unit (GOS) to be processed. That is, the three-dimensional data encoding device 100 does not refer to the second processing units (SPC) included in a first processing unit (GOS) different from the first processing unit (GOS) to be processed.
[0081] On the other hand, when the first processing unit (GOS) to be processed is an open GOS, the three-dimensional data encoding device 100 encodes the second processing unit (SPC) to be processed included in the first processing unit (GOS) to be processed with reference to other second processing units (SPC) included in the first processing unit (GOS) to be processed or the second processing units (SPC) included in a first processing unit (GOS) different from the first processing unit (GOS) to be processed.
[0082] Further, the three-dimensional data encoding device 100 selects one of a first type (I-SPC) that does not refer to other second processing units (SPCs), a second type (P-SPC) that refers to one other second processing unit (SPC), and a third type that refers to two other second processing units (SPCs) as the type of the second processing unit (SPC) to be processed, and encodes the second processing unit (SPC) to be processed according to the selected type.
[0083] Next, the configuration and operation flow of the three-dimensional data decoding device according to the present embodiment will be described. FIG. 8 is a block diagram of the blocks of the three-dimensional data decoding device 200 according to the present embodiment. FIG. 9 is a flowchart showing an operation example of the three-dimensional data decoding device 200.
[0084] The three-dimensional data decoding device 200 shown in FIG. 8 generates decoded three-dimensional data 212 by decoding the encoded three-dimensional data 211. Here, the encoded three-dimensional data 211 is, for example, the encoded three-dimensional data 112 generated by the three-dimensional data encoding device 100. This three-dimensional data decoding device 200 includes an acquisition unit 201, a decoding start GOS determination unit 202, a decoding SPC determination unit 203, and a decoding unit 204.
[0085] First, the acquisition unit 201 acquires the encoded three-dimensional data 211 (S201). Next, the decoding start GOS determination unit 202 determines the GOS to be decoded (S202). Specifically, the decoding start GOS determination unit 202 refers to the meta information stored within the encoded three-dimensional data 211 or separately from the encoded three-dimensional data, and determines the GOS including the SPC corresponding to the spatial position, object, or time at which decoding is to start as the GOS to be decoded.
[0086] Next, the decoding SPC determination unit 203 determines the type (I, P, B) of the SPC to be decoded within the GOS (S203). For example, the decoding SPC determination unit 203 determines whether to decode (1) only I-SPCs, (2) I-SPCs and P-SPCs, or (3) all types. Note that if the type of SPC to be decoded in advance, such as decoding all SPCs, has been determined, this step may not be performed.
[0087] Next, the decoding unit 204 acquires the address position where the first SPC to be decoded in decoding order (the same as the encoding order) within the GOS starts in the encoded three-dimensional data 211, acquires the encoded data of the first SPC from the address position, and sequentially decodes each SPC in order from the first SPC (S204). Note that the above address position is stored in meta information or the like.
[0088] In this way, the three-dimensional data decoding device 200 decodes the decoded three-dimensional data 212. Specifically, the three-dimensional data decoding device 200 generates the decoded three-dimensional data 212 of the first processing unit (GOS) by decoding each of the encoded three-dimensional data 211 of the first processing unit (GOS) which is a random access unit and each of which is associated with a three-dimensional coordinate. More specifically, the three-dimensional data decoding device 200 decodes each of the plurality of second processing units (SPCs) in each first processing unit (GOS). Further, the three-dimensional data decoding device 200 decodes each of the plurality of third processing units (VLM) in each second processing unit (SPC).
[0089] Hereinafter, the meta information for random access will be described. This meta information is generated by the three-dimensional data encoding device 100 and is included in the encoded three-dimensional data 112 (211).
[0090] In the random access in the conventional two-dimensional moving image, decoding starts from the first frame of the random access unit near the specified time. On the other hand, in the world, in addition to time, random access to space (coordinates or objects, etc.) is assumed.
[0091] Therefore, in order to realize random access to at least three elements: coordinates, objects, and time, a table is prepared to associate each element with the index number of the GOS. Further, the index number of the GOS is associated with the address of the first I-SPC of the GOS. FIG. 10 is a diagram showing an example of the table included in the meta information. Note that it is not necessary to use all the tables shown in FIG. 10, and at least one table may be used.
[0092] Hereinafter, as an example, random access starting from coordinates will be described. When accessing the coordinates (x2, y2, z2), first, by referring to the coordinate-GOS table, it can be found that the point with the coordinates (x2, y2, z2) is included in the second GOS. Next, by referring to the GOS address table, it is found that the address of the first I-SPC in the second GOS is addr(2). Therefore, the decoding unit 204 acquires data from this address and starts decoding.
[0093] Note that the address may be an address in the logical format or a physical address of the HDD or memory. Also, information specifying a file segment may be used instead of the address. For example, a file segment is a unit obtained by segmenting one or more GOSs, etc.
[0094] Also, when an object spans multiple GOSs, in the object-GOS table, multiple GOSs to which the object belongs may be shown. If the multiple GOSs are closed GOSs, the encoding device and the decoding device can perform encoding or decoding in parallel. On the other hand, if the multiple GOSs are open GOSs, the compression efficiency can be improved by the multiple GOSs referring to each other.
[0095] Examples of objects include a human, an animal, a car, a bicycle, a signal, or a landmark building. For example, the three-dimensional data encoding device 100 can extract feature points unique to an object from a three-dimensional point cloud or the like during world encoding, detect the object based on the feature points, and set the detected object as a random access point.
[0096] In this way, the three-dimensional data encoding device 100 generates first information indicating a plurality of first processing units (GOS) and three-dimensional coordinates associated with each of the plurality of first processing units (GOS). Further, the encoded three-dimensional data 112(211) includes this first information. Further, the first information further indicates at least one of an object, a time, and a data storage destination associated with each of the plurality of first processing units (GOS).
[0097] The three-dimensional data decoding device 200 acquires the first information from the encoded three-dimensional data 211, uses the first information to identify the encoded three-dimensional data 211 of the first processing unit corresponding to the specified three-dimensional coordinates, object, or time, and decodes the encoded three-dimensional data 211.
[0098] Hereinafter, examples of other meta information will be described. In addition to the meta information for random access, the three-dimensional data encoding device 100 may generate and store the following meta information. Further, the three-dimensional data decoding device 200 may use this meta information during decoding.
[0099] When using three-dimensional data as map information, etc., a profile is defined according to the application, and information indicating the profile may be included in the meta information. For example, profiles for urban areas or suburbs, or for flying objects are defined, and the maximum or minimum size of the world, SPC, or VLM, etc. is defined in each case. For example, for urban areas, more detailed information is required than for suburbs, so the minimum size of the VLM is set smaller.
[0100] The meta information may include a tag value indicating the type of the object. This tag value is associated with the VLM, SPC, or GOS that constitutes the object. For example, the tag value "0" indicates "person", the tag value "1" indicates "car", the tag value "2" indicates "traffic signal", etc. The tag value may be set for each type of object. Alternatively, when the type of the object is difficult to determine or there is no need to determine it, a tag value indicating a property such as the size or whether the object is a dynamic object or a static object may be used.
[0101] In addition, the meta information may include information indicating the range of the space area occupied by the world.
[0102] In addition, the meta information may store the size of the SPC or VXL as common header information for the entire stream of encoded data or for a plurality of SPCs such as the SPC within the GOS.
[0103] In addition, the meta information may include identification information of a distance sensor or a camera used for generating the point cloud, or information indicating the positional accuracy of the point group within the point cloud.
[0104] In addition, the meta information may include information indicating whether the world is composed only of static objects or includes dynamic objects.
[0105] Hereinafter, a modification of the present embodiment will be described.
[0106] The encoding device or the decoding device may encode or decode two or more different SPCs or GOSs in parallel. The GOSs to be encoded or decoded in parallel can be determined based on meta information indicating the spatial position of the GOSs.
[0107] In a case where the three-dimensional data is used as a spatial map when a car or a flying object moves, or such a spatial map is generated, the encoding device or the decoding device may encode or decode the GOS or SPC included in the space specified based on GPS, route information, or zoom ratio, etc.
[0108] Also, the decoding device may perform decoding in order from a space close to its own position or the traveling route. The encoding device or the decoding device may perform encoding or decoding on a space far from its own position or the traveling route with a lower priority than a close space. Here, lowering the priority means lowering the processing order, lowering the resolution (subsampling for processing), or lowering the image quality (increasing the encoding efficiency. For example, increasing the quantization step).
[0109] Also, when decoding encoded data hierarchically encoded in a space, the decoding device may decode only the lower layer.
[0110] Also, the decoding device may preferentially decode from the lower layer according to the zoom ratio or use of the map.
[0111] Also, in applications such as self-position estimation or object recognition performed during autonomous driving of a vehicle or a robot, the encoding device or the decoding device may perform encoding or decoding with a reduced resolution outside a region within a specific height from the road surface (the region where recognition is performed).
[0112] Also, the encoding device may separately encode point clouds representing the spatial shapes of the indoor and outdoor areas. For example, by separating the GOS representing the indoor area (indoor GOS) and the GOS representing the outdoor area (outdoor GOS), the decoding device can select the GOS to be decoded according to the viewpoint position when using the encoded data.
[0113] Also, the encoding device may encode an indoor GOS and an outdoor GOS with close coordinates so that they are adjacent in the encoding stream. For example, the encoding device associates the identifiers of both and stores information indicating the associated identifiers in the encoding stream or in separately stored meta information. Thereby, the decoding device can identify the indoor GOS and the outdoor GOS with close coordinates by referring to the information in the meta information.
[0114] In addition, the encoding device may switch the size of GOS or SPC between indoor GOS and outdoor GOS. For example, the encoding device sets the size of GOS to be smaller indoors than outdoors. Also, the encoding device may change the accuracy when extracting feature points from the point cloud, or the accuracy of object detection, etc., between indoor GOS and outdoor GOS.
[0115] In addition, the encoding device may add information for the decoding device to distinguish dynamic objects from static objects to the encoded data. As a result, the decoding device can display the dynamic object together with a red frame or explanatory text, etc. Note that the decoding device may display only the red frame or explanatory text instead of the dynamic object. Also, the decoding device may display more detailed object types. For example, a red frame may be used for a vehicle, and a yellow frame may be used for a person.
[0116] Also, the encoding device or the decoding device may determine whether to encode or decode dynamic objects and static objects as different SPCs or GOSs according to the appearance frequency of the dynamic objects, or the ratio between static objects and dynamic objects, etc. For example, when the appearance frequency or ratio of the dynamic objects exceeds a threshold, an SPC or GOS in which dynamic objects and static objects are mixed is allowed, and when the appearance frequency or ratio of the dynamic objects does not exceed the threshold, an SPC or GOS in which dynamic objects and static objects are mixed is not allowed.
[0117] When detecting dynamic objects from the two-dimensional image information of the camera instead of the point cloud, the encoding device may separately acquire information (such as a frame or text) for identifying the detection result and the object position, and encode these information as part of the three-dimensional encoded data. In this case, the decoding device superimposes and displays auxiliary information (a frame or text) indicating the dynamic object on the decoding result of the static object.
[0118] Further, the encoding device may change the coarseness of VXL or VLM in SPC according to, for example, the complexity of the shape of the static object. For example, the more complex the shape of the static object is, the denser the VXL or VLM is set. Further, the encoding device may determine, according to the coarseness of VXL or VLM, quantization steps when quantizing spatial position or color information. For example, the encoding device sets a smaller quantization step as VXL or VLM is denser.
[0119] As described above, the encoding device or decoding device according to the present embodiment performs spatial encoding or decoding in space units having coordinate information.
[0120] Further, the encoding device and the decoding device perform encoding or decoding in volume units within the space. A volume includes voxels which are the minimum units to which position information is associated.
[0121] Further, the encoding device and the decoding device perform encoding or decoding by associating arbitrary elements with each other using a table that associates each element of spatial information including coordinates, objects, time, etc. with a GOP, or a table that associates between each element. Further, the decoding device determines coordinates using the value of the selected element, specifies a volume, voxel or space from the coordinates, and decodes the space including the volume or voxel, or the specified space.
[0122] Further, the encoding device determines a volume, voxel or space that can be selected by an element by feature point extraction or object recognition, and encodes it as a volume, voxel or space that can be randomly accessed.
[0123] Spaces are classified into three types: I-SPC that can be encoded or decoded in the space unit itself, P-SPC that is encoded or decoded with reference to any one processed space, and B-SPC that is encoded or decoded with reference to any two processed spaces.
[0124] One or more volumes correspond to static objects or dynamic objects. The space including static objects and the space including dynamic objects are encoded or decoded as different GOSs. That is, the SPC including static objects and the SPC including dynamic objects are assigned different GOSs.
[0125] Dynamic objects are encoded or decoded for each object and associated with one or more spaces including static objects. That is, a plurality of dynamic objects are encoded individually, and the encoded data of the obtained plurality of dynamic objects is associated with the SPC including static objects.
[0126] The encoding device and the decoding device perform encoding or decoding by increasing the priority of the I-SPC in the GOS. For example, the encoding device performs encoding so that the deterioration of the I-SPC is reduced (so that the original three-dimensional data is reproduced more faithfully after decoding). Also, the decoding device decodes only the I-SPC, for example.
[0127] The encoding device may perform encoding by changing the frequency of using the I-SPC according to the density or number (quantity) of objects in the world. That is, the encoding device changes the frequency of selecting the I-SPC according to the number or density of objects included in the three-dimensional data. For example, the encoding device increases the frequency of using the I-space as the objects in the world are denser.
[0128] Also, the encoding device sets random access points in units of GOS and stores information indicating the space area corresponding to the GOS in the header information.
[0129] The encoding device uses, for example, a default value as the space size of the GOS. Note that the encoding device may change the size of the GOS according to the number (quantity) or density of objects or dynamic objects. For example, the encoding device reduces the space size of the GOS as the objects or dynamic objects are denser or the number is larger.
[0130] In addition, the space or volume includes a feature point group derived using information obtained from sensors such as a depth sensor, a gyroscope, or a camera. The coordinates of the feature points are set at the center positions of the voxels. Also, high-precision position information can be realized by subdividing the voxels.
[0131] The feature point group is derived using a plurality of pictures. The plurality of pictures have at least two types of time information, namely actual time information and the same time information (for example, the encoding time used for rate control, etc.) in a plurality of pictures associated with the space.
[0132] Also, encoding or decoding is performed in GOS units including one or more spaces.
[0133] The encoding device and the decoding device predict the P space or the B space in the GOS to be processed by referring to the spaces in the processed GOS.
[0134] Alternatively, the encoding device and the decoding device predict the P space or the B space in the GOS to be processed using the processed spaces in the GOS to be processed without referring to different GOSs.
[0135] Also, the encoding device and the decoding device transmit or receive an encoded stream in world units including one or more GOSs.
[0136] Also, the GOS has at least a layer structure in one direction within the world, and the encoding device and the decoding device perform encoding or decoding from the lower layer. For example, a randomly accessible GOS belongs to the lowest layer. The GOSs belonging to the upper layer refer to the GOSs belonging to the same layer or below. That is, the GOS is spatially divided in a predetermined direction and includes a plurality of layers each containing one or more SPCs. The encoding device and the decoding device encode or decode each SPC by referring to the SPCs included in the same layer or a lower layer than the SPC.
[0137] Also, the encoding device and the decoding device encode or decode GOSs continuously within a world unit including a plurality of GOSs. The encoding device and the decoding device write or read, as metadata, information indicating the order (direction) of encoding or decoding. That is, the encoded data includes information indicating the encoding order of a plurality of GOSs.
[0138] Also, the encoding device and the decoding device encode or decode two or more different spaces or GOSs in parallel.
[0139] Also, the encoding device and the decoding device encode or decode the spatial information (coordinates, size, etc.) of a space or GOS.
[0140] Also, the encoding device and the decoding device encode or decode a space or GOS included in a specific space specified based on external information such as GPS, route information, or magnification regarding their own position or / and region size.
[0141] The encoding device or the decoding device encodes or decodes a space far from its own position with a lower priority compared to a nearby space.
[0142] The encoding device sets one direction of the world according to magnification or use, and encodes a GOS having a layer structure in that direction. Also, the decoding device preferentially decodes a GOS having a layer structure in one direction of the world set according to magnification or use, starting from the lower layer.
[0143] The encoding device changes the extraction of feature points included in a space, the accuracy of object recognition, or the spatial region size, etc. between indoors and outdoors. However, the encoding device and the decoding device encode or decode adjacent indoor GOSs and outdoor GOSs with close coordinates within the world, and also encode or decode their identifiers in association.
[0144] (Embodiment 2) When using the encoded data of the point cloud in an actual device or service, it is desirable to transmit and receive the information necessary according to the application in order to suppress the network bandwidth. However, until now, such a function has not existed in the encoding structure of three-dimensional data, nor has there been an encoding method therefor.
[0145] In the present embodiment, a three-dimensional data encoding method, a three-dimensional data encoding apparatus for providing a function of transmitting and receiving only the information necessary according to the application in the encoded data of the three-dimensional point cloud, and a three-dimensional data decoding method and a three-dimensional data decoding apparatus for decoding the encoded data will be described.
[0146] A voxel (VXL) having a feature amount equal to or more than a certain value is defined as a feature voxel (FVXL), and a world (WLD) composed of FVXLs is defined as a sparse world (SWLD). FIG. 11 is a diagram showing a configuration example of the sparse world and the world. The SWLD includes a FGOS that is a GOS composed of FVXLs, a FSPC that is an SPC composed of FVXLs, and a FVLM that is a VLM composed of FVXLs. The data structure and prediction structure of the FGOS, FSPC, and FVLM may be the same as those of the GOS, SPC, and VLM.
[0147] The feature amount is a feature amount representing the three-dimensional position information of the VXL or the visible light information of the VXL position, and is a feature amount that is particularly frequently detected at the corners and edges of three-dimensional objects. Specifically, this feature amount is a three-dimensional feature amount or a visible light feature amount as described below, but any feature amount that represents the position, luminance, or color information of the VXL may be used.
[0148] As the three-dimensional feature amount, a SHOT feature amount (Signature of Histograms of OrienTations), a PFH feature amount (Point Feature Histograms), or a PPF feature amount (Point Pair Feature) is used.
[0149] The SHOT feature quantity is obtained by dividing the periphery of the VXL, calculating the inner product of the reference point and the normal vector of the divided region, and then creating a histogram. This SHOT feature quantity has the characteristics of a high number of dimensions and high feature representation ability.
[0150] The PFH feature quantity is obtained by selecting a large number of point pairs in the vicinity of the VXL, calculating the normal vector and the like from those two points, and then creating a histogram. Since this PFH feature quantity is a histogram feature, it has robustness against some disturbances and also has the feature of high feature representation ability.
[0151] The PPF feature quantity is a feature quantity calculated using the normal vector and the like for each pair of two VXLs. Since all VXLs are used for this PPF feature quantity, it has robustness against occlusion.
[0152] Also, as feature quantities of visible light, SIFT (Scale-Invariant Feature Transform), SURF (Speeded Up Robust Features), or HOG (Histogram of Oriented Gradients) using information such as the luminance gradient information of the image can be used.
[0153] SWLD is generated by calculating the above-mentioned feature quantity from each VXL of the WLD and extracting the FVXL. Here, SWLD may be updated every time the WLD is updated, or it may be updated periodically after a certain period of time regardless of the update timing of the WLD.
[0154] SWLD may be generated for each feature quantity. For example, separate SWLDs may be generated for each feature quantity, such as SWLD1 based on the SHOT feature quantity and SWLD2 based on the SIFT feature quantity, and the SWLDs may be used appropriately according to the application. Also, the feature quantity of each calculated FVXL may be held in each FVXL as feature quantity information.
[0155] Next, the method of using the sparse world (SWLD) will be described. Since SWLD only includes feature voxels (FVXL), it generally has a smaller data size compared to the WLD that includes all voxels.
[0156] In an application that achieves some purpose using feature quantities, by using the information of SWLD instead of WLD, the read time from the hard disk, as well as the bandwidth and transfer time during network transfer, can be suppressed. For example, as map information, both WLD and SWLD are held in the server, and by switching the map information to be transmitted to WLD or SWLD according to the request from the client, the network bandwidth and transfer time can be suppressed. Specific examples will be shown below.
[0157] Figures 12 and 13 are diagrams showing usage examples of SWLD and WLD. As shown in Figure 12, when the client 1, which is an in-vehicle device, needs map information for self-position determination, the client 1 sends a request to the server to acquire map data for self-position estimation (S301). The server transmits SWLD to the client 1 according to the acquisition request (S302). The client 1 performs self-position determination using the received SWLD (S303). At this time, the client 1 acquires VXL information around the client 1 by various methods such as a distance sensor such as a range finder, a stereo camera, or a combination of a plurality of monocular cameras, and estimates the self-position information from the obtained VXL information and SWLD. Here, the self-position information includes the three-dimensional position information and orientation of the client 1.
[0158] As shown in FIG. 13, when the client 2, which is an in-vehicle device, needs map information for map drawing applications such as a three-dimensional map, the client 2 sends a request to the server to obtain map data for map drawing (S311). The server transmits the WLD to the client 2 in response to the acquisition request (S312). The client 2 performs map drawing using the received WLD (S313). At this time, the client 2 creates a rendering image using, for example, an image captured by its own visible light camera or the like and the WLD obtained from the server, and draws the created image on a screen such as a car navigation system.
[0159] As described above, the server transmits the SWLD to the client for applications that mainly require feature amounts of each VXL such as self-position estimation, and transmits the WLD to the client when detailed VXL information is required such as map drawing. This enables efficient transmission and reception of map data.
[0160] Note that the client may determine which of the SWLD and the WLD is required by itself and request the server to transmit the SWLD or the WLD. Also, the server may determine which of the SWLD or the WLD should be transmitted according to the client or network situation.
[0161] Next, a method for switching the transmission and reception of the sparse world (SWLD) and the world (WLD) will be described.
[0162] It may be possible to switch to receive WLD or SWLD according to the network bandwidth. FIG. 14 is a diagram showing an operation example in this case. For example, when a low-speed network with limited available network bandwidth such as in an LTE (Long Term Evolution) environment is used, the client accesses the server via the low-speed network (S321) and acquires SWLD as map information from the server (S322). On the other hand, when a high-speed network with sufficient network bandwidth such as in a Wi-Fi (registered trademark) environment is used, the client accesses the server via the high-speed network (S323) and acquires WLD from the server (S324). Thereby, the client can acquire appropriate map information according to the network bandwidth of the client.
[0163] Specifically, the client receives SWLD via LTE outdoors and acquires WLD via Wi-Fi (registered trademark) when entering indoors such as in a facility. Thereby, the client can acquire more detailed map information indoors.
[0164] In this way, the client may request WLD or SWLD from the server according to the bandwidth of the network it uses. Or, the client may send information indicating the bandwidth of the network it uses to the server, and the server may send data (WLD or SWLD) suitable for the client according to the information. Or, the server may determine the network bandwidth of the client and send data (WLD or SWLD) suitable for the client.
[0165] Alternatively, it may be possible to switch whether to receive the WLD or SWLD according to the moving speed. FIG. 15 is a diagram showing an operation example in this case. For example, when the client is moving at high speed (S331), the client receives the SWLD from the server (S332). On the other hand, when the client is moving at low speed (S333), the client receives the WLD from the server (S334). Thereby, the client can acquire map information suitable for the speed while suppressing the network bandwidth. Specifically, while driving on a highway, the client can update the rough map information at an appropriate speed by receiving the SWLD with a small data amount. On the other hand, while driving on a general road, the client can acquire more detailed map information by receiving the WLD.
[0166] In this way, the client may request the WLD or SWLD from the server according to its own moving speed. Alternatively, the client may transmit information indicating its own moving speed to the server, and the server may transmit data (WLD or SWLD) suitable for the client according to the information. Alternatively, the server may determine the moving speed of the client and transmit data (WLD or SWLD) suitable for the client.
[0167] Also, the client may first acquire the SWLD from the server and then acquire the WLD of important areas therein. For example, when acquiring map data, the client first acquires rough map information with the SWLD, then narrows down the areas where features such as buildings, signs, or people appear frequently, and later acquires the WLD of the narrowed-down areas. Thereby, the client can acquire detailed information of necessary areas while suppressing the amount of received data from the server.
[0168] In addition, the server may create separate SWLDs for each object from the WLD, and the client may receive each of them according to the application. This can suppress the network bandwidth. For example, the server recognizes a person or a vehicle in advance from the WLD and creates a person's SWLD and a vehicle's SWLD. The client receives the person's SWLD when it wants to obtain information about the surrounding people, and receives the vehicle's SWLD when it wants to obtain vehicle information. Also, such types of SWLDs may be distinguished by information (such as a flag or type) added to the header or the like.
[0169] Next, the configuration and operation flow of the three-dimensional data encoding device (for example, a server) according to the present embodiment will be described. FIG. 16 is a block diagram of a three-dimensional data encoding device 400 according to the present embodiment. FIG. 17 is a flowchart of the three-dimensional data encoding process by the three-dimensional data encoding device 400.
[0170] The three-dimensional data encoding device 400 shown in FIG. 16 generates encoded three-dimensional data 413 and 414, which are encoded streams, by encoding the input three-dimensional data 411. Here, the encoded three-dimensional data 413 is the encoded three-dimensional data corresponding to the WLD, and the encoded three-dimensional data 414 is the encoded three-dimensional data corresponding to the SWLD. This three-dimensional data encoding device 400 includes an acquisition unit 401, an encoding region determination unit 402, an SWLD extraction unit 403, a WLD encoding unit 404, and an SWLD encoding unit 405.
[0171] As shown in FIG. 17, first, the acquisition unit 401 acquires the input three-dimensional data 411, which is point cloud data in the three-dimensional space (S401).
[0172] Next, the encoding region determination unit 402 determines the spatial region to be encoded based on the spatial region where the point cloud data exists (S402).
[0173] Next, the SWLD extraction unit 403 defines the spatial region to be encoded as the WLD, and calculates feature amounts from each VXL included in the WLD. Then, the SWLD extraction unit 403 extracts VXLs whose feature amounts are equal to or greater than a predetermined threshold value, defines the extracted VXLs as FVXLs, and generates the extracted three-dimensional data 412 by adding the FVXLs to the SWLD (S403). That is, the extracted three-dimensional data 412 whose feature amounts are equal to or greater than the threshold value is extracted from the input three-dimensional data 411.
[0174] Next, the WLD encoding unit 404 generates the encoded three-dimensional data 413 corresponding to the WLD by encoding the input three-dimensional data 411 corresponding to the WLD (S404). At this time, the WLD encoding unit 404 adds information for distinguishing that the encoded three-dimensional data 413 is a stream including the WLD to the header of the encoded three-dimensional data 413.
[0175] Also, the SWLD encoding unit 405 generates the encoded three-dimensional data 414 corresponding to the SWLD by encoding the extracted three-dimensional data 412 corresponding to the SWLD (S405). At this time, the SWLD encoding unit 405 adds information for distinguishing that the encoded three-dimensional data 414 is a stream including the SWLD to the header of the encoded three-dimensional data 414.
[0176] Note that the processing order of the process of generating the encoded three-dimensional data 413 and the process of generating the encoded three-dimensional data 414 may be reversed from the above. Also, some or all of these processes may be performed in parallel.
[0177] As information attached to the headers of the encoded three-dimensional data 413 and 414, for example, a parameter called "world_type" is defined. When world_type = 0, it indicates that the stream includes WLD, and when world_type = 1, it indicates that the stream includes SWLD. When defining a number of other types, the numerical value assigned like world_type = 2 can be increased as well. Also, a specific flag may be included in one of the encoded three-dimensional data 413 and 414. For example, a flag indicating that the stream includes SWLD may be attached to the encoded three-dimensional data 414. In this case, the decoding device can determine whether the stream includes WLD or SWLD based on the presence or absence of the flag.
[0178] Also, the encoding method used when the WLD encoding unit 404 encodes WLD and the encoding method used when the SWLD encoding unit 405 encodes SWLD may be different.
[0179] For example, since data is decimated in SWLD, the correlation with surrounding data may be lower than that in WLD. Therefore, in the encoding method used for SWLD, inter prediction among intra prediction and inter prediction may be prioritized over the encoding method used for WLD.
[0180] Also, the method of expressing the three-dimensional position may be different between the encoding method used for SWLD and the encoding method used for WLD. For example, in SWLD, the three-dimensional position of FVXL may be expressed by three-dimensional coordinates, and in WLD, the three-dimensional position may be expressed by an octree described later, or vice versa.
[0181] Also, the SWLD encoding unit 405 performs encoding so that the data size of the encoded three-dimensional data 414 of SWLD is smaller than the data size of the encoded three-dimensional data 413 of WLD. For example, as described above, SWLD may have lower correlation between data compared to WLD. As a result, the encoding efficiency decreases, and the data size of the encoded three-dimensional data 414 may become larger than the data size of the encoded three-dimensional data 413 of WLD. Therefore, when the data size of the obtained encoded three-dimensional data 414 is larger than the data size of the encoded three-dimensional data 413 of WLD, the SWLD encoding unit 405 regenerates the encoded three-dimensional data 414 with a reduced data size by performing re-encoding.
[0182] For example, the SWLD extraction unit 403 regenerates the extracted three-dimensional data 412 with a reduced number of feature points to be extracted, and the SWLD encoding unit 405 encodes the extracted three-dimensional data 412. Alternatively, the degree of quantization in the SWLD encoding unit 405 may be made coarser. For example, in the octree structure described later, the degree of quantization can be made coarser by rounding the data in the bottom layer.
[0183] Also, when the SWLD encoding unit 405 cannot make the data size of the encoded three-dimensional data 414 of SWLD smaller than the data size of the encoded three-dimensional data 413 of WLD, it may not be necessary to generate the encoded three-dimensional data 414 of SWLD. Alternatively, the encoded three-dimensional data 413 of WLD may be copied to the encoded three-dimensional data 414 of SWLD. That is, the encoded three-dimensional data 413 of WLD may be used as it is as the encoded three-dimensional data 414 of SWLD.
[0184] Next, the configuration and operation flow of the three-dimensional data decoding device (for example, a client) according to the present embodiment will be described. FIG. 18 is a block diagram of the three-dimensional data decoding device 500 according to the present embodiment. FIG. 19 is a flowchart of the three-dimensional data decoding process by the three-dimensional data decoding device 500.
[0185] The three-dimensional data decoding device 500 shown in FIG. 18 generates decoded three-dimensional data 512 or 513 by decoding the encoded three-dimensional data 511. Here, the encoded three-dimensional data 511 is, for example, the encoded three-dimensional data 413 or 414 generated by the three-dimensional data encoding device 400.
[0186] This three-dimensional data decoding device 500 includes an acquisition unit 501, a header analysis unit 502, a WLD decoding unit 503, and a SWLD decoding unit 504.
[0187] As shown in FIG. 19, first, the acquisition unit 501 acquires the encoded three-dimensional data 511 (S501). Next, the header analysis unit 502 analyzes the header of the encoded three-dimensional data 511 to determine whether the encoded three-dimensional data 511 is a stream including WLD or a stream including SWLD (S502). For example, the above-described parameter of world_type is referred to for the determination.
[0188] When the encoded three-dimensional data 511 is a stream including WLD (Yes in S503), the WLD decoding unit 503 generates the decoded three-dimensional data 512 of WLD by decoding the encoded three-dimensional data 511 (S504). On the other hand, when the encoded three-dimensional data 511 is a stream including SWLD (No in S503), the SWLD decoding unit 504 generates the decoded three-dimensional data 513 of SWLD by decoding the encoded three-dimensional data 511 (S505).
[0189] Also, similar to the encoding device, the decoding method used when the WLD decoding unit 503 decodes WLD and the decoding method used when the SWLD decoding unit 504 decodes SWLD may be different. For example, in the decoding method used for SWLD, inter prediction among intra prediction and inter prediction may be prioritized over the decoding method used for WLD.
[0190] Also, the method of decoding used in SWLD and the method of decoding used in WLD may differ in the method of expressing the three-dimensional position. For example, in SWLD, the three-dimensional position of FVXL may be expressed by three-dimensional coordinates, and in WLD, the three-dimensional position may be expressed by an octree described later, or vice versa.
[0191] Next, the octree representation, which is a method of expressing the three-dimensional position, will be described. The VXL data included in the three-dimensional data is converted into an octree structure and then encoded. FIG. 20 is a diagram showing an example of VXL of WLD. FIG. 21 is a diagram showing the octree structure of the WLD shown in FIG. 20. In the example shown in FIG. 20, there are three VXLs 1 to 3 that are VXLs (hereinafter, valid VXLs) including a point group. As shown in FIG. 21, the octree structure is composed of nodes and leaves. Each node has up to eight nodes or leaves. Each leaf has VXL information. Here, among the leaves shown in FIG. 21, leaves 1, 2, and 3 represent VXLs 1, 2, and 3 shown in FIG. 20, respectively.
[0192] Specifically, each node and leaf corresponds to a three-dimensional position. Node 1 corresponds to the entire block shown in FIG. 20. The block corresponding to Node 1 is divided into eight blocks. Among the eight blocks, the blocks containing valid VXL are set as nodes, and the other blocks are set as leaves. The block corresponding to the node is further divided into eight nodes or leaves, and this process is repeated for the hierarchy of the tree structure. Also, all the blocks in the bottom layer are set as leaves.
[0193] Further, FIG. 22 is a diagram showing an example of the SWLD generated from the WLD shown in FIG. 20. VXL1 and VXL2 shown in FIG. 20 are determined as FVXL1 and FVXL2 as a result of feature extraction and are added to the SWLD. On the other hand, VXL3 is not determined as FVXL and is not included in the SWLD. FIG. 23 is a diagram showing the octree structure of the SWLD shown in FIG. 22. In the octree structure shown in FIG. 23, leaf 3 corresponding to VXL3 shown in FIG. 21 is deleted. As a result, node 3 shown in FIG. 21 no longer has valid VXLs and is changed to a leaf. In general, the number of leaves of the SWLD is thus smaller than the number of leaves of the WLD, and the encoded three-dimensional data of the SWLD is also smaller than the encoded three-dimensional data of the WLD.
[0194] Next, a modification of this embodiment will be described.
[0195] For example, when a client such as an in-vehicle device performs self-position estimation, it receives the SWLD from the server and uses the SWLD to perform self-position estimation. When performing obstacle detection, the client may perform obstacle detection based on three-dimensional information of the surroundings obtained by itself using various methods such as a distance sensor such as a range finder, a stereo camera, or a combination of a plurality of monocular cameras.
[0196] Also, generally, the SWLD is unlikely to include VXL data of flat areas. Therefore, the server may hold a sub-sampled world (subWLD) obtained by sub-sampling the WLD for detecting static obstacles, and transmit the SWLD and the subWLD to the client. Thereby, while suppressing the network bandwidth, self-position estimation and obstacle detection can be performed on the client side.
[0197] Also, when a client renders three-dimensional map data at high speed, it may be more convenient if the map information has a mesh structure. Therefore, the server may generate a mesh from the WLD and hold it in advance as a Mesh World (MWLD). For example, the client may receive the MWLD when it needs rough three-dimensional rendering, and receive the WLD when it needs detailed three-dimensional rendering. This can suppress the network bandwidth.
[0198] Also, among each VXL, the server has set the VXL whose feature amount is equal to or greater than the threshold as the FVXL, but the FVXL may be calculated by different methods. For example, the server may determine that the VXL, VLM, SPC, or GOS that make up a signal or an intersection is necessary for self-position estimation, driving assistance, or autonomous driving, etc., and include them in the SWLD as FVXL, FVLM, FSPC, FGOS. Also, the above determination may be made manually. In addition, the FVXL obtained by the above method may be added to the FVXL etc. set based on the feature amount. That is, the SWLD extraction unit 403 may further extract, as the extracted three-dimensional data 412, data corresponding to an object having a predetermined attribute from the input three-dimensional data 411.
[0199] Also, it may be labeled separately from the feature amount to indicate the necessity for those applications. Also, the server may separately hold the FVXL necessary for self-position estimation, driving assistance, or autonomous driving, etc., such as signals or intersections, as the upper layer of the SWLD (for example, the lane world).
[0200] Also, the server may add attributes to the VXLs in the WLD for each random access unit or predetermined unit. The attributes include, for example, information indicating whether it is necessary or unnecessary for self-position estimation, or information indicating whether it is important as traffic information such as signals or intersections. Also, the attributes may include the correspondence relationship with Features (such as intersections or roads) in lane information (such as GDF: Geographic Data Files).
[0201] Also, the following methods may be used as the method for updating the WLD or SWLD.
[0202] Update information indicating changes in people, construction, or trees (for trucks), etc. is uploaded to the server as point clouds or metadata. Based on the upload, the server updates the WLD, and then updates the SWLD using the updated WLD.
[0203] Also, when the client detects an inconsistency between the three-dimensional information generated by itself during self-position estimation and the three-dimensional information received from the server, the client may send the three-dimensional information generated by itself to the server together with an update notification. In this case, the server updates the SWLD using the WLD. If the SWLD is not updated, the server determines that the WLD itself is old.
[0204] Also, although information for distinguishing between the WLD and the SWLD is added as the header information of the encoded stream, for example, when there are multiple types of worlds such as a mesh world or a lane world, information for distinguishing them may be added to the header information. Also, when there are a large number of SWLDs with different feature amounts, information for distinguishing each of them may be added to the header information.
[0205] Also, although the SWLD is assumed to be composed of FVXLs, it may include VXLs that are not determined to be FVXLs. For example, the SWLD may include adjacent VXLs used when calculating the feature amounts of the FVXLs. Thereby, even when feature amount information is not added to each FVXL of the SWLD, the client can calculate the feature amounts of the FVXLs when receiving the SWLD. In that case, the SWLD may include information for distinguishing whether each VXL is an FVXL or a VXL.
[0206] As described above, the three-dimensional data encoding device 400 extracts the extracted three-dimensional data 412 (second three-dimensional data) whose feature amount is equal to or greater than the threshold value from the input three-dimensional data 411 (first three-dimensional data), and generates the encoded three-dimensional data 414 (first encoded three-dimensional data) by encoding the extracted three-dimensional data 412.
[0207] According to this, the three-dimensional data encoding device 400 generates the encoded three-dimensional data 414 obtained by encoding the data whose feature amount is equal to or greater than the threshold value. Thereby, the data amount can be reduced as compared with the case where the input three-dimensional data 411 is encoded as it is. Therefore, the three-dimensional data encoding device 400 can reduce the data amount to be transmitted.
[0208] In addition, the three-dimensional data encoding device 400 further generates encoded three-dimensional data 413 (second encoded three-dimensional data) by encoding the input three-dimensional data 411.
[0209] According to this, the three-dimensional data encoding device 400 can selectively transmit the encoded three-dimensional data 413 and the encoded three-dimensional data 414, for example, according to the usage purpose and the like.
[0210] In addition, the extracted three-dimensional data 412 is encoded by the first encoding method, and the input three-dimensional data 411 is encoded by a second encoding method different from the first encoding method.
[0211] According to this, the three-dimensional data encoding device 400 can use an encoding method suitable for each of the input three-dimensional data 411 and the extracted three-dimensional data 412.
[0212] In addition, in the first encoding method, inter prediction is prioritized over intra prediction among the intra prediction and the inter prediction as compared with the second encoding method.
[0213] According to this, the three-dimensional data encoding device 400 can increase the priority of inter prediction for the extracted three-dimensional data 412 in which the correlation between adjacent data is likely to be low.
[0214] In addition, the first encoding method and the second encoding method differ in the method of expressing three-dimensional positions. For example, in the second encoding method, the three-dimensional position is expressed by an octree, and in the first encoding method, the three-dimensional position is expressed by three-dimensional coordinates.
[0215] According to this, the three-dimensional data encoding device 400 can use a more suitable method of expressing three-dimensional positions for three-dimensional data with different numbers of data (the number of VXLs or FVXLs).
[0216] In addition, at least one of the encoded three-dimensional data 413 and 414 includes an identifier indicating whether the encoded three-dimensional data is the encoded three-dimensional data obtained by encoding the input three-dimensional data 411 or the encoded three-dimensional data obtained by encoding a part of the input three-dimensional data 411. That is, the identifier indicates whether the encoded three-dimensional data is the encoded three-dimensional data 413 of the WLD or the encoded three-dimensional data 414 of the SWLD.
[0217] According to this, the decoding device can easily determine whether the acquired encoded three-dimensional data is the encoded three-dimensional data 413 or the encoded three-dimensional data 414.
[0218] In addition, the three-dimensional data encoding device 400 encodes the extracted three-dimensional data 412 so that the data amount of the encoded three-dimensional data 414 is smaller than the data amount of the encoded three-dimensional data 413.
[0219] According to this, the three-dimensional data encoding device 400 can make the data amount of the encoded three-dimensional data 414 smaller than the data amount of the encoded three-dimensional data 413.
[0220] In addition, the three-dimensional data encoding device 400 further extracts, as the extracted three-dimensional data 412, data corresponding to an object having a predetermined attribute from the input three-dimensional data 411. For example, an object having a predetermined attribute is an object necessary for self-position estimation, driving assistance, or autonomous driving, such as a signal or an intersection.
[0221] According to this, the three-dimensional data encoding device 400 can generate encoded three-dimensional data 414 including data required by the decoding device.
[0222] Further, the three-dimensional data encoding device 400 (server) further transmits one of the encoded three-dimensional data 413 and 414 to the client according to the state of the client.
[0223] According to this, the three-dimensional data encoding device 400 can transmit appropriate data according to the state of the client.
[0224] Further, the state of the client includes the communication status of the client (for example, network bandwidth) or the moving speed of the client.
[0225] Further, the three-dimensional data encoding device 400 further transmits one of the encoded three-dimensional data 413 and 414 to the client according to the request of the client.
[0226] According to this, the three-dimensional data encoding device 400 can transmit appropriate data according to the request of the client.
[0227] Further, the three-dimensional data decoding device 500 according to the present embodiment decodes the encoded three-dimensional data 413 or 414 generated by the three-dimensional data encoding device 400.
[0228] That is, the three-dimensional data decoding device 500 decodes the encoded three-dimensional data 414 obtained by encoding the extracted three-dimensional data 412 whose feature amount extracted from the input three-dimensional data 411 is equal to or greater than the threshold value by the first decoding method. Further, the three-dimensional data decoding device 500 decodes the encoded three-dimensional data 413 obtained by encoding the input three-dimensional data 411 by a second decoding method different from the first decoding method.
[0229] According to this, the three-dimensional data decoding device 500 can selectively receive the encoded three-dimensional data 414 obtained by encoding data with a feature amount equal to or greater than a threshold value and the encoded three-dimensional data 413, for example, according to the usage purpose or the like. Thereby, the three-dimensional data decoding device 500 can reduce the amount of data to be transmitted. Further, the three-dimensional data decoding device 500 can use decoding methods suitable for the input three-dimensional data 411 and the extracted three-dimensional data 412, respectively.
[0230] Also, in the first decoding method, inter prediction among intra prediction and inter prediction is prioritized over the second decoding method.
[0231] According to this, the three-dimensional data decoding device 500 can increase the priority of inter prediction for the extracted three-dimensional data in which the correlation between adjacent data is likely to be low.
[0232] Also, in the first decoding method and the second decoding method, the expression method of the three-dimensional position is different. For example, in the second decoding method, the three-dimensional position is expressed by an octree, and in the first decoding method, the three-dimensional position is expressed by three-dimensional coordinates.
[0233] According to this, the three-dimensional data decoding device 500 can use a more suitable expression method of the three-dimensional position for three-dimensional data with different numbers of data (the number of VXLs or FVXLs).
[0234] Further, at least one of the encoded three-dimensional data 413 and 414 includes an identifier indicating whether the encoded three-dimensional data is the encoded three-dimensional data obtained by encoding the input three-dimensional data 411 or the encoded three-dimensional data obtained by encoding a part of the input three-dimensional data 411. The three-dimensional data decoding device 500 identifies the encoded three-dimensional data 413 and 414 with reference to the identifier.
[0235] According to this, the three-dimensional data decoding device 500 can easily determine whether the acquired encoded three-dimensional data is the encoded three-dimensional data 413 or the encoded three-dimensional data 414.
[0236] Further, the three-dimensional data decoding device 500 further notifies the server of the state of the client (the three-dimensional data decoding device 500). The three-dimensional data decoding device 500 receives one of the encoded three-dimensional data 413 and 414 transmitted from the server according to the state of the client.
[0237] According to this, the three-dimensional data decoding device 500 can receive appropriate data according to the state of the client.
[0238] Also, the state of the client includes the communication status of the client (for example, network bandwidth) or the moving speed of the client.
[0239] Further, the three-dimensional data decoding device 500 further requests one of the encoded three-dimensional data 413 and 414 from the server, and receives one of the encoded three-dimensional data 413 and 414 transmitted from the server in response to the request.
[0240] According to this, the three-dimensional data decoding device 500 can receive appropriate data according to the application.
[0241] (Embodiment 3) In this embodiment, a method for transmitting and receiving three-dimensional data between vehicles will be described. For example, three-dimensional data is transmitted and received between the host vehicle and surrounding vehicles.
[0242] FIG. 24 is a block diagram of a three-dimensional data creation device 620 according to this embodiment. This three-dimensional data creation device 620 is included in, for example, the host vehicle, and creates denser third three-dimensional data 636 by synthesizing the received second three-dimensional data 635 with the first three-dimensional data 632 created by the three-dimensional data creation device 620.
[0243] This three-dimensional data creation device 620 includes a three-dimensional data creation unit 621, a request range determination unit 622, a search unit 623, a reception unit 624, a decoding unit 625, and a synthesis unit 626.
[0244] First, the three-dimensional data creation unit 621 creates first three-dimensional data 632 using sensor information 631 detected by sensors provided in the host vehicle. Next, the required range determination unit 622 determines a required range, which is a three-dimensional space range lacking data, from among the created first three-dimensional data 632.
[0245] Next, the search unit 623 searches for surrounding vehicles that possess three-dimensional data within the required range, and transmits required range information 633 indicating the required range to the surrounding vehicles identified by the search. Next, the reception unit 624 receives encoded three-dimensional data 634, which is an encoded stream of the required range, from the surrounding vehicles (S624). Note that the search unit 623 may issue requests indiscriminately to all vehicles existing within a specific range, and receive the encoded three-dimensional data 634 from the party that responds. Further, the search unit 623 may issue requests not only to vehicles, but also to objects such as traffic signals or signs, and receive the encoded three-dimensional data 634 from the object.
[0246] Next, the decoding unit 625 obtains second three-dimensional data 635 by decoding the received encoded three-dimensional data 634. Next, the combining unit 626 creates denser third three-dimensional data 636 by combining the first three-dimensional data 632 and the second three-dimensional data 635.
[0247] Next, the configuration and operation of the three-dimensional data transmission device 640 according to the present embodiment will be described. FIG. 25 is a block diagram of the three-dimensional data transmission device 640.
[0248] The three-dimensional data transmission device 640 is included in, for example, the surrounding vehicles described above, processes fifth three-dimensional data 652 created by the surrounding vehicles into sixth three-dimensional data 654 required by the host vehicle, generates encoded three-dimensional data 634 by encoding the sixth three-dimensional data 654, and transmits the encoded three-dimensional data 634 to the host vehicle.
[0249] The three-dimensional data transmission device 640 includes a three-dimensional data creation unit 641, a reception unit 642, an extraction unit 643, an encoding unit 644, and a transmission unit 645.
[0250] First, the three-dimensional data creation unit 641 creates fifth three-dimensional data 652 using the sensor information 651 detected by sensors provided in surrounding vehicles. Next, the reception unit 642 receives the requested range information 633 transmitted from the host vehicle.
[0251] Next, the extraction unit 643 processes the fifth three-dimensional data 652 into sixth three-dimensional data 654 by extracting the three-dimensional data within the requested range indicated by the requested range information 633 from the fifth three-dimensional data 652. Next, the encoding unit 644 encodes the sixth three-dimensional data 654 to generate encoded three-dimensional data 634 which is an encoded stream. Then, the transmission unit 645 transmits the encoded three-dimensional data 634 to the host vehicle.
[0252] Here, an example is described in which the host vehicle includes a three-dimensional data creation device 620 and surrounding vehicles include a three-dimensional data transmission device 640, but each vehicle may have the functions of the three-dimensional data creation device 620 and the three-dimensional data transmission device 640.
[0253] (Embodiment 4) In the present embodiment, abnormal operations in self-position estimation based on a three-dimensional map will be described.
[0254] It is expected that applications such as automatic driving of vehicles, or autonomous movement of moving bodies such as robots or flying objects such as drones will expand in the future. As an example of a means for realizing such autonomous movement, there is a method in which a moving body travels according to a map while estimating its own position (self-position estimation) within a three-dimensional map.
[0255] Self-position estimation can be realized by matching a three-dimensional map with three-dimensional information around the host vehicle (hereinafter, host vehicle detection three-dimensional data) acquired by sensors such as a range finder (such as LiDAR) or a stereo camera mounted on the host vehicle, and estimating the position of the host vehicle within the three-dimensional map.
[0256] The three-dimensional map may include not only three-dimensional point clouds, but also two-dimensional map data such as the shape information of roads and intersections, or information that changes in real time such as traffic jams and accidents, like the HD map proposed by HERE. The three-dimensional map is composed of multiple layers such as three-dimensional data, two-dimensional data, and metadata that changes in real time, and the device can also acquire or refer to only the necessary data.
[0257] The data of the point cloud may be the above-mentioned SWLD, or may include point cloud data that is not feature points. Also, the transmission and reception of the data of the point cloud are performed based on one or more random access units.
[0258] The following method can be used as a method for matching the three-dimensional map and the three-dimensional data of the host vehicle detection. For example, the device compares the shapes of the point clouds in the respective point clouds, and determines that the part with a high similarity between the feature points is the same position. Also, when the three-dimensional map is composed of SWLD, the device performs matching by comparing the feature points that make up the SWLD with the three-dimensional feature points extracted from the three-dimensional data of the host vehicle detection.
[0259] Here, in order to perform highly accurate self-position estimation, it is necessary that (A) the three-dimensional map and the three-dimensional data of the host vehicle detection can be acquired, and (B) their accuracy meets a predetermined standard. However, in the following abnormal cases, (A) or (B) cannot be satisfied.
[0260] (1) The three-dimensional map cannot be acquired via communication.
[0261] (2) The three-dimensional map does not exist, or the three-dimensional map has been acquired but is damaged.
[0262] (3) The sensors of the host vehicle are malfunctioning, or due to bad weather, the generation accuracy of the three-dimensional data of the host vehicle detection is not sufficient.
[0263] The operations for handling these abnormal cases will be described below. In the following, the operations will be described by taking a vehicle as an example. However, the following method can be applied to all moving objects that move autonomously, such as robots or drones.
[0264] Next, the configuration and operations of the three-dimensional information processing apparatus according to the present embodiment for coping with abnormal cases in the three-dimensional map or the ego-vehicle detection three-dimensional data will be described. FIG. 26 is a block diagram showing a configuration example of the three-dimensional information processing apparatus 700 according to the present embodiment.
[0265] The three-dimensional information processing apparatus 700 is mounted on a moving object such as an automobile, for example. As shown in FIG. 26, the three-dimensional information processing apparatus 700 includes a three-dimensional map acquisition unit 701, an ego-vehicle detection data acquisition unit 702, an abnormal case determination unit 703, a coping operation determination unit 704, and an operation control unit 705.
[0266] Note that the three-dimensional information processing apparatus 700 may include a two-dimensional or one-dimensional sensor (not shown) for detecting structures or moving objects around the ego-vehicle, such as a camera that acquires two-dimensional images, or a sensor for one-dimensional data using ultrasonic waves or lasers. Further, the three-dimensional information processing apparatus 700 may include a communication unit (not shown) for acquiring the three-dimensional map via a mobile communication network such as 4G or 5G, or vehicle-to-vehicle communication or road-to-vehicle communication.
[0267] The three-dimensional map acquisition unit 701 acquires a three-dimensional map 711 in the vicinity of the travel route. For example, the three-dimensional map acquisition unit 701 acquires the three-dimensional map 711 via a mobile communication network, or vehicle-to-vehicle communication or road-to-vehicle communication.
[0268] Next, the ego-vehicle detection data acquisition unit 702 acquires ego-vehicle detection three-dimensional data 712 based on sensor information. For example, the ego-vehicle detection data acquisition unit 702 generates the ego-vehicle detection three-dimensional data 712 based on the sensor information acquired by the sensors provided in the ego-vehicle.
[0269] Next, the abnormal case determination unit 703 detects an abnormal case by performing a predetermined check on at least one of the acquired three-dimensional map 711 and the host vehicle detection three-dimensional data 712. That is, the abnormal case determination unit 703 determines whether at least one of the acquired three-dimensional map 711 and the host vehicle detection three-dimensional data 712 is abnormal.
[0270] When an abnormal case is detected, the countermeasure operation determination unit 704 determines a countermeasure operation for the abnormal case. Next, the operation control unit 705 controls the operations of each processing unit necessary for implementing the countermeasure operation, such as the three-dimensional map acquisition unit 701.
[0271] On the other hand, when no abnormal case is detected, the three-dimensional information processing device 700 ends the processing.
[0272] In addition, the three-dimensional information processing device 700 estimates the self-position of the vehicle having the three-dimensional information processing device 700 by using the three-dimensional map 711 and the host vehicle detection three-dimensional data 712. Next, the three-dimensional information processing device 700 automatically drives the vehicle by using the result of the self-position estimation.
[0273] In this way, the three-dimensional information processing device 700 acquires map data (three-dimensional map 711) including the first three-dimensional position information via a communication path. For example, the first three-dimensional position information is encoded with a partial space having three-dimensional coordinate information as a unit, each of which is an aggregate of one or more partial spaces, and includes a plurality of random access units that can be independently decoded. For example, the first three-dimensional position information is data (SWLD) in which feature points where three-dimensional feature amounts are equal to or greater than a predetermined threshold are encoded.
[0274] In addition, the three-dimensional information processing device 700 generates second three-dimensional position information (host vehicle detection three-dimensional data 712) from the information detected by the sensor. Next, the three-dimensional information processing device 700 determines whether the first three-dimensional position information or the second three-dimensional position information is abnormal by performing an abnormality determination process on the first three-dimensional position information or the second three-dimensional position information.
[0275] When it is determined that the first three-dimensional position information or the second three-dimensional position information is abnormal, the three-dimensional information processing device 700 determines a countermeasure operation for the abnormality. Next, the three-dimensional information processing device 700 performs control necessary for implementing the countermeasure operation.
[0276] Thereby, the three-dimensional information processing device 700 can detect an abnormality in the first three-dimensional position information or the second three-dimensional position information and perform a countermeasure operation.
[0277] (Embodiment 5) In the present embodiment, a method for transmitting three-dimensional data to a following vehicle and the like will be described.
[0278] FIG. 27 is a block diagram showing a configuration example of a three-dimensional data creation device 810 according to the present embodiment. This three-dimensional data creation device 810 is mounted on a vehicle, for example. The three-dimensional data creation device 810 transmits and receives three-dimensional data to and from an external traffic monitoring cloud, a preceding vehicle, or a following vehicle, and creates and accumulates three-dimensional data.
[0279] The three-dimensional data creation device 810 includes a data reception unit 811, a communication unit 812, a reception control unit 813, a format conversion unit 814, a plurality of sensors 815, a three-dimensional data creation unit 816, a three-dimensional data synthesis unit 817, a three-dimensional data storage unit 818, a communication unit 819, a transmission control unit 820, a format conversion unit 821, and a data transmission unit 822.
[0280] The data reception unit 811 receives three-dimensional data 831 from a traffic monitoring cloud or a preceding vehicle. The three-dimensional data 831 includes information such as point cloud, visible light video, depth information, sensor position information, or speed information, which includes areas that cannot be detected by the sensors 815 of the host vehicle, for example.
[0281] The communication unit 812 communicates with a traffic monitoring cloud or a preceding vehicle and transmits a data transmission request or the like to the traffic monitoring cloud or the preceding vehicle.
[0282] The reception control unit 813 exchanges information such as the corresponding format with the communication destination via the communication unit 812 and establishes communication with the communication destination.
[0283] The format conversion unit 814 generates three-dimensional data 832 by performing format conversion or the like on the three-dimensional data 831 received by the data reception unit 811. Further, when the three-dimensional data 831 is compressed or encoded, the format conversion unit 814 performs decompression or decoding processing.
[0284] The plurality of sensors 815 are a group of sensors that acquire information outside the vehicle, such as a LiDAR, a visible light camera, or an infrared camera, and generate sensor information 833. For example, when the sensor 815 is a laser sensor such as a LiDAR, the sensor information 833 is three-dimensional data such as a point cloud (point group data). Note that the number of sensors 815 does not have to be plural.
[0285] The three-dimensional data creation unit 816 generates three-dimensional data 834 from the sensor information 833. The three-dimensional data 834 includes information such as, for example, a point cloud, a visible light video, depth information, sensor position information, or speed information.
[0286] The three-dimensional data synthesis unit 817 synthesizes the three-dimensional data 832 created by the traffic monitoring cloud or the preceding vehicle or the like with the three-dimensional data 834 created based on the sensor information 833 of the host vehicle, thereby constructing three-dimensional data 835 that includes the space in front of the preceding vehicle that cannot be detected by the sensor 815 of the host vehicle.
[0287] The three-dimensional data storage unit 818 stores the generated three-dimensional data 835 and the like.
[0288] The communication unit 819 communicates with the traffic monitoring cloud or the following vehicle and transmits a data transmission request or the like to the traffic monitoring cloud or the following vehicle.
[0289] The transmission control unit 820 exchanges information such as the corresponding format with the communication destination via the communication unit 819 and establishes communication with the communication destination. Further, the transmission control unit 820 determines a transmission area, which is the space of the three-dimensional data to be transmitted, based on the three-dimensional data construction information of the three-dimensional data 832 generated by the three-dimensional data synthesis unit 817 and the data transmission request from the communication destination.
[0290] Specifically, the transmission control unit 820 determines a transmission area that includes the space in front of the host vehicle that cannot be detected by the sensors of the following vehicle in response to a data transmission request from the traffic monitoring cloud or the following vehicle. Further, the transmission control unit 820 determines the transmission area by judging, based on the three-dimensional data construction information, whether there is an update to the space that can be transmitted or the transmitted space. For example, the transmission control unit 820 determines the transmission area as the area specified in the data transmission request and in which the corresponding three-dimensional data 835 exists. Then, the transmission control unit 820 notifies the format conversion unit 821 of the format corresponding to the communication destination and the transmission area.
[0291] The format conversion unit 821 generates the three-dimensional data 837 by converting the three-dimensional data 836 in the transmission area among the three-dimensional data 835 stored in the three-dimensional data storage unit 818 into the format corresponding to the receiving side. Note that the format conversion unit 821 may reduce the data amount by compressing or encoding the three-dimensional data 837.
[0292] The data transmission unit 822 transmits the three-dimensional data 837 to the traffic monitoring cloud or the following vehicle. This three-dimensional data 837 includes information such as the point cloud in front of the host vehicle, visible light video, depth information, or sensor position information, which includes the area that becomes a blind spot for the following vehicle.
[0293] Here, an example in which format conversion and the like are performed by the format conversion units 814 and 821 has been described, but the format conversion may not be performed.
[0294] With such a configuration, the three-dimensional data creation device 810 acquires three-dimensional data 831 of an area that cannot be detected by the sensor 815 of the host vehicle from the outside, and generates three-dimensional data 835 by synthesizing the three-dimensional data 831 and three-dimensional data 834 based on the sensor information 833 detected by the sensor 815 of the host vehicle. Thereby, the three-dimensional data creation device 810 can generate three-dimensional data in a range that cannot be detected by the sensor 815 of the host vehicle.
[0295] Further, the three-dimensional data creation device 810 can transmit three-dimensional data including the space in front of the host vehicle that cannot be detected by the sensor of the following vehicle to the traffic monitoring cloud or the following vehicle or the like in response to a data transmission request from the traffic monitoring cloud or the following vehicle.
[0296] (Embodiment 6) In Embodiment 5, an example in which a client device such as a vehicle transmits three-dimensional data to a server such as another vehicle or a traffic monitoring cloud has been described. In this embodiment, the client device transmits sensor information obtained by the sensor to the server or another client device.
[0297] First, the configuration of the system according to this embodiment will be described. FIG. 28 is a diagram showing the configuration of a three-dimensional map and sensor information transmission / reception system according to this embodiment. This system includes a server 901 and client devices 902A and 902B. When the client devices 902A and 902B are not particularly distinguished, they are also referred to as the client device 902.
[0298] The client device 902 is, for example, an in-vehicle device mounted on a moving body such as a vehicle. The server 901 is, for example, a traffic monitoring cloud or the like and can communicate with a plurality of client devices 902.
[0299] The server 901 transmits a three-dimensional map composed of point clouds to the client device 902. Note that the configuration of the three-dimensional map is not limited to point clouds, and may represent other three-dimensional data such as a mesh structure.
[0300] The client device 902 transmits the sensor information acquired by the client device 902 to the server 901. The sensor information includes, for example, at least one of LiDAR acquisition information, visible light image, infrared image, depth image, sensor position information, and speed information.
[0301] The data transmitted and received between the server 901 and the client device 902 may be compressed for data reduction, or may remain uncompressed to maintain the accuracy of the data. When compressing the data, for example, a three-dimensional compression method based on an octree structure can be used for the point cloud. Also, a two-dimensional image compression method can be used for the visible light image, infrared image, and depth image. The two-dimensional image compression method is, for example, MPEG-4 AVC or HEVC standardized by MPEG.
[0302] In addition, the server 901 transmits the three-dimensional map managed by the server 901 to the client device 902 in response to a transmission request for the three-dimensional map from the client device 902. Note that the server 901 may transmit the three-dimensional map without waiting for a transmission request for the three-dimensional map from the client device 902. For example, the server 901 may broadcast the three-dimensional map to one or more client devices 902 in a predetermined space. Also, the server 901 may transmit a three-dimensional map suitable for the position of the client device 902 to the client device 902 that has received a transmission request at regular intervals. Further, the server 901 may transmit the three-dimensional map to the client device 902 each time the three-dimensional map managed by the server 901 is updated.
[0303] The client device 902 issues a transmission request for the three-dimensional map to the server 901. For example, when the client device 902 wants to perform self-position estimation during driving, the client device 902 transmits a transmission request for the three-dimensional map to the server 901.
[0304] In addition, in the following cases, the client device 902 may send a request to the server 901 to transmit the three-dimensional map. When the three-dimensional map held by the client device 902 is old, the client device 902 may send a request to the server 901 to transmit the three-dimensional map. For example, when a certain period of time has elapsed since the client device 902 acquired the three-dimensional map, the client device 902 may send a request to the server 901 to transmit the three-dimensional map.
[0305] Before a certain time when the client device 902 exits from the space represented by the three-dimensional map held by the client device 902, the client device 902 may send a request to the server 901 to transmit the three-dimensional map. For example, when the client device 902 exists within a predetermined distance from the boundary of the space represented by the three-dimensional map held by the client device 902, the client device 902 may send a request to the server 901 to transmit the three-dimensional map. Further, when the movement path and movement speed of the client device 902 can be grasped, based on these, the time when the client device 902 exits from the space represented by the three-dimensional map held by the client device 902 may be predicted.
[0306] When the error during alignment between the three-dimensional data created by the client device 902 from sensor information and the three-dimensional map is a certain value or more, the client device 902 may send a request to the server 901 to transmit the three-dimensional map.
[0307] The client device 902 transmits sensor information to the server 901 in response to a transmission request for sensor information sent from the server 901. Note that the client device 902 may send the sensor information to the server 901 without waiting for a transmission request for sensor information from the server 901. For example, when the client device 902 once obtains a transmission request for sensor information from the server 901, it may periodically transmit the sensor information to the server 901 for a certain period of time. Also, when the error at the time of alignment between the three-dimensional data created by the client device 902 based on the sensor information and the three-dimensional map obtained from the server 901 is equal to or greater than a certain value, the client device 902 determines that there may be a change in the three-dimensional map around the client device 902, and may transmit that fact and the sensor information to the server 901.
[0308] The server 901 issues a transmission request for sensor information to the client device 902. For example, the server 901 receives position information of the client device 902 such as GPS from the client device 902. When the server 901 determines based on the position information of the client device 902 that the client device 902 is approaching a space with little information in the three-dimensional map managed by the server 901, the server 901 issues a transmission request for sensor information to the client device 902 to generate a new three-dimensional map. Also, when the server 901 wants to update the three-dimensional map, when it wants to check the road conditions such as during snow accumulation or a disaster, or when it wants to check the traffic congestion situation, or an accident situation, etc., it may issue a transmission request for sensor information.
[0309] Also, the client device 902 may set the data volume of the sensor information to be transmitted to the server 901 according to the communication state or bandwidth at the time of receiving a transmission request for sensor information received from the server 901. Setting the data volume of the sensor information to be transmitted to the server 901 means, for example, increasing or decreasing the data itself, or appropriately selecting a compression method.
[0310] FIG. 29 is a block diagram showing a configuration example of the client device 902. The client device 902 receives a three-dimensional map composed of a point cloud or the like from the server 901, and estimates its own position of the client device 902 from the three-dimensional data created based on the sensor information of the client device 902. Further, the client device 902 transmits the acquired sensor information to the server 901.
[0311] The client device 902 includes a data reception unit 1011, a communication unit 1012, a reception control unit 1013, a format conversion unit 1014, a plurality of sensors 1015, a three-dimensional data creation unit 1016, a three-dimensional image processing unit 1017, a three-dimensional data storage unit 1018, a format conversion unit 1019, a communication unit 1020, a transmission control unit 1021, and a data transmission unit 1022.
[0312] The data reception unit 1011 receives the three-dimensional map 1031 from the server 901. The three-dimensional map 1031 is data including a point cloud such as WLD or SWLD. The three-dimensional map 1031 may include either compressed data or uncompressed data.
[0313] The communication unit 1012 communicates with the server 901 and transmits a data transmission request (for example, a transmission request for a three-dimensional map) to the server 901.
[0314] The reception control unit 1013 exchanges information such as the corresponding format with the communication destination via the communication unit 1012 and establishes communication with the communication destination.
[0315] The format conversion unit 1014 generates the three-dimensional map 1032 by performing format conversion or the like on the three-dimensional map 1031 received by the data reception unit 1011. Further, when the three-dimensional map 1031 is compressed or encoded, the format conversion unit 1014 performs decompression or decoding processing. Note that the format conversion unit 1014 does not perform decompression or decoding processing if the three-dimensional map 1031 is uncompressed data.
[0316] The plurality of sensors 1015 are a group of sensors that acquire information outside the vehicle on which the client device 902 is mounted, such as LiDAR, a visible light camera, an infrared camera, or a depth sensor, and generate sensor information 1033. For example, when the sensor 1015 is a laser sensor such as LiDAR, the sensor information 1033 is three-dimensional data such as point cloud (point group data). Note that the number of sensors 1015 does not have to be plural.
[0317] The three-dimensional data creation unit 1016 creates three-dimensional data 1034 around the host vehicle based on the sensor information 1033. For example, the three-dimensional data creation unit 1016 creates point cloud data with color information around the host vehicle using the information acquired by LiDAR and the visible light video obtained by the visible light camera.
[0318] The three-dimensional image processing unit 1017 performs self-position estimation processing of the host vehicle using the received three-dimensional map 1032 such as point cloud and the three-dimensional data 1034 around the host vehicle generated from the sensor information 1033. Note that the three-dimensional image processing unit 1017 may create three-dimensional data 1035 around the host vehicle by synthesizing the three-dimensional map 1032 and the three-dimensional data 1034, and perform self-position estimation processing using the created three-dimensional data 1035.
[0319] The three-dimensional data storage unit 1018 stores the three-dimensional map 1032, the three-dimensional data 1034, the three-dimensional data 1035, and the like.
[0320] The format conversion unit 1019 generates sensor information 1037 by converting the sensor information 1033 into a format supported by the receiving side. Note that the format conversion unit 1019 may reduce the data amount by compressing or encoding the sensor information 1037. Also, the format conversion unit 1019 may omit the process when format conversion is not necessary. Further, the format conversion unit 1019 may control the data amount to be transmitted according to the specification of the transmission range.
[0321] The communication unit 1020 communicates with the server 901 and receives from the server 901 a data transmission request (a request for transmitting sensor information) and the like.
[0322] The transmission control unit 1021 exchanges information such as a corresponding format with the communication destination via the communication unit 1020 and establishes communication.
[0323] The data transmission unit 1022 transmits the sensor information 1037 to the server 901. The sensor information 1037 includes information acquired by a plurality of sensors 1015 such as, for example, information acquired by a LiDAR, a luminance image acquired by a visible light camera, an infrared image acquired by an infrared camera, a depth image acquired by a depth sensor, sensor position information, and speed information.
[0324] Next, the configuration of the server 901 will be described. FIG. 30 is a block diagram showing a configuration example of the server 901. The server 901 receives the sensor information transmitted from the client device 902 and creates three-dimensional data based on the received sensor information. The server 901 updates the three-dimensional map managed by the server 901 using the created three-dimensional data. Further, the server 901 transmits the updated three-dimensional map to the client device 902 in response to a transmission request for the three-dimensional map from the client device 902.
[0325] The server 901 includes a data reception unit 1111, a communication unit 1112, a reception control unit 1113, a format conversion unit 1114, a three-dimensional data creation unit 1116, a three-dimensional data synthesis unit 1117, a three-dimensional data storage unit 1118, a format conversion unit 1119, a communication unit 1120, a transmission control unit 1121, and a data transmission unit 1122.
[0326] The data reception unit 1111 receives the sensor information 1037 from the client device 902. The sensor information 1037 includes, for example, information acquired by a LiDAR, a luminance image acquired by a visible light camera, an infrared image acquired by an infrared camera, a depth image acquired by a depth sensor, sensor position information, and speed information.
[0327] The communication unit 1112 communicates with the client device 902 and transmits data transmission requests (e.g., a request to transmit sensor information) to the client device 902.
[0328] The reception control unit 1113 exchanges information such as the corresponding format with the communication destination via the communication unit 1112 to establish communication.
[0329] When the received sensor information 1037 is compressed or encoded, the format conversion unit 1114 generates the sensor information 1132 by performing decompression or decoding processing. Note that if the sensor information 1037 is uncompressed data, the format conversion unit 1114 does not perform decompression or decoding processing.
[0330] The three-dimensional data creation unit 1116 creates three-dimensional data 1134 around the client device 902 based on the sensor information 1132. For example, the three-dimensional data creation unit 1116 creates point cloud data with color information around the client device 902 using the information acquired by LiDAR and the visible light video obtained by the visible light camera.
[0331] The three-dimensional data synthesis unit 1117 updates the three-dimensional map 1135 by synthesizing the three-dimensional data 1134 created based on the sensor information 1132 with the three-dimensional map 1135 managed by the server 901.
[0332] The three-dimensional data storage unit 1118 stores the three-dimensional map 1135 and the like.
[0333] The format conversion unit 1119 generates the three-dimensional map 1031 by converting the three-dimensional map 1135 into a format that the receiving side can handle. Note that the format conversion unit 1119 may reduce the data amount by compressing or encoding the three-dimensional map 1135. Also, if there is no need to perform format conversion, the format conversion unit 1119 may omit the processing. Further, the format conversion unit 1119 may control the data amount to be transmitted according to the specified transmission range.
[0334] The communication unit 1120 communicates with the client device 902 and receives from the client device 902 a data transmission request (a request for transmitting a three-dimensional map) or the like.
[0335] The transmission control unit 1121 exchanges information such as a corresponding format with the communication destination via the communication unit 1120 and establishes communication.
[0336] The data transmission unit 1122 transmits the three-dimensional map 1031 to the client device 902. The three-dimensional map 1031 is data including a point cloud such as a WLD or an SWLD. The three-dimensional map 1031 may include either compressed data or uncompressed data.
[0337] Next, the operation flow of the client device 902 will be described. FIG. 31 is a flowchart showing the operation when the client device 902 acquires a three-dimensional map.
[0338] First, the client device 902 requests the server 901 to transmit a three-dimensional map (point cloud or the like) (S1001). At this time, the client device 902 may request the server 901 to transmit a three-dimensional map related to the position information by transmitting the position information of the client device 902 obtained by GPS or the like together.
[0339] Next, the client device 902 receives a three-dimensional map from the server 901 (S1002). If the received three-dimensional map is compressed data, the client device 902 decrypts the received three-dimensional map to generate an uncompressed three-dimensional map (S1003).
[0340] Next, the client device 902 creates three-dimensional data 1034 around the client device 902 from the sensor information 1033 obtained by the plurality of sensors 1015 (S1004). Next, the client device 902 estimates its own position using the three-dimensional map 1032 received from the server 901 and the three-dimensional data 1034 created from the sensor information 1033 (S1005).
[0341] FIG. 32 is a flowchart showing the operation when the client device 902 transmits sensor information. First, the client device 902 receives a sensor information transmission request from the server 901 (S1011). The client device 902 that has received the transmission request transmits the sensor information 1037 to the server 901 (S1012). Note that when the sensor information 1033 includes a plurality of pieces of information obtained by the plurality of sensors 1015, the client device 902 may generate the sensor information 1037 by compressing each piece of information using a compression method suitable for each piece of information.
[0342] Next, the operation flow of the server 901 will be described. FIG. 33 is a flowchart showing the operation when the server 901 acquires sensor information. First, the server 901 requests the client device 902 to transmit sensor information (S1021). Next, the server 901 receives the sensor information 1037 transmitted from the client device 902 in response to the request (S1022). Next, the server 901 creates three-dimensional data 1134 using the received sensor information 1037 (S1023). Next, the server 901 reflects the created three-dimensional data 1134 in the three-dimensional map 1135 (S1024).
[0343] FIG. 34 is a flowchart showing the operation at the time of transmission of the three-dimensional map by the server 901. First, the server 901 receives a transmission request for the three-dimensional map from the client device 902 (S1031). The server 901 that has received the transmission request for the three-dimensional map transmits the three-dimensional map 1031 to the client device 902 (S1032). At this time, the server 901 may extract the three-dimensional map in the vicinity according to the position information of the client device 902 and transmit the extracted three-dimensional map. Further, the server 901 may compress the three-dimensional map composed of the point cloud using, for example, a compression method based on an octree structure, and transmit the compressed three-dimensional map.
[0344] Hereinafter, a modification of the present embodiment will be described.
[0345] The server 901 creates three-dimensional data 1134 near the position of the client device 902 using the sensor information 1037 received from the client device 902. Next, the server 901 calculates the difference between the three-dimensional data 1134 and the three-dimensional map 1135 by performing matching between the created three-dimensional data 1134 and the three-dimensional map 1135 of the same area managed by the server 901. When the difference is equal to or greater than a predetermined threshold, the server 901 determines that some abnormality has occurred around the client device 902. For example, when ground subsidence or the like occurs due to a natural disaster such as an earthquake, a large difference may occur between the three-dimensional map 1135 managed by the server 901 and the three-dimensional data 1134 created based on the sensor information 1037.
[0346] The sensor information 1037 may include information indicating at least one of the type of the sensor, the performance of the sensor, and the model number of the sensor. Also, a class ID or the like corresponding to the performance of the sensor may be added to the sensor information 1037. For example, when the sensor information 1037 is information acquired by LiDAR, it is conceivable to assign an identifier to the performance of the sensor such that a sensor capable of acquiring information with an accuracy of several millimeters is class 1, a sensor capable of acquiring information with an accuracy of several centimeters is class 2, and a sensor capable of acquiring information with an accuracy of several meters is class 3. Further, the server 901 may estimate the performance information of the sensor, etc., from the model number of the client device 902. For example, when the client device 902 is mounted on a vehicle, the server 901 may determine the specification information of the sensor from the vehicle type of the vehicle. In this case, the server 901 may have acquired the information of the vehicle type of the vehicle in advance, or the information may be included in the sensor information. Also, the server 901 may use the acquired sensor information 1037 to switch the degree of correction for the three-dimensional data 1134 created using the sensor information 1037. For example, when the sensor performance is high precision (class 1), the server 901 does not correct the three-dimensional data 1134. When the sensor performance is low precision (class 3), the server 901 applies correction corresponding to the accuracy of the sensor to the three-dimensional data 1134. For example, the server 901 increases the degree of correction (strength) as the accuracy of the sensor decreases.
[0347] The server 901 may simultaneously send a request for transmitting sensor information to a plurality of client devices 902 in a certain space. When the server 901 receives a plurality of sensor information from a plurality of client devices 902, it is not necessary to use all the sensor information for creating the three-dimensional data 1134. For example, the server 901 may select the sensor information to be used according to the performance of the sensor. For example, when updating the three-dimensional map 1135, the server 901 may select high-precision sensor information (class 1) from the plurality of received sensor information and create the three-dimensional data 1134 using the selected sensor information.
[0348] The server 901 is not limited to only servers such as traffic monitoring clouds, and may be other client devices (in-vehicle). FIG. 35 is a diagram showing the system configuration in this case.
[0349] For example, the client device 902C sends a transmission request for sensor information to the client device 902A nearby and acquires the sensor information from the client device 902A. Then, the client device 902C creates three-dimensional data using the acquired sensor information of the client device 902A and updates the three-dimensional map of the client device 902C. Thereby, the client device 902C can generate a three-dimensional map of the space that can be acquired from the client device 902A, taking advantage of the performance of the client device 902C. For example, such a case is considered to occur when the performance of the client device 902C is high.
[0350] Also, in this case, the client device 902A that provided the sensor information is given the right to acquire the high-precision three-dimensional map generated by the client device 902C. The client device 902A receives the high-precision three-dimensional map from the client device 902C according to that right.
[0351] Also, the client device 902C may send a transmission request for sensor information to a plurality of nearby client devices 902 (client device 902A and client device 902B). When the sensors of the client device 902A or the client device 902B are high-performance, the client device 902C can create three-dimensional data using the sensor information obtained by this high-performance sensor.
[0352] FIG. 36 is a block diagram showing the functional configuration of the server 901 and the client device 902. The server 901 includes, for example, a three-dimensional map compression / decompression processing unit 1201 that compresses and decompresses a three-dimensional map, and a sensor information compression / decompression processing unit 1202 that compresses and decompresses sensor information.
[0353] The client device 902 includes a three-dimensional map decoding processing unit 1211 and a sensor information compression processing unit 1212. The three-dimensional map decoding processing unit 1211 receives the encoded data of the compressed three-dimensional map, decodes the encoded data, and acquires the three-dimensional map. The sensor information compression processing unit 1212 compresses the sensor information itself instead of the three-dimensional data created from the acquired sensor information, and transmits the encoded data of the compressed sensor information to the server 901. With this configuration, the client device 902 only needs to hold internally a processing unit (device or LSI) that performs the process of decoding the three-dimensional map (point cloud, etc.), and does not need to hold internally a processing unit that performs the process of compressing the three-dimensional data of the three-dimensional map (point cloud, etc.). Thereby, the cost and power consumption of the client device 902 can be suppressed.
[0354] As described above, the client device 902 according to the present embodiment is mounted on a moving body, and creates three-dimensional data 1034 of the periphery of the moving body from sensor information 1033 indicating the peripheral situation of the moving body obtained by a sensor 1015 mounted on the moving body. The client device 902 estimates its own position of the moving body using the created three-dimensional data 1034. The client device 902 transmits the acquired sensor information 1033 to the server 901 or another moving body 902.
[0355] According to this, the client device 902 transmits the sensor information 1033 to the server 901 or the like. Thereby, there is a possibility that the data amount of the transmitted data can be reduced as compared with the case of transmitting the three-dimensional data. Further, since it is not necessary for the client device 902 to perform processes such as compression or encoding of the three-dimensional data, the processing amount of the client device 902 can be reduced. Therefore, the client device 902 can realize reduction of the data amount to be transmitted or simplification of the device configuration.
[0356] Further, the client device 902 further sends a transmission request for the three-dimensional map to the server 901 and receives the three-dimensional map 1031 from the server 901. The client device 902 estimates its own position using the three-dimensional data 1034 and the three-dimensional map 1032 in the estimation of its own position.
[0357] Also, the sensor information 1033 includes at least one of information obtained by a laser sensor, a luminance image, an infrared image, a depth image, the position information of the sensor, and the speed information of the sensor.
[0358] Also, the sensor information 1033 includes information indicating the performance of the sensor.
[0359] Also, the client device 902 encodes or compresses the sensor information 1033 and transmits the encoded or compressed sensor information 1037 to the server 901 or another moving body 902 in the transmission of the sensor information. According to this, the client device 902 can reduce the amount of data to be transmitted.
[0360] For example, the client device 902 includes a processor and a memory, and the processor performs the above processing using the memory.
[0361] Also, the server 901 according to the present embodiment is communicable with the client device 902 mounted on the moving body, and receives the sensor information 1037 indicating the surrounding situation of the moving body obtained by the sensor 1015 mounted on the moving body from the client device 902. The server 901 creates three-dimensional data 1134 around the moving body from the received sensor information 1037.
[0362] According to this, the server 901 creates three-dimensional data 1134 using the sensor information 1037 transmitted from the client device 902. Thereby, compared with the case where the client device 902 transmits three-dimensional data, there is a possibility of reducing the data amount of the transmitted data. Also, since it is not necessary for the client device 902 to perform processing such as compression or encoding of the three-dimensional data, the processing amount of the client device 902 can be reduced. Therefore, the server 901 can achieve reduction of the data amount to be transmitted or simplification of the device configuration.
[0363] Also, the server 901 further transmits a transmission request for sensor information to the client device 902.
[0364] Also, the server 901 further updates the three-dimensional map 1135 using the created three-dimensional data 1134, and transmits the three-dimensional map 1135 to the client device 902 in response to a transmission request for the three-dimensional map 1135 from the client device 902.
[0365] Also, the sensor information 1037 includes at least one of information obtained by a laser sensor, a luminance image, an infrared image, a depth image, sensor position information, and sensor speed information.
[0366] Also, the sensor information 1037 includes information indicating the performance of the sensor.
[0367] Also, the server 901 further corrects the three-dimensional data according to the performance of the sensor. According to this, the three-dimensional data creation method can improve the quality of the three-dimensional data.
[0368] Also, when receiving the sensor information, the server 901 receives a plurality of sensor information 1037 from a plurality of client devices 902, and selects the sensor information 1037 to be used for creating the three-dimensional data 1134 based on a plurality of information indicating the performance of the sensors included in the plurality of sensor information 1037. According to this, the server 901 can improve the quality of the three-dimensional data 1134.
[0369] In addition, the server 901 decrypts or decompresses the received sensor information 1037, and creates three-dimensional data 1134 from the decrypted or decompressed sensor information 1132. According to this, the server 901 can reduce the amount of data to be transmitted.
[0370] For example, the server 901 includes a processor and a memory, and the processor performs the above processing using the memory.
[0371] (Embodiment 7) In this embodiment, an encoding method and a decoding method for three-dimensional data using inter prediction processing will be described.
[0372] FIG. 37 is a block diagram of a three-dimensional data encoding apparatus 1300 according to this embodiment. This three-dimensional data encoding apparatus 1300 generates an encoded bit stream (hereinafter, also simply referred to as a bit stream), which is an encoded signal, by encoding three-dimensional data. As shown in FIG. 37, the three-dimensional data encoding apparatus 1300 includes a division unit 1301, a subtraction unit 1302, a conversion unit 1303, a quantization unit 1304, an inverse quantization unit 1305, an inverse conversion unit 1306, an addition unit 1307, a reference volume memory 1308, an intra prediction unit 1309, a reference space memory 1310, an inter prediction unit 1311, a prediction control unit 1312, and an entropy encoding unit 1313.
[0373] The division unit 1301 divides each space (SPC) included in the three-dimensional data into a plurality of volumes (VLM) that are encoding units. In addition, the division unit 1301 octree-expresses (Octree-expresses) the voxels within each volume. Note that the division unit 1301 may set the space and the volume to the same size and octree-express the space. Further, the division unit 1301 may add information (such as depth information) necessary for octree conversion to the header of the bit stream or the like.
[0374] The subtraction unit 1302 calculates the difference between the volume (the volume to be encoded) output from the division unit 1301 and the predicted volume generated by intra prediction or inter prediction described later, and outputs the calculated difference as a prediction residual to the conversion unit 1303. FIG. 38 is a diagram showing an example of calculating the prediction residual. Note that the bit sequences of the volume to be encoded and the predicted volume shown here are, for example, position information indicating the positions of three-dimensional points (e.g., point cloud) included in the volume.
[0375] Hereinafter, the octree representation and the voxel scan order will be described. After the volume is converted into an octree structure (octree conversion), it is encoded. The octree structure is composed of nodes and leaves. Each node has eight nodes or leaves, and each leaf has voxel (VXL) information. FIG. 39 is a diagram showing an example of the structure of a volume including a plurality of voxels. FIG. 40 is a diagram showing an example of converting the volume shown in FIG. 39 into an octree structure. Here, among the leaves shown in FIG. 40, leaves 1, 2, and 3 represent voxels VXL1, VXL2, and VXL3 shown in FIG. 39, respectively, and represent VXLs (hereinafter, valid VXLs) including point clouds.
[0376] The octree is represented by, for example, a binary sequence of 0 and 1. For example, if a node or a valid VXL is set to value 1 and the others are set to value 0, the binary sequences shown in FIG. 40 are assigned to each node and leaf. Then, according to the breadth-first or depth-first scan order, this binary sequence is scanned. For example, when scanned in breadth-first order, the binary sequence shown in A of FIG. 41 is obtained. When scanned in depth-first order, the binary sequence shown in B of FIG. 41 is obtained. The binary sequence obtained by this scan is encoded by entropy encoding to reduce the amount of information.
[0377] Next, the depth information in the octree representation will be described. The depth in the octree representation is used to control up to what granularity the point cloud information contained in the volume is retained. When the depth is set large, the point cloud information can be reproduced to a finer level, but the amount of data for representing nodes and leaves increases. Conversely, when the depth is set small, the amount of data decreases, but since point cloud information at multiple different positions and with different colors is regarded as being at the same position and having the same color, the information possessed by the original point cloud information is lost.
[0378] For example, FIG. 42 is a diagram showing an example in which the octree with a depth = 2 shown in FIG. 40 is represented by an octree with a depth = 1. The octree shown in FIG. 42 has less data volume than the octree shown in FIG. 40. That is, the octree shown in FIG. 42 has fewer bits after binarization than the octree shown in FIG. 42. Here, leaf 1 and leaf 2 shown in FIG. 40 will be represented by leaf 1 shown in FIG. 41. That is, the information that leaf 1 and leaf 2 shown in FIG. 40 are at different positions is lost.
[0379] FIG. 43 is a diagram showing the volume corresponding to the octree shown in FIG. 42. VXL1 and VXL2 shown in FIG. 39 correspond to VXL12 shown in FIG. 43. In this case, the three-dimensional data encoding device 1300 generates the color information of VXL12 shown in FIG. 43 from the color information of VXL1 and VXL2 shown in FIG. 39. For example, the three-dimensional data encoding device 1300 calculates the average value, median value, or weighted average value of the color information of VXL1 and VXL2 as the color information of VXL12. In this way, the three-dimensional data encoding device 1300 may control the reduction of the data volume by changing the depth of the octree.
[0380] The three-dimensional data encoding device 1300 may set the depth information of the octree in any unit of world unit, space unit, or volume unit. At this time, the three-dimensional data encoding device 1300 may add the depth information to the world header information, space header information, or volume header information. Also, the same value may be used for the depth information in all worlds, spaces, and volumes with different times. In this case, the three-dimensional data encoding device 1300 may add the depth information to the header information that manages the worlds of all times.
[0381] When the voxel contains color information, the conversion unit 1303 applies a frequency conversion such as an orthogonal conversion to the prediction residual of the color information of the voxels in the volume. For example, the conversion unit 1303 creates a one-dimensional array by scanning the prediction residuals in a certain scan order. Then, the conversion unit 1303 converts the created one-dimensional array into the frequency domain by applying a one-dimensional orthogonal conversion to the one-dimensional array. As a result, when the values of the prediction residuals in the volume are close, the values of the low-frequency components become large and the values of the high-frequency components become small. Therefore, the quantization unit 1304 can more efficiently reduce the amount of code.
[0382] Also, the conversion unit 1303 may use an orthogonal conversion of two or more dimensions instead of one dimension. For example, the conversion unit 1303 maps the prediction residuals to a two-dimensional array in a certain scan order and applies a two-dimensional orthogonal conversion to the obtained two-dimensional array. Also, the conversion unit 1303 may select the orthogonal conversion method to be used from a plurality of orthogonal conversion methods. In this case, the three-dimensional data encoding device 1300 adds information indicating which orthogonal conversion method is used to the bitstream. Also, the conversion unit 1303 may select the orthogonal conversion method to be used from a plurality of orthogonal conversion methods with different dimensions. In this case, the three-dimensional data encoding device 1300 adds which-dimensional orthogonal conversion method is used to the bitstream.
[0383] For example, the conversion unit 1303 adjusts the scan order of the prediction residual to match the scan order (such as breadth-first or depth-first) in the octree within the volume. As a result, there is no need to add information indicating the scan order of the prediction residual to the bitstream, so the overhead can be reduced. Further, the conversion unit 1303 may apply a scan order different from the scan order of the octree. In this case, the three-dimensional data encoding device 1300 adds information indicating the scan order of the prediction residual to the bitstream. Thereby, the three-dimensional data encoding device 1300 can efficiently encode the prediction residual. Further, the three-dimensional data encoding device 1300 may add information (such as a flag) indicating whether to apply the scan order of the octree to the bitstream, and when not applying the scan order of the octree, may also add information indicating the scan order of the prediction residual to the bitstream.
[0384] The conversion unit 1303 may convert not only the prediction residual of the color information but also other attribute information that the voxel has. For example, the conversion unit 1303 may convert and encode information such as reflectance obtained when acquiring a point cloud with LiDAR or the like.
[0385] When the space does not have attribute information such as color information, the conversion unit 1303 may skip the process. Further, the three-dimensional data encoding device 1300 may add information (flag) indicating whether to skip the process of the conversion unit 1303 to the bitstream.
[0386] The quantization unit 1304 generates quantization coefficients by performing quantization on the frequency components of the prediction residual generated by the conversion unit 1303 using quantization control parameters. As a result, the amount of information is reduced. The generated quantization coefficients are output to the entropy encoding unit 1313. The quantization unit 1304 may control the quantization control parameters in world units, space units, or volume units. In that case, the three-dimensional data encoding device 1300 adds the quantization control parameters to respective header information and the like. Also, the quantization unit 1304 may perform quantization control with different weights for each frequency component of the prediction residual. For example, the quantization unit 1304 may perform fine quantization on low-frequency components and coarse quantization on high-frequency components. In this case, the three-dimensional data encoding device 1300 may add a parameter representing the weight of each frequency component to the header.
[0387] When the space does not have attribute information such as color information, the quantization unit 1304 may skip the process. Also, the three-dimensional data encoding device 1300 may add information (flag) indicating whether to skip the process of the quantization unit 1304 to the bit stream.
[0388] The inverse quantization unit 1305 generates inverse quantization coefficients of the prediction residual by performing inverse quantization on the quantization coefficients generated by the quantization unit 1304 using the quantization control parameters, and outputs the generated inverse quantization coefficients to the inverse conversion unit 1306.
[0389] The inverse conversion unit 1306 generates a prediction residual after inverse conversion application by applying inverse conversion to the inverse quantization coefficients generated by the inverse quantization unit 1305. Since this prediction residual after inverse conversion application is the prediction residual generated after quantization, it does not have to exactly match the prediction residual output by the conversion unit 1303.
[0390] The addition unit 1307 adds the post-inverse transformation prediction residual generated by the inverse transformation unit 1306 and the prediction volume generated by intra prediction or inter prediction (to be described later) used for generating the prediction residual before quantization to generate a reconstruction volume. This reconstruction volume is stored in the reference volume memory 1308 or the reference space memory 1310.
[0391] The intra prediction unit 1309 generates a prediction volume of the volume to be encoded using the attribute information of the adjacent volume stored in the reference volume memory 1308. The attribute information includes the color information or reflectance of the voxel. The intra prediction unit 1309 generates a predicted value of the color information or reflectance of the volume to be encoded.
[0392] FIG. 44 is a diagram for explaining the operation of the intra prediction unit 1309. For example, the intra prediction unit 1309 generates a predicted volume of the volume to be encoded (volume idx = 3) shown in FIG. 44 from an adjacent volume (volume idx = 0). Here, the volume idx is identifier information added to the volumes in the space, and different values are assigned to each volume. The assignment order of the volume idx may be the same as the encoding order or different from the encoding order. For example, the intra prediction unit 1309 uses the average value of the color information of the voxels included in the adjacent volume volume idx = 0 as the predicted value of the color information of the volume to be encoded shown in FIG. 44. In this case, a prediction residual is generated by subtracting the predicted value of the color information from the color information of each voxel included in the volume to be encoded. The processing after the conversion unit 1303 is performed on this prediction residual. Also, in this case, the three-dimensional data encoding device 1300 adds the adjacent volume information and the prediction mode information to the bitstream. Here, the adjacent volume information is information indicating the adjacent volume used for prediction, and for example, indicates the volume idx of the adjacent volume used for prediction. Also, the prediction mode information indicates the mode used for generating the predicted volume. The mode is, for example, an average value mode that generates a predicted value from the average value of the voxels in the adjacent volume, or a median value mode that generates a predicted value from the median value of the voxels in the adjacent volume, etc.
[0393] The intra prediction unit 1309 may generate a predicted volume from a plurality of adjacent volumes. For example, in the configuration shown in FIG. 44, the intra prediction unit 1309 generates a predicted volume 0 from the volume of volume idx = 0 and generates a predicted volume 1 from the volume of volume idx = 1. Then, the intra prediction unit 1309 generates the average of the predicted volume 0 and the predicted volume 1 as the final predicted volume. In this case, the three-dimensional data encoding device 1300 may add the plurality of volume idxs of the plurality of volumes used for generating the predicted volume to the bitstream.
[0394] (Embodiment 8) In this embodiment, a method for representing three-dimensional points (point clouds) in the encoding of three-dimensional data will be described.
[0395] FIG. 45 is a block diagram showing the configuration of a three-dimensional data distribution system according to this embodiment. The distribution system shown in FIG. 45 includes a server 1501 and a plurality of clients 1502.
[0396] The server 1501 includes a storage unit 1511 and a control unit 1512. The storage unit 1511 stores an encoded three-dimensional map 1513, which is encoded three-dimensional data.
[0397] FIG. 46 is a diagram showing a configuration example of the bit stream of the encoded three-dimensional map 1513. The three-dimensional map is divided into a plurality of sub-maps, and each sub-map is encoded. A random access header (RA) including sub-coordinate information is added to each sub-map. The sub-coordinate information is used to improve the encoding efficiency of the sub-map. This sub-coordinate information indicates the sub-coordinate of the sub-map. The sub-coordinate is the coordinate of the sub-map with respect to a reference coordinate. Note that a three-dimensional map including a plurality of sub-maps is called an overall map. Also, a coordinate (for example, the origin) serving as a reference in the overall map is called a reference coordinate. That is, the sub-coordinate is the coordinate of the sub-map in the coordinate system of the overall map. In other words, the sub-coordinate indicates the offset between the coordinate system of the overall map and the coordinate system of the sub-map. Also, a coordinate in the coordinate system of the overall map with respect to the reference coordinate is called an overall coordinate. A coordinate in the coordinate system of the sub-map with respect to the sub-coordinate is called a differential coordinate.
[0398] Client 1502 sends a message to server 1501. This message includes the location information of client 1502. The control unit 1512 included in server 1501 obtains the bit stream of the sub-map at the position closest to the position of client 1502 based on the location information included in the received message. The bit stream of the sub-map includes sub-coordinate information and is sent to client 1502. The decoder 1521 included in client 1502 uses this sub-coordinate information to obtain the overall coordinates of the sub-map with respect to the reference coordinates. The application 1522 included in client 1502 executes an application related to its own position using the obtained overall coordinates of the sub-map.
[0399] Also, the sub-map indicates a partial area of the overall map. The sub-coordinates are the coordinates at which the sub-map is located in the reference coordinate space of the overall map. For example, assume that in the overall map of A, there are sub-map A of AA and sub-map B of AB. When the vehicle wants to refer to the map of AA, it starts decoding from sub-map A, and when it wants to refer to the map of AB, it starts decoding from sub-map B. Here, the sub-map is a random access point. Specifically, A is Osaka Prefecture, AA is Osaka City, AB is Takatsuki City, etc.
[0400] Each sub-map is sent to the client together with the sub-coordinate information. The sub-coordinate information is included in the header information of each sub-map, or in the transmission packet, etc.
[0401] The reference coordinates that serve as the reference for the sub-coordinate information of each sub-map may be added to the header information of the space higher than the sub-map, such as the header information of the overall map.
[0402] The sub-map may be composed of one space (SPC). Also, the sub-map may be composed of a plurality of SPCs.
[0403] In addition, the sub-map may include a GOS (Group of Space). Also, the sub-map may be composed of worlds. For example, when there are a plurality of objects in the sub-map, if the plurality of objects are assigned to separate SPCs, the sub-map is composed of a plurality of SPCs. Also, if the plurality of objects are assigned to one SPC, the sub-map is composed of one SPC.
[0404] Next, the improvement effect of the encoding efficiency when using sub-coordinate information will be described. FIG. 47 is a diagram for explaining this effect. For example, in order to encode a three-dimensional point A at a position far from the reference coordinates shown in FIG. 47, a large number of bits are required. Here, the distance between the sub-coordinates and the three-dimensional point A is shorter than the distance between the reference coordinates and the three-dimensional point A. Therefore, the encoding efficiency can be improved by encoding the coordinates of the three-dimensional point A with respect to the sub-coordinates rather than encoding the coordinates of the three-dimensional point A with respect to the reference coordinates. Also, the bit stream of the sub-map includes sub-coordinate information. By sending the bit stream of the sub-map and the reference coordinates to the decoding side (client), the overall coordinates of the sub-map can be restored on the decoding side.
[0405] FIG. 48 is a flowchart of the processing by the server 1501 which is the transmission side of the sub-map.
[0406] First, the server 1501 receives a message including the position information of the client 1502 from the client 1502 (S1501). The control unit 1512 acquires the encoded bit stream of the sub-map based on the position information of the client from the storage unit 1511 (S1502). Then, the server 1501 transmits the encoded bit stream of the sub-map and the reference coordinates to the client 1502 (S1503).
[0407] FIG. 49 is a flowchart of the processing by the client 1502 which is the reception side of the sub-map.
[0408] First, the client 1502 receives the encoded bitstream of the submap and the reference coordinates transmitted from the server 1501 (S1511). Next, the client 1502 obtains the submap and the subcoordinate information by decoding the encoded bitstream (S1512). Next, the client 1502 restores the differential coordinates in the submap to the global coordinates using the reference coordinates and the subcoordinates (S1513).
[0409] Next, an example of the syntax of the information regarding the submap will be described. In the encoding of the submap, the three-dimensional data encoding device calculates the differential coordinates by subtracting the subcoordinates from the coordinates of each point cloud (three-dimensional point). Then, the three-dimensional data encoding device encodes the differential coordinates into a bitstream as the value of each point cloud. Also, the encoding device encodes the subcoordinate information indicating the subcoordinates as the header information of the bitstream. Thereby, the three-dimensional data decoding device can obtain the global coordinates of each point cloud. For example, the three-dimensional data encoding device is included in the server 1501, and the three-dimensional data decoding device is included in the client 1502.
[0410] FIG. 50 is a diagram showing an example of the syntax of the submap. NumOfPoint shown in FIG. 50 indicates the number of point clouds included in the submap. sub_coordinate_x, sub_coordinate_y, and sub_coordinate_z are the subcoordinate information. sub_coordinate_x indicates the x coordinate of the subcoordinates. sub_coordinate_y indicates the y coordinate of the subcoordinates. sub_coordinate_z indicates the z coordinate of the subcoordinates.
[0411] Also, diff_x[i], diff_y[i], and diff_z[i] are the differential coordinates of the i-th point cloud within the sub-map. diff_x[i] indicates the difference value between the x-coordinate of the i-th point cloud within the sub-map and the x-coordinate of the sub-coordinates. diff_y[i] indicates the difference value between the y-coordinate of the i-th point cloud within the sub-map and the y-coordinate of the sub-coordinates. diff_z[i] indicates the difference value between the z-coordinate of the i-th point cloud within the sub-map and the z-coordinate of the sub-coordinates.
[0412] The three-dimensional data decoding device decodes point_cloud[i]_x, point_cloud[i]_y, and point_cloud[i]_z, which are the overall coordinates of the i-th point cloud, using the following formula. point_cloud[i]_x is the x-coordinate of the overall coordinates of the i-th point cloud. point_cloud[i]_y is the y-coordinate of the overall coordinates of the i-th point cloud. point_cloud[i]_z is the z-coordinate of the overall coordinates of the i-th point cloud.
[0413] point_cloud[i]_x = sub_coordinate_x + diff_x[i] point_cloud[i]_y = sub_coordinate_y + diff_y[i] point_cloud[i]_z = sub_coordinate_z + diff_z[i]
[0414] Next, the switching process for applying octree encoding will be described. When encoding submaps, the three-dimensional data encoder selects whether to use octree representation to encode each point cloud (hereinafter referred to as octree encoding) or to encode the difference value from the subcoordinates (hereinafter referred to as non-octree encoding). FIG. 51 is a diagram schematically showing this operation. For example, when the number of point clouds in a submap is equal to or greater than a predetermined threshold, the three-dimensional data encoder applies octree encoding to the submap. When the number of point clouds in a submap is less than the threshold, the three-dimensional data encoder applies non-octree encoding to the submap. As a result, the three-dimensional data encoder can appropriately select whether to use octree encoding or non-octree encoding according to the shape and density of the objects included in the submap, thereby improving the encoding efficiency.
[0415] In addition, the three-dimensional data encoder adds information indicating which of octree encoding and non-octree encoding is applied to the submap (hereinafter referred to as octree encoding application information) to the header of the submap or the like. As a result, the three-dimensional data decoder can determine whether the bitstream is a bitstream obtained by octree encoding the submap or a bitstream obtained by non-octree encoding the submap.
[0416] Moreover, the three-dimensional data encoder calculates the encoding efficiency when applying octree encoding and non-octree encoding to the same point cloud, and may apply the encoding method with good encoding efficiency to the submap.
[0417] FIG. 52 is a diagram showing an example of the syntax of a sub-map when performing this switching. The coding_type shown in FIG. 52 is information indicating the coding type and is the above-mentioned octree coding application information. coding_type = 00 indicates that octree coding has been applied. coding_type = 01 indicates that non-octree coding has been applied. coding_type = 10 or 11 indicates that another coding method other than the above has been applied.
[0418] When the coding type is non-octree coding, the sub-map includes NumOfPoint and sub-coordinate information (sub_coordinate_x, sub_coordinate_y, and sub_coordinate_z).
[0419] When the coding type is octree coding, the sub-map includes octree_info. octree_info is information necessary for octree coding and includes, for example, depth information.
[0420] When the coding type is non-octree coding, the sub-map includes differential coordinates (diff_x[i], diff_y[i], and diff_z[i]).
[0421] When the coding type is octree coding, the sub-map includes octree_data, which is coding data related to octree coding.
[0422] Here, an example using the xyz coordinate system as the coordinate system of the point cloud is shown, but a polar coordinate system may also be used.
[0423] FIG. 53 is a flowchart of the three-dimensional data encoding process by the three-dimensional data encoder. First, the three-dimensional data encoder calculates the number of point clouds in the target sub-map, which is the sub-map to be processed (S1521). Next, the three-dimensional data encoder determines whether the calculated number of point clouds is equal to or greater than a predetermined threshold (S1522).
[0424] When the number of point clouds is equal to or greater than the threshold (Yes in S1522), the three-dimensional data encoder applies octree encoding to the target sub-map (S1523). Also, the three-dimensional point data encoder adds octree encoding application information indicating that octree encoding has been applied to the target sub-map to the header of the bitstream (S1525).
[0425] On the other hand, when the number of point clouds is less than the threshold (No in S1522), the three-dimensional data encoder applies non-octree encoding to the target sub-map (S1524). Also, the three-dimensional point data encoder adds octree encoding application information indicating that non-octree encoding has been applied to the target sub-map to the header of the bitstream (S1525).
[0426] FIG. 54 is a flowchart of the three-dimensional data decoding process by the three-dimensional data decoder. First, the three-dimensional data decoder decodes the octree encoding application information from the header of the bitstream (S1531). Next, the three-dimensional data decoder determines whether the encoding type applied to the target sub-map is octree encoding based on the decoded octree encoding application information (S1532).
[0427] When the encoding type indicated by the octree encoding application information is octree encoding (Yes in S1532), the three-dimensional data decoder decodes the target sub-map by octree decoding (S1533). On the other hand, when the encoding type indicated by the octree encoding application information is non-octree encoding (No in S1532), the three-dimensional data decoder decodes the target sub-map by non-octree decoding (S1534).
[0428] Hereinafter, a modification example of the present embodiment will be described. FIGS. 55 to 57 are diagrams schematically showing the operation of a modification example of the encoding type switching process.
[0429] As shown in FIG. 55, the three-dimensional data encoding device may select whether to apply octree encoding or non-octree encoding for each space. In this case, the three-dimensional data encoding device adds octree encoding application information to the header of the space. Thereby, the three-dimensional data decoding device can determine whether octree encoding has been applied for each space. Also, in this case, the three-dimensional data encoding device sets sub-coordinates for each space and encodes the difference value obtained by subtracting the value of the sub-coordinates from the coordinates of each point cloud within the space.
[0430] Thereby, the three-dimensional data encoding device can appropriately switch whether to apply octree encoding according to the shape of the object or the number of point clouds within the space, so that the encoding efficiency can be improved.
[0431] Also, as shown in FIG. 56, the three-dimensional data encoding device may select whether to apply octree encoding or non-octree encoding for each volume. In this case, the three-dimensional data encoding device adds octree encoding application information to the header of the volume. Thereby, the three-dimensional data decoding device can determine whether octree encoding has been applied for each volume. Also, in this case, the three-dimensional data encoding device sets sub-coordinates for each volume and encodes the difference value obtained by subtracting the value of the sub-coordinates from the coordinates of each point cloud within the volume.
[0432] Thereby, the three-dimensional data encoding device can appropriately switch whether to apply octree encoding according to the shape of the object or the number of point clouds within the volume, so that the encoding efficiency can be improved.
[0433] In the above description, as an example of non-octree encoding, an example of encoding the difference obtained by subtracting the sub-coordinates from the coordinates of each point cloud was shown. However, it is not necessarily limited to this, and any encoding method other than octree encoding may be used for encoding. For example, as shown in FIG. 57, as non-octree encoding, the three-dimensional data encoding device may use a method of encoding the values of the point cloud itself in the sub-map, space, or volume (hereinafter referred to as original coordinate encoding) instead of the difference from the sub-coordinates.
[0434] In that case, the three-dimensional data encoding device stores information indicating that the original coordinate encoding has been applied to the target space (sub-map, space, or volume) in the header. Thereby, the three-dimensional data decoding device can determine whether the original coordinate encoding has been applied to the target space.
[0435] Also, when applying the original coordinate encoding, the three-dimensional data encoding device may perform encoding without applying quantization and arithmetic encoding to the original coordinates. Further, the three-dimensional data encoding device may encode the original coordinates with a predetermined fixed bit length. Thereby, the three-dimensional data encoding device can generate a stream of a certain bit length at a certain timing.
[0436] In the above description, as an example of non-octree encoding, an example of encoding the difference obtained by subtracting the sub-coordinates from the coordinates of each point cloud was shown. However, it is not necessarily limited to this.
[0437] For example, the three-dimensional data encoding device may sequentially encode the difference values between the coordinates of each point cloud. FIG. 58 is a diagram for explaining the operation in this case. For example, in the example shown in FIG. 58, when encoding the point cloud PA, the three-dimensional data encoding device uses the sub-coordinates as the predicted coordinates and encodes the difference value between the coordinates of the point cloud PA and the predicted coordinates. Also, when encoding the point cloud PB, the three-dimensional data encoding device uses the coordinates of the point cloud PA as the predicted coordinates and encodes the difference value between the point cloud PB and the predicted coordinates. Further, when encoding the point cloud PC, the three-dimensional data encoding device uses the point cloud PB as the predicted coordinates and encodes the difference value between the point cloud PB and the predicted coordinates. In this way, the three-dimensional data encoding device may set the scan order for a plurality of point clouds and encode the difference value between the coordinates of the target point cloud to be processed and the coordinates of the immediately preceding point cloud in the scan order with respect to the target point cloud.
[0438] Also, in the above description, the sub-coordinates were the coordinates of the lower left front corner of the sub-map, but the position of the sub-coordinates is not limited to this. FIGS. 59 to 61 are diagrams showing another example of the position of the sub-coordinates. The sub-coordinates may be set at any coordinates within the target space (sub-map, space, or volume). That is, the sub-coordinates may be the coordinates of the lower left front corner of the target space as described above. As shown in FIG. 59, the sub-coordinates may be the coordinates of the center of the target space. As shown in FIG. 60, the sub-coordinates may be the coordinates of the upper right back corner of the target space. Also, the sub-coordinates are not limited to the coordinates of the lower left front or upper right back corner of the target space, and may be the coordinates of any corner of the target space.
[0439] Also, the set position of the sub-coordinates may be the same as the coordinates of a certain point cloud within the target space (sub-map, space, or volume). For example, in the example shown in FIG. 61, the coordinates of the sub-coordinates match the coordinates of the point cloud PD.
[0440] In addition, in this embodiment, an example of switching between applying octree encoding and non-octree encoding is shown, but it is not necessarily limited to this. For example, the three-dimensional data encoding device may switch between applying another tree structure other than the octree and applying a non-tree structure other than the tree structure. For example, another tree structure is a kd-tree that performs division using a plane perpendicular to one of the coordinate axes. Note that any method may be used as another tree structure.
[0441] In addition, in this embodiment, an example of encoding the coordinate information of the point cloud is shown, but it is not necessarily limited to this. The three-dimensional data encoding device may encode, for example, color information, three-dimensional feature amounts, or feature amounts of visible light in the same way as the coordinate information. For example, the three-dimensional data encoding device may set the average value of the color information of each point cloud in the submap as sub-color information, and encode the difference between the color information of each point cloud and the sub-color information.
[0442] In addition, in this embodiment, an example of selecting an encoding method (octree encoding or non-octree encoding) with good encoding efficiency according to the number of point clouds, etc. is shown, but it is not necessarily limited to this. For example, the three-dimensional data encoding device on the server side holds the bitstream of the point cloud encoded by octree encoding, the bitstream of the point cloud encoded by non-octree encoding, and the bitstream of the point cloud encoded by both, and switches the bitstream to be transmitted to the three-dimensional data decoding device according to the communication environment or the processing capacity of the three-dimensional data decoding device.
[0443] FIG. 62 is a diagram showing a syntax example of volume when switching the application of octree encoding. The syntax shown in FIG. 62 is basically the same as the syntax shown in FIG. 52, but the difference is that each piece of information is information in volume units. Specifically, NumOfPoint indicates the number of point clouds included in the volume. sub_coordinate_x, sub_coordinate_y, and sub_coordinate_z are sub-coordinate information of the volume.
[0444] Also, diff_x[i], diff_y[i], and diff_z[i] are the differential coordinates of the i-th point cloud in the volume. diff_x[i] indicates the difference value between the x coordinate of the i-th point cloud in the volume and the x coordinate of the sub-coordinate. diff_y[i] indicates the difference value between the y coordinate of the i-th point cloud in the volume and the y coordinate of the sub-coordinate. diff_z[i] indicates the difference value between the z coordinate of the i-th point cloud in the volume and the z coordinate of the sub-coordinate.
[0445] Note that when the relative position of the volume in the space can be calculated, the three-dimensional data encoding device may not include the sub-coordinate information in the header of the volume. That is, the three-dimensional data encoding device may calculate the relative position of the volume in the space without including the sub-coordinate information in the header, and use the calculated position as the sub-coordinate of each volume.
[0446] As described above, the three-dimensional data encoding device according to the present embodiment determines whether to encode a target space unit among a plurality of space units (for example, sub-maps, spaces, or volumes) included in the three-dimensional data in an octree structure (for example, S1522 in FIG. 53). For example, when the number of three-dimensional points included in the target space unit is more than a predetermined threshold, the three-dimensional data encoding device determines to encode the target space unit in an octree structure. Also, when the number of three-dimensional points included in the target space unit is less than or equal to the above threshold, the three-dimensional data encoding device determines not to encode the target space unit in an octree structure.
[0447] When it is determined that the target space unit is to be encoded in an octree structure (Yes in S1522), the three-dimensional data encoding device encodes the target space unit using the octree structure (S1523). Also, when it is determined that the target space unit is not to be encoded in the octree structure (No in S1522), the three-dimensional data encoding device encodes the target space unit in a manner different from the octree structure (S1524). For example, in a different manner, the three-dimensional data encoding device encodes the coordinates of the three-dimensional points included in the target space unit. Specifically, in a different manner, the three-dimensional data encoding device encodes the difference between the reference coordinates of the target space unit and the coordinates of the three-dimensional points included in the target space unit.
[0448] Next, the three-dimensional data encoding device adds information indicating whether the target space unit has been encoded in the octree structure to the bitstream (S1525).
[0449] According to this, the three-dimensional data encoding device can reduce the data amount of the encoded signal, so the encoding efficiency can be improved.
[0450] For example, the three-dimensional data encoding device includes a processor and a memory, and the processor performs the above processing using the memory.
[0451] Also, the three-dimensional data decoding device according to the present embodiment decodes from the bitstream information indicating whether to decode a target space unit (for example, a submap, a space, or a volume) among a plurality of target space units included in the three-dimensional data in an octree structure (for example, S1531 in FIG. 54). When it is indicated by the above information that the target space unit is to be decoded in the octree structure (Yes in S1532), the three-dimensional data decoding device decodes the target space unit using the octree structure (S1533).
[0452] When it is shown that the target space unit is not decoded in the octree structure based on the above information (No in S1532), the three-dimensional data decoding device decodes the target space unit in a method different from the octree structure (S1534). For example, the three-dimensional data decoding device decodes the coordinates of the three-dimensional points included in the target space unit in a different method. Specifically, the three-dimensional data decoding device decodes the difference between the reference coordinates of the target space unit and the coordinates of the three-dimensional points included in the target space unit in a different method.
[0453] According to this, the three-dimensional data decoding device can reduce the data amount of the encoded signal, so the encoding efficiency can be improved.
[0454] For example, the three-dimensional data decoding device includes a processor and a memory, and the processor performs the above processing using the memory.
[0455] (Embodiment 9) In this embodiment, another example of the encoding method of the tree structure such as the octree structure will be described. FIG. 63 is a diagram showing an example of the tree structure according to this embodiment. Note that FIG. 63 shows an example of a quadtree structure.
[0456] A leaf containing three-dimensional points is called an effective leaf, and a leaf not containing three-dimensional points is called an invalid leaf. A branch with the number of effective leaves greater than or equal to the threshold is called a dense branch. A branch with the number of effective leaves less than the threshold is called a sparse branch.
[0457] The three-dimensional data encoding device calculates the number of three-dimensional points (that is, the number of effective leaves) included in each branch in a certain layer of the tree structure. FIG. 63 shows an example when the threshold is 5. In this example, there are two branches in layer 1. Since the left branch contains 7 three-dimensional points, the left branch is determined to be a dense branch. Since the right branch contains 2 three-dimensional points, the right branch is determined to be a sparse branch.
[0458] FIG. 64 is a diagram showing an example of the number of valid leaves (3D points) that each branch of layer 5 has. The horizontal axis of FIG. 64 indicates an index that is the identification number of the branches of layer 5. As shown in FIG. 64, a specific branch contains significantly more three-dimensional points than other branches. Occupancy encoding is more effective for such dense branches than for sparse branches.
[0459] Hereinafter, the application methods of occupancy encoding and location encoding will be described. FIG. 65 is a diagram showing the relationship between the number of three-dimensional points (number of valid leaves) contained in each branch in layer 5 and the encoding method to be applied. As shown in FIG. 65, the three-dimensional data encoding device applies occupancy encoding to dense branches and location encoding to sparse branches. Thereby, the encoding efficiency can be improved.
[0460] FIG. 66 is a diagram showing an example of a dense branch region in LiDAR data. As shown in FIG. 66, depending on the region, the density of the three-dimensional points calculated from the number of three-dimensional points contained in each branch is different.
[0461] Also, separating dense three-dimensional points (branches) and sparse three-dimensional points (branches) has the following advantages. The density of the three-dimensional points increases as the distance from the LiDAR sensor decreases. Therefore, by separating the branches according to density, it becomes possible to partition in the distance direction. Such partitioning is effective in specific applications. Also, for sparse branches, it is effective to use a method other than occupancy encoding.
[0462] In the present embodiment, the three-dimensional data encoding device separates the input three-dimensional point cloud into two or more sub-three-dimensional point clouds and applies different encoding methods to each sub-three-dimensional point cloud.
[0463] For example, the three-dimensional data encoding device separates the input three-dimensional point cloud into a sub-three-dimensional point cloud A (dense cloud) including dense branches and a sub-three-dimensional point cloud B (sparse cloud) including sparse branches. FIG. 67 is a diagram showing an example of the sub-three-dimensional point cloud A (dense three-dimensional point cloud) including dense branches separated from the tree structure shown in FIG. 63. FIG. 68 is a diagram showing an example of the sub-three-dimensional point cloud B (sparse three-dimensional point cloud) including sparse branches separated from the tree structure shown in FIG. 63.
[0464] Next, the three-dimensional data encoding device encodes the sub-three-dimensional point cloud A by occupancy encoding and encodes the sub-three-dimensional point cloud B by location encoding.
[0465] Here, as different encoding methods, examples of applying different encoding schemes (occupancy encoding and location encoding) are shown. However, for example, the three-dimensional data encoding device may use the same encoding scheme for the sub-three-dimensional point cloud A and the sub-three-dimensional point cloud B, and may vary the parameters used for encoding between the sub-three-dimensional point cloud A and the sub-three-dimensional point cloud B.
[0466] Hereinafter, the flow of the three-dimensional data encoding process by the three-dimensional data encoding device will be described. FIG. 69 is a flowchart of the three-dimensional data encoding process by the three-dimensional data encoding device according to the present embodiment.
[0467] First, the three-dimensional data encoding device separates the input three-dimensional point cloud into sub-three-dimensional point clouds (S1701). The three-dimensional data encoding device may perform this separation automatically or based on information input by the user. For example, the range of the sub-three-dimensional point cloud may be specified by the user. As an example of automatic performance, for example, when the input data is LiDAR data, the three-dimensional data encoding device performs the separation using the distance information to each point cloud. Specifically, the three-dimensional data encoding device separates the point cloud within a certain range from the measurement point and the point cloud outside the range. Also, the three-dimensional data encoding device may perform the separation using information such as an important area and a non-important area.
[0468] Next, the three-dimensional data encoding device generates encoded data (encoded bit stream) by encoding the sub-three-dimensional point cloud A with method A (S1702). Further, the three-dimensional data encoding device generates encoded data by encoding the sub-three-dimensional point cloud B with method B (S1703). Note that the three-dimensional data encoding device may encode the sub-three-dimensional point cloud B with method A. In this case, the three-dimensional data encoding device encodes the sub-three-dimensional point cloud B using parameters different from the encoding parameters used for encoding the sub-three-dimensional point cloud A. For example, this parameter may be a quantization parameter. For example, the three-dimensional data encoding device encodes the sub-three-dimensional point cloud B using a quantization parameter larger than the quantization parameter used for encoding the sub-three-dimensional point cloud A. In this case, the three-dimensional data encoding device may add information indicating the quantization parameter used for encoding the sub-three-dimensional point cloud to the header of the encoded data of each sub-three-dimensional point cloud.
[0469] Next, the three-dimensional data encoding device generates a bit stream by combining the encoded data obtained in step S1702 and the encoded data obtained in step S1703 (S1704).
[0470] Further, the three-dimensional data encoding device may encode information for decoding each sub-three-dimensional point cloud as header information of the bit stream. For example, the three-dimensional data encoding device may encode the following information.
[0471] The header information may include information indicating the number of encoded sub-three-dimensional points. In this example, this information indicates 2.
[0472] The header information may include information indicating the number of three-dimensional points included in each sub-three-dimensional point cloud and the encoding method. In this example, this information indicates the number of three-dimensional points included in the sub-three-dimensional point cloud A, the encoding method (method A) applied to the sub-three-dimensional point cloud A, the number of three-dimensional points included in the sub-three-dimensional point cloud B, and the encoding method (method B) applied to the sub-three-dimensional point cloud B.
[0473] The header information may include information for identifying the start position or the end position of the encoded data of each sub-three-dimensional point cloud.
[0474] Also, the three-dimensional data encoding device may encode the sub-three-dimensional point cloud A and the sub-three-dimensional point cloud B in parallel. Alternatively, the three-dimensional data encoding device may encode the sub-three-dimensional point cloud A and the sub-three-dimensional point cloud B in order.
[0475] Also, the method of separating into sub-three-dimensional point clouds is not limited to the above. For example, the three-dimensional data encoding device changes the separation method, performs encoding using each of a plurality of separation methods, and calculates the encoding efficiency of the encoded data obtained using each separation method. Then, the three-dimensional data encoding device selects the separation method with the highest encoding efficiency. For example, the three-dimensional data encoding device separates the three-dimensional point cloud in each of a plurality of layers, calculates the encoding efficiency in each case, selects the separation method (that is, the layer for performing separation) with the highest encoding efficiency, and generates and encodes the sub-three-dimensional point cloud using the selected separation method.
[0476] Also, when combining the encoded data, the three-dimensional data encoding device may arrange the encoding information of the important sub-three-dimensional point cloud closer to the head of the bit stream. As a result, the three-dimensional data decoding device can obtain important information by simply decoding the head bit stream, so that it becomes possible to obtain important information earlier.
[0477] Next, the flow of the three-dimensional data decoding process by the three-dimensional data decoding device will be described. FIG. 70 is a flowchart of the three-dimensional data decoding process by the three-dimensional data decoding device according to the present embodiment.
[0478] First, the three-dimensional data decoding device acquires, for example, the bitstream generated by the above three-dimensional data encoding device. Next, the three-dimensional data decoding device separates the encoded data of the sub-three-dimensional point cloud A and the encoded data of the sub-three-dimensional point cloud B from the acquired bitstream (S1711). Specifically, the three-dimensional data decoding device decodes information for decoding each sub-three-dimensional point cloud from the header information of the bitstream, and separates the encoded data of each sub-three-dimensional point cloud using the information.
[0479] Next, the three-dimensional data decoding device obtains the sub-three-dimensional point cloud A by decoding the encoded data of the sub-three-dimensional point cloud A using Method A (S1712). Also, the three-dimensional data decoding device obtains the sub-three-dimensional point cloud B by decoding the encoded data of the sub-three-dimensional point cloud B using Method B (S1713). Next, the three-dimensional data decoding device combines the sub-three-dimensional point cloud A and the sub-three-dimensional point cloud B (S1714).
[0480] Note that the three-dimensional data decoding device may decode the sub-three-dimensional point cloud A and the sub-three-dimensional point cloud B in parallel. Or, the three-dimensional data decoding device may decode the sub-three-dimensional point cloud A and the sub-three-dimensional point cloud B in sequence.
[0481] Also, the three-dimensional data decoding device may decode the necessary sub-three-dimensional point clouds. For example, the three-dimensional data decoding device may decode the sub-three-dimensional point cloud A and not decode the sub-three-dimensional point cloud B. For example, when the sub-three-dimensional point cloud A is a three-dimensional point cloud included in an important area of LiDAR data, the three-dimensional data decoding device decodes the three-dimensional point cloud of the important area. Self-position estimation in a vehicle or the like is performed using the three-dimensional point cloud of this important area.
[0482] Next, a specific example of the encoding process according to this embodiment will be described. FIG. 71 is a flowchart of the three-dimensional data encoding process by the three-dimensional data encoding device according to this embodiment.
[0483] First, the three-dimensional data encoding device separates the input three-dimensional points into a sparse three-dimensional point group and a dense three-dimensional point group (S1721). Specifically, the three-dimensional data encoding device counts the number of valid leaves of the branches of a layer with an octree structure. The three-dimensional data encoding device sets each branch as a dense branch or a sparse branch according to the number of valid leaves of each branch. Then, the three-dimensional data encoding device generates a sub-three-dimensional point group (dense three-dimensional point group) formed by gathering dense branches and a sub-three-dimensional point group (sparse three-dimensional point group) formed by gathering sparse branches.
[0484] Next, the three-dimensional data encoding device generates encoded data by encoding the sparse three-dimensional point group (S1722). For example, the three-dimensional data encoding device encodes the sparse three-dimensional point group using location encoding.
[0485] Also, the three-dimensional data encoding device generates encoded data by encoding the dense three-dimensional point group (S1723). For example, the three-dimensional data encoding device encodes the dense three-dimensional point group using occupancy encoding.
[0486] Next, the three-dimensional data encoding device generates a bitstream by combining the encoded data of the sparse three-dimensional point group obtained in step S1722 and the encoded data of the dense three-dimensional point group obtained in step S1723 (S1724).
[0487] Also, the three-dimensional data encoding device may encode information for decoding the sparse three-dimensional point group and the dense three-dimensional point group as the header information of the bitstream. For example, the three-dimensional data encoding device may encode the following information.
[0488] The header information may include information indicating the number of encoded sub-three-dimensional point groups. In this example, this information indicates 2.
[0489] The header information may include information indicating the number of three-dimensional points included in each sub-three-dimensional point group and the encoding method. In this example, this information indicates the number of three-dimensional points included in the sparse three-dimensional point group, the encoding method (location encoding) applied to the sparse three-dimensional point group, the number of three-dimensional points included in the dense three-dimensional point group, and the encoding method (occupancy encoding) applied to the dense three-dimensional point group.
[0490] The header information may include information for identifying the start position or end position of the encoded data of each sub-three-dimensional point group. In this example, this information indicates at least one of the start position and end position of the encoded data of the sparse three-dimensional point group and the start position and end position of the encoded data of the dense three-dimensional point group.
[0491] Also, the three-dimensional data encoding device may encode the sparse three-dimensional point group and the dense three-dimensional point group in parallel. Alternatively, the three-dimensional data encoding device may encode the sparse three-dimensional point group and the dense three-dimensional point group in sequence.
[0492] Next, a specific example of the three-dimensional data decoding process will be described. FIG. 72 is a flowchart of the three-dimensional data decoding process by the three-dimensional data decoding device according to the present embodiment.
[0493] First, the three-dimensional data decoding device acquires, for example, the bitstream generated by the above three-dimensional data encoding device. Next, the three-dimensional data decoding device separates the encoded data of the sparse three-dimensional point group and the encoded data of the dense three-dimensional point group from the acquired bitstream (S1731). Specifically, the three-dimensional data decoding device decodes the information for decoding each sub-three-dimensional point group from the header information of the bitstream, and separates the encoded data of each sub-three-dimensional point group using the information. In this example, the three-dimensional data decoding device separates the encoded data of the sparse three-dimensional point group and the dense three-dimensional point group from the bitstream using the header information.
[0494] Next, the three-dimensional data decoding device obtains a sparse three-dimensional point cloud by decoding the encoded data of the sparse three-dimensional point cloud (S1732). For example, the three-dimensional data decoding device decodes the sparse three-dimensional point cloud using location decoding for decoding the location-encoded encoded data.
[0495] Also, the three-dimensional data decoding device obtains a dense three-dimensional point cloud by decoding the encoded data of the dense three-dimensional point cloud (S1733). For example, the three-dimensional data decoding device decodes the dense three-dimensional point cloud using occupancy decoding for decoding the occupancy-encoded encoded data.
[0496] Next, the three-dimensional data decoding device combines the sparse three-dimensional point cloud obtained in step S1732 and the dense three-dimensional point cloud obtained in step S1733 (S1734).
[0497] Note that the three-dimensional data decoding device may decode the sparse three-dimensional point cloud and the dense three-dimensional point cloud in parallel. Or, the three-dimensional data decoding device may decode the sparse three-dimensional point cloud and the dense three-dimensional point cloud in order.
[0498] Also, the three-dimensional data decoding device may decode some necessary sub-three-dimensional point clouds. For example, the three-dimensional data decoding device may decode the dense three-dimensional point cloud and not decode the sparse three-dimensional data. For example, when the dense three-dimensional point cloud is a three-dimensional point cloud included in an important area of LiDAR data, the three-dimensional data decoding device decodes the three-dimensional point cloud of the important area. Self-position estimation in a vehicle or the like is performed using the three-dimensional point cloud of this important area.
[0499] FIG. 73 is a flowchart of the encoding process according to the present embodiment. First, the three-dimensional data encoding device generates a sparse three-dimensional point cloud and a dense three-dimensional point cloud by separating the input three-dimensional point cloud into a sparse three-dimensional point cloud and a dense three-dimensional point cloud (S1741).
[0500] Next, the three-dimensional data encoding device generates encoded data by encoding a dense three-dimensional point cloud (S1742). Also, the three-dimensional data encoding device generates encoded data by encoding a sparse three-dimensional point cloud (S1743). Finally, the three-dimensional data encoding device generates a bitstream by combining the encoded data of the sparse three-dimensional point cloud obtained in step S1742 and the encoded data of the dense three-dimensional point cloud obtained in step S1743 (S1744).
[0501] FIG. 74 is a flowchart of the decoding process according to the present embodiment. First, the three-dimensional data decoding device extracts the encoded data of the dense three-dimensional point cloud and the encoded data of the sparse three-dimensional point cloud from the bitstream (S1751). Next, the three-dimensional data decoding device obtains the decoded data of the dense three-dimensional point cloud by decoding the encoded data of the dense three-dimensional point cloud (S1752). Also, the three-dimensional data decoding device obtains the decoded data of the sparse three-dimensional point cloud by decoding the encoded data of the sparse three-dimensional point cloud (S1753). Next, the three-dimensional data decoding device generates a three-dimensional point cloud by combining the decoded data of the dense three-dimensional point cloud obtained in step S1752 and the decoded data of the sparse three-dimensional point cloud obtained in step S1753 (S1754).
[0502] Note that the three-dimensional data encoding device and the three-dimensional data decoding device may encode or decode either the dense three-dimensional point cloud or the sparse three-dimensional point cloud first. Also, the encoding process or the decoding process may be performed in parallel by a plurality of processors or the like.
[0503] Further, the three-dimensional data encoding device may encode either the dense three-dimensional point cloud or the sparse three-dimensional point cloud. For example, when important information is included in the dense three-dimensional point cloud, the three-dimensional data encoding device extracts the dense three-dimensional point cloud and the sparse three-dimensional point cloud from the input three-dimensional point cloud, encodes the dense three-dimensional point cloud, and does not encode the sparse three-dimensional point cloud. Thereby, the three-dimensional data encoding device can add important information to the stream while suppressing the amount of bits. For example, between a server and a client, when the server receives a transmission request for three-dimensional point cloud information around the client from the client, the server encodes important information around the client as a dense three-dimensional point cloud and transmits it to the client. Thereby, the server can transmit the information required by the client while suppressing the network bandwidth.
[0504] Further, the three-dimensional data decoding device may decode either the dense three-dimensional point cloud or the sparse three-dimensional point cloud. For example, when important information is included in the dense three-dimensional point cloud, the three-dimensional data decoding device decodes the dense three-dimensional point cloud and does not decode the sparse three-dimensional point cloud. Thereby, the three-dimensional data decoding device can obtain the necessary information while suppressing the processing load of the decoding process.
[0505] FIG. 75 is a flowchart of the three-dimensional point separation process (S1741) shown in FIG. 73. First, the three-dimensional data encoding device sets the layer L and the threshold TH (S1761). Note that the three-dimensional data encoding device may add information indicating the set layer L and threshold TH to the bit stream. That is, the three-dimensional data encoding device may generate a bit stream including information indicating the set layer L and threshold TH.
[0506] Next, the three-dimensional data encoding device moves the position of the processing target from the root of the octree to the first branch of layer L. That is, the three-dimensional data encoding device selects the first branch of layer L as the branch to be processed (S1762).
[0507] Next, the three-dimensional data encoding device counts the number of valid leaves of the branch to be processed in layer L (S1763). If the number of valid leaves of the branch to be processed is more than the threshold TH (Yes in S1764), the three-dimensional data encoding device registers the branch to be processed as a dense branch in the dense three-dimensional point cloud (S1765). On the other hand, if the number of valid leaves of the branch to be processed is less than or equal to the threshold TH (No in S1764), the three-dimensional data encoding device registers the branch to be processed as a sparse branch in the sparse three-dimensional point cloud (S1766).
[0508] If the processing of all the branches in layer L has not been completed (No in S1767), the three-dimensional data encoding device moves the position to be processed to the next branch in layer L. That is, the three-dimensional data encoding device selects the next branch in layer L as the branch to be processed (S1768). Then, the three-dimensional data encoding device performs the processing from step S1763 onwards on the selected next branch to be processed.
[0509] The above processing is repeated until the processing of all the branches in layer L is completed (Yes in S1767).
[0510] Note that in the above description, layer L and threshold TH are set in advance, but it is not necessarily limited to this. For example, the three-dimensional data encoding device sets a plurality of patterns of the pair of layer L and threshold TH, generates a dense three-dimensional point cloud and a sparse three-dimensional point cloud using each pair, and encodes each of them. The three-dimensional data encoding device finally encodes the dense three-dimensional point cloud and the sparse three-dimensional point cloud with the pair of layer L and threshold TH for which the encoding efficiency of the generated encoded data is the highest among the plurality of pairs. Thereby, the encoding efficiency can be improved. Also, the three-dimensional data encoding device may, for example, calculate layer L and threshold TH. For example, the three-dimensional data encoding device may set the value that is half of the maximum value of the layers included in the tree structure as layer L. Also, the three-dimensional data encoding device may set the value that is half of the total number of the plurality of three-dimensional points included in the tree structure as threshold TH.
[0511] In the above description, an example of classifying the input three-dimensional point cloud into two types, a dense three-dimensional point cloud and a sparse three-dimensional point cloud, was described. However, the three-dimensional data encoding device may classify the input three-dimensional point cloud into three or more types of three-dimensional point clouds. For example, when the number of valid leaves of the branch to be processed is equal to or greater than the threshold TH1, the three-dimensional data encoding device classifies the branch to be processed into the first dense three-dimensional point cloud. When the number of valid leaves of the branch to be processed is less than the first threshold TH1 and equal to or greater than the second threshold TH2, the three-dimensional data encoding device classifies the branch to be processed into the second dense three-dimensional point cloud. When the number of valid leaves of the branch to be processed is less than the second threshold TH2 and equal to or greater than the third threshold TH3, the three-dimensional data encoding device classifies the branch to be processed into the first sparse three-dimensional point cloud. When the number of valid leaves of the branch to be processed is less than the threshold TH3, the three-dimensional data encoding device classifies the branch to be processed into the second sparse three-dimensional point cloud.
[0512] Hereinafter, a syntax example of the encoded data of the three-dimensional point cloud according to the present embodiment will be described. FIG. 76 is a diagram showing this syntax example. pc_header() is, for example, the header information of a plurality of input three-dimensional points.
[0513] Num_sub_pc shown in FIG. 76 indicates the number of sub three-dimensional point clouds. numPoint[i] indicates the number of three-dimensional points included in the i-th sub three-dimensional point cloud. coding_type[i] is coding type information indicating the coding type (encoding method) applied to the i-th sub three-dimensional point cloud. For example, coding_type = 00 indicates that location encoding is applied. coding_type = 01 indicates that occupancy encoding is applied. coding_type = 10 or 11 indicates that another encoding method is applied.
[0514] data_sub_cloud() is the encoded data of the i-th sub three-dimensional point cloud. coding_type_00_data is the encoded data to which the encoding type with coding_type being 00 is applied, for example, the encoded data to which location encoding is applied. coding_type_01_data is the encoded data to which the encoding type with coding_type being 01 is applied, for example, the encoded data to which occupancy encoding is applied.
[0515] end_of_data is the end information indicating the end of the encoded data. For example, a fixed bit string not used for the encoded data is assigned to this end_of_data. Thereby, the three-dimensional data decoding device can skip the decoding process of the encoded data that does not need to be decoded, for example, by searching for the bit string of end_of_data from the bit stream.
[0516] Note that the three-dimensional data encoding device may entropy-encode the encoded data generated by the above method. For example, the three-dimensional data encoding device binarizes each value and then calculates and encodes it.
[0517] Also, in this embodiment, examples of a quadtree structure or an octree structure are shown, but it is not necessarily limited to this. The above method may be applied to an N-ary tree (N is an integer of 2 or more) such as a binary tree, a 16-ary tree, or other tree structures.
[0518] [Modification Example] In the above description, as shown in FIGS. 68 and 69, a tree structure including a dense branch and its upper layer (the tree structure from the root of the entire tree structure to the root of the dense branch) is encoded, and a tree structure including a sparse branch and its upper layer (the tree structure from the root of the entire tree structure to the root of the sparse branch) is encoded. In this modification example, the three-dimensional data encoding device separates the dense branch and the sparse branch and encodes the dense branch and the sparse branch. That is, the upper layer tree structure is not included in the tree structure to be encoded. For example, the three-dimensional data encoding device applies occupancy encoding to the dense branch and location encoding to the sparse branch.
[0519] FIG. 77 is a diagram showing an example of a dense branch separated from the tree structure shown in FIG. 63. FIG. 78 is a diagram showing an example of a sparse branch separated from the tree structure shown in FIG. 63. In this modification example, the tree structures shown in FIGS. 77 and 78 are each encoded.
[0520] Further, instead of encoding the upper tree structure, the three-dimensional data encoding device encodes information indicating the position of a branch. For example, this information indicates the position of the root of the branch.
[0521] For example, the three-dimensional data encoding device encodes layer information indicating the layer in which a dense branch is generated and branch information indicating which branch in the layer the dense branch is as the encoded data of the dense branch. Thereby, the three-dimensional data decoding device decodes the layer information and the branch information from the bit stream, and using these layer information and branch information, can grasp which layer and which branch's three-dimensional point group the decoded dense branch is. Similarly, the three-dimensional data encoding device encodes layer information indicating the layer in which a sparse branch is generated and branch information indicating which branch in the layer the sparse branch is as the encoded data of the sparse branch.
[0522] Thereby, the three-dimensional data decoding device decodes the layer information and the branch information from the bit stream, and using these layer information and branch information, can grasp which layer and which branch's three-dimensional point group the decoded sparse branch is. Thereby, since the overhead due to encoding the information of the layer above the dense branch and the sparse branch can be reduced, the encoding efficiency can be improved.
[0523] Note that the branch information may indicate values assigned to each branch in the layer indicated by the layer information. Also, the branch information may indicate values assigned to each node starting from the root of the octree. In this case, the layer information may not be encoded. Also, the three-dimensional data encoding device may generate a plurality of dense branches and a plurality of sparse branches respectively.
[0524] FIG. 79 is a flowchart of the encoding process in this modified example. First, the three-dimensional data encoding device generates one or more sparse branches and one or more dense branches from the input three-dimensional point cloud (S1771).
[0525] Next, the three-dimensional data encoding device generates encoded data by encoding the dense branches (S1772). Next, the three-dimensional data encoding device determines whether the encoding of all the dense branches generated in step S1771 has been completed (S1773).
[0526] If the encoding of all the dense branches has not been completed (No in S1773), the three-dimensional data encoding device selects the next dense branch (S1774) and generates encoded data by encoding the selected dense branch (S1772).
[0527] On the other hand, if the encoding of all the dense branches has been completed (Yes in S1773), the three-dimensional data encoding device generates encoded data by encoding the sparse branches (S1775). Next, the three-dimensional data encoding device determines whether the encoding of all the sparse branches generated in step S1771 has been completed (S1776).
[0528] If the encoding of all the sparse branches has not been completed (No in S1776), the three-dimensional data encoding device selects the next sparse branch (S1777) and generates encoded data by encoding the selected sparse branch (S1775).
[0529] On the other hand, if the encoding of all the sparse branches has been completed (Yes in S1776), the three-dimensional data encoding device combines the encoded data generated in steps S1772 and S1775 to generate a bit stream (S1778).
[0530] Figure 79 is a flowchart of the decoding process in this modified example. First, the three-dimensional data decoding device extracts one or more encoded data of dense branches and one or more encoded data of sparse branches from the bit stream (S1781). Next, the three-dimensional data decoding device obtains decoded data of the dense branches by decoding the encoded data of the dense branches (S1782).
[0531] Next, the three-dimensional data decoding device determines whether or not the decoding of all the encoded data of the dense branches extracted in step S1781 has been completed (S1783). If the decoding of all the encoded data of the dense branches has not been completed (No in S1783), the three-dimensional data decoding device selects the encoded data of the next dense branch (S1784) and obtains decoded data of the dense branches by decoding the selected encoded data of the dense branches (S1782).
[0532] On the other hand, if the decoding of all the encoded data of the dense branches has been completed (Yes in S1783), the three-dimensional data decoding device obtains decoded data of the sparse branches by decoding the encoded data of the sparse branches (S1785).
[0533] Next, the three-dimensional data decoding device determines whether or not the decoding of all the encoded data of the sparse branches extracted in step S1781 has been completed (S1786). If the decoding of all the encoded data of the sparse branches has not been completed (No in S1786), the three-dimensional data decoding device selects the encoded data of the next sparse branch (S1787) and obtains decoded data of the sparse branches by decoding the selected encoded data of the sparse branches (S1785).
[0534] On the other hand, if the decoding of all the encoded data of the sparse branches has been completed (Yes in S1786), the three-dimensional data decoding device generates a three-dimensional point cloud by combining the decoded data obtained in steps S1782 and S1785 (S1788).
[0535] Note that the three-dimensional data encoding device and the three-dimensional data decoding device may encode or decode either the dense branches or the sparse branches first. Also, the encoding process or the decoding process may be performed in parallel by a plurality of processors or the like.
[0536] Further, the three-dimensional data encoding device may encode either the dense branches or the sparse branches. Further, the three-dimensional data encoding device may encode a part of a plurality of dense branches. For example, when important information is included in a specific dense branch, the three-dimensional data encoding device extracts the dense branches and the sparse branches from the input three-dimensional point cloud, encodes the dense branch containing the important information, and does not encode the other dense branches and sparse branches. Thereby, the three-dimensional data encoding device can add important information to the stream while suppressing the amount of bits. For example, between a server and a client, when the server receives a transmission request for three-dimensional point cloud information around the client from the client, the server encodes important information around the client as dense branches and transmits it to the client. Thereby, the server can transmit the information required by the client while suppressing the network bandwidth.
[0537] Further, the three-dimensional data decoding device may decode either the dense branches or the sparse branches. Further, the three-dimensional data decoding device may decode a part of a plurality of dense branches. For example, when important information is included in a specific dense branch, the three-dimensional data decoding device decodes the specific dense branch and does not decode the other dense branches and sparse branches. Thereby, the three-dimensional data decoding device can obtain the necessary information while suppressing the processing load of the decoding process.
[0538] FIG. 81 is a flowchart of the three-dimensional point separation process (S1771) shown in FIG. 79. First, the three-dimensional data encoding device sets a layer L and a threshold TH (S1761). Note that the three-dimensional data encoding device may add information indicating the set layer L and threshold TH to the bit stream.
[0539] Next, the three-dimensional data encoding device selects the branch at the head of layer L as the branch to be processed (S1762). Next, the three-dimensional data encoding device counts the number of valid leaves of the branch to be processed in layer L (S1763). If the number of valid leaves of the branch to be processed is greater than the threshold TH (Yes in S1764), the three-dimensional data encoding device sets the branch to be processed as a dense branch and adds layer information and branch information to the bit stream (S1765A). On the other hand, if the number of valid leaves of the branch to be processed is less than or equal to the threshold TH (No in S1764), the three-dimensional data encoding device sets the branch to be processed as a sparse branch and adds layer information and branch information to the bit stream (S1766A).
[0540] If the processing of all branches in layer L has not been completed (No in S1767), the three-dimensional data encoding device selects the next branch in layer L as the branch to be processed (S1768). Then, the three-dimensional data encoding device performs the processing from step S1763 and later on the selected next branch to be processed. The above processing is repeated until the processing of all branches in layer L is completed (Yes in S1767).
[0541] Note that in the above description, layer L and threshold TH are set in advance, but it is not necessarily limited to this. For example, the three-dimensional data encoding device sets a plurality of patterns of the combination of layer L and threshold TH, generates dense branches and sparse branches using each combination, and encodes each of them. The three-dimensional data encoding device finally encodes the dense branches and sparse branches with the combination of layer L and threshold TH in which the encoding efficiency of the generated encoded data is the highest among the plurality of combinations. Thereby, the encoding efficiency can be improved. Also, the three-dimensional data encoding device may, for example, calculate layer L and threshold TH. For example, the three-dimensional data encoding device may set the value that is half of the maximum value of the layers included in the tree structure as layer L. Also, the three-dimensional data encoding device may set the value that is half of the total number of the plurality of three-dimensional points included in the tree structure as threshold TH.
[0542] Next, a syntax example of the encoded data of the three-dimensional point group according to this modification example will be described. FIG. 82 is a diagram showing this syntax example. In the syntax example shown in FIG. 82, layer_id[i] which is layer information and branch_id[i] which is branch information are added to the syntax example shown in FIG. 76.
[0543] layer_id[i] indicates the layer number to which the i-th sub three-dimensional point group belongs. branch_id[i] indicates the branch number within layer_id[i] of the i-th sub three-dimensional point group.
[0544] layer_id[i] and branch_id[i] are layer information and branch information representing the location of branches in, for example, an octree. For example, layer_id[i]=2 and branch_id[i]=5 indicate that the i-th branch is the 5th branch in layer 2.
[0545] Note that the three-dimensional data encoding device may entropy-encode the encoded data generated by the above method. For example, the three-dimensional data encoding device binarizes each value and then calculates and encodes it.
[0546] Also, in this modification example, an example of a quadtree structure or an octree structure is shown, but it is not necessarily limited to this. The above method may be applied to an N-ary tree (N is an integer of 2 or more) such as a binary tree or a 16-ary tree, or other tree structures.
[0547] As described above, the three-dimensional data encoding device according to this embodiment performs the processing shown in FIG. 83.
[0548] First, the three-dimensional data encoding device generates an N-ary tree structure (N is an integer of 2 or more) of a plurality of three-dimensional points included in the three-dimensional data (S1801).
[0549] Next, the three-dimensional data encoding device generates first encoded data by encoding a first branch having a first node included in a first layer which is any one of a plurality of layers included in the N-ary tree structure with a first encoding process (S1802).
[0550] Further, the three-dimensional data encoding device generates second encoded data by encoding a second branch rooted at a second node different from the first node included in the first layer with a second encoding process different from the first encoding process (S1803).
[0551] Next, the three-dimensional data encoding device generates a bit stream including the first encoded data and the second encoded data (S1804).
[0552] According to this, the three-dimensional data encoding device can apply an encoding process suitable for each branch included in the N-ary tree structure, so that the encoding efficiency can be improved.
[0553] For example, the number of three-dimensional points included in the first branch is less than a predetermined threshold, and the number of three-dimensional points included in the second branch is more than the threshold. That is, when the number of three-dimensional points included in the branch to be processed is less than the threshold, the three-dimensional data encoding device sets the branch to be processed as the first branch, and when the number of three-dimensional points included in the branch to be processed is more than the threshold, the three-dimensional data encoding device sets the branch to be processed as the second branch.
[0554] For example, the first encoded data includes first information representing a first N-ary tree structure of a plurality of first three-dimensional points included in the first branch in a first method. The second encoded data includes second information representing a second N-ary tree structure of a plurality of second three-dimensional points included in the second branch in a second method. That is, the first encoding process and the second encoding process differ in the encoding method.
[0555] For example, location encoding is used in the first encoding process, and occupancy encoding is used in the second encoding process. That is, the first information includes three-dimensional point information corresponding to each of a plurality of first three-dimensional points. Each three-dimensional point information includes indexes corresponding to each of a plurality of layers in the first N-ary tree structure. Each index indicates the sub-block to which the corresponding first three-dimensional point belongs among the N sub-blocks belonging to the corresponding layer. The second information corresponds to each of a plurality of sub-blocks belonging to a plurality of layers in the second N-ary tree structure, and includes a plurality of 1-bit information indicating whether or not a three-dimensional point exists in the corresponding sub-block.
[0556] For example, the quantization parameter used in the second encoding process is different from the quantization parameter used in the first encoding process. That is, the first encoding process and the second encoding process have the same encoding method but different parameters used.
[0557] For example, as shown in FIGS. 67 and 68, in the encoding of the first branch, the three-dimensional data encoding apparatus encodes, by the first encoding process, a tree structure including a tree structure from the root of the N-ary tree structure to the first node and the first branch. In the encoding of the second branch, the three-dimensional data encoding apparatus encodes, by the second encoding process, a tree structure including a tree structure from the root of the N-ary tree structure to the second node and the second branch.
[0558] For example, the first encoded data includes the encoded data of the first branch and third information indicating the position of the first node in the N-ary tree structure. The second encoded data includes the encoded data of the second branch and fourth information indicating the position of the second node in the N-ary tree structure.
[0559] For example, the third information includes information indicating the first layer (layer information) and information indicating which node among the nodes included in the first layer the first node is (branch information). The fourth information includes information indicating the first layer (layer information) and information indicating which node among the nodes included in the first layer the second node is (branch information).
[0560] For example, the first encoded data includes information (numPoint) indicating the number of three-dimensional points included in the first branch, and the second encoded data includes information (numPoint) indicating the number of three-dimensional points included in the second branch.
[0561] For example, the three-dimensional data encoding device includes a processor and a memory, and the processor performs the above processing using the memory.
[0562] Also, the three-dimensional data decoding device according to the embodiment performs the processing shown in FIG. 84.
[0563] First, the three-dimensional data decoding device obtains, from the bit stream, first encoded data obtained by encoding a first branch having as a root a first node included in a first layer that is any one of a plurality of layers included in an N-ary tree structure (N is an integer of 2 or more) of a plurality of three-dimensional points, and second encoded data obtained by encoding a second branch having as a root a second node different from the first node and included in the first layer (S1811).
[0564] Next, the three-dimensional data decoding device generates first decoded data of the first branch by decoding the first encoded data with a first decoding process (S1812).
[0565] Also, the three-dimensional data decoding device generates second decoded data of the second branch by decoding the second encoded data with a second decoding process different from the first decoding process (S1813).
[0566] Next, the three-dimensional data decoding device restores a plurality of three-dimensional points using the first decoded data and the second decoded data (S1814). For example, these three-dimensional points include a plurality of three-dimensional points indicated by the first decoded data and a plurality of three-dimensional points indicated by the second decoded data.
[0567] According to this, the three-dimensional data decoding device can decode a bit stream with improved encoding efficiency.
[0568] For example, the number of three-dimensional points included in the first branch is less than a predetermined threshold value, and the number of three-dimensional points included in the second branch is more than the threshold value.
[0569] For example, the first encoded data includes first information representing the first N-ary tree structure of a plurality of first three-dimensional points included in the first branch in a first manner. The second encoded data includes second information representing the second N-ary tree structure of a plurality of second three-dimensional points included in the second branch in a second manner. That is, the first decoding process and the second decoding process differ in the encoding method (decoding method).
[0570] For example, location encoding is used for the first encoded data, and occupancy encoding is used for the second encoded data. That is, the first information includes three-dimensional point information corresponding to each of the plurality of first three-dimensional points. Each three-dimensional point information includes indexes corresponding to each of a plurality of layers in the first N-ary tree structure. Each index indicates the sub-block to which the corresponding first three-dimensional point belongs among the N sub-blocks belonging to the corresponding layer. The second information corresponds to each of a plurality of sub-blocks belonging to a plurality of layers in the second N-ary tree structure, and includes a plurality of 1-bit information indicating whether or not a three-dimensional point exists in the corresponding sub-block.
[0571] For example, the quantization parameter used in the second decoding process is different from the quantization parameter used in the first decoding process. That is, the first decoding process and the second decoding process have the same encoding method (decoding method), but different parameters are used.
[0572] For example, as shown in FIGS. 67 and 68, in decoding the first branch, the three-dimensional data decoding device decodes, by a first decoding process, a tree structure including a tree structure from the root of the N-ary tree structure to the first node and the first branch. In decoding the second branch, the three-dimensional data decoding device decodes, by a second decoding process, a tree structure including a tree structure from the root of the N-ary tree structure to the second node and the second branch.
[0573] For example, the first encoded data includes the encoded data of the first branch and third information indicating the position of the first node in the N-ary tree structure. The second encoded data includes the encoded data of the second branch and fourth information indicating the position of the second node in the N-ary tree structure.
[0574] For example, the third information includes information indicating the first layer (layer information) and information indicating which node among the nodes included in the first layer the first node is (branch information). The fourth information includes information indicating the first layer (layer information) and information indicating which node among the nodes included in the first layer the second node is (branch information).
[0575] For example, the first encoded data includes information (numPoint) indicating the number of three-dimensional points included in the first branch, and the second encoded data includes information (numPoint) indicating the number of three-dimensional points included in the second branch.
[0576] For example, the three-dimensional data decoding device includes a processor and a memory, and the processor performs the above processing using the memory.
[0577] (Embodiment 10) In this embodiment, a method for controlling reference during encoding of the occupancy code will be described. Note that hereinafter, the operation of the three-dimensional data encoding device will be mainly described, but the same processing may also be performed in the three-dimensional data decoding device.
[0578] FIGS. 85 and 86 are diagrams showing the reference relationship according to this embodiment. FIG. 85 is a diagram showing the reference relationship on an octree structure, and FIG. 86 is a diagram showing the reference relationship in a spatial region.
[0579] In this embodiment, when encoding the encoding information of a node to be encoded (hereinafter referred to as a target node), the three-dimensional data encoding device refers to the encoding information of each node in the parent node to which the target node belongs (hereinafter referred to as a parent node). However, it does not refer to the encoding information of each node in other nodes in the same layer as the parent node (hereinafter referred to as parent adjacent nodes). That is, the three-dimensional data encoding device sets the reference to the parent adjacent nodes to be unavailable or prohibits the reference.
[0580] Note that the three-dimensional data encoding device may permit the reference to the encoding information in the parent node to which the parent node belongs (hereinafter referred to as a grandparent node). That is, the three-dimensional data encoding device may encode the encoding information of the target node by referring to the encoding information of the parent node and the grandparent node to which the target node belongs.
[0581] Here, the encoding information is, for example, an occupancy code. When encoding the occupancy code of the target node, the three-dimensional data encoding device refers to information indicating whether a point cloud is included in each node in the parent node to which the target node belongs (hereinafter referred to as occupancy information). In other words, when encoding the occupancy code of the target node, the three-dimensional data encoding device refers to the occupancy code of the parent node. On the other hand, the three-dimensional data encoding device does not refer to the occupancy information of each node in the parent adjacent nodes. That is, the three-dimensional data encoding device does not refer to the occupancy code of the parent adjacent nodes. Also, the three-dimensional data encoding device may refer to the occupancy information of each node in the grandparent node. That is, the three-dimensional data encoding device may refer to the occupancy information of the parent node and the parent adjacent nodes.
[0582] For example, when encoding the occupancy code of a target node, the three-dimensional data encoding device switches the encoding table used when entropy-encoding the occupancy code of the target node using the occupancy code of the parent node or grandparent node to which the target node belongs. The details will be described later. At this time, the three-dimensional data encoding device does not necessarily need to refer to the occupancy code of the parent adjacent node. As a result, when encoding the occupancy code of the target node, the three-dimensional data encoding device can appropriately switch the encoding table according to the information of the occupancy code of the parent node or grandparent node, so that the encoding efficiency can be improved. In addition, by not referring to the parent adjacent node, the three-dimensional data encoding device can suppress the confirmation process of the information of the parent adjacent node and the memory capacity for storing them. Also, it becomes easy to scan and encode the occupancy codes of the nodes of the octree in depth-first order.
[0583] Hereinafter, Modification Example 1 of the present embodiment will be described. FIG. 87 is a diagram showing the reference relationship in this modification example. In the above embodiment, the three-dimensional data encoding device does not refer to the occupancy code of the parent adjacent node, but whether to refer to the occupancy encoding of the parent adjacent node may be switched according to a specific condition.
[0584] For example, when the three-dimensional data encoding device performs encoding while scanning the octree in breadth-first order, it refers to the occupancy information of the nodes in the parent adjacent node to encode the occupancy code of the target node. On the other hand, when the three-dimensional data encoding device performs encoding while scanning the octree in depth-first order, it prohibits referring to the occupancy information of the nodes in the parent adjacent node. By appropriately switching the nodes that can be referred to according to the scan order (encoding order) of the nodes of the octree in this way, it is possible to improve the encoding efficiency and suppress the processing load.
[0585] In addition, the three-dimensional data encoding device may add information such as whether the octree is encoded in a breadth-first manner or a depth-first manner to the header of the bit stream. FIG. 88 is a diagram showing an example of the syntax of the header information in this case. The octree_scan_order shown in FIG. 88 is encoding order information (encoding order flag) indicating the encoding order of the octree. For example, when octree_scan_order is 0, it indicates breadth-first, and when it is 1, it indicates depth-first. Thereby, the three-dimensional data decoding device can know whether the bit stream is encoded in a breadth-first or depth-first manner by referring to octree_scan_order, so that the bit stream can be decoded appropriately.
[0586] Further, the three-dimensional data encoding device may add information indicating whether to prohibit the reference to the parent adjacent node to the header information of the bit stream. FIG. 89 is a diagram showing an example of the syntax of the header information in this case. The limit_refer_flag is prohibition switching information (prohibition switching flag) indicating whether to prohibit the reference to the parent adjacent node. For example, when limit_refer_flag is 1, it indicates that the reference to the parent adjacent node is prohibited, and when it is 0, it indicates no reference restriction (permission to reference the parent adjacent node).
[0587] That is, the three-dimensional data encoding device determines whether to prohibit the reference to the parent adjacent node, and based on the result of the above determination, switches whether to prohibit or permit the reference to the parent adjacent node. Further, the three-dimensional data encoding device generates a bit stream including prohibition switching information indicating whether to prohibit the reference to the parent adjacent node, which is the result of the above determination.
[0588] In addition, the three-dimensional data decoding device acquires the prohibition switching information indicating whether to prohibit the reference to the parent adjacent node from the bit stream, and based on the prohibition switching information, switches whether to prohibit or permit the reference to the parent adjacent node.
[0589] As a result, the three-dimensional data encoding device can generate a bit stream by controlling the reference to the parent adjacent node. Also, the three-dimensional data decoding device can obtain information indicating whether or not the reference to the parent adjacent node is prohibited from the header of the bit stream.
[0590] In addition, in this embodiment, although the encoding process of the occupancy code is described as an example of the encoding process for prohibiting the reference to the parent adjacent node, it is not necessarily limited to this. For example, the same method can be applied when encoding other information of the nodes of the octree. For example, when encoding other attribute information such as the color, normal vector, or reflectance added to the node, the method of this embodiment may be applied. Also, the same method can be applied when encoding the encoding table or the predicted value.
[0591] Next, Modification Example 2 of this embodiment will be described. In the above description, an example in which three reference adjacent nodes are used is shown, but four or more reference adjacent nodes may be used. FIG. 90 is a diagram showing an example of a target node and reference adjacent nodes.
[0592] For example, the three-dimensional data encoding device calculates an encoding table for entropy encoding the occupancy code of the target node shown in FIG. 90 by, for example, the following formula.
[0593] CodingTable=(FlagX0<<3)+(FlagX1<<2)+(FlagY<<1)+(FlagZ)
[0594] Here, CodingTable indicates an encoding table for the occupancy code of the target node, and indicates any value from 0 to 15. FlagXN is the occupancy information of the adjacent node XN (N = 0..1), and indicates 1 if the adjacent node XN contains a point cloud (occupied), and 0 otherwise. FlagY is the occupancy information of the adjacent node Y, and indicates 1 if the adjacent node Y contains a point cloud (occupied), and 0 otherwise. FlagZ is the occupancy information of the adjacent node Z, and indicates 1 if the adjacent node Z contains a point cloud (occupied), and 0 otherwise.
[0595] At this time, if an adjacent node, for example, the adjacent node X0 in FIG. 90, is not accessible (reference prohibited), the three-dimensional data encoding device may use a fixed value such as 1 (occupied) or 0 (unoccupied) as an alternative value.
[0596] FIG. 91 is a diagram showing an example of a target node and adjacent nodes. As shown in FIG. 91, when an adjacent node is not accessible (reference prohibited), the occupancy information of the adjacent node may be calculated by referring to the occupancy code of the grandfather node of the target node. For example, instead of the adjacent node X0 shown in FIG. 91, the three-dimensional data encoding device may calculate FlagX0 in the above formula using the occupancy information of the adjacent node G0, and determine the value in the encoding table using the calculated FlagX0. Note that the adjacent node G0 shown in FIG. 91 is an adjacent node whose occupancy can be determined by the occupancy code of the grandfather node. The adjacent node X1 is an adjacent node whose occupancy can be determined by the occupancy code of the parent node.
[0597] Hereinafter, Modification Example 3 of the present embodiment will be described. FIGS. 92 and 93 are diagrams showing the reference relationship according to this modification example. FIG. 92 is a diagram showing the reference relationship on an octree structure, and FIG. 93 is a diagram showing the reference relationship in a spatial region.
[0598] In this modified example, when the three-dimensional data encoding device encodes the encoding information of a node to be encoded (hereinafter referred to as the target node 2), it refers to the encoding information of each node in the parent node to which the target node 2 belongs. That is, the three-dimensional data encoding device permits the reference of information (for example, occupancy information) of the child nodes of the first node where the target node and the parent node are the same among a plurality of adjacent nodes. For example, when the three-dimensional data encoding device encodes the occupancy code of the target node 2 shown in FIG. 92, it refers to the nodes existing in the parent node to which the target node 2 belongs, for example, the occupancy code of the target node shown in FIG. 92. As shown in FIG. 93, the occupancy code of the target node shown in FIG. 92 represents, for example, whether each node in the target nodes adjacent to the target node 2 is occupied or not. Therefore, since the three-dimensional data encoding device can switch the encoding table of the occupancy code of the target node 2 according to the finer shape of the target node, the encoding efficiency can be improved.
[0599] When entropy-encoding the occupancy code of the target node 2, the three-dimensional data encoding device may calculate the encoding table by, for example, the following formula.
[0600] CodingTable=(FlagX1<<5)+(FlagX2<<4)+(FlagX3<<3)+(FlagX4<<2)+(FlagY<<1)+(FlagZ)
[0601] Here, CodingTable indicates the encoding table for the occupancy code of the target node 2 and indicates any value from 0 to 63. FlagXN is the occupancy information of the adjacent node XN (N = 1..4), and indicates 1 if the adjacent node XN includes a point cloud (occupied), and 0 otherwise. FlagY is the occupancy information of the adjacent node Y, and indicates 1 if the adjacent node Y includes a point cloud (occupied), and 0 otherwise. FlagZ is the occupancy information of the adjacent node Y, and indicates 1 if the adjacent node Z includes a point cloud (occupied), and 0 otherwise.
[0602] Note that the three-dimensional data encoding device may change the calculation method of the encoding table according to the node position of the target node 2 in the parent node.
[0603] Also, when the reference to the parent adjacent node is not prohibited, the three-dimensional data encoding device may refer to the encoding information of each node in the parent adjacent node. For example, when the reference to the parent adjacent node is not prohibited, the reference to the information (for example, occupancy information) of the child node of the third node different from the target node and the parent node is permitted. For example, in the example shown in FIG. 91, the three-dimensional data encoding device refers to the occupancy code of the adjacent node X0 different from the target node and the parent node, and acquires the occupancy information of the child node of the adjacent node X0. The three-dimensional data encoding device switches the encoding table used for the entropy encoding of the occupancy code of the target node based on the acquired occupancy information of the child node of the adjacent node X0.
[0604] As described above, the three-dimensional data encoding device according to the present embodiment encodes the information (for example, occupancy code) of the target node included in the N-ary tree structure (where N is an integer of 2 or more) of a plurality of three-dimensional points included in the three-dimensional data. As shown in FIGS. 85 and 86, in the above encoding, the three-dimensional data encoding device permits the reference to the information (for example, occupancy information) of the first node in which the target node and the parent node are the same among the plurality of adjacent nodes spatially adjacent to the target node, and prohibits the reference to the information (for example, occupancy information) of the second node in which the target node and the parent node are different. In other words, in the above encoding, the three-dimensional data encoding device permits the reference to the information (for example, occupancy code) of the parent node, and prohibits the reference to the information (for example, occupancy code) of other nodes (parent adjacent nodes) in the same layer as the parent node.
[0605] According to this, the three-dimensional data encoding device can improve the encoding efficiency by referring to the information of the first node where the target node and the parent node are the same among a plurality of adjacent nodes that are spatially adjacent to the target node. Further, the three-dimensional data encoding device can reduce the processing amount by not referring to the information of the second node where the target node and the parent node are different among the plurality of adjacent nodes. Thus, the three-dimensional data encoding device can improve the encoding efficiency and reduce the processing amount.
[0606] For example, the three-dimensional data encoding device further determines whether to prohibit referring to the information of the second node, and in the above encoding, based on the result of the determination, switches whether to prohibit or permit referring to the information of the second node. The three-dimensional data encoding device further generates a bit stream including prohibition switching information (for example, limit_refer_flag shown in FIG. 89) indicating whether to prohibit referring to the information of the second node, which is the result of the above determination.
[0607] According to this, the three-dimensional data encoding device can switch whether to prohibit referring to the information of the second node. Further, the three-dimensional data decoding device can appropriately perform decoding processing using the prohibition switching information.
[0608] For example, the information of the target node is information (for example, occupancy code) indicating whether a three-dimensional point exists in each of the child nodes belonging to the target node, the information of the first node is information (occupancy information of the first node) indicating whether a three-dimensional point exists in the first node, and the information of the second node is information (occupancy information of the second node) indicating whether a three-dimensional point exists in the second node.
[0609] For example, in the above encoding, the three-dimensional data encoding device selects an encoding table based on whether a three-dimensional point exists in the first node, and uses the selected encoding table to perform entropy encoding on the information of the target node (for example, occupancy code).
[0610] For example, in the above encoding, the three-dimensional data encoding device permits reference to information (e.g., occupancy information) of child nodes of a first node among a plurality of adjacent nodes, as shown in FIGS. 92 and 93.
[0611] According to this, the three-dimensional data encoding device can refer to more detailed information of adjacent nodes, so the encoding efficiency can be improved.
[0612] For example, in the above encoding, the three-dimensional data encoding device switches the adjacent node to be referred to among a plurality of adjacent nodes according to the spatial position within the parent node of the target node.
[0613] According to this, the three-dimensional data encoding device can refer to an appropriate adjacent node according to the spatial position within the parent node of the target node.
[0614] For example, the three-dimensional data encoding device includes a processor and a memory, and the processor performs the above processing using the memory.
[0615] Further, the three-dimensional data decoding device according to the present embodiment decodes information (e.g., occupancy code) of a target node included in an N-ary tree structure (where N is an integer of 2 or more) of a plurality of three-dimensional points included in the three-dimensional data. As shown in FIGS. 85 and 86, in the above decoding, the three-dimensional data decoding device permits reference to information (e.g., occupancy information) of a first node in which the target node and the parent node are the same among a plurality of adjacent nodes that are spatially adjacent to the target node, and prohibits reference to information (e.g., occupancy information) of a second node in which the target node and the parent node are different. In other words, in the above decoding, the three-dimensional data decoding device permits reference to information (e.g., occupancy code) of the parent node and prohibits reference to information (e.g., occupancy code) of other nodes (parent adjacent nodes) in the same layer as the parent node.
[0616] According to this, the three-dimensional data decoding device can improve the encoding efficiency by referring to the information of the first node where the target node and the parent node are the same among a plurality of adjacent nodes that are spatially adjacent to the target node. Further, the three-dimensional data decoding device can reduce the processing amount by not referring to the information of the second node where the target node and the parent node are different among the plurality of adjacent nodes. In this way, the three-dimensional data decoding device can improve the encoding efficiency and reduce the processing amount.
[0617] For example, the three-dimensional data decoding device further obtains prohibition switching information (for example, limit_refer_flag shown in FIG. 89) indicating whether to prohibit referring to the information of the second node from the bit stream, and in the above decoding, based on the prohibition switching information, switches whether to prohibit or permit referring to the information of the second node.
[0618] According to this, the three-dimensional data decoding device can appropriately perform the decoding process using the prohibition switching information.
[0619] For example, the information of the target node is information (for example, occupancy code) indicating whether there is a three-dimensional point in each of the child nodes belonging to the target node, the information of the first node is information (occupancy information of the first node) indicating whether there is a three-dimensional point in the first node, and the information of the second node is information (occupancy information of the second node) indicating whether there is a three-dimensional point in the second node.
[0620] For example, in the above decoding, the three-dimensional data decoding device selects an encoding table based on whether there is a three-dimensional point in the first node, and uses the selected encoding table to perform entropy decoding on the information (for example, occupancy code) of the target node.
[0621] For example, in the above decoding, as shown in FIGS. 92 and 93, the three-dimensional data decoding device permits referring to the information (for example, occupancy information) of the child nodes of the first node among the plurality of adjacent nodes.
[0622] According to this, since the three-dimensional data decoding device can refer to more detailed information of adjacent nodes, the encoding efficiency can be improved.
[0623] For example, in the above decoding, the three-dimensional data decoding device switches the adjacent nodes to be referred to among a plurality of adjacent nodes according to the spatial position in the parent node of the target node.
[0624] According to this, the three-dimensional data decoding device can refer to appropriate adjacent nodes according to the spatial position in the parent node of the target node.
[0625] For example, the three-dimensional data decoding device includes a processor and a memory, and the processor performs the above processing using the memory.
[0626] (Embodiment 11) In the present embodiment, the three-dimensional data encoding device separates the input three-dimensional point cloud into two or more sub-three-dimensional point clouds, and encodes each sub-three-dimensional point cloud so that no dependency occurs among the plurality of sub-three-dimensional point clouds. Thereby, the three-dimensional data encoding device can encode a plurality of sub-three-dimensional point clouds in parallel. For example, the three-dimensional data encoding device separates the input three-dimensional point cloud into a sub-three-dimensional point cloud A and a sub-three-dimensional point cloud B, and encodes the sub-three-dimensional point cloud A and the sub-three-dimensional point cloud B in parallel.
[0627] Note that as a separation method, when the three-dimensional data encoding device performs encoding using, for example, an octree structure, it encodes the eight child nodes divided into the octree in parallel. For example, the three-dimensional data encoding device encodes a plurality of tree structures with each child node as a root in parallel.
[0628] Note that the three-dimensional data encoding device does not necessarily have to encode a plurality of sub-three-dimensional point clouds in parallel, and may encode the plurality of sub-three-dimensional point clouds sequentially so that no dependency occurs. Also, the method of this embodiment may be applied not only to octrees but also to N-ary trees (N is an integer of 2 or more) such as quaternary trees or 16-ary trees. Further, the three-dimensional data encoding device may perform division using attribute information such as the color, reflectance, or normal vector of the point cloud. Also, as described with reference to FIGS. 67 to 68 of Embodiment 9, the three-dimensional data encoding device may perform division based on the difference in the density of the point cloud.
[0629] Further, the three-dimensional data encoding device may combine the encoded data of a plurality of sub-three-dimensional point clouds into one bit stream. At this time, the three-dimensional data encoding device may include the start position of each encoded data of each sub-three-dimensional point cloud in the header of the bit stream or the like. For example, the three-dimensional data encoding device may include the address (bit position, number of bytes, etc.) from the start of the bit stream in the header or the like. Thereby, the three-dimensional data decoding device can know the start position of the encoded data of each sub-three-dimensional point cloud by decoding the start of the bit stream. Also, since the three-dimensional data decoding device can decode the encoded data of a plurality of sub-three-dimensional point clouds in parallel, the processing time can be reduced.
[0630] Note that the three-dimensional data encoding device may add a flag indicating that the plurality of sub-three-dimensional point clouds have been encoded so that no dependency occurs among the plurality of sub-three-dimensional point clouds, or indicating that the plurality of sub-three-dimensional point clouds have been encoded in parallel, to the header of the bit stream. Thereby, the three-dimensional data decoding device can determine whether it is possible to decode the encoded data of the plurality of three-dimensional point clouds in parallel by decoding the header.
[0631] Here, the fact that no dependency occurs among a plurality of sub-three-dimensional point groups means, for example, having an encoding table (such as a probability table used for entropy encoding) for encoding occupancy codes or leaf information of a plurality of nodes of a plurality of sub-three-dimensional point groups independently for each sub-three-dimensional point group. For example, in order to encode sub-three-dimensional point group A and sub-three-dimensional point group B without generating a dependency, a three-dimensional data encoding device uses different encoding tables for sub-three-dimensional point group A and sub-three-dimensional point group B. Or, when the three-dimensional data encoding device processes sub-three-dimensional point group A and sub-three-dimensional point group B sequentially, the encoding table is initialized after encoding sub-three-dimensional point group A and before encoding sub-three-dimensional point group B so that no dependency occurs between sub-three-dimensional point group A and sub-three-dimensional point group B. In this way, the three-dimensional data encoding device can encode a plurality of sub-three-dimensional point groups so that no dependency occurs among the plurality of sub-three-dimensional point groups by having an encoding table for each sub-three-dimensional point group independently or by initializing the encoding table before encoding. Similarly, the three-dimensional data decoding device can also appropriately decode each sub-three-dimensional point group by having an encoding table (decoding table) for each sub-three-dimensional point group independently or by initializing the encoding table before decoding each sub-three-dimensional point group.
[0632] Also, the fact that no dependency occurs among a plurality of sub-three-dimensional point groups may mean, for example, prohibiting reference between sub-three-dimensional point groups when encoding occupancy codes or leaf information of a plurality of nodes of a plurality of sub-three-dimensional point groups. For example, when encoding the occupancy code of a target node to be encoded, a three-dimensional data encoding device performs encoding using information of adjacent nodes in an octree. In this case, when an adjacent node is included in another sub-three-dimensional point group, the three-dimensional data encoding device encodes the target node without referring to the adjacent node. In this case, the three-dimensional data encoding device may perform encoding assuming that the adjacent node does not exist, or may encode the target node under the condition that the adjacent node exists but the adjacent node is included in another sub-three-dimensional point group.
[0633] Similarly, when decrypting the occupancy codes or leaf information of multiple nodes of a plurality of sub-three-dimensional point clouds, for example, the three-dimensional data decrypting device prohibits references between sub-three-dimensional point clouds. For example, when decrypting the occupancy code of a target node to be decrypted, the three-dimensional data decrypting device performs decryption using the information of adjacent nodes in the octree. In this case, when an adjacent node is included in another sub-three-dimensional point cloud, the three-dimensional data decrypting device decrypts the target node without referring to the adjacent node. In this case, the three-dimensional data decrypting device may perform decryption assuming that the adjacent node does not exist, or may decrypt the target node under the condition that the adjacent node exists but the adjacent node is included in another sub-three-dimensional point cloud.
[0634] In addition, when encoding the three-dimensional position information and attribute information (such as color, reflectivity, or normal vector) of a plurality of sub-three-dimensional point clouds, the three-dimensional data encoding device may perform encoding so that no dependency occurs for one of them, and may perform encoding so that there is a dependency for the other. For example, the three-dimensional data encoding device may encode the three-dimensional position information so that no dependency occurs, and may encode the attribute information so that there is a dependency. Thereby, the three-dimensional data encoding device can reduce the processing time by encoding the three-dimensional position information in parallel, and can reduce the encoding amount by encoding the attribute information sequentially. Note that the three-dimensional data encoding device may add both information indicating whether the three-dimensional position information is encoded without dependency and information indicating whether the attribute information is encoded without dependency to the header. Thereby, the three-dimensional data decrypting device can determine, by decrypting the header, whether the three-dimensional position information can be decrypted without dependency and whether the attribute information can be decrypted without dependency. Thereby, the three-dimensional data decrypting device can perform decryption in parallel when there is no dependency. For example, when the three-dimensional position information is encoded so that no dependency occurs and the attribute information is encoded so that there is a dependency, the three-dimensional data decrypting device reduces the processing time by decrypting the three-dimensional position information in parallel and decrypts the attribute information sequentially.
[0635] FIG. 94 is a diagram showing an example of a tree structure. In FIG. 94, an example of a quadtree is shown, but other tree structures such as an octree may be used. The three-dimensional data encoding device divides the tree structure shown in FIG. 94 into, for example, a sub-three-dimensional point group A shown in FIG. 95 and a sub-three-dimensional group B shown in FIG. 96. In this example, the division is performed at the valid nodes of layer 1. That is, in the case of a quadtree, up to 4 sub-three-dimensional point groups are generated, and in the case of an octree, up to 8 sub-three-dimensional point groups are generated. Further, the three-dimensional data encoding device may perform the division using attribute information or information such as point group density.
[0636] The three-dimensional data encoding device performs encoding so that no dependency occurs between the sub-three-dimensional point group A and the sub-three-dimensional point group B. For example, the three-dimensional data encoding device switches the encoding table used for entropy encoding of the occupancy code for each sub-three-dimensional point group. Alternatively, the three-dimensional data encoding device initializes the encoding table before encoding each sub-three-dimensional point group. Alternatively, when calculating the adjacent information of the nodes, the three-dimensional data encoding device prohibits the reference to the adjacent node when the adjacent node is included in a different sub-three-dimensional point group.
[0637] FIG. 97 is a diagram showing a configuration example of a bit stream according to the present embodiment. As shown in FIG. 97, the bit stream includes a header, encoded data of the sub-three-dimensional point group A, and encoded data of the sub-three-dimensional point group B. The header includes point group number information, dependency information, start address information A, and start address information B.
[0638] The point group number information indicates the number of sub-three-dimensional point groups included in the bit stream. Note that the number may be indicated by an occupancy code as the point group number information. For example, in the example shown in FIG. 94, the occupancy code "1010" of layer 0 is used, and the number of sub-three-dimensional point groups is indicated by the number of "1"s included in the occupancy code.
[0639] The dependency relationship information indicates whether the sub-three-dimensional point cloud is encoded without dependency. For example, the three-dimensional data decoding device determines whether to decode the sub-three-dimensional point cloud in parallel based on this dependency relationship information.
[0640] The start address information A indicates the start address of the encoded data of the sub-three-dimensional point cloud A. The start address information B indicates the start address of the encoded data of the sub-three-dimensional point cloud B.
[0641] Hereinafter, the effect of parallel encoding will be described. In the octree data of the three-dimensional point cloud (point cloud), by dividing the geometric information (three-dimensional position information) or attribute information and performing parallel encoding, the processing time can be reduced. Parallel encoding can be realized when the node is independent of other nodes in the hierarchy of the parent node. That is, it is not necessary to refer to adjacent parent nodes. This condition must also be satisfied for all of the child nodes and grandchild nodes.
[0642] FIG. 98 is a diagram showing an example of a tree structure. In the example shown in FIG. 98, when depth-first encoding is used, node A is independent of node C from layer 1. Also, node C is independent of node D from layer 2. Node A is independent of node B from layer 3.
[0643] The three-dimensional data encoding device selects a parallel encoding method to be used from two types of parallel encoding methods based on the independent information of each node, based on the type of hardware, user settings, algorithm, or adaptability of the data, etc.
[0644] These two methods are full parallel encoding and incremental parallel encoding.
[0645] First, full parallel encoding will be described. In parallel processing or parallel programming, since it is necessary to process a large amount of data simultaneously, the processing is very heavy.
[0646] The number of parallel - processable nodes is determined using the number of processing units (PUs) included in a GPU (Graphics Processing Unit), the number of cores included in a CPU, or the number of threads in a software implementation.
[0647] Here, generally, the number of nodes included in an octree is larger than the number of available PUs. The three - dimensional data encoding device determines whether the number of nodes included in a layer is the optimal number corresponding to the number of available PUs using information indicating the number of encoded nodes included in the layer, and starts full - parallel encoding triggered by the fact that the number of nodes included in the layer has reached the optimal number. In parallel processing, breadth - first or depth - first processing can be used.
[0648] The three - dimensional data encoding device may store information indicating the nodes (layers) for which parallel encoding processing has been started in the header of the bitstream. Thereby, the three - dimensional data decoding device can perform parallel decoding processing using this information if necessary. Note that the format of the information indicating the nodes for which parallel encoding processing has been started may be arbitrary. For example, a location code may be used.
[0649] Also, the three - dimensional data encoding device prepares an encoding table (probability table) for each node (sub - three - dimensional point group) that performs parallel encoding. This encoding table is initialized with an initial value or values different for each node. For example, the values different for each node are values based on the occupancy code of the parent node. This full - parallel encoding has the advantage that the initialization of the GPU needs to be done only once.
[0650] FIG. 99 is a diagram for explaining full - parallel encoding and shows an example of a tree structure. FIG. 100 is a diagram spatially showing the sub - three - dimensional point groups to be processed in parallel. The three - dimensional data encoding device starts parallel processing triggered by the fact that the number of nodes correlated with the number of PUs or threads has reached the optimal point.
[0651] In the example shown in FIG. 99, in layer 3, the number of occupied nodes included in the layer is 9, which exceeds the optimal number. Therefore, the three-dimensional data encoding device divides the three-dimensional points (nodes) below layer 3 into a plurality of sub-three-dimensional point groups with each occupied node in layer 3 as the root, and processes each sub-three-dimensional point group in parallel. For example, in the example shown in FIG. 99, nine sub-three-dimensional point groups are generated.
[0652] The three-dimensional data encoding device may encode layer information indicating the layer at which parallel processing starts. Further, the three-dimensional data encoding device may encode information indicating the number of occupied nodes at the start of parallel processing (9 in the example of FIG. 99).
[0653] Further, the three-dimensional data encoding device encodes, for example, a plurality of sub-three-dimensional groups while prohibiting mutual reference. Further, the three-dimensional data encoding device initializes, for example, an encoding table (probability table, etc.) used for entropy encoding before encoding each sub-three-dimensional point group.
[0654] FIG. 101 is a diagram showing a configuration example of a bit stream according to the present embodiment. As shown in FIG. 101, the bit stream includes a header, upper layer encoded data, a sub-header, encoded data of sub-three-dimensional point group A, and encoded data of sub-three-dimensional point group B.
[0655] The header includes spatial maximum size information and parallel start layer information. The spatial size information indicates the first three-dimensional space that divides the three-dimensional point group into an octree. For example, the spatial size information indicates the maximum coordinates (x, y, z) of the first three-dimensional space.
[0656] The parallel start layer information indicates the parallel start layer, which is a layer at which parallel processing can start. Here, the parallel start layer information indicates, for example, layer N.
[0657] The upper layer encoded data is encoded data up to layer N before starting parallel processing, and is node information up to layer N. For example, the upper layer encoded data includes occupancy codes of nodes up to layer N.
[0658] The sub-header contains information necessary for decoding subsequent to layer N. For example, the sub-header indicates the starting address of the encoded data of each sub-three-dimensional point group, etc. In the example shown in FIG. 101, the sub-header includes starting address information A and starting address information B. The starting address information A indicates the starting address of the encoded data of the sub-three-dimensional point group A. The starting address information B indicates the starting address of the encoded data of the sub-three-dimensional point group B.
[0659] Note that the three-dimensional data encoding device may store the starting address information A and the starting address information B in the header. Thereby, the three-dimensional data decoding device can decode the encoded data of the sub-three-dimensional point groups in parallel prior to the upper-layer encoded data. In this case, the sub-header may include information indicating the space of each sub-three-dimensional point group. This information indicates the maximum coordinates (x, y, z) of the space of each sub-three-dimensional point group.
[0660] FIG. 102 is a diagram for explaining the parallel decoding process. As shown in FIG. 102, the three-dimensional data decoding device decodes the encoded data of the sub-three-dimensional point group A and the encoded data of the sub-three-dimensional point group B in parallel, and generates the decoded data of the sub-three-dimensional point A and the decoded data of the sub-three-dimensional point group B...
Claims
1. A bit stream is generated that includes first control information common to a plurality of subspaces included in a target space of three-dimensional data and a plurality of encoded data each corresponding to the plurality of subspaces, the first control information includes information of the plurality of subspaces associated with a plurality of identifiers respectively assigned to the plurality of subspaces; The header of each of the plurality of encoded data includes information identifying the corresponding subspace. A method for generating three-dimensional data.
2. The plurality of identifiers are included in the first control information. The three-dimensional data generating method according to claim 1.
3. In the bitstream, the first control information is placed before the plurality of encoded data.
3. The three-dimensional data generating method according to claim 1 or 2.
4. The first control information includes position information of each of the plurality of subspaces.
3. The three-dimensional data generating method according to claim 1 or 2.
5. The first control information includes size information of each of the plurality of subspaces.
3. The three-dimensional data generating method according to claim 1 or 2.
6. Obtaining a bit stream including first control information common to a plurality of subspaces included in a target space of three-dimensional data and a plurality of encoded data each corresponding to the plurality of subspaces; the first control information includes information of the plurality of subspaces associated with each of a plurality of identifiers assigned to the plurality of subspaces, A header of each of the plurality of encoded data includes information identifying a corresponding subspace. Three-dimensional data acquisition methods.
7. The plurality of identifiers are included in the first control information. The three-dimensional data acquisition method according to claim 6.
8. In the bitstream, the first control information is placed before the plurality of encoded data. The three-dimensional data acquisition method according to claim 6 or 7.
9. The first control information includes position information of each of the plurality of subspaces. The three-dimensional data acquisition method according to claim 6 or 7.
10. The first control information includes size information of each of the plurality of subspaces. The three-dimensional data acquisition method according to claim 6 or 7.
11. A processor; A memory. The processor uses the memory to: generating a bit stream including first control information common to a plurality of subspaces included in a target space of the three-dimensional data and a plurality of pieces of encoded data each corresponding to the plurality of subspaces; the first control information includes information of the plurality of subspaces associated with a plurality of identifiers respectively assigned to the plurality of subspaces; A header of each of the plurality of encoded data includes information identifying a corresponding subspace. Three-dimensional data generation device.
12. A processor; A memory. The processor uses the memory to: obtaining a bit stream including first control information common to a plurality of subspaces included in a target space of the three-dimensional data and a plurality of pieces of encoded data each corresponding to the plurality of subspaces; the first control information includes information of the plurality of subspaces associated with each of a plurality of identifiers assigned to the plurality of subspaces, A header of each of the plurality of encoded data includes information identifying a corresponding subspace. Three-dimensional data acquisition device.
Citation Information
Patent Citations
Progressive three-dimensional mesh information coding / decoding method, and apparatus therefor
JP2006187015A
Adaptive entropy coding method for tree structures
JP2014527735A
Hierarchical entropy coding and decoding
JP2014529950A
Location coding based on spatial trees with overlapping points
JP2015505389A
Map display device
WO2014020663A1