Three-dimensional data encoding method, three-dimensional data decoding method, three-dimensional data encoding device, and three-dimensional data decoding device

By encoding and decoding the N forktree structure information using the same encoding style as the octree structure information, the problem of low encoding efficiency of three-dimensional data in the prior art is solved, and more efficient data compression and processing is achieved.

CN112805751BActive Publication Date: 2025-05-20PANASONIC INTELLECTUAL PROPERTY CORP OF AMERICA
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN201980066145.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2018-10-12
Filing Date
2019-10-11
Publication Date
2025-05-20
Estimated Expiration
2039-10-11

AI Technical Summary

Technical Problem

In the prior art, the encoding efficiency of three-dimensional data is low, and it is difficult to effectively compress large-scale point cloud data.

Method used

The method of encoding and decoding the information of the N forktree structure and the octree structure is adopted, and the N forktree structure information is encoded and decoded using the same encoding style as the octree structure information to reduce the processing load.

Benefits of technology

It improves the coding efficiency of three-dimensional data, reduces processing load, and achieves more efficient data compression.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN112805751B_ABST
    Figure CN112805751B_ABST
Patent Text Reader

Abstract

The three-dimensional data encoding method encodes (S5903, S5904) the first information of the first object node contained in the N (N is 2 or 4) octree structure of multiple first three-dimensional points of the first three-dimensional point group, or the second information of the second object node contained in the octree structure of multiple second three-dimensional points of the second three-dimensional point group. In the encoding, the first information is encoded using a first encoding style that is the same as the second encoding style used in the encoding of the second information.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to a three-dimensional data encoding method, a three-dimensional data decoding method, a three-dimensional data encoding apparatus, and a three-dimensional data decoding apparatus. Background Art

[0002] In large fields such as computer vision, map information, monitoring, infrastructure inspection, or video distribution for autonomous operation of automobiles or robots, devices or services that make flexible use of three-dimensional data will become popular in the future. Three-dimensional data is obtained by various methods such as distance sensors such as rangefinders, stereo cameras, or combinations of multiple monocular cameras.

[0003] As a representation method of three-dimensional data, there is a representation method called point cloud, which represents the shape of a three-dimensional structure by a point group in a three-dimensional space (for example, refer to Non-Patent Document 1). In the point cloud, the positions and colors of the point group are stored. Although it is expected that the point cloud will become the mainstream as a representation method of three-dimensional data, the data volume of the point group is very large. Therefore, in the storage or transmission of three-dimensional data, like two-dimensional moving images (as an example, MPEG-4 AVC or HEVC standardized as MPEG), data volume compression needs to be performed by encoding.

[0004] Furthermore, for the compression of point clouds, part of it is supported by publicly available libraries (PointCloud Library) that perform point cloud correlation processing.

[0005] Furthermore, there is a well-known technology that uses three-dimensional map data to retrieve and display facilities around a vehicle (for example, refer to Patent Document 1).

[0006] Prior Art Documents

[0007] Patent Documents

[0008] Patent Document 1 International Publication No. 2014 / 020663 Summary of the Invention

[0009] Problems to be Solved by the Invention

[0010] It is desired to improve the encoding efficiency in the encoding process of three-dimensional data.

[0011] An object of the present disclosure is to provide a three-dimensional data encoding method, a three-dimensional data decoding method, a three-dimensional data encoding apparatus, or a three-dimensional data decoding apparatus that can improve the encoding efficiency.

[0012] Means for Solving the Problems

[0013] A three-dimensional data encoding method according to one aspect of the present disclosure encodes first information of a first object node included in an N-ary tree structure (where N is 2 or 4) of a plurality of first three-dimensional points in a first three-dimensional point group, or second information of a second object node included in an octree structure of a plurality of second three-dimensional points in a second three-dimensional point group. In the encoding, the first information is encoded using a first encoding style common to a second encoding style used in the encoding of the second information.

[0014] A three-dimensional data decoding method according to one aspect of the present disclosure decodes first information of a first object node included in an N-ary tree structure (where N is 2 or 4) of a plurality of first three-dimensional points in a first three-dimensional point group, or second information of a second object node included in an octree structure of a plurality of second three-dimensional points in a second three-dimensional point group. In the decoding, the first information is decoded using a first decoding style common to a second decoding style used in the decoding of the second information.

[0015] Advantageous Effects of the Invention

[0016] The present disclosure can provide a three-dimensional data encoding method, a three-dimensional data decoding method, a three-dimensional data encoding device, or a three-dimensional data decoding device that can improve encoding efficiency. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] Figure 1 FIG. is a diagram showing the configuration of encoded three-dimensional data according to Embodiment 1.

[0018] Figure 2 FIG. is a diagram showing an example of a prediction structure between SPCs belonging to the lowest layer of the GOS according to Embodiment 1.

[0019] Figure 3 FIG. is a diagram showing an example of an inter-layer prediction structure according to Embodiment 1.

[0020] Figure 4 FIG. is a diagram showing an example of the encoding order of the GOS according to Embodiment 1.

[0021] Figure 5 FIG. is a diagram showing an example of the encoding order of the GOS according to Embodiment 1.

[0022] Figure 6 FIG. is a block diagram of a three-dimensional data encoding device according to Embodiment 1.

[0023] Figure 7 FIG. is a flowchart of the encoding process according to Embodiment 1.

[0024] Figure 8 FIG. is a block diagram of a three-dimensional data decoding device according to Embodiment 1.

[0025] Figure 9 is a flowchart of the decoding process according to Embodiment 1.

[0026] Figure 10 is a diagram showing an example of the meta information according to Embodiment 1.

[0027] Figure 11 is a diagram showing a configuration example of the SWLD according to Embodiment 2.

[0028] Figure 12 is a diagram showing an operation example of the server and the client according to Embodiment 2.

[0029] Figure 13 is a diagram showing an operation example of the server and the client according to Embodiment 2.

[0030] Figure 14 is a diagram showing an operation example of the server and the client according to Embodiment 2.

[0031] Figure 15 is a diagram showing an operation example of the server and the client according to Embodiment 2.

[0032] Figure 16 is a block diagram of the three-dimensional data encoding device according to Embodiment 2.

[0033] Figure 17 is a flowchart of the encoding process according to Embodiment 2.

[0034] Figure 18 is a block diagram of the three-dimensional data decoding device according to Embodiment 2.

[0035] Figure 19 is a flowchart of the decoding process according to Embodiment 2.

[0036] Figure 20 is a diagram showing a configuration example of the WLD according to Embodiment 2.

[0037] Figure 21 is a diagram showing an example of the octree structure of the WLD according to Embodiment 2.

[0038] Figure 22 is a diagram showing a configuration example of the SWLD according to Embodiment 2.

[0039] Figure 23 is a diagram showing an example of the octree structure of the SWLD according to Embodiment 2.

[0040] Figure 24 is a schematic diagram showing the state of transmission and reception of three-dimensional data between vehicles according to Embodiment 3.

[0041] Figure 25 This is a diagram showing an example of three-dimensional data transmitted between vehicles according to Embodiment 3.

[0042] Figure 26 This is a block diagram of a three-dimensional data creation device according to Embodiment 3.

[0043] Figure 27 This is a flowchart of a three-dimensional data creation process according to Embodiment 3.

[0044] Figure 28 This is a block diagram of a three-dimensional data transmission device according to Embodiment 3.

[0045] Figure 29 This is a flowchart of a three-dimensional data transmission process according to Embodiment 3.

[0046] Figure 30 This is a block diagram of a three-dimensional data creation device according to Embodiment 3.

[0047] Figure 31 This is a flowchart of a three-dimensional data creation process according to Embodiment 3.

[0048] Figure 32 This is a block diagram of a three-dimensional data transmission device according to Embodiment 3.

[0049] Figure 33 This is a flowchart of a three-dimensional data transmission process according to Embodiment 3.

[0050] Figure 34 This is a block diagram of a three-dimensional information processing device according to Embodiment 4.

[0051] Figure 35 This is a flowchart of a three-dimensional information processing method according to Embodiment 4.

[0052] Figure 36 This is a flowchart of a three-dimensional information processing method according to Embodiment 4.

[0053] Figure 37 This is a diagram for explaining the transmission process of three-dimensional data according to Embodiment 5.

[0054] Figure 38 This is a block diagram of a three-dimensional data creation device according to Embodiment 5.

[0055] Figure 39 This is a flowchart of a three-dimensional data creation method according to Embodiment 5.

[0056] Figure 40 This is a flowchart of a three-dimensional data creation method according to Embodiment 5.

[0057] Figure 41 It is a flowchart of the display method according to Embodiment 6.

[0058] Figure 42 It is a diagram showing an example of the surrounding environment seen through the windshield according to Embodiment 6.

[0059] Figure 43 It is a diagram showing a display example of the head-up display according to Embodiment 6.

[0060] Figure 44 It is a diagram showing a display example of the adjusted head-up display according to Embodiment 6.

[0061] Figure 45 It is a diagram showing the configuration of the system according to Embodiment 7.

[0062] Figure 46 It is a block diagram of the client device according to Embodiment 7.

[0063] Figure 47 It is a block diagram of the server according to Embodiment 7.

[0064] Figure 48 It is a flowchart of the three-dimensional data creation process performed by the client device according to Embodiment 7.

[0065] Figure 49 It is a flowchart of the sensor information transmission process performed by the client device according to Embodiment 7.

[0066] Figure 50 It is a flowchart of the three-dimensional data creation process performed by the server according to Embodiment 7.

[0067] Figure 51 It is a flowchart of the three-dimensional map transmission process performed by the server according to Embodiment 7.

[0068] Figure 52 It is a diagram showing the configuration of a modified example of the system according to Embodiment 7.

[0069] Figure 53 It is a diagram showing the configuration of the server and the client device according to Embodiment 7.

[0070] Figure 54 It is a block diagram of the three-dimensional data encoding device according to Embodiment 8.

[0071] Figure 55 It is a diagram showing an example of the prediction residual according to Embodiment 8.

[0072] Figure 56This is a diagram showing an example of the volume related to Embodiment 8.

[0073] Figure 57 This is a diagram showing an example of the octree representation of the volume related to Embodiment 8.

[0074] Figure 58 This is a diagram showing an example of the bit string of the volume related to Embodiment 8.

[0075] Figure 59 This is a diagram showing an example of the octree representation of the volume related to Embodiment 8.

[0076] Figure 60 This is a diagram showing an example of the volume related to Embodiment 8.

[0077] Figure 61 This is a diagram for explaining the intra prediction process related to Embodiment 8.

[0078] Figure 62 This is a diagram for explaining the rotation and translation processes related to Embodiment 8.

[0079] Figure 63 This is a diagram showing a syntax example of the RT application flag and RT information related to Embodiment 8.

[0080] Figure 64 This is a diagram for explaining the inter prediction process related to Embodiment 8.

[0081] Figure 65 This is a block diagram of the 3D data decoding device related to Embodiment 8.

[0082] Figure 66 This is a flowchart of the 3D data encoding process performed by the 3D data encoding device related to Embodiment 8.

[0083] Figure 67 This is a flowchart of the 3D data decoding process performed by the 3D data decoding device related to Embodiment 8.

[0084] Figure 68 This is a diagram showing an example of the tree structure related to Embodiment 9.

[0085] Figure 69 This is a diagram showing an example of the occupancy rate encoding related to Embodiment 9.

[0086] Figure 70 This is a diagram schematically showing the operation of the 3D data encoding device related to Embodiment 9.

[0087] Figure 71 This is a diagram showing an example of the geometric information related to Embodiment 9.

[0088] Figure 72 It is a diagram showing a selection example of an encoding table using geometric information related to Embodiment 9.

[0089] Figure 73 It is a diagram showing a selection example of an encoding table using structural information related to Embodiment 9.

[0090] Figure 74 It is a diagram showing a selection example of an encoding table using attribute information related to Embodiment 9.

[0091] Figure 75 It is a diagram showing a selection example of an encoding table using attribute information related to Embodiment 9.

[0092] Figure 76 It is a diagram showing a structural example of a bitstream related to Embodiment 9.

[0093] Figure 77 It is a diagram showing an example of an encoding table related to Embodiment 9.

[0094] Figure 78 It is a diagram showing an example of an encoding table related to Embodiment 9.

[0095] Figure 79 It is a diagram showing a structural example of a bitstream related to Embodiment 9.

[0096] Figure 80 It is a diagram showing an example of an encoding table related to Embodiment 9.

[0097] Figure 81 It is a diagram showing an example of an encoding table related to Embodiment 9.

[0098] Figure 82 It is a diagram showing an example of the bit number of occupancy encoding related to Embodiment 9.

[0099] Figure 83 It is a flowchart of the encoding process using geometric information related to Embodiment 9.

[0100] Figure 84 It is a flowchart of the decoding process using geometric information related to Embodiment 9.

[0101] Figure 85 It is a flowchart of the encoding process using structural information related to Embodiment 9.

[0102] Figure 86 It is a flowchart of the decoding process using structural information related to Embodiment 9.

[0103] Figure 87 It is a flowchart of the encoding process using attribute information according to Embodiment 9.

[0104] Figure 88 It is a flowchart of the decoding process using attribute information according to Embodiment 9.

[0105] Figure 89 It is a flowchart of the encoding table selection process using geometric information according to Embodiment 9.

[0106] Figure 90 It is a flowchart of the encoding table selection process using structure information according to Embodiment 9.

[0107] Figure 91 It is a flowchart of the encoding table selection process using attribute information according to Embodiment 9.

[0108] Figure 92 It is a block diagram of a three-dimensional data encoding device according to Embodiment 9.

[0109] Figure 93 It is a block diagram of a three-dimensional data decoding device according to Embodiment 9.

[0110] Figure 94 It is a diagram showing the reference relationship in the octree structure according to Embodiment 10.

[0111] Figure 95 It is a diagram showing the reference relationship in the spatial region according to Embodiment 10.

[0112] Figure 96 It is a diagram showing an example of adjacent reference nodes according to Embodiment 10.

[0113] Figure 97 It is a diagram showing the relationship between a parent node and a node according to Embodiment 10.

[0114] Figure 98 It is a diagram showing an example of occupancy rate encoding of a parent node according to Embodiment 10.

[0115] Figure 99 It is a block diagram of a three-dimensional data encoding device according to Embodiment 10.

[0116] Figure 100 It is a block diagram of a three-dimensional data decoding device according to Embodiment 10.

[0117] Figure 101 It is a flowchart of the three-dimensional data encoding process according to Embodiment 10.

[0118] Figure 102It is a flowchart showing the three-dimensional data decoding process according to Embodiment 10.

[0119] Figure 103 It is a diagram showing an example of switching of the coding table according to Embodiment 10.

[0120] Figure 104 It is a diagram showing the reference relationship in the spatial region according to Modification 1 of Embodiment 10.

[0121] Figure 105 It is a diagram showing a syntactic example of the header information according to Modification 1 of Embodiment 10.

[0122] Figure 106 It is a diagram showing a syntactic example of the header information according to Modification 1 of Embodiment 10.

[0123] Figure 107 It is a diagram showing an example of adjacent reference nodes according to Modification 2 of Embodiment 10.

[0124] Figure 108 It is a diagram showing examples of an object node and adjacent nodes according to Modification 2 of Embodiment 10.

[0125] Figure 109 It is a diagram showing the reference relationship in the octree structure according to Modification 3 of Embodiment 10.

[0126] Figure 110 It is a diagram showing the reference relationship in the spatial region according to Modification 3 of Embodiment 10.

[0127] Figure 111 It is a diagram for schematically explaining the three-dimensional data encoding method according to Embodiment 11.

[0128] Figure 112 It is a diagram for explaining a transformation method for transforming a plane detected to be inclined according to Embodiment 11 into an X-Y plane.

[0129] Figure 113 It is a diagram showing the relationship between the plane according to Embodiment 11 and the point group selected in each method.

[0130] Figure 114 It is a diagram showing the frequency distribution of the quantified distance between the plane detected from the three-dimensional point group and the point group around the plane (candidate for the plane point group) in the first method according to Embodiment 11.

[0131] Figure 115 It is a diagram showing the frequency distribution of the quantified distance between the plane detected from the three-dimensional point group and the point group around the plane in the second method according to Embodiment 11.

[0132] Figure 116 This is a diagram showing an example in which the two-dimensional space related to Embodiment 11 is divided into four sub-spaces.

[0133] Figure 117 This is a diagram showing an example in which the four sub-spaces in the two-dimensional space related to Embodiment 11 are applied to eight sub-spaces in a three-dimensional space.

[0134] Figure 118 This is a diagram showing the adjacency relationship of three-dimensional points of the first three-dimensional point group arranged on a plane related to Embodiment 11.

[0135] Figure 119 This is a diagram showing the adjacency relationship of three-dimensional points of the three-dimensional point group arranged in a three-dimensional space related to Embodiment 11.

[0136] Figure 120 This is a block diagram showing the structure of the three-dimensional data encoding device related to Embodiment 11.

[0137] Figure 121 This is a block diagram showing the detailed structure of the quadtree encoding unit using the first method related to Embodiment 11.

[0138] Figure 122 This is a block diagram showing the detailed structure of the quadtree encoding unit using the second method related to Embodiment 11.

[0139] Figure 123 This is a block diagram showing the structure of the three-dimensional data decoding device related to Embodiment 11.

[0140] Figure 124 This is a block diagram showing the detailed structure of the quadtree decoding unit using the first method related to Embodiment 11.

[0141] Figure 125 This is a block diagram showing the detailed structure of the quadtree decoding unit using the second method related to Embodiment 11.

[0142] Figure 126 This is a flowchart of the three-dimensional data encoding method related to Embodiment 11.

[0143] Figure 127 This is a flowchart of the three-dimensional data decoding method related to Embodiment 11.

[0144] Figure 128 This is a flowchart of the quadtree encoding process related to Embodiment 11.

[0145] Figure 129 This is a flowchart of the octree encoding process related to Embodiment 11.

[0146] Figure 130 It is a flowchart of the quadtree decoding process according to Embodiment 11.

[0147] Figure 131 It is a flowchart of the octree decoding process according to Embodiment 11. Specific Embodiment

[0148] A 3D data encoding method according to one aspect of the present disclosure encodes the first information of the first object nodes included in the N (N is 2 or 4) - ary tree structure of a plurality of first 3D points of the first 3D point group, or the second information of the second object nodes included in the octree structure of a plurality of second 3D points of the second 3D point group. In the encoding, the first information is encoded using a first encoding style common to the second encoding style used in the encoding of the second information.

[0149] Thus, by encoding the information of the N - ary tree structure using an encoding style common to the encoding of the information of the octree structure, this 3D data encoding method can reduce the processing load.

[0150] For example, it may be that the first encoding style is an encoding style for selecting an encoding table used in the encoding of the first information, the second encoding style is an encoding style for selecting an encoding table used in the encoding of the second information, and in the encoding, the first encoding style is generated based on the first adjacent information of a plurality of first adjacent nodes that are spatially adjacent to the first object node in a plurality of directions, and the second encoding style is generated based on the second adjacent information of a plurality of second adjacent nodes that are spatially adjacent to the second object node in the plurality of directions.

[0151] For example, it may be that in the generation of the first encoding style, the first encoding style including a 6 - bit third bit pattern is generated, the third bit pattern is composed of a first bit pattern and a second bit pattern, the first bit pattern is composed of one or more bits, the one or more bits represent one or more first adjacent nodes that are spatially adjacent to the first object node in a specified direction among the plurality of directions and indicate that they are not occupied by the point group, the second bit pattern is composed of a plurality of bits, the plurality of bits represent a plurality of second adjacent nodes that are spatially adjacent to the first object node in directions other than the specified direction among the plurality of directions. In the generation of the second encoding style, the second encoding style including a 6 - bit fourth bit pattern is generated, the fourth bit pattern is composed of a plurality of bits, and the plurality of bits represent a plurality of third adjacent nodes that are spatially adjacent to the second object node in the plurality of directions.

[0152] For example, it can also be that, in the encoding, based on the first encoding style, a first encoding table is selected, and using the selected first encoding table, entropy encoding is performed on the first information; based on the second encoding style, a second encoding table is selected, and using the selected second encoding table, entropy encoding is performed on the second information.

[0153] For example, it can also be that, in the encoding, by encoding the first information, a bitstream including a third bit string composed of 8 bits is generated, the first information represents whether each of N first subspaces obtained by N-dividing the first object node contains the first three-dimensional point, the third bit string is composed of a first bit string and an invalid second bit string, the first bit string is composed of N bits corresponding to the first information, and the second bit string is composed of (8 - N) bits.

[0154] For example, it can also be that a bitstream including identification information indicating whether the object of the encoding is the first information or the second information is further generated.

[0155] For example, it can also be that the first three-dimensional point group is a point group arranged on a plane, and the second three-dimensional point group is a point group arranged around the plane.

[0156] The three-dimensional data decoding method according to one aspect of the present disclosure decodes the first information of the first object node included in the N (N is 2 or 4)-ary tree structure of multiple first three-dimensional points of the first three-dimensional point group, or the second information of the second object node included in the octree structure of multiple second three-dimensional points of the second three-dimensional point group. In the decoding, the first information is decoded using the first decoding style common to the second decoding style used in the decoding of the second information.

[0157] Thereby, the three-dimensional data decoding method can reduce the processing load by decoding the information of the N-ary tree structure using the decoding style common to the decoding of the information of the octree structure.

[0158] For example, it can also be that the first decoding style is a decoding style for selecting a decoding table used in the decoding of the first information, the second decoding style is a decoding style for selecting a decoding table used in the decoding of the second information, and in the decoding, the first decoding style is generated according to the first adjacent information of multiple first adjacent nodes that are spatially adjacent to the first object node in multiple directions, and the second decoding style is generated according to the second adjacent information of multiple second adjacent nodes that are spatially adjacent to the second object node in the multiple directions.

[0159] For example, it may also be that in the generation of the first decoding pattern, the first decoding pattern including a third bit pattern of 6 bits is generated. The third bit pattern is composed of a first bit pattern and a second bit pattern. The first bit pattern is composed of one or more bits, and the one or more bits represent one or more first adjacent nodes that are spatially adjacent to the first object node in a specified direction among a plurality of directions and are each not occupied by a point group. The second bit pattern is composed of a plurality of bits, and the plurality of bits represent a plurality of second adjacent nodes that are spatially adjacent to the first object node in directions other than the specified direction among the plurality of directions. In the generation of the second decoding pattern, the second decoding pattern is generated. The second decoding pattern includes a fourth bit pattern of 6 bits. The fourth bit pattern is composed of a plurality of bits, and the plurality of bits represent a plurality of third adjacent nodes that are spatially adjacent to the second object node in the plurality of directions.

[0160] For example, it may also be that in the decoding, a first decoding table is selected based on the first decoding pattern, and the first information is entropy decoded using the selected first decoding table. A second decoding table is selected based on the second decoding pattern, and the second information is entropy decoded using the selected second decoding table.

[0161] For example, it may also be that in the decoding, a bit stream including a third bit string composed of 8 bits is obtained. The third bit string is composed of a first bit string of N bits and an invalid second bit string of (8 - N) bits. The first information is decoded from the first bit string of the bit stream. The first information indicates whether each of N first subspaces obtained by dividing the first object node into N parts contains the first three-dimensional point.

[0162] For example, it may also be that the bit stream includes identification information indicating whether the coding object is the first information or the second information. In the decoding, when the identification information indicates that the coding object is the first information, the first bit string in the bit stream is decoded.

[0163] For example, it may also be that the first three-dimensional point group is a point group arranged on a plane, and the second three-dimensional point group is a point group arranged around the plane.

[0164] In addition, a three-dimensional data encoding apparatus according to an aspect of the present disclosure includes a processor and a memory. The processor uses the memory to encode first information of a first object node included in an N-ary tree structure (N is 2 or 4) of a plurality of first three-dimensional points of a first three-dimensional point group, or second information of a second object node included in an octree structure of a plurality of second three-dimensional points of a second three-dimensional point group. In the encoding, the first information is encoded using a first encoding style common to a second encoding style used in the encoding of the second information.

[0165] Accordingly, this three-dimensional data encoding method can reduce the processing load by encoding the information of the N-ary tree structure using an encoding style common to the encoding of the information of the octree structure.

[0166] A three-dimensional data decoding apparatus according to an aspect of the present disclosure includes a processor and a memory. The processor uses the memory to decode first information of a first object node included in an N-ary tree structure (N is 2 or 4) of a plurality of first three-dimensional points of a first three-dimensional point group, or second information of a second object node included in an octree structure of a plurality of second three-dimensional points of a second three-dimensional point group. In the decoding, the first information is decoded using a first decoding style common to a second decoding style used in the decoding of the second information.

[0167] Accordingly, this three-dimensional data decoding method can reduce the processing load by decoding the information of the N-ary tree structure using a decoding style common to the decoding of the information of the octree structure.

[0168] In addition, these general or specific forms can be implemented by a system, a method, an integrated circuit, a computer program, or a recording medium such as a computer-readable CD-ROM, and can be implemented by any combination of a system, a method, an integrated circuit, a computer program, and a recording medium.

[0169] Hereinafter, embodiments will be specifically described with reference to the drawings. In addition, the embodiments to be described below are all specific examples showing the present disclosure. The numerical values, shapes, materials, constituent elements, arrangement positions and connection forms of the constituent elements, steps, and the order of steps shown in the following embodiments are all examples, and the gist thereof is not to limit the present disclosure. In addition, among the constituent elements of the following embodiments, the constituent elements not described in the technical solution showing the uppermost concept are described as arbitrary constituent elements.

[0170] (Embodiment 1)

[0171] First, the data structure of the encoded three-dimensional data (hereinafter also referred to as encoded data) according to the present embodiment will be described. Figure 1This is a diagram showing the composition of the three-dimensional data encoding according to the present embodiment.

[0172] In the present embodiment, the three-dimensional space is divided into spaces (SPCs) corresponding to pictures in the encoding of moving images, and the three-dimensional data is encoded in units of space. The space is further divided into volumes (VLMs) corresponding to macroblocks and the like in the moving image encoding, and prediction and transformation are performed in units of VLM. A volume includes a plurality of voxels (VXLs), which are the smallest units corresponding to position coordinates. In addition, prediction means, similar to the prediction performed in two-dimensional images, referring to other processing units, generating predicted three-dimensional data similar to the processing unit of the object to be processed, and encoding the difference between the predicted three-dimensional data and the processing unit of the object to be processed. And this prediction includes not only spatial prediction referring to other prediction units at the same time, but also temporal prediction referring to prediction units at different times.

[0173] For example, when a three-dimensional data encoding device (hereinafter also referred to as an encoding device) encodes a three-dimensional space represented by point cloud data such as a point cloud, it encodes each point of the point cloud or a plurality of points included in a voxel together according to the size of the voxel. If the voxel is subdivided, the three-dimensional shape of the point cloud can be represented with high accuracy, and if the size of the voxel is increased, the three-dimensional shape of the point cloud can be represented roughly.

[0174] In addition, although the case where the three-dimensional data is a point cloud is described below as an example, the three-dimensional data is not limited to the point cloud and can be three-dimensional data in any form.

[0175] Moreover, hierarchical voxels can be used. In this case, in the n-th hierarchy, it is possible to sequentially show whether there are sampling points in the hierarchies below the (n - 1)-th hierarchy (the lower layer of the n-th hierarchy). For example, when only decoding the n-th hierarchy, when there are sampling points in the hierarchies below the (n - 1)-th hierarchy, it can be regarded that there are sampling points at the center of the voxels in the n-th hierarchy for decoding.

[0176] And the encoding device obtains point cloud data through a distance sensor, a stereo camera, a monocular camera, a gyroscope, or an inertial sensor, etc.

[0177] Regarding the space, similar to the encoding of moving images, it is at least classified into any one of the following three prediction structures: an intra-frame space (I-SPC) that can be decoded independently, a prediction space (P-SPC) that can only be referred to unidirectionally, and a bidirectional space (B-SPC) that can be referred to bidirectionally. And the space has two types of time information: a decoding time and a display time.

[0178] And, as Figure 1As shown, as a processing unit including multiple spaces, there is a GOS (Group Of Space) which is a random access unit. Moreover, as a processing unit including multiple GOSs, there is a World Space (WLD).

[0179] The space area occupied by the world space is associated with the absolute position on the earth through GPS or latitude and longitude information, etc. This position information is stored as meta information. In addition, the meta information can be included in the encoded data or transmitted separately from the encoded data.

[0180] Also, within a GOS, all SPCs can be three-dimensionally adjacent, or there can be SPCs that are not three-dimensionally adjacent to other SPCs.

[0181] In addition, hereinafter, processes such as encoding, decoding, or referencing corresponding to the three-dimensional data included in processing units such as GOS, SPC, or VLM will also be simply referred to as encoding, decoding, or referencing the processing unit. And the three-dimensional data included in the processing unit includes at least one group of spatial positions such as three-dimensional coordinates and characteristic values such as color information.

[0182] Next, the prediction structure of SPCs in a GOS will be described. Multiple SPCs within the same GOS, or multiple VLMs within the same SPC, although they occupy different spaces from each other, hold the same time information (decoding time and display time).

[0183] Also, within a GOS, the SPC that is the first in the decoding order is an I-SPC. And there are two types of GOSs in a GOS: a closed GOS and an open GOS. A closed GOS is a GOS that can decode all SPCs within the GOS when starting to decode from the first I-SPC. In an open GOS, within the GOS, a part of the SPCs whose display time is earlier than that of the first I-SPC refer to different GOSs and can only be decoded in that GOS.

[0184] In addition, in the encoded data such as map information, there is a case of decoding the WLD in the direction opposite to the encoding order. If there is a dependency between GOSs, it is difficult to perform reverse regeneration. Therefore, in this case, a closed GOS is basically adopted.

[0185] Also, a GOS has a layer structure in the height direction, and encoding or decoding is performed sequentially starting from the SPCs in the bottom layer.

[0186] Figure 2 is a diagram showing an example of the prediction structure between SPCs in the layer that represents the bottom layer belonging to a GOS. Figure 3 is a diagram showing an example of the inter-layer prediction structure.

[0187] There is more than one I-SPC in the GOS. Although there are objects such as people, animals, cars, bicycles, traffic lights, or buildings that serve as land marks in the three-dimensional space, it is particularly effective when encoding small-sized objects as I-SPCs. For example, when a three-dimensional data decoding device (hereinafter also referred to as the decoding device) decodes the GOS with a low processing volume or at high speed, it only decodes the I-SPCs in the GOS.

[0188] Moreover, the encoding device can switch the encoding interval or the occurrence frequency of the I-SPCs according to the density of the objects in the WLD.

[0189] Moreover, in Figure 3 In the configuration shown, the encoding device or the decoding device encodes or decodes multiple layers sequentially starting from the lower layer (layer 1). Accordingly, for example, for an automatically moving vehicle or the like, it is possible to increase the priority of the data near the ground where there is a large amount of information.

[0190] In addition, in the encoded data used in a drone or the like, in the GOS, encoding or decoding can be performed sequentially starting from the SPC of the layer above in the height direction.

[0191] Moreover, the encoding device or the decoding device can also encode or decode multiple layers in such a way that the decoding device generally grasps the GOS and can gradually increase the resolution. For example, the encoding device or the decoding device can perform encoding or decoding in the order of layer 3, 8, 1, 9...

[0192] Next, a method for corresponding static objects and dynamic objects will be described.

[0193] In the three-dimensional space, there are static objects or scenes such as buildings or roads (hereinafter collectively referred to as static objects), and dynamic objects such as vehicles or people (hereinafter referred to as dynamic objects). Detection of the objects can be performed separately by extracting feature points from the data of the point cloud, or the images captured by a stereo camera or the like. Here, an example of the encoding method for dynamic objects will be described.

[0194] The first method is a method of encoding without distinguishing between static objects and dynamic objects. The second method is a method of distinguishing between static objects and dynamic objects by identification information.

[0195] For example, the GOS is used as an identification unit. In this case, the GOS including the SPCs constituting the static object and the GOS including the SPCs constituting the dynamic object are distinguished in the encoded data, or by the identification information stored separately from the encoded data.

[0196] Alternatively, the SPC is used as an identification unit. In this case, only the SPCs that constitute the VLM of the static object and the SPCs that include the VLM of the dynamic object are distinguished by the above-mentioned identification information.

[0197] Alternatively, the VLM or VXL can be used as an identification unit. In this case, the VLM or VXL that includes the static object and the VLM or VXL that includes the dynamic object are distinguished by the above-mentioned identification information.

[0198] Moreover, the encoding device can encode the dynamic object as one or more VLMs or SPCs, and encode the VLM or SPC that includes the static object and the SPC that includes the dynamic object as different GOSs. Moreover, when the size of the GOS becomes variable according to the size of the dynamic object, the encoding device stores the size of the GOS separately as meta-information.

[0199] Moreover, the encoding device encodes the static object and the dynamic object independently of each other, and for the world space composed of the static object, the dynamic object can be overlapped. At this time, the dynamic object is composed of one or more SPCs, and each SPC corresponds to one or more SPCs of the static object that overlaps the SPC. In addition, the dynamic object may not be represented by the SPC, but may be represented by one or more VLMs or VXLs.

[0200] Moreover, the encoding device can encode the static object and the dynamic object as different streams.

[0201] Moreover, the encoding device can also generate a GOS that includes one or more SPCs that constitute the dynamic object. Moreover, the encoding device can set the GOS (GOS_M) that includes the dynamic object and the GOS of the static object corresponding to the spatial region of GOS_M to have the same size (occupy the same spatial region). In this way, the overlapping process can be performed in units of GOS.

[0202] The P-SPC or B-SPC that constitutes the dynamic object can also refer to the SPCs included in different encoded GOSs. The position of the dynamic object changes over time. In the case where the same dynamic object is encoded as GOSs at different times, the reference across GOSs is effective from the viewpoint of the compression ratio.

[0203] Moreover, the above-mentioned first method and second method can also be switched according to the use of the encoded data. For example, when the three-dimensional data is encoded and applied as a map, since it is desired to be separated from the dynamic object, the encoding device adopts the second method. In addition, when the encoding device encodes the three-dimensional data of an event such as a concert or a sports event, if there is no need to separate the dynamic object, the first method is adopted.

[0204] Moreover, the decoding time and display time of GOS or SPC can be stored in the encoded data or as meta-information. Also, the time information of static objects can all be the same. In this case, the actual decoding time and display time can be determined by the decoding device. Alternatively, as the decoding time, different values can be assigned for each GOS or SPC, and as the display time, the same value can be assigned to all of them. Moreover, as shown in the decoder mode in video coding such as the HRD (Hypothetical Reference Decoder) of HEVC, the decoder has a buffer of a specified size. As long as the bitstream is read at a specified bit rate according to the decoding time, a model that will not be damaged and can be guaranteed to be decoded can be imported.

[0205] Next, the configuration of GOS in the world space will be described. The coordinates of the three-dimensional space in the world space are represented by three mutually orthogonal coordinate axes (x-axis, y-axis, z-axis). By setting a specified rule in the encoding order of GOS, GOS that are adjacent in space can be encoded continuously in the encoded data. For example, in Figure 4 the example shown, the GOS in the xz plane are encoded continuously. After the encoding of all GOS in one xz plane is completed, the value of the y-axis is updated. That is, as encoding continues, the world space extends in the y-axis direction. Also, the index number of GOS is set as the encoding order.

[0206] Here, the three-dimensional space of the world space corresponds one-to-one with GPS or geographical absolute coordinates such as latitude and longitude. Alternatively, the three-dimensional space can be represented by the relative position with respect to a preset reference position. The directions of the x-axis, y-axis, and z-axis of the three-dimensional space are represented as direction vectors determined based on latitude and longitude, etc., and this direction vector is stored together with the encoded data as meta-information.

[0207] Moreover, the size of GOS is set to be fixed, and the encoding device stores this size as meta-information. Also, the size of GOS can be switched, for example, according to whether it is in the city or indoors or outdoors. That is, the size of GOS can be switched according to the quantity or nature of the object having information value. Alternatively, the encoding device can appropriately switch the size of GOS or the interval of I-SPC within GOS in the same world space according to the density of the object, etc. For example, when the density of the object is higher, the encoding device sets the size of GOS to be smaller and the interval of I-SPC within GOS to be shorter.

[0208] In Figure 5In the example, in the region from the 3rd to the 10th GOS, since the density of the objects is high, in order to achieve random access with a fine granularity, the GOS is subdivided. Also, the 7th to 10th GOSs are respectively located on the back of the 3rd to 6th GOSs.

[0209] Next, the configuration and operation flow of the three-dimensional data encoding device according to the present embodiment will be described. Figure 6 is a block diagram of the three-dimensional data encoding device 100 according to the present embodiment. Figure 7 is a flowchart showing an operation example of the three-dimensional data encoding device 100.

[0210] Figure 6 The three-dimensional data encoding device 100 shown generates encoded three-dimensional data 112 by encoding three-dimensional data 111. This three-dimensional data encoding device 100 includes: an acquisition unit 101, an encoding region determination unit 102, a division unit 103, and an encoding unit 104.

[0211] As Figure 7 shown, first, the acquisition unit 101 acquires three-dimensional data 111 as point cloud data (S101).

[0212] Next, the encoding region determination unit 102 determines the region to be encoded from the spatial region corresponding to the acquired point cloud data (S102). For example, the encoding region determination unit 102 determines the spatial region around the position of the user or vehicle as the region to be encoded according to the position.

[0213] Next, the division unit 103 divides the point cloud data included in the region to be encoded into respective processing units. Here, the processing units are the above-mentioned GOS and SPC, etc. And this region to be encoded corresponds to the above-mentioned world space, for example. Specifically, the division unit 103 divides the point cloud data into processing units according to the size of the GOS set in advance, the presence or absence or size of dynamic objects (S103). And the division unit 103 determines the start position of the SPC that becomes the head in the encoding order in each GOS.

[0214] Next, the encoding unit 104 generates encoded three-dimensional data 112 by sequentially encoding a plurality of SPCs in each GOS (S104).

[0215] In addition, here, after dividing the region to be encoded into GOS and SPC, although an example of encoding each GOS is shown, the order of processing is not limited to the above. For example, after determining the configuration of one GOS, the GOS can be encoded, and then the order of determining the configuration of the GOS, etc.

[0216] In this way, the three-dimensional data encoding device 100 generates the encoded three-dimensional data 112 by encoding the three-dimensional data 111. Specifically, the three-dimensional data encoding device 100 divides the three-dimensional data into random access units, that is, divides it into first processing units (GOS) corresponding to three-dimensional coordinates respectively, divides the first processing units (GOS) into a plurality of second processing units (SPC), and divides the second processing units (SPC) into a plurality of third processing units (VLM). Moreover, the third processing unit (VLM) includes one or more voxels (VXL), and the voxel (VXL) is the smallest unit corresponding to the position information.

[0217] Next, the three-dimensional data encoding device 100 generates the encoded three-dimensional data 112 by encoding each of the plurality of first processing units (GOS). Specifically, the three-dimensional data encoding device 100 encodes each of the plurality of second processing units (SPC) in each first processing unit (GOS). Moreover, the three-dimensional data encoding device 100 encodes each of the plurality of third processing units (VLM) in each second processing unit (SPC).

[0218] For example, when the first processing unit (GOS) of the processing object is a closed GOS, for the second processing unit (SPC) of the processing object included in the first processing unit (GOS) of the processing object, encoding is performed with reference to other second processing units (SPC) included in the first processing unit (GOS) of the processing object. That is, the three-dimensional data encoding device 100 does not refer to the second processing units (SPC) included in the first processing units (GOS) different from the first processing unit (GOS) of the processing object.

[0219] Moreover, when the first processing unit (GOS) of the processing object is an open GOS, for the second processing unit (SPC) of the processing object included in the first processing unit (GOS) of the processing object, encoding is performed with reference to other second processing units (SPC) included in the first processing unit (GOS) of the processing object or the second processing units (SPC) included in the first processing units (GOS) different from the first processing unit (GOS) of the processing object.

[0220] Moreover, the three-dimensional data encoding device 100 selects one from the first type (I-SPC) that never refers to other second processing units (SPC), the second type (P-SPC) that refers to one other second processing unit (SPC), and the third type that refers to two other second processing units (SPC) as the type of the second processing unit (SPC) of the processing object, and encodes the second processing unit (SPC) of the processing object according to the selected type.

[0221] Next, the configuration and operation flow of the three-dimensional data decoding device according to this embodiment will be described. Figure 8 is a block diagram of the three-dimensional data decoding device 200 according to this embodiment. Figure 9 is a flowchart showing an operation example of the three-dimensional data decoding device 200.

[0222] Figure 8 The three-dimensional data decoding device 200 shown generates decoded three-dimensional data 212 by decoding the encoded three-dimensional data 211. Here, the encoded three-dimensional data 211 is, for example, the encoded three-dimensional data 112 generated by the three-dimensional data encoding device 100. The three-dimensional data decoding device 200 includes: an acquisition unit 201, a decoding start GOS determination unit 202, a decoding SPC determination unit 203, and a decoding unit 204.

[0223] First, the acquisition unit 201 acquires the encoded three-dimensional data 211 (S201). Next, the decoding start GOS determination unit 202 determines the GOS to be decoded (S202). Specifically, the decoding start GOS determination unit 202 refers to the meta-information in the encoded three-dimensional data 211 or stored separately from the encoded three-dimensional data, and determines the GOS including the spatial position, object, or SPC corresponding to the time at which decoding starts as the GOS to be decoded.

[0224] Next, the decoding SPC determination unit 203 determines the type (I, P, B) of the SPC to be decoded within the GOS (S203). For example, the decoding SPC determination unit 203 determines (1) whether to decode only I-SPC, (2) whether to decode I-SPC and P-SPC, and (3) whether to decode all types. In addition, in the case where the type of SPC to be decoded is predetermined such as decoding all SPCs, this step may not be performed.

[0225] Next, the decoding unit 204 acquires the address position in the encoded three-dimensional data 211 where the SPC at the beginning in the decoding order (the same as the encoding order) within the GOS is located, acquires the encoded data of the beginning SPC from this address position, and decodes each SPC in sequence from this beginning SPC (S204). And the above address position is stored in meta-information or the like.

[0226] In this way, the three-dimensional data decoding device 200 decodes the decoded three-dimensional data 212. Specifically, the three-dimensional data decoding device 200 generates the decoded three-dimensional data 212 of the first processing unit (GOS) as a random access unit by decoding each of the encoded three-dimensional data 211 of the first processing unit (GOS) corresponding to the three-dimensional coordinates respectively. More specifically, the three-dimensional data decoding device 200 decodes each of the plurality of second processing units (SPC) in each of the first processing units (GOS). Further, the three-dimensional data decoding device 200 decodes each of the plurality of third processing units (VLM) in each of the second processing units (SPC).

[0227] The meta information for random access will be described below. This meta information is generated by the three-dimensional data encoding device 100 and is included in the encoded three-dimensional data 112 (211).

[0228] In the random access of conventional two-dimensional moving images, decoding starts from the first frame of a random access unit near the specified time. However, in the world space, random access is envisioned not only for time but also for (coordinates or objects, etc.).

[0229] Therefore, in order to implement random access to at least the three elements of coordinates, objects, and time, a table in which the index numbers of each element are associated with the GOS is prepared. Moreover, the index number of the GOS is associated with the address of the I-SPC that is the start of the GOS. Figure 10 This is a diagram showing an example of the table included in the meta information. Additionally, it is not necessary to use Figure 10 all of the tables shown. At least one table can be used.

[0230] As an example below, random access starting from coordinates will be described. When accessing the coordinates (x2, y2, z2), first referring to the coordinate-GOS table, it can be known that the location with the coordinates (x2, y2, z2) is included in the second GOS. Then, referring to the GOS address table, since it can be known that the address of the I-SPC at the start of the second GOS is addr(2), the decoding unit 204 obtains data from this address and starts decoding.

[0231] In addition, the address can be an address in a logical format or a physical address of an HDD or a memory. Also, information for determining a file segment can be used instead of the address. For example, a file segment is a unit obtained by segmenting one or more GOSs, etc.

[0232] Also, in the case where the object spans multiple GOSs, the GOSs to which the multiple objects belong can also be shown in the object GOS table. If the multiple GOSs are closed GOSs, the encoding device and the decoding device can perform encoding or decoding in parallel. Additionally, if the multiple GOSs are open GOSs, by cross-referencing between the multiple GOSs, the compression efficiency can be further improved.

[0233] Examples of the object include a person, an animal, a car, a bicycle, a traffic signal, or a building serving as a land mark. For example, when the three-dimensional data encoding device 100 encodes in the world space, it extracts the feature points unique to the object from a three-dimensional point cloud or the like, detects the object based on the feature points, and can set the detected object as a random access point.

[0234] In this way, the three-dimensional data encoding device 100 generates the first information, which shows multiple first processing units (GOSs) and the three-dimensional coordinates corresponding to each of the multiple first processing units (GOSs). And the encoded three-dimensional data 112(211) includes this first information. And the first information further shows at least one of the object, time, and data storage destination corresponding to each of the multiple first processing units (GOSs).

[0235] The three-dimensional data decoding device 200 obtains the first information from the encoded three-dimensional data 211, uses the first information to determine the encoded three-dimensional data 211 of the first processing unit corresponding to the specified three-dimensional coordinates, object, or time, and decodes the encoded three-dimensional data 211.

[0236] Examples of other meta-information are described below. In addition to the meta-information for random access, the three-dimensional data encoding device 100 can also generate and store the following meta-information. And the three-dimensional data decoding device 200 can also use this meta-information during decoding.

[0237] In the case of using the three-dimensional data as map information, etc., a profile is defined according to the usage, and the information showing the profile can be included in the meta-information. For example, a profile for urban areas or suburbs is defined, or a profile for flying objects is defined, and the maximum or minimum size of the world space, SPC, or VLM is defined respectively. For example, in the profile for urban areas, more detailed information is required than in the suburbs, so the minimum size of the VLM is set smaller.

[0238] The meta-information may also include a tag value indicating the type of the object. This tag value corresponds to the VLM, SPC, or GOS that constitutes the object. The tag value can be set according to the type of the object, etc. For example, the tag value "0" represents "person", the tag value "1" represents "car", and the tag value "2" represents "signal lamp". Alternatively, in a case where the type of the object is difficult to determine or does not need to be determined, a tag value indicating a property such as size, or whether it is a dynamic object or a static object may also be used.

[0239] Furthermore, the meta-information may also include information indicating the range of the spatial region occupied by the world space.

[0240] Furthermore, the meta-information may store the size of the SPC or VXL as the entire stream of encoded data, or as header information shared by multiple SPCs such as SPCs within the GOS.

[0241] Furthermore, the meta-information may also include identification information such as a distance sensor or a camera used in the generation of the point cloud, or information indicating the position accuracy of the point group within the point cloud.

[0242] Furthermore, the meta-information may include information indicating whether the world space is composed only of static objects or contains dynamic objects.

[0243] A modification example of the present embodiment will be described below.

[0244] The encoding device or the decoding device may encode or decode two or more different SPCs or GOSs in parallel. The GOSs encoded or decoded in parallel can be determined based on the meta-information indicating the spatial position of the GOSs, etc.

[0245] In a case where the three-dimensional data is used as a spatial map when a vehicle or a flying object moves, or in a case where such a spatial map is generated, etc., the encoding device or the decoding device may encode or decode the GOS or SPC included in the space determined based on GPS, path information, or zoom ratio, etc.

[0246] Furthermore, the decoding device may also start decoding sequentially from the space close to its own position or the walking path. The encoding device or the decoding device may also perform encoding or decoding by making the priority of the space far from its own position or the walking path lower than that of the close space. Here, reducing the priority means reducing the processing order, reducing the resolution (post-screening processing), or reducing the image quality (improving the encoding efficiency. For example, increasing the quantization step size), etc.

[0247] Furthermore, when decoding the encoded data hierarchically encoded within the space, the decoding device may also decode only the lower hierarchy.

[0248] Also, the decoding device can also start decoding from the lower layer according to the zoom ratio or use of the map.

[0249] Also, in applications such as self-position estimation or object recognition during the automatic driving of a vehicle or a robot, the encoding device or the decoding device can also reduce the resolution of areas outside the area within a specified height from the road surface (the area to be recognized) for encoding or decoding.

[0250] Also, the encoding device can also encode the point clouds representing the spatial shapes of the indoor and outdoor spaces independently. For example, by separating the GOS representing the indoor (indoor GOS) from the GOS representing the outdoor (outdoor GOS), the decoding device can select the GOS to be decoded according to the viewpoint position when using the encoded data.

[0251] Also, the encoding device can make the indoor GOS and the outdoor GOS with adjacent coordinates adjacent in the encoding stream for encoding. For example, the encoding device corresponds their identifiers and stores the information showing the identifiers corresponding in the encoding stream or in the meta-information stored separately. Accordingly, the decoding device can identify the indoor GOS and the outdoor GOS with adjacent coordinates by referring to the information in the meta-information.

[0252] Also, the encoding device can also switch the size of the GOS or SPC between the indoor GOS and the outdoor GOS. For example, the encoding device sets the size of the GOS to be smaller indoors than outdoors. Also, the encoding device can also change the accuracy when extracting feature points from the point cloud or the accuracy of object detection, etc., between the indoor GOS and the outdoor GOS.

[0253] Also, the encoding device can attach information for the decoding device to distinguish and display dynamic objects from static objects to the encoded data. Accordingly, the decoding device can combine the dynamic objects with a red frame or explanatory text, etc., for display. In addition, the decoding device can also use only a red frame or explanatory text instead of the dynamic objects for display. And the decoding device can represent more detailed object categories. For example, a car can use a red frame and a person can use a yellow frame.

[0254] Also, the encoding device or the decoding device can determine whether to encode or decode the dynamic objects and the static objects as different SPCs or GOSs according to the appearance frequency of the dynamic objects, or the ratio of the static objects to the dynamic objects, etc. For example, when the appearance frequency or ratio of the dynamic objects exceeds the threshold, the SPC or GOS in which the dynamic objects and the static objects are mixed is allowed, and when the appearance frequency or ratio of the dynamic objects does not exceed the threshold, the SPC or GOS in which the dynamic objects and the static objects are mixed is not allowed.

[0255] When the dynamic object is detected from the two-dimensional image information of the camera instead of the point cloud, the encoding device can separately obtain the information (such as a frame or text) for identifying the detection result and the object position, and encode these information as part of the three-dimensional encoded data. In this case, the decoding device overlays and displays the auxiliary information (frame or text) representing the dynamic object on the decoding result of the static object.

[0256] Moreover, the encoding device can change the density of VXL or VLM according to the complexity of the shape of the static object, etc. For example, the more complex the shape of the static object is, the denser the encoding device sets VXL or VLM. Also, the encoding device can determine the quantization step size, etc. when quantifying the spatial position or color information according to the density of VXL or VLM. For example, the denser VXL or VLM is, the smaller the encoding device sets the quantization step size.

[0257] As described above, the encoding device or decoding device according to the present embodiment performs spatial encoding or decoding in a spatial unit having coordinate information.

[0258] Moreover, the encoding device and the decoding device perform encoding or decoding in a volume unit within the space. The volume includes voxels, which are the smallest units corresponding to the position information.

[0259] Moreover, the encoding device and the decoding device establish correspondences between any elements by using a table in which each element of the spatial information including coordinates, objects, and time, etc. is associated with a GOP, or a table corresponding between the elements, and perform encoding or decoding. And the decoding device judges the coordinates by using the value of the selected element, determines the volume, voxel, or space according to the coordinates, and decodes the space including the volume or voxel, or the determined space.

[0260] Moreover, the encoding device judges the volume, voxel, or space that can be selected by the element through feature point extraction or object recognition, and encodes it as a volume, voxel, or space that can be randomly accessed.

[0261] The space is divided into three types, namely: I-SPC that can be encoded or decoded by the space alone, P-SPC that encodes or decodes with reference to any one processed space, and B-SPC that encodes or decodes with reference to any two processed spaces.

[0262] One or more volumes correspond to static objects or dynamic objects. The space containing the static object and the space containing the dynamic object are encoded or decoded as different GOSs. That is, the SPC containing the static object and the SPC containing the dynamic object are assigned to different GOSs.

[0263] Dynamic objects are encoded or decoded on a per-object basis and correspond to more than one space that contains only static objects. That is, multiple dynamic objects are encoded separately, and the encoded data of the multiple dynamic objects corresponds to the SPC that contains only static objects.

[0264] The encoding device and the decoding device improve the priority of the I-SPC in the GOS to perform encoding or decoding. For example, the encoding device performs encoding in a manner that reduces the degradation of the I-SPC (after decoding, the original three-dimensional data can be reproduced more faithfully). Also, the decoding device decodes only the I-SPC, for example.

[0265] The encoding device can change the frequency of using the I-SPC according to the density or value (quantity) of the objects in the world space to perform encoding. That is, the encoding device changes the frequency of selecting the I-SPC according to the quantity or density of the objects included in the three-dimensional data. For example, the encoding device increases the usage frequency of the I-space as the density of the objects in the world space increases.

[0266] Also, the encoding device sets random access points in units of GOS and stores information indicating the spatial region corresponding to the GOS in the header information.

[0267] The encoding device uses a default value as the spatial size of the GOS, for example. Additionally, the encoding device can also change the size of the GOS according to the value (quantity) or density of the objects or dynamic objects. For example, the encoding device sets the spatial size of the GOS to be smaller as the objects or dynamic objects are denser or more numerous.

[0268] Also, the space or volume includes a feature point group derived using information obtained from sensors such as depth sensors, gyroscopes, or cameras. The coordinates of the feature points are set as the center positions of the voxels. And through the subdivision of the voxels, high-precision position information can be achieved.

[0269] The feature point group is derived using multiple pictures. The multiple pictures have at least the following two types of time information, namely: actual time information, and the same time information in the multiple pictures corresponding to the space (for example, the encoding time for rate control, etc.).

[0270] Also, encoding or decoding is performed in units of GOS that include more than one space.

[0271] The encoding device and the decoding device predict the P space or B space in the GOS of the processing object with reference to the space in the processed GOS.

[0272] Alternatively, the encoding device and the decoding device do not refer to different GOSs, and use the processed space within the GOS of the object to be processed to predict the P space or the B space within the GOS of the object to be processed.

[0273] Furthermore, the encoding device and the decoding device send or receive an encoded stream in units of a world space including one or more GOSs.

[0274] Moreover, the GOS has a layer structure in at least one direction within the world space, and the encoding device and the decoding device perform encoding or decoding starting from the lower layer. For example, a GOS that can be randomly accessed belongs to the lowest layer. A GOS belonging to an upper layer only refers to GOSs belonging to layers below the same layer. That is, the GOS is spatially divided in a predetermined direction and includes multiple layers each having one or more SPCs. The encoding device and the decoding device perform encoding or decoding for each SPC by referring to SPCs included in the same layer as or a lower layer than the SPC.

[0275] In addition, the encoding device and the decoding device continuously perform encoding or decoding on GOSs within a world space unit including multiple GOSs. The encoding device and the decoding device write or read information indicating the order (direction) of encoding or decoding as metadata. That is, the encoded data includes information indicating the encoding order of multiple GOSs.

[0276] Also, the encoding device and the decoding device perform encoding or decoding on two or more different spaces or GOSs in parallel.

[0277] Moreover, the encoding device and the decoding device encode or decode the spatial information (coordinates, size, etc.) of a space or GOS.

[0278] Furthermore, the encoding device and the decoding device encode or decode a space or GOS included in a specific space determined based on external information such as GPS, path information, or magnification related to its own position or / and region size.

[0279] The encoding device or the decoding device performs encoding or decoding by making the priority of a space far from its own position lower than that of a space close to its own position.

[0280] The encoding device sets a direction in the world space according to magnification or use, and encodes a GOS having a layer structure in that direction. And the decoding device preferentially performs decoding starting from the lower layer for a GOS having a layer structure in one direction of the world space set according to magnification or use.

[0281] The encoding device changes the extraction of feature points, the accuracy of object recognition, or the size of the spatial region, etc. included in the indoor and outdoor spaces. However, the encoding device and the decoding device encode or decode the indoor GOS and the outdoor GOS with adjacent coordinates in the world space, and also encode or decode these identifiers in correspondence with each other.

[0282] (Embodiment 2)

[0283] When using the encoded data of the point cloud for an actual device or service, in order to suppress the network bandwidth, it is desired to transmit and receive the required information according to the usage. However, such a function does not exist in the encoding structure of the three-dimensional data so far, and thus there is no encoding method corresponding thereto.

[0284] In this embodiment, a three-dimensional data encoding method, a three-dimensional data encoding device for providing a function of transmitting and receiving the required information according to the usage in the encoded data of the three-dimensional point cloud, a three-dimensional data decoding method for decoding the encoded data, and a three-dimensional data decoding device will be described.

[0285] A voxel (VXL) having a feature amount equal to or more than a certain level is defined as a feature voxel (FVXL), and a world space (WLD) composed of FVXLs is defined as a sparse world space (SWLD). Figure 11 It is a diagram showing a configuration example of the sparse world space and the world space. In the SWLD, there are included: FGOS, which is a GOS composed of FVXLs; FSPC, which is an SPC composed of FVXLs; and FVLM, which is a VLM composed of FVXLs. The data structure and prediction structure of FGOS, FSPC, and FVLM can be the same as those of GOS, SPC, and VLM.

[0286] The feature amount refers to a feature amount representing the three-dimensional position information of the VXL or the visible light information at the VXL position, and in particular, a larger number of feature amounts can be detected at the corners and edges of a three-dimensional object. Specifically, although this feature amount is the three-dimensional feature amount or the visible light feature amount described below, as long as it is a feature amount representing the position, brightness, or color information, etc. of the VXL, it can be any feature amount.

[0287] As the three-dimensional feature amount, a SHOT feature amount (Signature of Histograms of OrienTations), a PFH feature amount (Point Feature Histograms), or a PPF feature amount (Point Pair Feature) is adopted.

[0288] The SHOT feature quantity is obtained by segmenting the periphery of the VXL, calculating the inner product of the normal vector between the reference point and the segmented region, and performing histogramming. This SHOT feature quantity has the characteristics of high dimensionality and high feature expressiveness.

[0289] The PFH feature quantity is obtained by selecting multiple two-point groups near the VXL, calculating the normal vector, etc. based on these two points, and performing histogramming. Since this PFH feature quantity is a histogram feature, it is robust against a small amount of interference and has the characteristic of high feature expressiveness.

[0290] The PPF feature quantity is a feature quantity calculated using the normal vector, etc. according to two VXLs. In this PPF feature quantity, since all VXLs are used, it is robust against occlusion.

[0291] Moreover, as feature quantities of visible light, SIFT (Scale-Invariant Feature Transform), SURF (Speeded Up Robust Features), or HOG (Histogram of Oriented Gradients), etc. that adopt information such as the luminance gradient information of the image can be used.

[0292] SWLD is generated by calculating the above-mentioned feature quantities from each VXL of the WLD and extracting the FVXL. Here, SWLD can be updated each time the WLD is updated, or it can be updated regularly after a certain period of time regardless of the update timing of the WLD.

[0293] SWLD can be generated for each feature quantity. For example, as shown by SWLD1 based on the SHOT feature quantity and SWLD2 based on the SIFT feature quantity, SWLD can be generated separately for each feature quantity, and SWLD can be distinguished and used according to the purpose. Also, the feature quantities of each calculated FVXL can be held as feature quantity information in each FVXL.

[0294] Next, the method of using the sparse world space (SWLD) will be described. Since SWLD only contains feature voxels (FVXL), generally, the data size is smaller compared to the WLD that includes all VXLs.

[0295] In an application that uses feature quantities to achieve a certain purpose, by using the information of SWLD instead of WLD, the read time from the hard disk can be suppressed, and the bandwidth and transmission time during network transmission can be suppressed. For example, as map information, WLD and SWLD are stored in the server in advance, and by switching the transmitted map information to WLD or SWLD according to the requirements from the client, the network bandwidth and transmission time can be suppressed. The following shows specific examples.

[0296] Figure 12 And Figure 13 is a diagram showing usage examples of SWLD and WLD. As Figure 12 shown, when the client 1 as a vehicle-mounted device needs map information for its own position judgment, the client 1 sends a request for obtaining map data for its own position estimation to the server (S301). The server sends SWLD to the client 1 according to this acquisition request (S302). The client 1 uses the received SWLD to judge its own position (S303). At this time, the client 1 obtains VXL information around the client 1 by various methods such as a distance sensor such as a rangefinder, a stereo camera, or a combination of multiple monocular cameras, and estimates its own position information based on the obtained VXL information and SWLD. Here, the own position information includes the three-dimensional position information and the orientation of the client 1, etc.

[0297] As Figure 13 shown, when the client 2 as a vehicle-mounted device needs map information for map drawing such as a three-dimensional map, the client 2 sends a request for obtaining map data for map drawing to the server (S311). The server sends WLD to the client 2 according to this acquisition request (S312). The client 2 uses the received WLD to perform map drawing (S313). At this time, the client 2, for example, uses an image taken by its own visible light camera, etc., and the WLD obtained from the server to create a conceptual image, and depicts the created image on a screen such as a car navigation.

[0298] As shown above, the server sends SWLD to the client in applications that mainly require the feature quantities of each VXL for its own position estimation, and sends WLD to the client when detailed VXL information is required, such as map drawing. Accordingly, the map data can be efficiently transmitted and received.

[0299] In addition, the client can judge which of SWLD and WLD it needs and request the server to send SWLD or WLD. And the server can judge which of SWLD or WLD should be sent according to the status of the client or the network.

[0300] Next, a method for switching the reception and transmission between the Sparse World Space (SWLD) and the World Space (WLD) will be described.

[0301] The reception of the WLD or SWLD can be switched according to the network bandwidth. Figure 14 This is a diagram showing an example of the operation in this case. For example, when a low-speed network with a network bandwidth such as in an LTE (Long Term Evolution) environment is used, when the client accesses the server via the low-speed network (S321), the client obtains the SWLD as map information from the server (S322). Additionally, when a high-speed network with a surplus network bandwidth such as in a WiFi environment is used, the client accesses the server via the high-speed network (S323) and obtains the WLD from the server (S324). Accordingly, the client can obtain appropriate map information according to the network bandwidth of the client.

[0302] Specifically, the client receives the SWLD via LTE outdoors, and when entering indoors such as in a facility, obtains the WLD via WiFi. Accordingly, the client can obtain more detailed map information of the interior.

[0303] In this way, the client can request the WLD or SWLD from the server according to the frequency band of the network it uses. Alternatively, the client can send information indicating the frequency band of the network it uses to the server, and the server sends appropriate data (WLD or SWLD) to the client according to this information. Or, the server can determine the network bandwidth of the client and send appropriate data (WLD or SWLD) to the client.

[0304] Moreover, the reception of the WLD or SWLD can be switched according to the moving speed. Figure 15 This is a diagram showing an example of the operation in this case. For example, when the client is moving at a high speed (S331), the client receives the SWLD from the server (S332). Additionally, when the client is moving at a low speed (S333), the client receives the WLD from the server (S334). Accordingly, the client can both suppress the network bandwidth and obtain map information according to the speed. Specifically, when the client is driving on a highway, by receiving the SWLD with a small data volume, the map information can be updated at an appropriate speed approximately. Additionally, when the client is driving on an ordinary road, by receiving the WLD, more detailed map information can be obtained.

[0305] In this way, the client can request the WLD or SWLD from the server according to its own moving speed. Alternatively, the client can send the information indicating its own moving speed to the server, and the server sends appropriate data (WLD or SWLD) to the client according to this information. Alternatively, the server can determine the moving speed of the client and send appropriate data (WLD or SWLD) to the client.

[0306] Moreover, it can also be that the client first obtains the SWLD from the server and then obtains the WLD of the important regions therein. For example, when the client obtains map data, it first obtains the general map information in the form of SWLD, screens out the regions where features such as buildings, signs, or people appear more frequently from it, and then obtains the WLD of the screened regions. Accordingly, the client can both suppress the amount of received data from the server and obtain the detailed information of the required regions.

[0307] Moreover, it can also be that the server respectively creates the SWLD for each object according to the WLD, and the client receives them respectively according to the usage. Accordingly, the network bandwidth can be suppressed. For example, the server pre-identifies people or vehicles from the WLD and creates the SWLD for people and the SWLD for vehicles. When the client wants to obtain information about the people around it, it receives the SWLD for people, and when it wants to obtain information about vehicles, it receives the SWLD for vehicles. And the types of such SWLD can be distinguished according to the information (flags or types, etc.) attached to the header, etc.

[0308] Next, the configuration and operation process of the three-dimensional data encoding device (such as a server) according to this embodiment will be described. Figure 16 It is a block diagram of the three-dimensional data encoding device 400 according to this embodiment. Figure 17 It is a flowchart of the three-dimensional data encoding process performed by the three-dimensional data encoding device 400.

[0309] Figure 16 The three-dimensional data encoding device 400 shown generates encoded three-dimensional data 413 and 414 as an encoded stream by encoding the input three-dimensional data 411. Here, the encoded three-dimensional data 413 is the encoded three-dimensional data corresponding to the WLD, and the encoded three-dimensional data 414 is the encoded three-dimensional data corresponding to the SWLD. The three-dimensional data encoding device 400 includes: an acquisition unit 401, an encoding region determination unit 402, an SWLD extraction unit 403, a WLD encoding unit 404, and an SWLD encoding unit 405.

[0310] As Figure 17 shown, first, the acquisition unit 401 acquires the input three-dimensional data 411 as point cloud data in a three-dimensional space (S401).

[0311] Next, the coding area determination unit 402 determines the spatial area of the coding object based on the spatial area where the point group data exists (S402).

[0312] Next, the SWLD extraction unit 403 defines the spatial area of the coding object as the WLD, and calculates the feature amount based on each VXL included in the WLD. Further, the SWLD extraction unit 403 extracts the VXL whose feature amount is equal to or greater than a preset threshold value, defines the extracted VXL as the FVXL, and generates the extracted three-dimensional data 412 by adding the FVXL to the SWLD (S403). That is, the extracted three-dimensional data 412 whose feature amount is equal to or greater than the threshold value is extracted from the input three-dimensional data 411.

[0313] Next, the WLD coding unit 404 generates the coded three-dimensional data 413 corresponding to the WLD by coding the input three-dimensional data 411 corresponding to the WLD (S404). At this time, the WLD coding unit 404 attaches information for distinguishing that the coded three-dimensional data 413 is a stream including the WLD to the head of the coded three-dimensional data 413.

[0314] Further, the SWLD coding unit 405 generates the coded three-dimensional data 414 corresponding to the SWLD by coding the extracted three-dimensional data 412 corresponding to the SWLD (S405). At this time, the SWLD coding unit 405 attaches information for distinguishing that the coded three-dimensional data 414 is a stream including the SWLD to the head of the coded three-dimensional data 414.

[0315] Further, the processing order of the process of generating the coded three-dimensional data 413 and the process of generating the coded three-dimensional data 414 may be the reverse of the above. Further, a part or all of the above processes may be executed in parallel.

[0316] The information given to the heads of the coded three-dimensional data 413 and 414 is defined as a parameter such as "world_type", for example. When world_type = 0, it indicates that the stream includes the WLD, and when world_type = 1, it indicates that the stream includes the SWLD. When defining more other categories, the allocated value can be increased as in world_type = 2. Further, a specific flag may be included in one of the coded three-dimensional data 413 and 414. For example, the coded three-dimensional data 414 may be given a flag indicating that the stream includes the SWLD. In this case, the decoding device can determine whether the stream includes the WLD or the SWLD based on the presence or absence of the flag.

[0317] Further, the coding method used by the WLD coding unit 404 when coding the WLD may be different from the coding method used by the SWLD coding unit 405 when coding the SWLD.

[0318] For example, since SWLD data is selected, its correlation with surrounding data may be lower compared to WLD. Therefore, in the encoding method for SWLD, among intra prediction and inter prediction, inter prediction is prioritized compared to the encoding method for WLD.

[0319] Also, the representation methods of three-dimensional positions may be different between the encoding method for SWLD and the encoding method for WLD. For example, in FWLD, the three-dimensional position of FVXL may be represented by three-dimensional coordinates, and in WLD, the three-dimensional position may be represented by an octree described later, and vice versa.

[0320] Moreover, the SWLD encoding unit 405 encodes in such a way that the data size of the encoded three-dimensional data 414 of SWLD is smaller than the data size of the encoded three-dimensional data 413 of WLD. As described above, the correlation between data in SWLD may be lower compared to WLD. Accordingly, the encoding efficiency decreases, and the data size of the encoded three-dimensional data 414 may be larger than the data size of the encoded three-dimensional data 413 of WLD. Therefore, when the data size of the obtained encoded three-dimensional data 414 of SWLD is larger than the data size of the encoded three-dimensional data 413 of WLD, the SWLD encoding unit 405 performs re-encoding to regenerate the encoded three-dimensional data 414 with a reduced data size.

[0321] For example, the SWLD extraction unit 403 regenerates the extracted three-dimensional data 412 with a reduced number of extracted feature points, and the SWLD encoding unit 405 encodes the extracted three-dimensional data 412. Or, the quantization level in the SWLD encoding unit 405 can be made coarser. For example, in the octree structure described later, by rounding the data in the bottom layer, the quantization level can be made coarser.

[0322] Also, when the SWLD encoding unit 405 cannot make the data size of the encoded three-dimensional data 414 of SWLD smaller than the data size of the encoded three-dimensional data 413 of WLD, it may not generate the encoded three-dimensional data 414 of SWLD. Or, the encoded three-dimensional data 413 of WLD can be copied to the encoded three-dimensional data 414 of SWLD. That is, the encoded three-dimensional data 413 of WLD can be directly used as the encoded three-dimensional data 414 of SWLD.

[0323] Next, the configuration and operation process of the three-dimensional data decoding device (e.g., client) according to the present embodiment will be described. Figure 18 This is a block diagram of the three-dimensional data decoding device 500 according to the present embodiment. Figure 19 This is a flowchart of the three-dimensional data decoding process performed by the three-dimensional data decoding device 500.

[0324] Figure 18 The three-dimensional data decoding device 500 shown generates decoded three-dimensional data 512 or 513 by decoding the encoded three-dimensional data 511. Here, the encoded three-dimensional data 511 is, for example, the encoded three-dimensional data 413 or 414 generated by the three-dimensional data encoding device 400.

[0325] The three-dimensional data decoding device 500 includes: an acquisition unit 501, a header analysis unit 502, a WLD decoding unit 503, and a SWLD decoding unit 504.

[0326] As Figure 19 shown, first, the acquisition unit 501 acquires the encoded three-dimensional data 511 (S501). Next, the header analysis unit 502 analyzes the header of the encoded three-dimensional data 511 to determine whether the encoded three-dimensional data 511 is a stream containing WLD or a stream containing SWLD (S502). For example, the determination is made with reference to the above-mentioned world_type parameter.

[0327] In the case where the encoded three-dimensional data 511 is a stream containing WLD (Yes in S503), the WLD decoding unit 503 decodes the encoded three-dimensional data 511 to generate decoded three-dimensional data 512 of WLD (S504). In addition, in the case where the encoded three-dimensional data 511 is a stream containing SWLD (No in S503), the SWLD decoding unit 504 decodes the encoded three-dimensional data 511 to generate decoded three-dimensional data 513 of SWLD (S505).

[0328] And, similar to the encoding device, the decoding method used by the WLD decoding unit 503 when decoding WLD and the decoding method used by the SWLD decoding unit 504 when decoding SWLD can be different. For example, in the decoding method for SWLD, inter-frame prediction in intra-frame prediction and inter-frame prediction can be given priority compared to the decoding method for WLD.

[0329] And, in the decoding method for SWLD and the decoding method for WLD, the representation method of the three-dimensional position can be different. For example, in SWLD, the three-dimensional position of FVXL can be represented by three-dimensional coordinates, and in WLD, the three-dimensional position can be represented by an octree described later, and vice versa.

[0330] Next, the octree representation as a representation method of the three-dimensional position will be described. The VXL data included in the three-dimensional data is transformed into an octree structure and then encoded. Figure 20 is a diagram showing an example of VXL of WLD. Figure 21 is showing Figure 20 a diagram of the octree structure of the WLD shown. In Figure 20In the example shown, there are three VXLs (hereinafter, valid VXLs) that include point groups, namely VXL1 to VXL3. As Figure 21 shown, the octree structure is composed of nodes and leaves. Each node has a maximum of 8 nodes or leaves. Each leaf has VXL information. Here, Figure 21 among the leaves shown, leaves 1, 2, and 3 respectively represent Figure 20 the VXL1, VXL2, and VXL3 shown.

[0331] Specifically, each node and leaf correspond to a three-dimensional position. Node 1 corresponds to Figure 20 all the blocks shown. The block corresponding to Node 1 is divided into 8 blocks. Among the 8 blocks, the blocks including valid VXLs are set as nodes, and the other blocks are set as leaves. The blocks corresponding to the nodes are further divided into 8 nodes or leaves, and the number of times this process is repeated is the same as the number of levels in the tree structure. And all the blocks in the bottom layer are set as leaves.

[0332] And, Figure 22 is a diagram showing an example of the SWLD generated from the Figure 20 shown WLD. Figure 20 The results of feature quantity extraction of the VXL1 and VXL2 shown are judged as FVXL1 and FVXL2 and added to the SWLD. In addition, since VXL3 is not judged as FVXL, it is not included in the SWLD. Figure 23 is a diagram showing the octree structure of the Figure 22 shown SWLD. In the Figure 23 shown octree structure, Figure 21 leaf 3, which corresponds to VXL3, shown is deleted. Accordingly, Figure 21 node 3 shown has no valid VXL and is changed to a leaf. In general, the number of leaves in the SWLD is smaller than the number of leaves in the WLD, and the encoded three-dimensional data of the SWLD is also smaller than the encoded three-dimensional data of the WLD.

[0333] The following describes a modification example of the present embodiment.

[0334] For example, it can also be the case where, when a client such as a vehicle-mounted device estimates its own position, it receives the SWLD from the server, uses the SWLD for its own position estimation, and performs obstacle detection. In this case, various methods such as a distance sensor such as a rangefinder, a stereo camera, or a combination of multiple monocular cameras are used to perform obstacle detection based on the three-dimensional information of the surrounding area obtained by itself.

[0335] Also, generally speaking, it is difficult to include VXL data of flat areas in the SWLD. For this reason, the server maintains a downsampled world space (SubWLD) obtained by downsampling the WLD for detecting stationary obstacles, and can send the SWLD and the SubWLD to the client. Accordingly, both network bandwidth can be suppressed, and self-position estimation and obstacle detection can be performed on the client side.

[0336] Also, when quickly rendering three-dimensional map data on the client side, it may be convenient if the map information has a grid structure. Thus, the server can generate a grid based on the WLD and maintain it in advance as a grid world space (MWLD). For example, when the client needs to perform rough three-dimensional rendering, it receives the MWLD, and when it needs to perform detailed three-dimensional rendering, it receives the WLD. Accordingly, network bandwidth can be suppressed.

[0337] Also, although the server sets the VXLs with feature amounts above the threshold as FVXLs from each VXL, the FVXLs can also be calculated by different methods. For example, if the server determines that VXLs, VLMs, SPCs, or GOSs that make up signals or intersections are required for self-position estimation, driving assistance, or autonomous driving, etc., they can be included in the SWLD as FVXLs, FVLMs, FSPCs, or FGOSs. And the above determination can be made manually. In addition, the FVXLs obtained by the above method can be added to the FVXLs, etc. set based on the feature amount. That is, the SWLD extraction unit 403 can further extract data corresponding to an object having a predetermined attribute from the input three-dimensional data 411 as the extracted three-dimensional data 412.

[0338] Also, different labels can be assigned to the situations that are required for these uses and are different from the feature amounts. The server can separately maintain the FVXLs required for self-position estimation, driving assistance, or autonomous driving, such as signals or intersections, as an upper layer of the SWLD (for example, a lane world space).

[0339] Also, the server can attach attributes to the VXLs in the WLD according to a random access unit or a specified unit. The attributes include, for example, information indicating whether it is required or not required for self-position estimation, or information indicating whether it is important as traffic information such as a signal or an intersection. And the attributes can also include the correspondence relationship with Features (intersections or roads, etc.) in lane information (GDF: Geographic Data Files, etc.).

[0340] Also, as a method for updating the WLD or SWLD, the following method can be adopted.

[0341] Update information such as changes in people, construction, or street trees (facing the track) is loaded into the server as a point cloud or metadata. The server updates the WLD based on this load, and then updates the SWLD using the updated WLD.

[0342] Moreover, when the client detects a mismatch between the three-dimensional information generated by itself during self-position estimation and the three-dimensional information received from the server, it can send the three-dimensional information generated by itself to the server together with an update notification. In this case, the server updates the SWLD using the WLD. If the SWLD is not updated, the server determines that the WLD itself is old.

[0343] Also, as the header information of the encoded stream, information for distinguishing between WLD and SWLD is attached. For example, in the case where there are multiple world spaces such as a grid world space or a lane world space, information for distinguishing them can be attached to the header information. And when there are multiple SWLDs with different feature amounts, information for distinguishing them separately can also be attached to the header information.

[0344] Furthermore, although the SWLD is composed of FVXLs, it can also include VXLs that are not determined to be FVXLs. For example, the SWLD can include adjacent VXLs used when calculating the feature amounts of FVXLs. Accordingly, even when no feature amount information is attached to each FVXL of the SWLD, the client can calculate the feature amounts of FVXLs when receiving the SWLD. Additionally, at this time, the SWLD can include information for distinguishing whether each VXL is an FVXL or a VXL.

[0345] As described above, the three-dimensional data encoding device 400 extracts the extracted three-dimensional data 412 (second three-dimensional data) with a feature amount equal to or greater than the threshold from the input three-dimensional data 411 (first three-dimensional data), and generates the encoded three-dimensional data 414 (first encoded three-dimensional data) by encoding the extracted three-dimensional data 412.

[0346] Accordingly, the three-dimensional data encoding device 400 generates the encoded three-dimensional data 414 obtained by encoding the data with a feature amount equal to or greater than the threshold. In this way, compared with directly encoding the input three-dimensional data 411, the data amount can be reduced. Therefore, the three-dimensional data encoding device 400 can reduce the data amount during transmission.

[0347] Moreover, the three-dimensional data encoding device 400 further generates the encoded three-dimensional data 413 (second encoded three-dimensional data) by encoding the input three-dimensional data 411.

[0348] Accordingly, the three-dimensional data encoding device 400 can selectively transmit the encoded three-dimensional data 413 and the encoded three-dimensional data 414 according to the usage purpose or the like.

[0349] Moreover, the extracted three-dimensional data 412 is encoded by the first encoding method, and the input three-dimensional data 411 is encoded by a second encoding method different from the first encoding method.

[0350] Accordingly, the three-dimensional data encoding device 400 can adopt appropriate encoding methods for the input three-dimensional data 411 and the extracted three-dimensional data 412 respectively.

[0351] Moreover, in the first encoding method, among intra prediction and inter prediction, inter prediction is prioritized compared to the second encoding method.

[0352] Accordingly, the three-dimensional data encoding device 400 can increase the priority of inter prediction for the extracted three-dimensional data 412 where the correlation between adjacent data is likely to be low.

[0353] Moreover, in the first encoding method and the second encoding method, the representation methods of three-dimensional positions are different. For example, in the second encoding method, the three-dimensional position is represented by an octree, and in the first encoding method, the three-dimensional position is represented by three-dimensional coordinates.

[0354] Accordingly, the three-dimensional data encoding device 400 can adopt a more appropriate representation method of three-dimensional positions for three-dimensional data with different numbers of data (the number of VXL or FVXL).

[0355] Moreover, at least one of the encoded three-dimensional data 413 and 414 includes an identifier indicating whether the encoded three-dimensional data is the encoded three-dimensional data obtained by encoding the input three-dimensional data 411 or the encoded three-dimensional data obtained by encoding a part of the input three-dimensional data 411. That is, this identifier indicates whether the encoded three-dimensional data is the encoded three-dimensional data 413 of WLD or the encoded three-dimensional data 414 of SWLD.

[0356] Accordingly, the decoding device can easily determine whether the acquired encoded three-dimensional data is the encoded three-dimensional data 413 or the encoded three-dimensional data 414.

[0357] Moreover, the three-dimensional data encoding device 400 encodes the extracted three-dimensional data 412 in such a way that the data amount of the encoded three-dimensional data 414 is less than the data amount of the encoded three-dimensional data 413.

[0358] Accordingly, the three-dimensional data encoding device 400 can make the data amount of the encoded three-dimensional data 414 less than the data amount of the encoded three-dimensional data 413.

[0359] Further, the three-dimensional data encoding device 400 extracts data corresponding to an object having a predetermined attribute from the input three-dimensional data 411 as the extracted three-dimensional data 412. For example, an object having a predetermined attribute refers to an object required for self-position estimation, driving assistance, or autonomous driving, such as a signal or an intersection.

[0360] Accordingly, the three-dimensional data encoding device 400 can generate the encoded three-dimensional data 414 including the data required by the decoding device.

[0361] Further, the three-dimensional data encoding device 400 (server) sends either the encoded three-dimensional data 413 or 414 to the client according to the state of the client.

[0362] Accordingly, the three-dimensional data encoding device 400 can send appropriate data according to the state of the client.

[0363] The state of the client includes the communication status of the client (e.g., network bandwidth) or the moving speed of the client.

[0364] Further, the three-dimensional data encoding device 400 sends either the encoded three-dimensional data 413 or 414 to the client according to the request of the client.

[0365] Accordingly, the three-dimensional data encoding device 400 can send appropriate data according to the request of the client.

[0366] The three-dimensional data decoding device 500 according to the present embodiment decodes the encoded three-dimensional data 413 or 414 generated by the three-dimensional data encoding device 400.

[0367] That is, the three-dimensional data decoding device 500 decodes the encoded three-dimensional data 414 obtained by encoding the extracted three-dimensional data 412 whose feature amount extracted from the input three-dimensional data 411 is above the threshold by the first decoding method. And the three-dimensional data decoding device 500 decodes the encoded three-dimensional data 413 obtained by encoding the input three-dimensional data 411 using a second decoding method different from the first decoding method.

[0368] Accordingly, the three-dimensional data decoding device 500 can selectively receive, for example, according to the usage purpose, etc., the encoded three-dimensional data 414 and the encoded three-dimensional data 413 obtained by encoding the data whose feature amount is above the threshold. Accordingly, the three-dimensional data decoding device 500 can reduce the amount of data during transmission. Moreover, the three-dimensional data decoding device 500 can adopt appropriate decoding methods for the input three-dimensional data 411 and the extracted three-dimensional data 412 respectively.

[0369] Further, in the first decoding method, among intra prediction and inter prediction, inter prediction is prioritized compared to the second decoding method.

[0370] Accordingly, the three-dimensional data decoding device 500 can increase the priority of inter prediction for the extracted three-dimensional data where the correlation between adjacent data is likely to decrease.

[0371] Further, in the first decoding method and the second decoding method, the representation methods of three-dimensional positions are different. For example, in the second decoding method, the three-dimensional position is represented by an octree, and in the first decoding method, the three-dimensional position is represented by three-dimensional coordinates.

[0372] Accordingly, the three-dimensional data decoding device 500 can adopt a more appropriate representation method of three-dimensional positions for three-dimensional data with different numbers of data (the number of VXLs or FVXLs).

[0373] Further, at least one of the encoded three-dimensional data 413 and 414 includes an identifier indicating whether the encoded three-dimensional data is the encoded three-dimensional data obtained by encoding the input three-dimensional data 411 or the encoded three-dimensional data obtained by encoding a part of the input three-dimensional data 411. The three-dimensional data decoding device 500 identifies the encoded three-dimensional data 413 and 414 with reference to this identifier.

[0374] Accordingly, the three-dimensional data decoding device 500 can easily determine whether the obtained encoded three-dimensional data is the encoded three-dimensional data 413 or the encoded three-dimensional data 414.

[0375] Further, the three-dimensional data decoding device 500 further notifies the server of the state of the client (the three-dimensional data decoding device 500). The three-dimensional data decoding device 500 receives one of the encoded three-dimensional data 413 and 414 sent from the server according to the state of the client.

[0376] Accordingly, the three-dimensional data decoding device 500 can receive appropriate data according to the state of the client.

[0377] Further, the state of the client includes the communication status of the client (e.g., network bandwidth) or the moving speed of the client.

[0378] Further, the three-dimensional data decoding device 500 further requests one of the encoded three-dimensional data 413 and 414 from the server and receives one of the encoded three-dimensional data 413 and 414 sent from the server according to this request.

[0379] Accordingly, the three-dimensional data decoding device 500 can receive appropriate data corresponding to the usage.

[0380] (Embodiment 3)

[0381] In the present embodiment, a method for transmitting and receiving three-dimensional data between vehicles will be described.

[0382] Figure 24 It is a schematic diagram showing the state of transmission and reception of three-dimensional data 607 between the host vehicle 600 and surrounding vehicles 601.

[0383] When acquiring three-dimensional data by sensors (distance sensors such as rangefinders, stereo cameras, or a combination of multiple monocular cameras, etc.) mounted on the host vehicle 600, due to the influence of obstacles such as surrounding vehicles 601, although within the detection range 602 of the sensors of the host vehicle 600, there will be areas where three-dimensional data cannot be created (hereinafter referred to as occlusion areas 604). Also, although the larger the space for obtaining three-dimensional data, the higher the accuracy of autonomous operation, however, the detection range of only the sensors of the host vehicle 600 is limited.

[0384] The detection range 602 of the sensors of the host vehicle 600 includes an area 603 where three-dimensional data can be obtained and an occlusion area 604. The range in which the host vehicle 600 wants to obtain three-dimensional data includes the detection range 602 of the sensors of the host vehicle 600 and areas other than this. Also, the detection range 605 of the sensors of the surrounding vehicle 601 includes the occlusion area 604 and an area 606 not included in the detection range 602 of the sensors of the host vehicle 600.

[0385] The surrounding vehicle 601 transmits the information detected by the surrounding vehicle 601 to the host vehicle 600. The host vehicle 600 can obtain three-dimensional data 607 of the occlusion area 604 and an area 606 outside the detection range 602 of the sensors of the host vehicle 600 by obtaining the information detected by surrounding vehicles 601 such as the vehicle traveling ahead. The host vehicle 600 uses the information obtained by the surrounding vehicle 601 to supplement the three-dimensional data of the occlusion area 604 and the area 606 outside the sensor detection range.

[0386] The uses of three-dimensional data in the autonomous operation of vehicles or robots are for self-position estimation, detection of surrounding conditions, or both. For example, in self-position estimation, three-dimensional data generated by the host vehicle 600 is used based on the sensor information of the host vehicle 600. In the detection of surrounding conditions, in addition to the three-dimensional data generated by the host vehicle 600, three-dimensional data obtained by the surrounding vehicle 601 is also used.

[0387] Surrounding vehicle 601 for transmitting three-dimensional data 607 to its own vehicle 600 can be determined according to the state of its own vehicle 600. For example, the surrounding vehicle 601 is the vehicle traveling in front when its own vehicle 600 goes straight, the oncoming vehicle when its own vehicle 600 turns right, or the vehicle behind when its own vehicle 600 reverses. Also, it can be that the driver of its own vehicle 600 directly designates the surrounding vehicle 601 for transmitting the three-dimensional data 607 to its own vehicle 600.

[0388] Moreover, its own vehicle 600 can also search for the surrounding vehicle 601 that holds the three-dimensional data of the area that its own vehicle 600 cannot obtain within the space where it wants to obtain the three-dimensional data 607. The area that its own vehicle 600 cannot obtain refers to the occlusion area 604 or the area 606 outside the sensor detection range 602, etc.

[0389] Furthermore, its own vehicle 600 can also determine the occlusion area 604 based on the sensor information of its own vehicle 600. For example, its own vehicle 600 determines the area within the sensor detection range 602 of its own vehicle 600 that cannot generate three-dimensional data as the occlusion area 604.

[0390] The following describes an example of the operation when the vehicle traveling in front is the one transmitting the three-dimensional data 607. Figure 25 This is a diagram showing an example of the three-dimensional data to be transmitted in this case.

[0391] As Figure 25 shown, the three-dimensional data 607 transmitted from the vehicle traveling in front is, for example, the sparse world space (SWLD) of the point cloud. That is, the vehicle traveling in front creates the three-dimensional data (point cloud) of the WLD based on the information detected by the sensors of that vehicle traveling in front, and creates the three-dimensional data (point cloud) of the SWLD by extracting the data with a feature amount above the threshold from the three-dimensional data of the WLD. Then, the vehicle traveling in front transmits the created three-dimensional data of the SWLD to its own vehicle 600.

[0392] Its own vehicle 600 receives the SWLD and merges the received SWLD into the point cloud created by its own vehicle 600.

[0393] The transmitted SWLD has the information of the absolute coordinates (the position of the SWLD in the coordinate system of the three-dimensional map). Its own vehicle 600 superimposes the point cloud generated by its own vehicle 600 according to the absolute coordinates, thereby enabling the merging process.

[0394] The SWLD transmitted from the surrounding vehicle 601 can be the SWLD of the area 606 that is outside the sensor detection range 602 of the host vehicle 600 and within the sensor detection range 605 of the surrounding vehicle 601, or can be the SWLD of the occlusion area 604 relative to the host vehicle 600, or can be the SWLD of both. Also, the transmitted SWLD can be the SWLD of the area used by the surrounding vehicle 601 when detecting the surrounding situation among the above-mentioned SWLDs.

[0395] Also, the surrounding vehicle 601 can change the density of the transmitted point cloud according to the communicable time based on the speed difference between the host vehicle 600 and the surrounding vehicle 601. For example, when the speed difference is large and the communicable time is short, the surrounding vehicle 601 can reduce the density (data volume) of the point cloud by extracting three-dimensional points with large feature amounts from the SWLD.

[0396] Also, the detection of the surrounding situation means detecting whether there are people, vehicles, road construction equipment, etc., determining their types, and detecting their positions, moving directions, moving speeds, etc.

[0397] Also, it can be that the host vehicle 600 obtains the braking information of the surrounding vehicle 601 instead of the three-dimensional data 607 generated by the surrounding vehicle 601, or in addition to the three-dimensional data 607, also obtains the braking information of the surrounding vehicle 601. Here, the braking information of the surrounding vehicle 601 is, for example, information indicating that the accelerator or brake of the surrounding vehicle 601 is depressed, or indicating the degree of depression.

[0398] Also, for the point cloud generated by each vehicle, considering low-latency communication between vehicles, the three-dimensional space is subdivided into random access units. In addition, for the map data downloaded from the server, compared with the case of vehicle-to-vehicle communication, three-dimensional maps, etc. are divided into larger random access units.

[0399] Data of areas that are likely to be occlusion areas, such as the area in front of the vehicle traveling ahead or the area behind the vehicle traveling behind, are divided into small random access units as data for low latency.

[0400] When traveling at high speed, since the importance of the front increases, each vehicle, when traveling at high speed, creates the SWLD of a narrow viewing angle range with small random access units.

[0401] When the area where the host vehicle 600 can obtain the point cloud is included in the SWLD created by the vehicle traveling ahead for transmission, the vehicle traveling ahead can reduce the transmission amount by removing the point cloud in this area.

[0402] Next, the configuration and operation of the three-dimensional data creation device 620, which is the three-dimensional data receiving device according to the present embodiment, will be described.

[0403] Figure 26 FIG. 4 is a block diagram of the three-dimensional data creation device 620 according to the present embodiment. The three-dimensional data creation device 620 is included in the own vehicle 600 described above, for example, and creates a denser third three-dimensional data 636 by synthesizing the received second three-dimensional data 635 and the first three-dimensional data 632 created by the three-dimensional data creation device 620.

[0404] The three-dimensional data creation device 620 includes: a three-dimensional data creation unit 621, a request range determination unit 622, a search unit 623, a reception unit 624, a decoding unit 625, and a synthesis unit 626. Figure 27 FIG. 5 is a flowchart showing the operation of the three-dimensional data creation device 620.

[0405] First, the three-dimensional data creation unit 621 creates the first three-dimensional data 632 using the sensor information 631 detected by the sensors included in the own vehicle 600 (S621). Next, the request range determination unit 622 determines a request range, which is a three-dimensional space range in which the data in the created first three-dimensional data 632 is insufficient (S622).

[0406] Next, the search unit 623 searches for surrounding vehicles 601 that hold three-dimensional data within the request range, and transmits request range information 633 indicating the request range to the surrounding vehicles 601 determined by the search (S623). Next, the reception unit 624 receives the encoded three-dimensional data 634, which is an encoded stream within the request range, from the surrounding vehicles 601 (S624). In addition, the search unit 623 can issue requests to all vehicles existing within the determined range without discrimination, and receive the encoded three-dimensional data 634 from the responding parties. Further, the search unit 623 is not limited to vehicles, and can also issue requests to objects such as traffic lights or signs, and receive the encoded three-dimensional data 634 from such objects.

[0407] Next, the received encoded three-dimensional data 634 is decoded by the decoding unit 625 to obtain the second three-dimensional data 635 (S625). Next, the first three-dimensional data 632 and the second three-dimensional data 635 are synthesized by the synthesis unit 626 to create a denser third three-dimensional data 636 (S626).

[0408] Next, the configuration and operation of the three-dimensional data transmission device 640 according to the present embodiment will be described. Figure 28 FIG. 6 is a block diagram of the three-dimensional data transmission device 640.

[0409] The three-dimensional data transmission device 640 is included, for example, in the above-mentioned surrounding vehicle 601. The fifth three-dimensional data 652 created by the surrounding vehicle 601 is processed into the sixth three-dimensional data 654 requested by the host vehicle 600, and the encoded three-dimensional data 634 is generated by encoding the sixth three-dimensional data 654, and the encoded three-dimensional data 634 is transmitted to the host vehicle 600.

[0410] The three-dimensional data transmission device 640 includes: a three-dimensional data creation unit 641, a reception unit 642, an extraction unit 643, an encoding unit 644, and a transmission unit 645. Figure 29 It is a flowchart showing the operation of the three-dimensional data transmission device 640.

[0411] First, the three-dimensional data creation unit 641 creates the fifth three-dimensional data 652 using the sensor information 651 detected by the sensors included in the surrounding vehicle 601 (S641). Next, the reception unit 642 receives the requested range information 633 transmitted from the host vehicle 600 (S642).

[0412] Next, the extraction unit 643 extracts the three-dimensional data of the requested range indicated by the requested range information 633 from the fifth three-dimensional data 652, and processes the fifth three-dimensional data 652 into the sixth three-dimensional data 654 (S643). Next, the encoding unit 644 encodes the sixth three-dimensional data 654, thereby generating the encoded three-dimensional data 634 as an encoded stream (S644). Then, the transmission unit 645 transmits the encoded three-dimensional data 634 to the host vehicle 600 (S645).

[0413] In addition, here, although the example in which the host vehicle 600 is equipped with the three-dimensional data creation device 620 and the surrounding vehicle 601 is equipped with the three-dimensional data transmission device 640 has been described, each vehicle may also have the functions of the three-dimensional data creation device 620 and the three-dimensional data transmission device 640.

[0414] Hereinafter, the configuration and operation in the case where the three-dimensional data creation device 620 is a surrounding situation detection device for realizing the detection process of the surrounding situation of the host vehicle 600 will be described. Figure 30 It is a block diagram showing the configuration of the three-dimensional data creation device 620A in this case. Figure 30 The shown three-dimensional data creation device 620A in addition to Figure 26 In addition to the configuration of the shown three-dimensional data creation device 620, it further includes: a detection area determination unit 627, a surrounding situation detection unit 628, and an autonomous operation control unit 629. And the three-dimensional data creation device 620A is included in the host vehicle 600.

[0415] Figure 31This is a flowchart for the surrounding situation detection process of the host vehicle 600 by the 3D data production device 620A.

[0416] First, the 3D data production unit 621 produces the first 3D data 632 as a point cloud using the sensor information 631 within the detection range of the host vehicle 600 detected by the sensors equipped on the host vehicle 600 (S661). Additionally, the 3D data production device 620A can also use the sensor information 631 to estimate its own position.

[0417] Next, the detection area determination unit 627 determines the detection target range, which is the spatial area for which the surrounding situation is to be detected (S662). For example, the detection area determination unit 627 calculates the area required for safe autonomous operation in the surrounding situation detection based on the driving direction and speed of the host vehicle 600 and other autonomous operation (autopilot) conditions, and determines this area as the detection target range.

[0418] Next, the request range determination unit 622 determines the occlusion area 604 and the spatial area that is outside the detection range of the sensors of the host vehicle 600 but is required for the surrounding situation detection as the request range (S663).

[0419] When there is a request range determined in step S663 (Yes in S664), the search unit 623 searches for surrounding vehicles that hold information related to the request range. For example, the search unit 623 can ask the surrounding vehicles whether they hold information related to the request range, and based on the request range and the positions of the surrounding vehicles, determine whether the surrounding vehicles hold information related to the request range. Next, the search unit 623 sends a delegation signal 637 for commissioning the transmission of 3D data to the surrounding vehicle 601 identified through the search. And after the search unit 623 receives a permission signal sent from the surrounding vehicle 601 indicating acceptance of the commission of the delegation signal 637, it sends the request range information 633 showing the request range to the surrounding vehicle 601 (S665).

[0420] Next, the receiving unit 624 detects the transmission notification of the transmission data 638, which is the information related to the request range, and accepts the transmission data 638 (S666).

[0421] Alternatively, the 3D data production device 620A may not search for the recipient of the request, but instead issue requests to all vehicles within the determined range without discrimination, and receive the transmission data 638 from the party that responds as holding information related to the request range. And the search unit 623 is not limited to vehicles, and may also issue requests to objects such as traffic lights or signs, and receive the transmission data 638 from the object.

[0422] Further, the transmitted data 638 includes at least one of the encoded three-dimensional data 634 obtained by encoding the three-dimensional data of the requested range generated by the surrounding vehicle 601 and the surrounding condition detection result 639 of the requested range. The surrounding condition detection result 639 shows the positions, moving directions, moving speeds, etc. of the people and vehicles detected by the surrounding vehicle 601. Further, the transmitted data 638 may also include information such as the position and movement of the surrounding vehicle 601. For example, the transmitted data 638 may also include the braking information of the surrounding vehicle 601.

[0423] When the encoded three-dimensional data 634 is included in the received transmitted data 638 (Yes in S667), the encoded three-dimensional data 634 is decoded by the decoding unit 625 to obtain the second three-dimensional data 635 of the SWLD (S668). That is, the second three-dimensional data 635 is three-dimensional data (SWLD) generated by extracting data with a feature amount equal to or greater than the threshold value from the fourth three-dimensional data (WLD).

[0424] Next, the first three-dimensional data 632 and the second three-dimensional data 635 are synthesized by the synthesis unit 626 to generate the third three-dimensional data 636 (S669).

[0425] Next, the surrounding condition detection unit 628 uses the point cloud of the spatial region required for surrounding condition detection, that is, the third three-dimensional data 636, to detect the surrounding conditions of the host vehicle 600 (S670). Further, when the surrounding condition detection result 639 is included in the received transmitted data 638, the surrounding condition detection unit 628 uses the surrounding condition detection result 639 in addition to the third three-dimensional data 636 to detect the surrounding conditions of the host vehicle 600. Further, when the braking information of the surrounding vehicle 601 is included in the received transmitted data 638, the surrounding condition detection unit 628 uses the braking information in addition to the third three-dimensional data 636 to detect the surrounding conditions of the host vehicle 600.

[0426] Next, the autonomous operation control unit 629 controls the autonomous operation (automatic driving) of the host vehicle 600 based on the surrounding condition detection result obtained by the surrounding condition detection unit 628 (S671). Further, the surrounding condition detection result may be presented to the driver through a UI (user interface) or the like.

[0427] Also, when there is no request range in step S663 (the "No" in S664), that is, when all the spatial area information required for surrounding situation detection is created based on the sensor information 631, the surrounding situation detection unit 628 uses the point cloud of the spatial area required for surrounding situation detection, i.e., the first three-dimensional data 632, to detect the surrounding situation of the host vehicle 600 (S672). Then, the autonomous driving control unit 629 controls the autonomous driving (autonomous operation) of the host vehicle 600 according to the surrounding situation detection result by the surrounding situation detection unit 628 (S671).

[0428] Also, when the received transmission data 638 does not include the encoded three-dimensional data 634 (the "No" in S667), that is, when the transmission data 638 only includes the surrounding situation detection result 639 of the surrounding vehicle 601 or the braking information, the surrounding situation detection unit 628 uses the first three-dimensional data 632 and the surrounding situation detection result 639 or the braking information to detect the surrounding situation of the host vehicle 600 (S673). Then, the autonomous driving control unit 629 controls the autonomous driving (autonomous operation) of the host vehicle 600 according to the surrounding situation detection result by the surrounding situation detection unit 628 (S671).

[0429] Next, the three-dimensional data transmission device 640A that transmits the transmission data 638 to the above-mentioned three-dimensional data creation device 620A will be described. Figure 32 It is a block diagram of the three-dimensional data transmission device 640A.

[0430] Figure 32 The shown three-dimensional data transmission device 640A in addition to Figure 28 the configuration of the three-dimensional data transmission device 640 shown, further includes a transmission availability determination unit 646. And the three-dimensional data transmission device 640A is included in the surrounding vehicle 601.

[0431] Figure 33 It is a flowchart showing an operation example of the three-dimensional data transmission device 640A. First, the three-dimensional data creation unit 641 uses the sensor information 651 detected by the sensors equipped in the surrounding vehicle 601 to create the fifth three-dimensional data 652 (S681).

[0432] Next, the receiving unit 642 receives a delegation signal 637 for delegating a transmission request for three-dimensional data from its own vehicle 600 (S682). Next, the transmission permission determination unit 646 determines whether to respond to the delegation indicated by the delegation signal 637 (S683). For example, the transmission permission determination unit 646 determines whether to respond to the delegation according to the content preset by the user in advance. Alternatively, the receiving unit 642 may first receive a request from the other party such as a request range, and the transmission permission determination unit 646 determines whether to respond to the delegation according to this content. For example, the transmission permission determination unit 646 may determine to respond to the delegation when holding three-dimensional data within the request range, and determine not to respond to the delegation when not having three-dimensional data within the request range.

[0433] In the case of responding to the delegation (Yes in S683), the three-dimensional data transmission device 640A sends a permission signal to its own vehicle 600, and the receiving unit 642 receives request range information 633 indicating the request range (S684). Next, the extraction unit 643 extracts the point cloud within the request range from the fifth three-dimensional data 652 which is a point cloud, and creates transmission data 638 of the SWLD (the sixth three-dimensional data 654) including the extracted point cloud (S685).

[0434] That is, it may be that the three-dimensional data transmission device 640A creates the seventh three-dimensional data (WLD) according to the sensor information 651, and creates the fifth three-dimensional data 652 (SWLD) by extracting data with a feature amount above the threshold from the seventh three-dimensional data (WLD). And it may be that the three-dimensional data creation unit 641 pre-creates the three-dimensional data of the SWLD, and the extraction unit 643 extracts the three-dimensional data of the SWLD within the request range from the three-dimensional data of the SWLD. Alternatively, the extraction unit 643 may generate the three-dimensional data of the SWLD within the request range according to the three-dimensional data of the WLD within the request range.

[0435] Moreover, the transmission data 638 may include the surrounding condition detection result 639 within the request range performed by the surrounding vehicle 601 and the braking information of the surrounding vehicle 601. Also, the transmission data 638 may not include the sixth three-dimensional data 654, but may include at least one of the surrounding condition detection result 639 within the request range performed by the surrounding vehicle 601 and the braking information of the surrounding vehicle 601.

[0436] In the case where the transmission data 638 includes the sixth three-dimensional data 654 (Yes in S686), the encoding unit 644 encodes the sixth three-dimensional data 654 to generate encoded three-dimensional data 634 (S687).

[0437] Then, the transmission unit 645 sends the transmission data 638 including the encoded three-dimensional data 634 to its own vehicle 600 (S688).

[0438] Further, when the transmission data 638 does not include the sixth three-dimensional data 654 (No in S686), the transmission unit 645 transmits the transmission data 638 including at least one of the surrounding condition detection result 639 of the request range performed by the surrounding vehicle 601 and the braking information of the surrounding vehicle 601 to its own vehicle 600 (S688).

[0439] A modification example of the present embodiment will be described below.

[0440] For example, the information transmitted from the surrounding vehicle 601 may not be the three-dimensional data or the surrounding condition detection result created by the surrounding vehicle, but may be the correct feature point information of the surrounding vehicle 601 itself. The own vehicle 600 uses the feature point information of the surrounding vehicle 601 to correct the feature point information of the vehicle traveling ahead in the point cloud obtained by the own vehicle 600. Accordingly, the own vehicle 600 can improve the matching accuracy when estimating its own position.

[0441] And, the feature point information of the vehicle traveling ahead is, for example, three-dimensional point information composed of color information and coordinate information. Accordingly, even if the sensor of the own vehicle 600 is a laser sensor or a stereo camera, the feature point information of the vehicle traveling ahead can be used regardless of its type.

[0442] In addition, the own vehicle 600 is not limited by the time of transmission, and can use the point cloud of the SWLD when calculating the accuracy of its own position estimation. For example, when the sensor of the own vehicle 600 is an imaging device such as a stereo camera, two-dimensional points on the image captured by the camera of the own vehicle 600 are detected, and the own position is estimated using the two-dimensional points. And, while estimating its own position, the own vehicle 600 creates a point cloud of surrounding objects. The own vehicle 600 projects the three-dimensional points of the SWLD among them onto the two-dimensional image again, and evaluates the accuracy of its own position estimation based on the error between the detection points on the two-dimensional image and the re-projected points.

[0443] And, when the sensor of the own vehicle 600 is a laser sensor such as LIDAR, the own vehicle 600 evaluates the accuracy of its own position estimation based on the error calculated by the Iterative Closest Point algorithm between the SWLD of the created point cloud and the SWLD of the three-dimensional map.

[0444] And, when the communication state via a base station or a server such as 5G is poor, the own vehicle 600 can obtain a three-dimensional map from the surrounding vehicle 601.

[0445] Furthermore, for distant information that cannot be obtained from vehicles in the vicinity of the host vehicle 600, it can be obtained through vehicle-to-vehicle communication. For example, regarding traffic accident information that just occurred several hundred meters or several kilometers ahead, the host vehicle 600 can obtain it through communication when passing by oncoming vehicles or in the way of sequential transmission among surrounding vehicles. At this time, the data form of the transmitted data is transmitted as meta-information at the upper layer of the dynamic three-dimensional map.

[0446] Moreover, the detection results of the surrounding conditions and the information detected by the host vehicle 600 can be presented to the user through a user interface. For example, these information are presented by overlapping them onto the navigation screen or the windshield.

[0447] Also, it is possible not to support autonomous driving. In a vehicle with cruise control, when a surrounding vehicle traveling in the autonomous driving mode is detected, the host vehicle can track that surrounding vehicle.

[0448] In addition, when a three-dimensional map cannot be obtained or the host vehicle 600 cannot estimate its own position due to excessive occlusion areas or other reasons, the host vehicle 600 can switch its operation mode from the autonomous driving mode to the tracking mode of the surrounding vehicle.

[0449] Moreover, the vehicle being tracked can issue a warning to the user about the being-tracked situation, and a user interface that allows the user to specify whether to allow tracking can also be installed. At this time, advertising displays can be designed on the tracking vehicle, and payment rewards or the like can be designed on the side being tracked.

[0450] Furthermore, although the transmitted information is basically the SWLD as three-dimensional data, it can also be information corresponding to the request settings set in the host vehicle 600 or the public settings of the vehicle traveling ahead. For example, the transmitted information can be the WLD of a dense point cloud, or the detection results of the surrounding conditions by the vehicle traveling ahead, or the braking information of the vehicle traveling ahead.

[0451] The host vehicle 600 receives the WLD, visualizes the three-dimensional data of the WLD, and presents the visualized three-dimensional data to the driver using the GUI. At this time, the host vehicle 600 can adopt a method that allows the user to distinguish the point cloud made by the host vehicle 600 from the received point cloud, and use different colors to represent the information for presentation.

[0452] When the host vehicle 600 presents the information detected by the host vehicle 600 and the detection results of the surrounding vehicle 601 to the driver using the GUI, the information is presented by color-coding or the like in a way that allows the user to distinguish the information detected by the host vehicle 600 from the received detection results.

[0453] As described above, in the three-dimensional data production device 620 according to this embodiment, the three-dimensional data production unit 621 produces the first three-dimensional data 632 based on the sensor information 631 detected by the sensor. The receiving unit 624 receives the encoded three-dimensional data 634 obtained by encoding the second three-dimensional data 635. The decoding unit 625 decodes the received encoded three-dimensional data 634 to obtain the second three-dimensional data 635. The synthesizing unit 626 produces the third three-dimensional data 636 by synthesizing the first three-dimensional data 632 and the second three-dimensional data 635.

[0454] Accordingly, the three-dimensional data production device 620 can produce the detailed third three-dimensional data 636 by using the produced first three-dimensional data 632 and the received second three-dimensional data 635.

[0455] Moreover, the synthesizing unit 626 can produce the third three-dimensional data 636 with a higher density than the first three-dimensional data 632 and the second three-dimensional data 635 by synthesizing the first three-dimensional data 632 and the second three-dimensional data 635.

[0456] Moreover, the second three-dimensional data 635 (e.g., SWLD) is three-dimensional data generated by extracting data with a feature amount equal to or greater than a threshold value from the fourth three-dimensional data (e.g., WLD).

[0457] Accordingly, the three-dimensional data production device 620 can reduce the data amount of the transmitted three-dimensional data.

[0458] Moreover, the three-dimensional data production device 620 further includes a search unit 623 that searches for the transmitting device that is the source of the encoded three-dimensional data 634. The receiving unit 624 receives the encoded three-dimensional data 634 from the searched transmitting device.

[0459] Accordingly, the three-dimensional data production device 620 can, for example, determine the transmitting device that holds the required three-dimensional data by searching.

[0460] Moreover, the three-dimensional data production device further includes a request range determination unit 622 that determines the request range, which is the range of the three-dimensional space for which the three-dimensional data is requested. The search unit 623 sends the request range information 633 indicating the request range to the transmitting device. The second three-dimensional data 635 includes the three-dimensional data within the request range.

[0461] Accordingly, the three-dimensional data production device 620 can not only receive the required three-dimensional data, but also reduce the data amount of the transmitted three-dimensional data.

[0462] Moreover, the request range determination unit 622 determines the spatial range including the occlusion area 604 that cannot be detected by the sensor as the request range.

[0463] Further, in the three-dimensional data transmission device 640 according to the present embodiment, the three-dimensional data creation unit 641 creates the fifth three-dimensional data 652 based on the sensor information 651 detected by the sensor. The extraction unit 643 creates the sixth three-dimensional data 654 by extracting a part of the fifth three-dimensional data 652. The encoding unit 644 generates the encoded three-dimensional data 634 by encoding the sixth three-dimensional data 654. The transmission unit 645 transmits the encoded three-dimensional data 634.

[0464] Accordingly, the three-dimensional data transmission device 640 can not only transmit the three-dimensional data created by itself to other devices, but also reduce the amount of data of the three-dimensional data to be transmitted.

[0465] Further, the three-dimensional data creation unit 641 creates the seventh three-dimensional data (e.g., WLD) based on the sensor information 651 detected by the sensor, and creates the fifth three-dimensional data 652 (e.g., SWLD) by extracting data with a feature amount equal to or greater than a threshold from the seventh three-dimensional data.

[0466] Accordingly, the three-dimensional data transmission device 640 can reduce the amount of data of the three-dimensional data to be transmitted.

[0467] Further, the three-dimensional data transmission device 640 further includes a reception unit 642 that receives request range information 633 indicating a request range from a receiving device, where the request range is the range of the three-dimensional space for which three-dimensional data is requested. The extraction unit 643 creates the sixth three-dimensional data 654 by extracting the three-dimensional data of the request range from the fifth three-dimensional data 652. The transmission unit 645 transmits the encoded three-dimensional data 634 to the receiving device.

[0468] Accordingly, the three-dimensional data transmission device 640 can reduce the amount of data of the three-dimensional data to be transmitted.

[0469] (Embodiment 4)

[0470] In the present embodiment, operations related to abnormal conditions in the self-position estimation based on the three-dimensional map will be described.

[0471] Applications such as autonomous driving of motor vehicles, robots, or flying objects such as drones, etc., will be expanded in the future. As an example of a method for realizing such autonomous movement, there is a method in which a moving object estimates its own position in a three-dimensional map (self-position estimation) and travels according to the map.

[0472] Self-position estimation is achieved by matching a three-dimensional map with three-dimensional information around the own vehicle (hereinafter referred to as own vehicle detection three-dimensional data) obtained by a distance measuring instrument (LIDAR, etc.) or a stereo camera mounted on the own vehicle, and estimating the position of the own vehicle in the three-dimensional map.

[0473] The three-dimensional map, such as the HD map proposed by HERE, etc., is not only a three-dimensional point cloud, but may also include two-dimensional map data such as the shape information of roads and intersections, or information such as traffic jams and accidents that changes in real time. The three-dimensional map is composed of multiple levels such as three-dimensional data, two-dimensional data, and metadata that changes in real time. The device can obtain only the required data, or can also refer to the required data.

[0474] The data of the point cloud can be the above-mentioned SWLD, or can also include point group data that is not feature points. And, the transmission and reception of the data of the point cloud are basically performed in one or more random access units.

[0475] As a method for matching the three-dimensional map and the three-dimensional data of the own vehicle detection, the following method can be adopted. For example, the device compares the shapes of the point groups in the respective point clouds, and determines the part with a high similarity between the feature points as the same position. And, when the three-dimensional map is composed of SWLD, the device compares the feature points constituting the SWLD with the three-dimensional feature points extracted from the three-dimensional data of the own vehicle detection and performs matching.

[0476] Here, in order to estimate the own position with high accuracy, the following (A) and (B) need to be satisfied. (A) The three-dimensional map and the three-dimensional data of the own vehicle detection can already be obtained. (B) Their accuracy satisfies a predetermined standard. However, in the following abnormal situations, (A) or (B) cannot be satisfied.

[0477] (1) The three-dimensional map cannot be obtained through the communication path.

[0478] (2) There is no three-dimensional map, or the obtained three-dimensional map is damaged.

[0479] (3) The sensor of the own vehicle fails, or due to bad weather, the generation accuracy of the three-dimensional data of the own vehicle detection is insufficient.

[0480] The following describes the actions for coping with these abnormal situations. Although the following describes the actions taking a vehicle as an example, the following method can also be applied to all moving objects that perform autonomous movement such as robots and drones.

[0481] The following will describe the configuration and actions of the three-dimensional information processing device according to the present embodiment for coping with abnormal situations in the three-dimensional map or the three-dimensional data of the own vehicle detection. Figure 34 It is a block diagram showing a configuration example of the three-dimensional information processing device 700 according to the present embodiment. Figure 35 It is a flowchart of the three-dimensional information processing method performed by the three-dimensional information processing device 700.

[0482] The three-dimensional information processing device 700 is mounted on a moving object such as a motor vehicle, for example. As Figure 34 shown, the three-dimensional information processing device 700 includes: a three-dimensional map acquisition unit 701, an own vehicle detection data acquisition unit 702, an abnormal situation determination unit 703, a response action determination unit 704, and an action control unit 705.

[0483] In addition, the three-dimensional information processing device 700 may include a camera that acquires a two-dimensional image, or may include a two-dimensional or one-dimensional sensor (not shown) such as a sensor that uses ultrasonic waves or lasers to detect a structural object or a moving object around the own vehicle. Further, the three-dimensional information processing device 700 may include a communication unit (not shown) that acquires a three-dimensional map through a mobile communication network such as 4G or 5G, vehicle-to-vehicle communication, or road-to-vehicle communication.

[0484] As Figure 35 shown, the three-dimensional map acquisition unit 701 acquires a three-dimensional map 711 near the driving route (S701). For example, the three-dimensional map acquisition unit 701 acquires the three-dimensional map 711 through a mobile communication network, vehicle-to-vehicle communication, or road-to-vehicle communication.

[0485] Next, the own vehicle detection data acquisition unit 702 acquires own vehicle detection three-dimensional data 712 based on sensor information (S702). For example, the own vehicle detection data acquisition unit 702 generates own vehicle detection three-dimensional data 712 based on the sensor information acquired by the sensors provided in the own vehicle.

[0486] Next, the abnormal situation determination unit 703 detects an abnormal situation by performing a pre-determined check on at least one of the acquired three-dimensional map 711 and own vehicle detection three-dimensional data 712 (S703). That is, the abnormal situation determination unit 703 determines whether at least one of the acquired three-dimensional map 711 and own vehicle detection three-dimensional data 712 is abnormal.

[0487] In step S703, when an abnormal situation is detected (Yes in S704), the response action determination unit 704 determines a response action for the abnormal situation (S705). Next, the action control unit 705 controls the operations of the various processing units required for implementing the response action, such as the three-dimensional map acquisition unit 701 (S706).

[0488] In addition, in step S703, when no abnormal situation is detected (No in S704), the three-dimensional information processing device 700 ends the process.

[0489] Furthermore, the three-dimensional information processing device 700 estimates the position of the vehicle equipped with the three-dimensional information processing device 700 by using the three-dimensional map 711 and the three-dimensional data of the vehicle itself detected by the device. Subsequently, the three-dimensional information processing device 700 uses the result of the self-position estimation to cause the vehicle to perform autonomous driving.

[0490] Accordingly, the three-dimensional information processing device 700 obtains map data (three-dimensional map 711) including the first three-dimensional position information via a channel. For example, the first three-dimensional position information is encoded in units of partial spaces having three-dimensional coordinate information, the first three-dimensional position information includes a plurality of random access units, each of the plurality of random access units is an aggregate of one or more partial spaces, and can be independently decoded. For example, the first three-dimensional position information is data (SWLD) in which feature points where the three-dimensional feature amount exceeds a specified threshold are encoded.

[0491] In addition, the three-dimensional information processing device 700 generates second three-dimensional position information (three-dimensional data of the vehicle itself detected by the device) based on the information detected by the sensor. Subsequently, the three-dimensional information processing device 700 determines whether the first three-dimensional position information or the second three-dimensional position information is abnormal by performing an abnormality determination process on the first three-dimensional position information or the second three-dimensional position information.

[0492] When the three-dimensional information processing device 700 determines that the first three-dimensional position information or the second three-dimensional position information is abnormal, it determines a response action for the abnormality. Subsequently, the three-dimensional information processing device 700 executes the control required for the implementation of the response action.

[0493] Accordingly, the three-dimensional information processing device 700 can detect an abnormality in the first three-dimensional position information or the second three-dimensional position information and can perform a response action.

[0494] The response action in the case of abnormality situation 1, that is, when the three-dimensional map 711 cannot be obtained via communication, will be described below.

[0495] The three-dimensional map 711 is required for self-position estimation. However, when the vehicle has not previously obtained the three-dimensional map 711 corresponding to the route to the destination, it is necessary to obtain the three-dimensional map 711 via communication. However, due to channel congestion or deterioration of the radio wave reception state, etc., there may be a situation where the vehicle cannot obtain the three-dimensional map 711 on the driving route.

[0496] The abnormality determination unit 703 confirms whether the three-dimensional map 711 has been obtained for all sections on the path to the destination or for sections within a specified range from the current position. If it cannot be obtained, it is determined as abnormal situation 1. That is, the abnormality determination unit 703 determines whether the three-dimensional map 711 (the first three-dimensional position information) can be obtained via the channel. If the three-dimensional map 711 cannot be obtained via the channel, the three-dimensional map 711 is determined to be abnormal.

[0497] In the case where it is determined as abnormal situation 1, the response action decision unit 704 selects one of the following two response actions: (1) continue with the self-position estimation, and (2) stop the self-position estimation.

[0498] First, an example of the response action in the case of (1) continuing with the self-position estimation will be described. When continuing with the self-position estimation, the three-dimensional map 711 on the path to the destination is required.

[0499] For example, the vehicle determines the places where the channel can be used within the range where the three-dimensional map 711 has been obtained, moves to those places, and obtains the three-dimensional map 711. At this time, the vehicle can obtain all the three-dimensional maps 711 to the destination, or can obtain the three-dimensional map 711 for each random access unit within the upper limit size that can be stored in the memory or HDD of the vehicle itself.

[0500] Alternatively, the following action can also be taken. The vehicle additionally obtains the communication status on the path. In the case where it can be predicted that the communication status on the path will deteriorate, before reaching the section with poor communication status, the three-dimensional map 711 of that section is obtained in advance, or the three-dimensional map 711 of the maximum range that can be obtained is obtained in advance. That is, the three-dimensional information processing device 700 predicts whether the vehicle will enter a region with poor communication status. When the three-dimensional information processing device 700 predicts that the vehicle will enter a region with poor communication status, before the vehicle enters that region, the three-dimensional map 711 is obtained.

[0501] Furthermore, it can also be that the vehicle determines the random access unit that constitutes the minimum three-dimensional map 711 required for self-position estimation on the path, which is narrower than the normal range, and receives the determined random access unit. That is, when the three-dimensional information processing device 700 cannot obtain the three-dimensional map 711 (the first three-dimensional position information) via the channel, it can obtain the third three-dimensional position information that is narrower than the range of the first three-dimensional position information via the channel.

[0502] In addition, when the vehicle cannot access the distribution server of the 3D map 711, it can obtain the 3D map 711 from a moving body such as another vehicle traveling around its own vehicle. In this case, the other vehicle or the like is a moving body that has obtained the 3D map 711 on the path to the destination and can communicate with its own vehicle.

[0503] Next, a specific example of the response action in the case of (2) stopping the self-position estimation will be described. In this case, the 3D map 711 on the path to the destination is not required.

[0504] For example, the vehicle notifies the driver of the situation where functions such as autonomous driving based on self-position estimation cannot be continued, and transfers the operation mode to a manual operation mode performed by the driver.

[0505] Normally, when performing self-position estimation, although there are cases where the level varies according to the presence of a person, autonomous driving is executed. In addition, the result of self-position estimation may also utilize navigation or the like when a person is driving. Therefore, the result of self-position estimation is not necessarily required for autonomous driving.

[0506] In addition, when the vehicle cannot use a normally used channel such as a mobile communication network such as 4G or 5G, it is confirmed whether the 3D map 711 can be obtained via another communication path such as Wi-Fi (registered trademark) or millimeter wave communication between the vehicle and the road, or vehicle-to-vehicle communication, and the used channel can be switched to a channel capable of obtaining the 3D map 711.

[0507] In addition, when the vehicle cannot obtain the 3D map 711, it can also obtain a 2D map and continue autonomous driving by using the 2D map and the three-dimensional data 712 detected by its own vehicle. That is, when the 3D information processing device 700 cannot obtain the 3D map 711 via the channel, it can obtain map data (2D map) including 2D position information via the channel, and perform self-position estimation of the vehicle by using the 2D position information and the three-dimensional data 712 detected by its own vehicle.

[0508] Specifically, the vehicle uses the 2D map and the three-dimensional data 712 detected by its own vehicle in self-position estimation, and uses the three-dimensional data 712 detected by its own vehicle in the detection of surrounding vehicles, pedestrians, and obstacles.

[0509] Here, map data such as HD maps can include not only a three-dimensional map 711 composed of three-dimensional point clouds or the like, but also two-dimensional map data (two-dimensional map), a simplified version of map data in which characteristic information such as road shapes or intersections is extracted from the two-dimensional map data, and metadata showing real-time information such as traffic jams, accidents, or construction. For example, the map data has a layer structure in which three-dimensional data (three-dimensional map 711), two-dimensional data (two-dimensional map), and metadata are sequentially arranged from the lower layer.

[0510] Here, compared with three-dimensional data, two-dimensional data is smaller in size. Therefore, even when the communication state is poor, the vehicle can obtain a two-dimensional map. And the vehicle can obtain a large range of two-dimensional maps uniformly in an area with good communication state. Therefore, when the channel state is poor and it is difficult to obtain the three-dimensional map 711, the vehicle can receive the layer including the two-dimensional map instead of receiving the three-dimensional map 711. In addition, since the data size of the metadata is small, for example, the vehicle can receive the metadata frequently without being affected by the communication state.

[0511] In the method of estimating the own position that uses a two-dimensional map and the own vehicle detection three-dimensional data 712, for example, there are the following two methods.

[0512] The first method is a method of performing matching of two-dimensional feature quantities. Specifically, the vehicle extracts two-dimensional feature quantities from the own vehicle detection three-dimensional data 712 and performs matching of the extracted two-dimensional feature quantities with the two-dimensional map.

[0513] For example, the vehicle projects the own vehicle detection three-dimensional data 712 onto the same plane as the two-dimensional map and performs matching of the obtained two-dimensional data with the two-dimensional map. The matching is performed using two-dimensional image feature quantities extracted from both.

[0514] When the three-dimensional map 711 includes SWLD, the three-dimensional map 711 can store both three-dimensional feature quantities at feature points in the three-dimensional space and two-dimensional feature quantities in the same plane as the two-dimensional map. For example, identification information is given to the two-dimensional feature quantities. Or, the two-dimensional feature quantities are stored in a layer different from the three-dimensional data and the two-dimensional map, and the vehicle obtains the data of the two-dimensional feature quantities while obtaining the two-dimensional map.

[0515] When the two-dimensional map represents information of positions at different heights (not in the same plane) from the ground, such as white lines, guardrails, and buildings, inside the road, in the same map, the vehicle extracts feature quantities from multiple height data of the own vehicle detection three-dimensional data 712.

[0516] And the information showing the correspondence between the feature points in the two-dimensional map and the feature points in the three-dimensional map 711 can be stored as meta-information of the map data.

[0517] The second method is a method of matching three-dimensional feature quantities. Specifically, the vehicle obtains three-dimensional feature quantities corresponding to the feature points of the two-dimensional map, and matches the obtained three-dimensional feature quantities with the three-dimensional feature quantities of its own vehicle detection three-dimensional data 712.

[0518] Specifically, the three-dimensional feature quantities corresponding to the feature points of the two-dimensional map are stored in the map data. When the vehicle obtains the two-dimensional map, it also obtains the three-dimensional feature quantities. In addition, when the three-dimensional map 711 includes SWLD, by giving the information of the feature points corresponding to the feature points of the two-dimensional map in the feature points for identifying SWLD, the vehicle can determine the three-dimensional feature quantities obtained together with the two-dimensional map according to the identification information. In addition, in this case, since it is only necessary to represent the two-dimensional position, the data amount can be reduced compared with the case of representing the three-dimensional position. Figure 1 In addition, in this case, since it is only necessary to represent the two-dimensional position, the data amount can be reduced compared with the case of representing the three-dimensional position.

[0519] Moreover, when estimating the vehicle's own position using the two-dimensional map, the accuracy of the own position estimation is lower than that of the three-dimensional map 711. Therefore, the vehicle determines whether to continue with autonomous driving even when the estimation accuracy is reduced, and only continues with autonomous driving when it is determined that it can continue.

[0520] Whether autonomous driving can continue is affected by the following factors, namely whether the road on which the vehicle is traveling is an urban area, or a road such as a highway where there are fewer other vehicles or pedestrians entering, or the driving environment such as the road width and the degree of chaos of the road (the density of vehicles or pedestrians). Moreover, markers for enabling sensors such as cameras to identify can be arranged in enterprise sites, streets, or buildings. In these specific areas, since the markers can be identified with high accuracy by two-dimensional sensors, for example, by including the position information of the markers in the two-dimensional map, high-accuracy own position estimation can be performed.

[0521] In addition, by including identification information indicating whether each area is a specific area in the map, the vehicle can determine whether the vehicle exists in a specific area. When the vehicle exists in a specific area, it is determined to continue autonomous driving. In this way, the vehicle can determine whether it can continue autonomous driving based on the accuracy of its own position estimation when using the two-dimensional map or the driving environment of the vehicle.

[0522] In this way, the three-dimensional information processing device 700 can determine whether to perform autonomous driving of the vehicle based on the driving environment (the moving environment of the moving body) using the result of estimating the vehicle's own position by using the two-dimensional map and its own vehicle detection three-dimensional data 712.

[0523] Also, it is possible that the vehicle does not determine whether it can continue with autonomous driving, but instead switches the level (mode) of autonomous driving according to the accuracy of its own position estimation or the driving environment of the vehicle. Here, the switching of the level (mode) of autonomous driving refers to, for example, restricting the speed, increasing the amount of operation by the driver (reducing the automatic level of autonomous driving), obtaining the driving information of the vehicle ahead and switching the mode with reference to this information, obtaining the driving information of the vehicle set to the same destination and switching the mode of autonomous driving using this information, and so on.

[0524] Moreover, the map may include information corresponding to the position information and showing the recommended level of autonomous driving in the case of estimating its own position using a two-dimensional map. The recommended level may be metadata that changes dynamically according to traffic volume and the like. Accordingly, the vehicle can determine the level not by sequentially judging according to the surrounding environment and the like, but by simply obtaining the information within the map. And by multiple vehicles referring to the same map, the level of autonomous driving of each vehicle can be kept stable. In addition, the recommended level may not be a recommended level, but a level that must be complied with.

[0525] Also, the vehicle can switch the level of autonomous driving according to whether there is a driver (whether it is manned or unmanned). For example, the vehicle reduces the level of autonomous driving when there is a person, and stops when there is no person. The vehicle judges the position where it can safely stop by recognizing the surrounding pedestrians, vehicles, and traffic signs. Or, the map may include position information showing the positions where the vehicle can safely stop, and the vehicle can refer to this position information to judge the positions where it can safely stop.

[0526] Next, the response actions in the case of abnormal situation 2, that is, when the three-dimensional map 711 does not exist or the obtained three-dimensional map 711 is damaged, will be described.

[0527] The abnormal situation determination unit 703 confirms which of the following (1) and (2) cases it belongs to, and determines it as abnormal situation 2 when it belongs to one of them. (1) The three-dimensional map 711 in part or all of the intervals on the path to the destination does not exist in the distribution server or the like that is the access destination and cannot be obtained, or (2) part or all of the obtained three-dimensional map 711 is damaged. That is, the abnormal situation determination unit 703 determines whether the data of the three-dimensional map 711 is complete, and determines the three-dimensional map 711 as abnormal when the data of the three-dimensional map 711 is incomplete.

[0528] In the case where it is determined as abnormal situation 2, the following response actions are performed. First, an example of the response action in the case where the three-dimensional map 711 cannot be obtained will be described.

[0529] For example, the vehicle sets a route that does not pass through an area without the 3D map 711.

[0530] Moreover, when the vehicle cannot set an alternative route because there is no alternative route or even if there is an alternative route, the distance will be greatly increased, etc., the vehicle sets a route that includes an area without the 3D map 711. And the vehicle notifies the driver to switch the driving mode in this area to switch the driving mode to the manual operation mode.

[0531] (2) When part or all of the obtained 3D map 711 is damaged, the following countermeasures are taken.

[0532] The vehicle determines the damaged part in the 3D map 711, requests the data of the damaged part through communication, obtains the data of the damaged part, and updates the 3D map 711 with the obtained data. At this time, the vehicle can specify the damaged part by position information such as absolute coordinates or relative coordinates in the 3D map 711, or can also be specified by the index number of the random access unit constituting the damaged part. In this case, the vehicle replaces the random access unit including the damaged part with the obtained random access unit.

[0533] Next, the countermeasure in the case of abnormal situation 3, that is, when the sensor of the own vehicle cannot generate the own vehicle detection three-dimensional data 712 due to a failure or bad weather, will be described.

[0534] The abnormal situation determination unit 703 confirms whether the generation error of the own vehicle detection three-dimensional data 712 is within the allowable range. If it is not within the allowable range, it is determined as abnormal situation 3. That is, the abnormal situation determination unit 703 determines whether the generation accuracy of the data of the own vehicle detection three-dimensional data 712 is above the reference value. When the generation accuracy of the data of the own vehicle detection three-dimensional data 712 is not above the reference value, the own vehicle detection three-dimensional data 712 is determined as abnormal.

[0535] As a method for confirming whether the generation error of the own vehicle detection three-dimensional data 712 is within the allowable range, the following method can be adopted.

[0536] According to the resolution in the depth direction and the scanning direction of the three-dimensional sensor of the own vehicle such as a rangefinder or a stereo camera, or the density of the point cloud that can be generated, etc., the spatial resolution of the own vehicle detection three-dimensional data 712 during normal operation is determined in advance. And the vehicle obtains the spatial resolution of the 3D map 711 based on the meta information included in the 3D map 711.

[0537] The vehicle uses the spatial resolutions of both to estimate a reference value of a matching error when matching the self-vehicle detection three-dimensional data 712 and the three-dimensional map 711 based on three-dimensional feature quantities or the like. As the matching error, statistics such as the error of the three-dimensional feature quantity of each feature point, the average value of the errors of the three-dimensional feature quantities between multiple feature points, or the error of the spatial distance between multiple feature points can be used. A permissible range of deviation from the reference value is set in advance.

[0538] When the matching error between the self-vehicle detection three-dimensional data 712 and the three-dimensional map 711 generated by the vehicle before starting to drive or during driving is not within the permissible range, it is determined as an abnormal situation 3.

[0539] Alternatively, the vehicle may use a test pattern having a known three-dimensional shape for accuracy inspection to obtain the self-vehicle detection three-dimensional data 712 for the test pattern such as before starting to drive, and determine whether it is an abnormal situation 3 based on whether the shape error is within the permissible range.

[0540] For example, the vehicle makes the above determination each time before starting to drive. Alternatively, the vehicle obtains the time-series change of the matching error by making the above determination at regular time intervals or the like during driving. When the matching error has an increasing tendency, even if the error is within the permissible range, it may be determined as an abnormal situation 3. And, based on the time-series change, when it is possible to predict an abnormality, the vehicle can notify the user of the predicted abnormal situation by displaying a message urging inspection or repair or the like. And, the vehicle can determine, through the time-series change, the abnormality caused by temporary reasons such as bad weather and the abnormality caused by a sensor failure, and notify only the user of the abnormality caused by the sensor failure.

[0541] And, when the vehicle determines that it is an abnormal situation 3, it selects or selectively executes any one of the following three response actions: (1) operating an emergency substitute sensor (rescue mode), (2) switching the operation mode, (3) performing action correction of the three-dimensional sensor.

[0542] First, (1) the case of operating an emergency substitute sensor will be described. The vehicle operates an emergency substitute sensor different from the three-dimensional sensor used during normal operation. That is, when the generation accuracy of the data of the self-vehicle detection three-dimensional data 712 by the three-dimensional information processing device 700 is not above the reference value, the self-vehicle detection three-dimensional data 712 (fourth three-dimensional position information) is generated based on information detected by a substitute different from the normal sensor.

[0543] Specifically, when the vehicle uses multiple cameras or LIDARs to obtain the three-dimensional data 712 of its own vehicle detection, the vehicle determines the malfunctioning sensor based on the direction in which the matching error of the three-dimensional data 712 of its own vehicle detection exceeds the allowable range. Then, the vehicle activates the substitute sensor corresponding to the malfunctioning sensor.

[0544] The substitute sensor can be a three-dimensional sensor, a camera that obtains a two-dimensional image, or a one-dimensional sensor such as an ultrasonic sensor. When the substitute sensor is a sensor other than a three-dimensional sensor, since there may be a decrease in the accuracy of self-position estimation or the inability to perform self-position estimation, the vehicle can switch the autonomous driving mode according to the type of the substitute sensor.

[0545] For example, when the substitute sensor is a three-dimensional sensor, the vehicle continues the autonomous driving mode. And when the substitute sensor is a two-dimensional sensor, the vehicle changes the operation mode from full autonomous driving to semi-autonomous driving premised on human operation. And when the substitute sensor is a one-dimensional sensor, the vehicle switches the operation mode to a manual operation mode where automatic braking control cannot be performed.

[0546] Also, the vehicle can switch the autonomous driving mode according to the driving environment. For example, when the substitute sensor is a two-dimensional sensor, if the vehicle is driving on a highway, it continues the full autonomous driving mode, and if it is driving in an urban area, it switches the operation mode to semi-autonomous driving.

[0547] Moreover, when there is no substitute sensor, if the normally operating sensors alone can obtain a sufficient number of feature points, the vehicle can continue self-position estimation. However, since detection in a specific direction cannot be performed, the vehicle switches the operation mode to semi-autonomous driving or manual operation mode.

[0548] Next, (2) the response actions for switching the operation mode will be described. The vehicle switches the operation mode from the autonomous driving mode to the manual operation mode. Or, the vehicle can continue autonomous driving until a safe stopping shoulder or the like and then stop. And after stopping, the vehicle can switch the operation mode to the manual operation mode. In this way, when the generation accuracy of the three-dimensional data 712 of its own vehicle detection by the three-dimensional information processing device 700 is not above the reference value, the autonomous driving mode is switched.

[0549] Next, (3) will describe the countermeasures for correcting the operation of the 3D sensor. The vehicle determines the malfunctioning 3D sensor based on the direction in which the matching error occurs, etc., and calibrates the determined sensor. Specifically, when using multiple LIDARs or cameras as sensors, a part of the 3D space reconstructed by each sensor overlaps. That is, the data of the overlapping part is obtained by multiple sensors. The 3D point cloud data obtained for the overlapping part is different between the normal sensor and the malfunctioning sensor. Therefore, the vehicle adjusts the operation of a predetermined part by performing origin correction of the LIDAR or exposure and focusing of the camera in such a way that the malfunctioning sensor can obtain 3D point cloud data equivalent to that of the normal sensor.

[0550] After the adjustment, if the matching error can be within the allowable range, the vehicle continues the previous operation mode. In addition, after the adjustment, if the matching accuracy cannot be within the allowable range, the vehicle performs the countermeasure of (1) operating the emergency replacement sensor or (2) switching the operation mode.

[0551] In this way, when the generation accuracy of the 3D data 712 detected by the own vehicle by the 3D information processing device 700 is not above the reference value, the operation of the sensor is corrected.

[0552] The following describes the method for selecting the countermeasures. The countermeasures can be selected by a user such as the driver, or can be automatically selected by the vehicle without passing through the user.

[0553] Moreover, the vehicle can also switch the control according to whether the driver is on board. For example, when the driver is on board, the vehicle gives priority to switching to the manual operation mode. In addition, when the driver is not on board, the vehicle gives priority to the mode of moving to a safe place and stopping.

[0554] The information indicating the stop place can be included in the 3D map 711 as meta-information. Or the vehicle can send a response request for the stop place to the service that manages the operation information of the autonomous driving, so as to obtain the information indicating the stop place.

[0555] Moreover, when the vehicle is running on a prescribed route, etc., the operation mode of the vehicle can be shifted to the mode in which the operation of the vehicle is managed by an operator via a channel. In particular, in a vehicle traveling in the fully autonomous driving mode, the risk of abnormality of the own position estimation function is high. Therefore, when an abnormality is detected or when the detected abnormality cannot be corrected, the vehicle notifies the service that manages the operation information of the occurrence of the abnormality via the channel. This service can notify the presence of the abnormal vehicle to the vehicles traveling in the vicinity of the vehicle, etc., or issue an instruction to vacate the nearby stop place.

[0556] Also, when detecting an abnormal situation, the vehicle can travel at a slower speed than normal.

[0557] When the vehicle is an autonomous vehicle providing a vehicle dispatching service such as a taxi and an abnormal situation occurs to the vehicle, the vehicle contacts the operation management center to stop at a safe location. Also, an alternative vehicle can be dispatched for the vehicle dispatching service. Or it can be that the user of the vehicle dispatching service drives the vehicle. In these situations, a discount on fees or awarding special points, etc. can be used in combination.

[0558] Also, in the method for dealing with abnormal situation 1, although the method for estimating its own position based on a two-dimensional map has been described, even under normal circumstances, a two-dimensional map can be used for estimating its own position. Figure 36 It is a flowchart of the own position estimation process in this case.

[0559] First, the vehicle obtains a three-dimensional map 711 near the driving path (S711). Next, the vehicle obtains its own vehicle detection three-dimensional data 712 based on the sensor information (S712).

[0560] Next, the vehicle determines whether a three-dimensional map 711 is required when estimating its own position (S713). Specifically, the vehicle determines whether a three-dimensional map 711 is required based on the accuracy of the own position estimation when using the two-dimensional map and the driving environment. For example, the same method as the method for dealing with abnormal situation 1 described above is adopted.

[0561] In the case where it is determined that the three-dimensional map 711 is not required (the "no" in S714), the vehicle obtains a two-dimensional map (S715). At this time, the vehicle can obtain the additional information described in the method for dealing with abnormal situation 1 at the same time. Also, the vehicle can generate a two-dimensional map based on the three-dimensional map 711. For example, the vehicle can extract an arbitrary plane from the three-dimensional map 711 to generate a two-dimensional map.

[0562] Next, the vehicle uses its own vehicle detection three-dimensional data 712 and the two-dimensional map to estimate its own position (S716). In addition, the method for estimating the own position using the two-dimensional map is, for example, the same as the method described in the method for dealing with abnormal situation 1 above.

[0563] In addition, in the case where it is determined that the three-dimensional map 711 is required (the "yes" in S714), the vehicle obtains the three-dimensional map 711 (S717). Next, the vehicle uses its own vehicle detection three-dimensional data 712 and the three-dimensional map 711 to estimate its own position (S718).

[0564] In addition, the vehicle can switch between basically using a two-dimensional map and a three-dimensional map 711 according to the corresponding speed of its own communication device or the condition of the channel. For example, while receiving the three-dimensional map 711, a communication speed required during driving is preset in advance. When the communication speed during driving is equal to or lower than the set value, the vehicle basically uses the two-dimensional map. When the communication speed during driving is greater than the set value, the vehicle basically uses the three-dimensional map 711. In addition, the vehicle can also basically use the two-dimensional map without making a switching judgment on whether to use the two-dimensional map or the three-dimensional map.

[0565] (Embodiment 5)

[0566] In this embodiment, a method for sending three-dimensional data to a following vehicle and the like will be described. Figure 37 It is a diagram showing an example of the object space of the three-dimensional data sent to a following vehicle or the like.

[0567] The vehicle 801 sends three-dimensional data such as a point cloud (point group) included in a rectangular parallelepiped space 802 with a width W, a height H, and a depth D at a distance L in front of the vehicle 801 to a traffic cloud monitor for monitoring the road condition or a following vehicle at a time interval of Δt.

[0568] When a vehicle or a person enters the space 802 or the like from the outside, and thus the three-dimensional data included in the space 802 that has been sent in the past has changed, the vehicle 801 also sends the three-dimensional data of the changed space.

[0569] In addition, although Figure 37 shows an example in which the shape of the space 802 is a rectangular parallelepiped, the space 802 only needs to include the space on the front road that is a blind spot from the perspective of the following vehicle, and it is not necessarily a rectangular parallelepiped.

[0570] The distance L is preferably set to a distance at which the following vehicle that has received the three-dimensional data can stop safely. For example, the distance L is set to the sum of the following distances: the distance that the following vehicle moves during the reception of the three-dimensional data, the distance that the following vehicle moves until it starts to decelerate according to the received data, and the distance required for the following vehicle to stop safely in principle. Since these distances change according to the speed, as shown by L = a×V + b (a and b are constants), the distance L can change according to the speed V of the vehicle.

[0571] The width W is set to a value that is at least larger than the width of the lane in which the vehicle 801 is traveling. More preferably, the width W is set to a size that includes adjacent spaces such as the left and right traffic lanes or the curb strip.

[0572] Although the depth D can be a fixed value, it can also vary according to the vehicle speed V as shown by D = c×V + d (where c and d are constants). Also, by setting D such that D > V×Δt, the transmission space can be made to repeat the space that has been transmitted in the past. Accordingly, the vehicle 801 can more reliably and without omission transmit the space on the driving route to the following vehicle or the like.

[0573] In this way, by limiting the three-dimensional data transmitted by the vehicle 801 to the space that is useful for the following vehicle, it is possible to effectively reduce the capacity of the transmitted three-dimensional data, achieving low latency and low cost of communication.

[0574] Next, the configuration of the three-dimensional data production device 810 according to the present embodiment will be described. Figure 38 It is a block diagram showing a configuration example of the three-dimensional data production device 810 according to the present embodiment. This three-dimensional data production device 810 is mounted on the vehicle 801, for example. The three-dimensional data production device 810 performs transmission and reception of three-dimensional data with external traffic cloud monitoring, the vehicle ahead, or the vehicle behind, and at the same time produces and stores the three-dimensional data.

[0575] The three-dimensional data production device 810 includes: a data reception unit 811, a communication unit 812, a reception control unit 813, a format conversion unit 814, a plurality of sensors 815, a three-dimensional data production unit 816, a three-dimensional data synthesis unit 817, a three-dimensional data storage unit 818, a communication unit 819, a transmission control unit 820, a format conversion unit 821, and a data transmission unit 822.

[0576] The data reception unit 811 receives three-dimensional data 831 from traffic cloud monitoring or the vehicle ahead. The three-dimensional data 831 includes information on areas that cannot be detected by the sensors 815 of the own vehicle, for example, such as point cloud, visible light image, depth information, sensor position information, or speed information.

[0577] The communication unit 812 communicates with traffic cloud monitoring or the vehicle ahead, and sends a data transmission request or the like to traffic cloud monitoring or the vehicle ahead.

[0578] The reception control unit 813 exchanges information such as the corresponding format with the communication partner via the communication unit 812 to establish communication with the communication partner.

[0579] The format conversion unit 814 generates three-dimensional data 832 by performing format conversion or the like on the three-dimensional data 831 received by the data reception unit 811. Also, when the three-dimensional data 831 is compressed or encoded, the format conversion unit 814 performs decompression or decoding processing.

[0580] A plurality of sensors 815 are a group of sensors such as LIDAR, visible light cameras, or infrared cameras that obtain information about the outside of the vehicle 801 and generate sensor information 833. For example, when the sensor 815 is a laser sensor such as LIDAR, the sensor information 833 is three-dimensional data such as point clouds (point group data). Additionally, there may not be a plurality of sensors 815.

[0581] The three-dimensional data creation unit 816 generates three-dimensional data 834 based on the sensor information 833. The three-dimensional data 834 includes, for example, information such as point clouds, visible light images, depth information, sensor position information, or speed information.

[0582] The three-dimensional data synthesis unit 817 synthesizes the three-dimensional data 832 created by traffic cloud monitoring or the vehicle ahead, etc., into the three-dimensional data 834 created based on the sensor information 833 of its own vehicle, thereby enabling the construction of three-dimensional data 835 that includes the space in front of the vehicle ahead that cannot be detected by the sensors 815 of its own vehicle.

[0583] The three-dimensional data storage unit 818 stores the generated three-dimensional data 835, etc.

[0584] The communication unit 819 communicates with traffic cloud monitoring or the vehicle behind, and sends data transmission requests, etc., to traffic cloud monitoring or the vehicle behind.

[0585] The transmission control unit 820 exchanges information such as the corresponding format, etc., with the communication partner via the communication unit 819 to establish communication with the communication partner. And the transmission control unit 820 determines the transmission area of the space of the three-dimensional data to be transmitted based on the three-dimensional data construction information of the three-dimensional data 832 generated in the three-dimensional data synthesis unit 817 and the data transmission request from the communication partner.

[0586] Specifically, the transmission control unit 820 determines the transmission area that includes the space in front of its own vehicle that cannot be detected by the sensors of the vehicle behind according to the data transmission request from traffic cloud monitoring or the vehicle behind. And the transmission control unit 820 determines the transmission area by judging, according to the three-dimensional data construction information, whether there is an update of the space that can be transmitted or the already transmitted space, etc. For example, the transmission control unit 820 determines the area that is both specified by the data transmission request and where the corresponding three-dimensional data 835 exists as the transmission area. And the transmission control unit 820 notifies the format conversion unit 821 of the format corresponding to the communication partner and the transmission area.

[0587] The format conversion unit 821 generates the three-dimensional data 837 by converting the three-dimensional data 836 of the transmission area in the three-dimensional data 835 stored in the three-dimensional data storage unit 818 into a format corresponding to the receiving side. Additionally, the format conversion unit 821 may compress or encode the three-dimensional data 837 to reduce the data volume.

[0588] The data transmission unit 822 transmits the three-dimensional data 837 to the traffic cloud monitoring or the following vehicle. The three-dimensional data 837 includes, for example, information on an area that is a blind spot for the following vehicle, such as point cloud, visible light image, depth information, or sensor position information in front of the host vehicle.

[0589] In addition, although the format conversion units 814 and 821 are taken as examples for format conversion and the like, format conversion may not be performed.

[0590] With this configuration, the three-dimensional data production device 810 obtains the three-dimensional data 831 of an area that cannot be detected by the sensor 815 of the host vehicle from the outside, and generates the three-dimensional data 835 by synthesizing the three-dimensional data 831 and the three-dimensional data 834 based on the sensor information 833 detected by the sensor 815 of the host vehicle. Accordingly, the three-dimensional data production device 810 can generate the three-dimensional data of a range that cannot be detected by the sensor 815 of the host vehicle.

[0591] Moreover, the three-dimensional data production device 810 can transmit the three-dimensional data of the space in front of the host vehicle, which cannot be detected by the sensors of the following vehicle, to the traffic cloud monitoring or the following vehicle, etc., in accordance with a data transmission request from the traffic cloud monitoring or the following vehicle.

[0592] Next, the transmission order of the three-dimensional data to the following vehicle in the three-dimensional data production device 810 will be described. Figure 39 It is a flowchart showing an example of the order of transmitting the three-dimensional data from the three-dimensional data production device 810 to the traffic cloud monitoring or the following vehicle.

[0593] First, the three-dimensional data production device 810 generates and updates the three-dimensional data 835 of the space 802 on the road in front of the host vehicle 801 (S801). Specifically, the three-dimensional data production device 810 synthesizes the three-dimensional data 831 produced by the traffic cloud monitoring or the vehicle in front, etc., into the three-dimensional data 834 produced based on the sensor information 833 of the host vehicle 801, and constructs the three-dimensional data 835 that also includes the space in front of the vehicle in front that cannot be detected by the sensor 815 of the host vehicle through this synthesis.

[0594] Next, the three-dimensional data production device 810 determines whether the three-dimensional data 835 included in the transmitted space has changed (S802).

[0595] When the three-dimensional data 835 included in the space changes due to a vehicle or a person entering the sent space from the outside, etc. (Yes in S802), the three-dimensional data production device 810 sends the three-dimensional data including the three-dimensional data 835 of the space where the change has occurred to the traffic cloud monitoring or the following vehicle (S803).

[0596] In addition, although the three-dimensional data production device 810 can send the three-dimensional data of the space where the change has occurred according to the transmission timing of sending the three-dimensional data at a prescribed interval, it can also be sent immediately after the change is detected. That is, the three-dimensional data production device 810 can give priority to sending the three-dimensional data of the space where the change has occurred over the three-dimensional data sent at a prescribed interval.

[0597] Moreover, the three-dimensional data production device 810 can also send all the three-dimensional data of the space where the change has occurred as the three-dimensional data of the space where the change has occurred, or can also send only the difference of the three-dimensional data (for example, information of three-dimensional points that appear or disappear, or displacement information of three-dimensional points, etc.).

[0598] Also, it can be that the three-dimensional data production device 810 sends metadata related to the danger avoidance action of its own vehicle, such as an emergency brake alarm, to the following vehicle before sending the three-dimensional data of the space where the change has occurred. Accordingly, the following vehicle can confirm in advance the emergency brake, etc. of the vehicle ahead, and thus can start danger avoidance actions such as decelerating as soon as possible.

[0599] When the three-dimensional data 835 included in the sent space has not changed (No in S802) or after step S803, the three-dimensional data production device 810 sends the three-dimensional data included in the space of a prescribed shape on the front distance L of its own vehicle 801 to the traffic cloud monitoring or the following vehicle (S804).

[0600] And, for example, the processes of steps S801 to S804 are repeatedly executed at a prescribed time interval.

[0601] Moreover, when there is no difference between the three-dimensional data 835 of the currently targeted space 802 and the three-dimensional map, the three-dimensional data production device 810 may not send the three-dimensional data 837 of the space 802.

[0602] Figure 40 It is a flowchart showing the operation of the three-dimensional data production device 810 in this case.

[0603] First, the three-dimensional data production device 810 generates and updates the three-dimensional data 835 of the space including the space 802 on the front road of its own vehicle 801 (S811).

[0604] Next, the three-dimensional data production device 810 determines whether there is an update in the three-dimensional data 835 of the generated space 802 that is different from the three-dimensional map (S812). That is, the three-dimensional data production device 810 determines whether there is a difference between the three-dimensional data 835 of the generated space 802 and the three-dimensional map. Here, the three-dimensional map is three-dimensional map information managed by devices on the infrastructure side such as traffic cloud monitoring. For example, this three-dimensional map can be obtained as the three-dimensional data 831.

[0605] In the case of an update (Yes in S812), the three-dimensional data production device 810, as described above, sends the three-dimensional data included in the space 802 to traffic cloud monitoring or the following vehicle (S813).

[0606] In addition, in the case of no update (No in S812), the three-dimensional data production device 810 does not send the three-dimensional data included in the space 802 to traffic cloud monitoring and the following vehicle (S814). In addition, the three-dimensional data production device 810 can be controlled so that the three-dimensional data of the space 802 is not sent by setting the volume of the space 802 to zero. Also, the three-dimensional data production device 810 can send information indicating that there is no update in the space 802 to traffic cloud monitoring or the following vehicle.

[0607] As described above, for example, when there are no obstacles on the road, there is no difference between the generated three-dimensional data 835 and the three-dimensional map on the infrastructure side, so data is not sent. In this way, it is possible to suppress the transmission of unnecessary data.

[0608] In addition, although the three-dimensional data production device 810 is mounted on a vehicle as an example in the above description, the three-dimensional data production device 810 is not limited to being mounted on a vehicle and can be mounted on any moving body.

[0609] As shown above, the three-dimensional data production device 810 according to this embodiment is mounted on a moving body equipped with a sensor 815 and a communication unit (such as a data receiving unit 811 or a data sending unit 822) that transceives three-dimensional data with the outside. The three-dimensional data production device 810 produces three-dimensional data 835 (second three-dimensional data) based on the sensor information 833 detected by the sensor 815 and the three-dimensional data 831 (first three-dimensional data) received by the data receiving unit 811. The three-dimensional data production device 810 sends three-dimensional data 837, which is part of the three-dimensional data 835, to the outside.

[0610] Accordingly, the three-dimensional data creation device 810 can generate three-dimensional data of a range that its own vehicle cannot detect. Also, the three-dimensional data creation device 810 can send the three-dimensional data of a range that other vehicles and the like cannot detect to such other vehicles.

[0611] Also, the three-dimensional data creation device 810 repeatedly performs the creation of the three-dimensional data 835 and the sending of the three-dimensional data 837 at a prescribed interval. The three-dimensional data 837 is the three-dimensional data of a small space 802 having a prescribed size located at a prescribed distance L in the forward direction of the moving direction of the vehicle 801 from the position of the current vehicle 801.

[0612] Accordingly, since the range of the sent three-dimensional data 837 is restricted, the data amount of the sent three-dimensional data 837 can be reduced.

[0613] Also, the prescribed distance L changes according to the moving speed V of the vehicle 801. For example, the greater the moving speed V, the longer the prescribed distance L. Accordingly, the vehicle 801 can set an appropriate small space 802 according to the moving speed V of the vehicle 801 and can send the three-dimensional data 837 of the small space 802 to a following vehicle or the like.

[0614] Also, the prescribed size changes according to the moving speed V of the vehicle 801. For example, the greater the moving speed V, the larger the prescribed size. For example, the greater the moving speed V, the greater the depth D which is the length in the moving direction of the vehicle of the small space 802. Accordingly, the vehicle 801 can set an appropriate small space 802 according to the moving speed V of the vehicle 801 and can send the three-dimensional data 837 of the small space 802 to a following vehicle or the like.

[0615] Also, the three-dimensional data creation device 810 determines whether there is a change in the three-dimensional data 835 of the small space 802 corresponding to the sent three-dimensional data 837. When the three-dimensional data creation device 810 determines that there is a change, it sends the three-dimensional data 837 (the fourth three-dimensional data) which is at least a part of the three-dimensional data 835 as the change to an external following vehicle or the like.

[0616] Accordingly, the vehicle 801 can send the three-dimensional data 837 of the space where a change has occurred to a following vehicle or the like.

[0617] Further, the three-dimensional data creation device 810 transmits the changed three-dimensional data 837 (the fourth three-dimensional data) preferentially over the normal three-dimensional data 837 (the third three-dimensional data) that is regularly transmitted. Specifically, the three-dimensional data creation device 810 transmits the changed three-dimensional data 837 (the fourth three-dimensional data) before transmitting the normal three-dimensional data 837 (the third three-dimensional data) that is regularly transmitted. That is, the three-dimensional data creation device 810 does not wait for the transmission of the normal three-dimensional data 837 that is regularly transmitted, but transmits the changed three-dimensional data 837 (the fourth three-dimensional data) irregularly.

[0618] Accordingly, since the vehicle 801 can preferentially transmit the three-dimensional data 837 of the changed space to the following vehicle or the like, the following vehicle or the like can make a quick judgment based on the three-dimensional data.

[0619] Further, the changed three-dimensional data 837 (the fourth three-dimensional data) shows the difference between the three-dimensional data 835 of the small space 802 corresponding to the transmitted three-dimensional data 837 and the changed three-dimensional data 835. Accordingly, the data volume of the transmitted three-dimensional data 837 can be reduced.

[0620] Further, when there is no difference between the three-dimensional data 837 of the small space 802 and the three-dimensional data 831 of the small space 802, the three-dimensional data creation device 810 does not transmit the three-dimensional data 837 of the small space 802. Moreover, the three-dimensional data creation device 810 may also transmit information indicating that there is no difference between the three-dimensional data 837 of the small space 802 and the three-dimensional data 831 of the small space 802 to the outside.

[0621] Accordingly, since the transmission of unnecessary three-dimensional data 837 can be suppressed, the data volume of the transmitted three-dimensional data 837 can be reduced.

[0622] (Embodiment 6)

[0623] In this embodiment, a display device and a display method for displaying information obtained from a three-dimensional map or the like, and a storage device and a storage method for the three-dimensional map or the like will be described.

[0624] A moving body such as a vehicle or a robot flexibly uses a three-dimensional map obtained through communication with a server or another vehicle and two-dimensional images obtained from sensors mounted on its own vehicle or its own vehicle detection three-dimensional data for autonomous driving of the vehicle or autonomous movement of the robot. The data that the user wants to view or save among these data varies according to the situation. A display device that switches the display according to different situations will be described below.

[0625] Figure 41It is a flowchart showing an outline of a display method of a display device. The display device is mounted on a moving body such as a vehicle or a robot. In addition, hereinafter, an example in which the moving body is a vehicle (motor vehicle) will be described.

[0626] First, the display device determines which of two-dimensional surrounding information and three-dimensional surrounding information to display according to the driving condition of the vehicle (S901). In addition, the two-dimensional surrounding information corresponds to the first surrounding information in the technical solution, and the three-dimensional surrounding information corresponds to the second surrounding information in the technical solution. Here, the surrounding information is information showing the surroundings of the moving body, for example, an image when looking in a specified direction from the vehicle, or a map of the surroundings of the vehicle.

[0627] The two-dimensional surrounding information is information generated using two-dimensional data. Here, the two-dimensional data refers to two-dimensional map information or an image. For example, the two-dimensional surrounding information is a map of the vehicle surroundings obtained from a two-dimensional map, or an image obtained by a camera mounted on the vehicle. And, for example, the two-dimensional surrounding information does not include three-dimensional information. That is, in the case where the two-dimensional surrounding information is a map of the vehicle surroundings, the map does not include information in the height direction. And, in the case where the two-dimensional surrounding information is an image obtained by a camera, the image does not include information in the depth direction.

[0628] And, the three-dimensional surrounding information is information generated using three-dimensional data. Here, the three-dimensional data is, for example, a three-dimensional map. In addition, the three-dimensional data may be information showing the three-dimensional position or three-dimensional shape of an object around the vehicle obtained from another vehicle or a server, or detected by the own vehicle, etc. For example, the three-dimensional surrounding information is a two-dimensional or three-dimensional image or map of the vehicle surroundings generated using a three-dimensional map. And, for example, the three-dimensional surrounding information includes three-dimensional information. For example, in the case where the three-dimensional surrounding information is an image in front of the vehicle, the image includes information showing the distance to an object in the image. And, in this image, for example, a pedestrian blocked by a vehicle in front is displayed. And, in the three-dimensional surrounding information, information showing these distances or pedestrians, etc. may also be overlaid on an image obtained by a sensor mounted on the vehicle. And, in the three-dimensional surrounding information, information in the height direction may also be overlaid on a two-dimensional map.

[0629] And, the three-dimensional data can be three-dimensionally displayed, and a two-dimensional image or two-dimensional map obtained from the three-dimensional data can also be displayed on a two-dimensional display or the like.

[0630] In step S901, when it is determined to display three-dimensional surrounding information (Yes in S902), the display device displays the three-dimensional surrounding information (S903). Also, in step S901, when it is determined to display two-dimensional surrounding information (No in S902), the display device displays the two-dimensional surrounding information (S904). In this way, the display device displays the three-dimensional surrounding information or the two-dimensional surrounding information to be displayed determined in step S901.

[0631] A specific example will be described below. In the first example, the display device switches the surrounding information to be displayed according to whether the vehicle is in autonomous driving or manual driving. Specifically, in autonomous driving, since the driver does not need to know in detail the surrounding road information, the display device displays two-dimensional surrounding information (e.g., a two-dimensional map). In addition, in manual driving, three-dimensional surrounding information (e.g., a three-dimensional map) is displayed in order to drive safely so that the detailed situation of the surrounding road information can be known.

[0632] Also, in autonomous driving, since the information based on which the own vehicle drives is displayed to the user, the display device can display the information that affects the driving operation (e.g., SWLD, lanes, road signs, and surrounding condition detection results used in the own position estimation). For example, the display device can add this information to the two-dimensional map.

[0633] In addition, the surrounding information displayed during autonomous driving and manual driving described above is only an example. The display device can also display three-dimensional surrounding information during autonomous driving and two-dimensional surrounding information during manual driving. Also, in at least one of autonomous driving and manual driving, in addition to the two-dimensional or three-dimensional map or image, the display device can also display metadata or surrounding condition detection results, and can display metadata or surrounding condition detection results instead of the two-dimensional or three-dimensional map or image. Here, the metadata refers to information showing the three-dimensional position or three-dimensional shape of an object obtained from a server or other vehicle. And the surrounding condition detection result is information showing the three-dimensional position or three-dimensional shape of an object detected by the own vehicle.

[0634] In the second example, the display device switches the surrounding information to be displayed according to the driving environment. For example, the display device switches the surrounding information to be displayed according to the external brightness. Specifically, when the area around the own vehicle is bright, the display device displays a two-dimensional image obtained by a camera mounted on the own vehicle, or displays three-dimensional surrounding information created using the two-dimensional image. In addition, when the area around the own vehicle is dark, since the two-dimensional image obtained by the camera mounted on the own vehicle is dark and not easy to view, the display device displays three-dimensional surrounding information created using lidar or millimeter-wave radar.

[0635] In addition, the display device can also switch the surrounding information to be displayed according to the current area where the vehicle is located, i.e., the driving area. For example, the display device can display three-dimensional surrounding information in such a way that it can provide information about surrounding buildings, etc. in tourist areas, city centers, or near destinations to the user. Additionally, considering that detailed surrounding information is mostly not required in mountainous or suburban areas, etc., the display device can display two-dimensional surrounding information.

[0636] In addition, the display device can also switch the surrounding information to be displayed according to the weather condition. For example, in sunny weather, the display device displays three-dimensional surrounding information created using a camera or lidar. Additionally, in rainy or foggy weather, since the three-dimensional surrounding information created using a camera or lidar is likely to contain noise, the display device displays three-dimensional surrounding information created using a millimeter-wave radar.

[0637] In addition, these display switches can be executed automatically by the system or manually by the user.

[0638] In addition, the three-dimensional surrounding information is generated based on one or more of the following data, which are: dense point cloud data generated according to WLD, grid data generated according to MWLD, sparse data generated according to SWLD, lane data generated according to the lane world space, two-dimensional map data including three-dimensional shape information of roads and intersections, etc., and metadata or own vehicle detection results including three-dimensional position or three-dimensional shape information that changes according to the actual time.

[0639] In addition, the above-mentioned WLD is three-dimensional point cloud data, and SWLD is data obtained by extracting point clouds with a feature amount above a threshold from WLD. And MWLD is data having a grid structure generated from WLD. The lane world space is data obtained by extracting point clouds with a feature amount above a threshold and required for self-position estimation, driving support, or autonomous driving, etc. from WLD.

[0640] Here, compared with WLD, MWLD and SWLD have less data volume. Therefore, in cases where more detailed data is required, WLD can be used, and in other cases, by using MWLD or SWLD, the communication data volume and processing volume can be appropriately reduced. And compared with SWLD, the lane world space has less data volume. Therefore, by using the lane world space, the communication data volume and processing volume can be further reduced.

[0641] Also, although the above has been described by taking the switching between two-dimensional surrounding information and three-dimensional surrounding information as an example, the display device may also switch the types of data (WLD, SWLD, etc.) used in the generation of three-dimensional surrounding information according to the above conditions. That is, in the above description, in the example where the display device displays three-dimensional surrounding information, the three-dimensional surrounding information generated based on the first data (e.g., WLD or SWLD) with a larger data volume can be displayed. In the example where two-dimensional surrounding information is displayed, the three-dimensional surrounding information generated not based on the two-dimensional surrounding information but based on the second data (e.g., SWLD or lane world space), which is data with a smaller data volume than the above first data, can be displayed.

[0642] Also, the display device displays two-dimensional surrounding information or three-dimensional surrounding information, for example, on a two-dimensional display, a head-up display, or a head-mounted display mounted on its own vehicle. Alternatively, the display device may transmit the two-dimensional surrounding information or three-dimensional surrounding information to a mobile terminal such as a smart phone via wireless communication and display it. That is, the display device is not limited to being mounted on a moving body and can operate in cooperation with the moving body. For example, when a user holding a display device such as a smart phone rides in a moving body or drives a moving body, the information of the moving body such as the position of the moving body estimated based on the own position of the moving body is displayed on the display device, or these information are displayed on the display device together with the surrounding information.

[0643] Also, when the display device displays a three-dimensional map, it can render the three-dimensional map and display it as two-dimensional data, or it can use a three-dimensional display or three-dimensional holography to display it as three-dimensional data.

[0644] Next, a method for saving a three-dimensional map will be described. Since a moving body such as a vehicle or a robot is for the autonomous driving of the vehicle or the autonomous movement of the robot, three-dimensional maps obtained from communication with a server or other vehicles, two-dimensional images obtained from sensors mounted on its own vehicle, or three-dimensional data detected by its own vehicle are utilized. It can be considered that among these data, the data that the user wants to view or save will vary depending on the situation. The following describes a method for saving data corresponding to the situation.

[0645] The saving device is mounted on a moving body such as a vehicle or a robot. In addition, the following takes the example of the moving body being a vehicle (motor vehicle). Also, the saving device may be included in the above display device.

[0646] In the first example, the storage device determines whether to store the 3D map according to the region. Here, by storing the 3D map in the storage medium of its own vehicle, autonomous driving can be performed without communicating with the server within the stored space. However, since the storage capacity is limited, only limited data can be stored. Therefore, the storage device is limited to the storage regions shown below.

[0647] For example, the storage device preferentially stores the 3D maps of regions frequently passed through, such as the commuting route or the area around the user's home. Accordingly, data of frequently used regions do not need to be acquired each time, thus effectively reducing the communication data volume. Additionally, preferential storage means storing data with a higher priority within a pre-determined storage capacity. For example, when new data cannot be stored within the storage capacity, data with a lower priority than the new data is deleted.

[0648] Moreover, the storage device preferentially stores the 3D maps of regions with poor communication environments. Accordingly, in regions with poor communication environments, since data does not need to be obtained through communication, the occurrence of a situation where the 3D map cannot be obtained due to poor communication can be suppressed.

[0649] Alternatively, the storage device preferentially stores the 3D maps of regions with heavy traffic. Accordingly, the 3D maps of accident-prone regions can be preferentially stored. Thus, in such regions, a reduction in the accuracy of autonomous driving or driving support caused by the inability to obtain the 3D map due to poor communication can be suppressed.

[0650] Alternatively, the storage device preferentially stores the 3D maps of regions with light traffic. Here, in regions with light traffic, the possibility of not being able to use the autonomous driving mode of automatically following the vehicle ahead increases. Accordingly, there may be a situation where more detailed surrounding information is required. Therefore, by preferentially storing the 3D maps of regions with light traffic, the accuracy of autonomous driving or driving support in such regions can be improved.

[0651] In addition, the above-mentioned multiple storage methods can be combined. And the regions for preferentially storing these 3D maps can be automatically determined by the system or specified by the user.

[0652] Moreover, the storage device can delete the 3D maps after a specified period has elapsed since storage, or update them to the latest data. Accordingly, old map data will not be used. And when the storage device updates the map data, by comparing the old map with the new map, a differential space region with differences, i.e., a differential region, is detected. By adding the data of the differential region of the new map to the old map, or removing the data of the differential region from the old map, only the data of the changed region can be updated in this way.

[0653] Also, in this example, the saved three-dimensional map is used for autonomous driving. Therefore, by using the SWLD as this three-dimensional map, the amount of communication data and the like can be reduced. In addition, the three-dimensional map is not limited to the SWLD and can also be other types of data such as the WLD.

[0654] In the second example, the saving device saves the three-dimensional map according to an event.

[0655] For example, the saving device saves a special event encountered during driving as a three-dimensional map. Accordingly, the user can view the details of the event afterwards. Examples of events saved as three-dimensional maps are shown below. In addition, the saving device can also save three-dimensional surrounding information generated from the three-dimensional map.

[0656] For example, the saving device saves the three-dimensional map before and after a collision accident or when danger is sensed, etc.

[0657] Alternatively, the saving device stores three-dimensional maps of characteristic scenes such as beautiful scenery, places where people gather, or tourist attractions.

[0658] These events to be saved can be automatically determined by the system or specified in advance by the user. For example, as a method for judging these events, machine learning can be adopted.

[0659] Also, in this example, the saved three-dimensional map is used for viewing. Therefore, by using the WLD as this three-dimensional map, high-quality images can be provided. In addition, the three-dimensional map is not limited to the WLD and can also be other types of data such as the SWLD.

[0660] The method for the display device to control the display according to the user is described below. When the display device overlaps the surrounding condition detection result obtained through vehicle-to-vehicle communication onto the map and displays it, by representing the surrounding vehicles with wireframes or adding transparency to the surrounding vehicles, etc., the detected objects blocked by the surrounding vehicles can be seen. Alternatively, the display device can display an image seen from an overhead viewpoint, or enable the own vehicle, surrounding vehicles, and the surrounding condition detection result to be seen from a top view.

[0661] When using a head-up display to overlap the surrounding condition detection result or point cloud data onto the surrounding environment seen through the windshield as shown, due to differences in the user's posture, body type, or eye position, the position of the overlapping information will deviate. Figure 42 This is a diagram showing a display example of the head-up display when the position has deviated. Figure 43

[0662] ​To correct such deviation, the display device uses information from a camera inside the vehicle or sensors mounted on the seat to detect the user's posture, body shape, or eye position. The display device adjusts the position of the overlapping information according to the detected user's posture, body shape, or eye position. Figure 44 FIG. is a diagram showing a display example of the adjusted head-up display.

[0663] In addition, the adjustment of such overlapping position can be manually performed by the user using a control device mounted on the vehicle.

[0664] Moreover, the display device can display a safe place during a disaster on a map and prompt it to the user. Alternatively, the vehicle can convey the disaster content and a message to move to a safe place to the user and perform autonomous driving until it reaches a safe place.

[0665] For example, the vehicle sets a high-altitude area that will not be engulfed by a tsunami during an earthquake as the destination. At this time, the vehicle can obtain road information that is difficult to pass due to the earthquake through communication with the server and process it according to disaster content such as a route that avoids such roads.

[0666] Furthermore, the autonomous driving can include multiple modes such as a moving mode and a long-distance driving mode.

[0667] In the moving mode, the vehicle decides the route to the destination considering factors such as early time, low cost, short driving distance, short running distance, and low fuel consumption, and performs autonomous driving according to the decided route.

[0668] In the long-distance driving mode, the vehicle automatically decides the route in such a way as to reach the destination at the user-specified time. For example, when the user sets the destination and the arrival time, the vehicle decides the route in such a way that it can perform sightseeing around the area and reach the destination at the set time, and performs autonomous driving according to the decided route.

[0669] (Embodiment 7)

[0670] In the example to be described in Embodiment 5, a client device such as a vehicle sends three-dimensional data to other vehicles or a server such as a traffic cloud monitor. In the present embodiment, the client device sends sensor information obtained by sensors to the server or other client devices.

[0671] First, the configuration of the system according to the present embodiment will be described. Figure 45 FIG. is a diagram showing the configuration of a three-dimensional map and a sensor information transceiver system according to the present embodiment. The system includes a server 901, client devices 902A and 902B. In addition, without particularly distinguishing between the client devices 902A and 902B, they are also denoted as the client device 902.

[0672] The client device 902 is, for example, an in-vehicle device mounted on a moving body such as a vehicle. The server 901 is, for example, a traffic cloud monitor or the like and can communicate with a plurality of client devices 902.

[0673] The server 901 sends a three-dimensional map composed of point clouds to the client device 902. In addition, the composition of the three-dimensional map is not limited to point clouds and can also be represented by other three-dimensional data such as a mesh structure.

[0674] The client device 902 sends sensor information obtained by the client device 902 to the server 901. The sensor information includes, for example, at least one of LIDAR acquisition information, visible light image, infrared image, depth image, sensor position information, and speed information.

[0675] Regarding the data transmitted and received between the server 901 and the client device 902, it can be compressed when reducing data is desired, and can be not compressed when maintaining the accuracy of the data is desired. When compressing the data, for example, a three-dimensional compression method based on an octree can be adopted in the point cloud. And, a two-dimensional image compression method can be adopted in the visible light image, infrared image, and depth image. The two-dimensional image compression method is, for example, MPEG-4 AVC or HEVC standardized by MPEG.

[0676] In addition, the server 901 sends the three-dimensional map managed by the server 901 to the client device 902 in accordance with a transmission request for the three-dimensional map from the client device 902. In addition, the server 901 may also send the three-dimensional map without waiting for a transmission request for the three-dimensional map from the client device 902. For example, the server 901 may broadcast the three-dimensional map to one or more client devices 902 in a pre-specified space. And, the server 901 may also send a three-dimensional map suitable for the position of the client device 902 to the client device 902 that has received a transmission request once at regular intervals. And, the server 901 may also send a three-dimensional map to the client device 902 whenever the three-dimensional map managed by the server 901 is updated.

[0677] The client device 902 issues a transmission request for the three-dimensional map to the server 901. For example, when the client device 902 wants to estimate its own position while driving, the client device 902 sends a transmission request for the three-dimensional map to the server 901.

[0678] In addition, the client device 902 may also send a request to the server 901 to send a 3D map under the following circumstances. When the 3D map held by the client device 902 is relatively old, the client device 902 may also send a request to the server 901 to send a 3D map. For example, when a certain period has passed since the client device 902 obtained the 3D map, the client device 902 may also send a request to the server 901 to send a 3D map.

[0679] It may also be that the client device 902 sends a request to the server 901 to send a 3D map a certain moment before the client device 902 is about to leave the space shown in the 3D map held by the client device 902. For example, it may also be that when the client device 902 is within a pre-specified distance from the boundary of the space shown in the 3D map held by the client device 902, the client device 902 sends a request to the server 901 to send a 3D map. Moreover, when the movement path and movement speed of the client device 902 are known, the moment when the client device 902 leaves the space shown in the 3D map held by the client device 902 can be predicted based on the known movement path and movement speed.

[0680] When the error in the position comparison between the 3D data generated by the client device 902 based on sensor information and the 3D map is above a certain range, the client device 902 may send a request to the server 901 to send a 3D map.

[0681] The client device 902 sends the sensor information to the server 901 in accordance with the request from the server 901 to send the sensor information. In addition, the client device 902 may also send the sensor information to the server 901 without waiting for the request from the server 901 to send the sensor information. For example, when the client device 902 has received a request from the server 901 to send the sensor information once, it may regularly send the sensor information to the server 901 within a certain period. It may also be that when the error in the position comparison between the 3D data generated by the client device 902 based on sensor information and the 3D map obtained from the server 901 is above a certain range, the client device 902 determines that there is a possibility that the 3D map around the client device 902 has changed, and sends this judgment result together with the sensor information to the server 901.

[0682] The server 901 sends a request to the client device 902 to send sensor information. For example, the server 901 receives the location information of the client device 902 such as GPS from the client device 902. When the server 901 determines, based on the location information of the client device 902, that the client device 902 is approaching a space with less information in the three-dimensional map managed by the server 901, in order to regenerate the three-dimensional map, the server 901 sends a request to the client device 902 to send sensor information. Also, the server 901 may send a request to send sensor information when it wants to update the three-dimensional map, when it wants to confirm road conditions such as during snow accumulation or disasters, or when it wants to confirm congestion conditions or accident conditions.

[0683] Also, the client device 902 may set the amount of sensor information to be sent to the server 901 according to the communication state or frequency band at the time of receiving the request to send sensor information from the server 901. Setting the amount of sensor information to be sent to the server 901, for example, means increasing or decreasing the data itself, or selecting an appropriate compression method.

[0684] Figure 46 It is a block diagram showing a configuration example of the client device 902. The client device 902 receives a three-dimensional map composed of point clouds, etc. from the server 901, and estimates its own position based on the three-dimensional data created from the sensor information of the client device 902. Then, the client device 902 sends the acquired sensor information to the server 901.

[0685] The client device 902 includes: a data receiving unit 1011, a communication unit 1012, a receiving control unit 1013, a format conversion unit 1014, a plurality of sensors 1015, a three-dimensional data creation unit 1016, a three-dimensional image processing unit 1017, a three-dimensional data storage unit 1018, a format conversion unit 1019, a communication unit 1020, a sending control unit 1021, and a data sending unit 1022.

[0686] The data receiving unit 1011 receives the three-dimensional map 1031 from the server 901. The three-dimensional map 1031 is data including point clouds such as WLD or SWLD. The three-dimensional map 1031 may include either compressed data or uncompressed data.

[0687] The communication unit 1012 communicates with the server 901 and sends a data sending request (for example, a request to send a three-dimensional map) to the server 901.

[0688] The receiving control unit 1013 exchanges information such as the corresponding format with the communication partner via the communication unit 1012 to establish communication with the communication partner.

[0689] The format conversion unit 1014 generates a 3D map 1032 by performing format conversion and the like on the 3D map 1031 received by the data reception unit 1011. Moreover, when the 3D map 1031 is compressed or encoded, the format conversion unit 1014 performs decompression or decoding processing. In addition, when the 3D map 1031 is uncompressed data, the format conversion unit 1014 does not perform decompression or decoding processing.

[0690] The multiple sensors 1015 are a group of sensors mounted on the client device 902 such as a LIDAR, a visible light camera, an infrared camera, or a depth sensor, which are used to obtain information about the outside of the vehicle, and generate sensor information 1033. For example, when the sensor 1015 is a laser sensor such as a LIDAR, the sensor information 1033 is 3D data such as a point cloud (point group data). In addition, the sensor 1015 may not be multiple.

[0691] The 3D data production unit 1016 produces 3D data 1034 around its own vehicle according to the sensor information 1033. For example, the 3D data production unit 1016 uses the information obtained by the LIDAR and the visible light image obtained by the visible light camera to produce point cloud data with color information around its own vehicle.

[0692] The 3D image processing unit 1017 uses the received 3D map 1032 such as a point cloud and the 3D data 1034 around its own vehicle generated according to the sensor information 1033 to perform its own position estimation processing and the like of its own vehicle. In addition, it may be that the 3D image processing unit 1017 synthesizes the 3D map 1032 and the 3D data 1034 to produce 3D data 1035 around its own vehicle, and uses the produced 3D data 1035 to perform its own position estimation processing.

[0693] The 3D data storage unit 1018 stores the 3D map 1032, the 3D data 1034, the 3D data 1035, etc.

[0694] The format conversion unit 1019 generates sensor information 1037 by converting the sensor information 1033 into a format corresponding to the receiving side. In addition, the format conversion unit 1019 can reduce the data volume by compressing or encoding the sensor information 1037. Moreover, when format conversion is not required, the format conversion unit 1019 can omit the processing. And the format conversion unit 1019 can control the data volume of the data transmitted according to the specified transmission range.

[0695] The communication unit 1020 communicates with the server 901 and receives a data transmission request (a transmission request for sensor information) and the like from the server 901.

[0696] The transmission control unit 1021 exchanges information such as corresponding formats with the communication partner via the communication unit 1020 to establish communication.

[0697] The data transmission unit 1022 transmits the sensor information 1037 to the server 901. The sensor information 1037 includes, for example, information obtained by LIDAR, luminance images (visible light images) obtained by visible light cameras, infrared images obtained by infrared cameras, depth images obtained by depth sensors, sensor position information, and speed information, etc., which are obtained by multiple sensors 1015.

[0698] Next, the configuration of the server 901 will be described. Figure 47 It is a block diagram showing a configuration example of the server 901. The server 901 receives the sensor information sent from the client device 902 and creates three-dimensional data based on the received sensor information. The server 901 updates the three-dimensional map managed by the server 901 using the created three-dimensional data. And the server 901 sends the updated three-dimensional map to the client device 902 according to the transmission request of the three-dimensional map from the client device 902.

[0699] The server 901 includes: a data reception unit 1111, a communication unit 1112, a reception control unit 1113, a format conversion unit 1114, a three-dimensional data creation unit 1116, a three-dimensional data synthesis unit 1117, a three-dimensional data storage unit 1118, a format conversion unit 1119, a communication unit 1120, a transmission control unit 1121, and a data transmission unit 1122.

[0700] The data reception unit 1111 receives the sensor information 1037 from the client device 902. The sensor information 1037 includes, for example, information obtained by LIDAR, luminance images (visible light images) obtained by visible light cameras, infrared images obtained by infrared cameras, depth images obtained by depth sensors, sensor position information, and speed information, etc.

[0701] The communication unit 1112 communicates with the client device 902 and sends data transmission requests (for example, transmission requests for sensor information) etc. to the client device 902.

[0702] The reception control unit 1113 exchanges information such as corresponding formats with the communication partner via the communication unit 1112 to establish communication.

[0703] When the received sensor information 1037 is compressed or encoded, the format conversion unit 1114 generates the sensor information 1132 by performing decompression or decoding processing. In addition, when the sensor information 1037 is uncompressed data, the format conversion unit 1114 does not perform decompression or decoding processing.

[0704] The 3D data creation unit 1116 creates 3D data 1134 of the periphery of the client device 902 based on the sensor information 1132. For example, the 3D data creation unit 1116 uses the information obtained by LIDAR and the visible light images obtained by the visible light camera to create point cloud data with color information of the periphery of the client device 902.

[0705] The 3D data synthesis unit 1117 synthesizes the 3D data 1134 created based on the sensor information 1132 with the 3D map 1135 managed by the server 901, and thereby updates the 3D map 1135.

[0706] The 3D data storage unit 1118 stores the 3D map 1135 and the like.

[0707] The format conversion unit 1119 generates a 3D map 1031 by converting the 3D map 1135 into a format corresponding to the receiving side. In addition, the format conversion unit 1119 can also reduce the data volume by compressing or encoding the 3D map 1135. And when format conversion is not required, the format conversion unit 1119 can also omit the processing. And the format conversion unit 1119 can control the data volume to be sent according to the specification of the sending range.

[0708] The communication unit 1120 communicates with the client device 902 and receives a data transmission request (a transmission request for a 3D map) and the like from the client device 902.

[0709] The transmission control unit 1121 exchanges information such as the corresponding format with the communication partner via the communication unit 1120, thereby establishing communication.

[0710] The data transmission unit 1122 transmits the 3D map 1031 to the client device 902. The 3D map 1031 is data of point cloud including WLD or SWLD, etc. Either compressed data or uncompressed data may be included in the 3D map 1031.

[0711] Next, the operation flow of the client device 902 will be described. Figure 48 It is a flowchart showing the operation when the client device 902 obtains a 3D map.

[0712] First, the client device 902 requests the server 901 to transmit a 3D map (point cloud, etc.) (S1001). At this time, the client device 902 also transmits the position information of the client device 902 obtained by GPS or the like together, and thereby, it is possible to request the server 901 to transmit the 3D map related to the position information.

[0713] Next, the client device 902 receives a three-dimensional map from the server 901 (S1002). If the received three-dimensional map is compressed data, the client device 902 decodes the received three-dimensional map to generate an uncompressed three-dimensional map (S1003).

[0714] Next, the client device 902 creates three-dimensional data 1034 of the surroundings of the client device 902 based on the sensor information 1033 obtained from the plurality of sensors 1015 (S1004). Next, the client device 902 estimates its own position using the three-dimensional map 1032 received from the server 901 and the three-dimensional data 1034 created based on the sensor information 1033 (S1005).

[0715] Figure 49 This is a flowchart showing the operations when the client device 902 transmits sensor information. First, the client device 902 receives a sensor information transmission request from the server 901 (S1011). The client device 902 that has received the transmission request transmits the sensor information 1037 to the server 901 (S1012). Further, when the sensor information 1033 includes a plurality of pieces of information obtained from the plurality of sensors 1015, the client device 902 compresses each piece of information in a compression method suitable for each piece of information to generate the sensor information 1037.

[0716] Next, the operation flow of the server 901 will be described. Figure 50 This is a flowchart showing the operations when the server 901 obtains sensor information. First, the server 901 requests the client device 902 to transmit sensor information (S1021). Next, the server 901 receives the sensor information 1037 transmitted from the client device 902 in response to the request (S1022). Next, the server 901 creates three-dimensional data 1134 using the received sensor information 1037 (S1023). Next, the server 901 reflects the created three-dimensional data 1134 in the three-dimensional map 1135 (S1024).

[0717] Figure 51It is a flowchart showing the actions when the server 901 sends a 3D map. First, the server 901 receives a request to send a 3D map from the client device 902 (S1031). The server 901 that has received the request to send a 3D map sends the 3D map 1031 to the client device 902 (S1032). At this time, the server 901 can extract the 3D map in its vicinity corresponding to the location information of the client device 902 and send the extracted 3D map. And it can be that the server 901 compresses the 3D map composed of point clouds, for example, using a compression method such as an octree, and sends the compressed 3D map.

[0718] Hereinafter, a modified example of the present embodiment will be described.

[0719] The server 901 uses the sensor information 1037 received from the client device 902 to create 3D data 1134 near the location of the client device 902. Next, the server 901 matches the created 3D data 1134 with the 3D map 1135 of the same area managed by the server 901 and calculates the difference between the 3D data 1134 and the 3D map 1135. When the difference is equal to or greater than a predetermined threshold, the server 901 determines that some abnormality has occurred around the client device 902. For example, when the ground surface sinks due to natural disasters such as earthquakes, a large difference may be considered to occur between the 3D map 1135 managed by the server 901 and the 3D data 1134 created based on the sensor information 1037.

[0720] The sensor information 1037 may also include at least one of the type of the sensor, the performance of the sensor, and the model of the sensor. Also, it may be that a category ID corresponding to the performance of the sensor is attached to the sensor information 1037. For example, when the sensor information 1037 is information obtained by LIDAR, it is possible to consider allocating identifiers according to the performance of the sensor. For example, category 1 is allocated to a sensor that can obtain information with an accuracy of several millimeters, category 2 is allocated to a sensor that can obtain information with an accuracy of several centimeters, and category 3 is allocated to a sensor that can obtain information with an accuracy of several meters. Also, the server 901 may estimate the performance information of the sensor from the model of the client device 902. For example, when the client device 902 is mounted on a vehicle, the server 901 can determine the specification information of the sensor according to the model of the vehicle. In this case, the server 901 may obtain the information of the vehicle model in advance, or may include this information in the sensor information. Also, it may be that the server 901 uses the obtained sensor information 1037 to switch the degree of correction for the three-dimensional data 1134 created using the sensor information 1037. For example, when the sensor performance is high accuracy (category 1), the server 901 does not perform correction for the three-dimensional data 1134. When the sensor performance is low accuracy (category 3), the server 901 applies correction suitable for the accuracy of the sensor to the three-dimensional data 1134. For example, the server 901 increases the degree (intensity) of correction as the accuracy of the sensor becomes lower.

[0721] The server 901 may also send a request to send sensor information to multiple client devices 902 existing in a certain space at the same time. When the server 901 receives multiple sensor information from the multiple client devices 902, it is not necessary to use all the sensor information for creating the three-dimensional data 1134. For example, the sensor information to be used can be selected according to the performance of the sensor. For example, when updating the three-dimensional map 1135, the server 901 can select high-accuracy sensor information (category 1) from the received multiple sensor information and use the selected sensor information to create the three-dimensional data 1134.

[0722] The server 901 is not limited to servers such as traffic cloud monitoring, and may also be other client devices (in-vehicle). Figure 52 It is a diagram showing the system configuration in this case.

[0723] For example, the client device 902C sends a request to the nearby client device 902A to send sensor information, and obtains the sensor information from the client device 902A. Then, the client device 902C uses the obtained sensor information of the client device 902A to create three-dimensional data and update the three-dimensional map of the client device 902C. In this way, the client device 902C can utilize the performance of the client device 902C to generate a three-dimensional map of the space that can be obtained from the client device 902A. For example, this may occur when the performance of the client device 902C is high.

[0724] Moreover, in this case, the client device 902A that provided the sensor information is given the right to obtain the highly accurate three-dimensional map generated by the client device 902C. The client device 902A receives the highly accurate three-dimensional map from the client device 902C according to this right.

[0725] Alternatively, the client device 902C may send a request to send sensor information to a plurality of nearby client devices 902 (client device 902A and client device 902B). When the sensor of the client device 902A or the client device 902B has high performance, the client device 902C can use the sensor information obtained through this high-performance sensor to create three-dimensional data.

[0726] Figure 53 It is a block diagram showing the functional configurations of the server 901 and the client device 902. The server 901 includes, for example: a three-dimensional map compression / decoding processing unit 1201 that compresses and decodes a three-dimensional map, and a sensor information compression / decoding processing unit 1202 that compresses and decodes sensor information.

[0727] The client device 902 includes: a three-dimensional map decoding processing unit 1211 and a sensor information compression processing unit 1212. The three-dimensional map decoding processing unit 1211 receives the encoded data of the compressed three-dimensional map, decodes the encoded data, and obtains the three-dimensional map. The sensor information compression processing unit 1212 does not compress the three-dimensional data created from the obtained sensor information, but compresses the sensor information itself, and sends the encoded data of the compressed sensor information to the server 901. According to this configuration, the client device 902 can keep the processing unit (device or LSI) for decoding the three-dimensional map (point cloud, etc.) inside, without having to keep the processing unit for compressing the three-dimensional data of the three-dimensional map (point cloud, etc.) inside. In this way, the cost and power consumption of the client device 902 can be suppressed.

[0728] As described above, the client device 902 according to this embodiment is mounted on a moving body, and creates three-dimensional data 1034 of the surroundings of the moving body based on sensor information 1033 indicating the surroundings of the moving body obtained by a sensor 1015 mounted on the moving body. The client device 902 estimates the own position of the moving body by using the created three-dimensional data 1034. The client device 902 transmits the acquired sensor information 1033 to the server 901 or another moving body 902.

[0729] Accordingly, the client device 902 transmits the sensor information 1033 to the server 901 or the like. In this way, there is a possibility that the amount of data to be transmitted can be reduced as compared with the case of transmitting three-dimensional data. Also, since it is not necessary to perform processing such as compression or encoding of the three-dimensional data on the client device 902, the amount of processing of the client device 902 can be reduced. Therefore, the client device 902 can achieve a reduction in the amount of transmitted data or a simplification of the device configuration.

[0730] In addition, the client device 902 further transmits a transmission request for a three-dimensional map to the server 901, and receives a three-dimensional map 1031 from the server 901. The client device 902 estimates its own position by using the three-dimensional data 1034 and the three-dimensional map 1032 in the estimation of its own position.

[0731] In addition, the sensor information 1033 includes at least one of information obtained by a laser sensor, a luminance image (visible light image), an infrared image, a depth image, position information of the sensor, and speed information of the sensor.

[0732] In addition, the sensor information 1033 includes information indicating the performance of the sensor.

[0733] In addition, the client device 902 encodes or compresses the sensor information 1033, and transmits the encoded or compressed sensor information 1037 to the server 901 or another moving body 902 in the transmission of the sensor information. Accordingly, the client device 902 can reduce the amount of data to be transmitted.

[0734] For example, the client device 902 includes a processor and a memory, and the processor performs the above-described processing by using the memory.

[0735] In addition, the server 901 according to this embodiment can communicate with the client device 902 mounted on a moving body, and receives sensor information 1037 indicating the surroundings of the moving body obtained by a sensor 1015 mounted on the moving body from the client device 902. The server 901 creates three-dimensional data 1134 of the surroundings of the moving body based on the received sensor information 1037.

[0736] Accordingly, the server 901 uses the sensor information 1037 sent from the client device 902 to create three-dimensional data 1134. In this way, compared with the case where the client device 902 sends three-dimensional data, there is a possibility of reducing the amount of data to be sent. Also, since it is not necessary to perform processing such as compression or encoding of three-dimensional data on the client device 902, the processing amount of the client device 902 can be reduced. In this way, the server 901 can achieve a reduction in the amount of transmitted data or a simplification of the device configuration.

[0737] Furthermore, the server 901 further sends a transmission request for the sensor information to the client device 902.

[0738] Furthermore, the server 901 further uses the created three-dimensional data 1134 to update the three-dimensional map 1135, and sends the three-dimensional map 1135 to the client device 902 according to the transmission request of the three-dimensional map 1135 from the client device 902.

[0739] Moreover, the sensor information 1037 includes at least one of information obtained by a laser sensor, a luminance image (visible light image), an infrared image, a depth image, the position information of the sensor, and the speed information of the sensor.

[0740] Moreover, the sensor information 1037 includes information indicating the performance of the sensor.

[0741] Furthermore, the server 901 corrects the three-dimensional data according to the performance of the sensor. Accordingly, this three-dimensional data creation method can improve the quality of the three-dimensional data.

[0742] Moreover, in receiving the sensor information, the server 901 receives a plurality of sensor information 1037 from a plurality of client devices 902, and selects the sensor information 1037 to be used in the creation of the three-dimensional data 1134 according to a plurality of information indicating the performance of the sensor included in the plurality of sensor information 1037. Accordingly, the server 901 can improve the quality of the three-dimensional data 1134.

[0743] Moreover, the server 901 decodes or decompresses the received sensor information 1037, and creates the three-dimensional data 1134 according to the decoded or decompressed sensor information 1132. Accordingly, the server 901 can reduce the amount of transmitted data.

[0744] For example, the server 901 includes a processor and a memory, and the processor uses the memory to perform the above processing.

[0745] (Embodiment 8)

[0746] In this embodiment, a method for encoding and decoding three-dimensional data using inter-frame prediction processing will be described.

[0747] Figure 54 FIG. 4 is a block diagram of a three-dimensional data encoding apparatus 1300 according to this embodiment. The three-dimensional data encoding apparatus 1300 generates an encoded bitstream (hereinafter also simply referred to as a bitstream) as an encoded signal by encoding three-dimensional data. As Figure 54 shown, the three-dimensional data encoding apparatus 1300 includes: a segmentation unit 1301, a subtraction unit 1302, a transformation unit 1303, a quantization unit 1304, an inverse quantization unit 1305, an inverse transformation unit 1306, an addition unit 1307, a reference volume memory 1308, an intra-frame prediction unit 1309, a reference space memory 1310, an inter-frame prediction unit 1311, a prediction control unit 1312, and an entropy encoding unit 1313.

[0748] The segmentation unit 1301 divides each space (SPC) included in the three-dimensional data into a plurality of volumes (VLM) as encoding units. Further, the segmentation unit 1301 performs octree representation (Octree conversion) on the voxels within each volume. In addition, the segmentation unit 1301 may make the space and the volume the same size and perform octree representation on the space. Further, the segmentation unit 1301 may attach information (such as depth information) required for octree conversion to the head of the bitstream or the like.

[0749] The subtraction unit 1302 calculates the difference between the volume (encoding target volume) output from the segmentation unit 1301 and the predicted volume generated by intra-frame prediction or inter-frame prediction described later, and outputs the calculated difference as a prediction residual to the transformation unit 1303. Figure 55 FIG. 5 is a diagram showing an example of calculation of a prediction residual. In addition, the bit strings of the encoding target volume and the predicted volume shown here are, for example, position information indicating the positions of three-dimensional points (for example, point clouds) included in the volume.

[0750] Hereinafter, octree representation and voxel scanning order will be described. After the volume is transformed into an octree structure (octree conversion), it is encoded. The octree structure is composed of nodes and leaf nodes. Each node has eight nodes or leaf nodes, and each leaf node has voxel (VXL) information. Figure 56 FIG. 6 is a diagram showing a configuration example of a volume including a plurality of voxels. Figure 57 is a diagram showing Figure 56 an example of transforming the volume shown in Figure 57 into an octree structure. Here, Figure 56 among the leaf nodes shown in

[0751] An octree is represented by a binary sequence of 0s and 1s, for example. For example, when a node or a valid VXL is set to the value 1 and the rest are set to the value 0, the binary sequence shown is assigned to each node and leaf node. Figure 57 Then, according to the breadth-first or depth-first scanning order, this binary sequence is scanned. For example, when scanned in breadth-first order, the binary sequence shown in A is obtained. When scanned in depth-first order, the binary sequence shown in B is obtained. Figure 58 The binary sequence obtained by this scanning is encoded by entropy coding, thereby reducing the amount of information. Figure 58 Next, the depth information in the octree representation is described. The depth in the octree representation is used for controlling up to which granularity the point cloud information contained in the volume is maintained. If the depth is set large, the point cloud information can be reproduced at a finer level, but the data volume for representing nodes and leaf nodes will increase. On the contrary, if the depth is set small, although the data volume can be reduced, point cloud information at multiple different positions and with different colors will be regarded as the same position and the same color, so the information originally possessed by the point cloud information will be lost.

[0752] For example,

[0753] For example, Figure 59 is a diagram showing an example of representing an octree with a depth of 2 shown in Figure 57 as an octree with a depth of 1. Figure 59 The octree shown in Figure 57 has a smaller data volume than the octree shown in Figure 59 That is, the octree shown in Figure 59 has fewer bits after binary serialization compared to the octree shown in Figure 57 Here, the leaf node 1 and leaf node 2 shown in Figure 58 become represented by the leaf node 1 shown in Figure 57 That is, the information that the leaf node 1 and leaf node 2 shown in

[0754] Figure 60 is a diagram showing the volume corresponding to the octree shown in Figure 59 are different positions is lost. Figure 56 The VXL1 and VXL2 shown in Figure 60 correspond to the VXL12 shown in Figure 56 In this case, the three-dimensional data encoding device 1300 generates Figure 60The color information of VXL12 shown. For example, the three-dimensional data encoding device 1300 calculates the average value, median value, weighted average value, etc. of the color information of VXL1 and VXL2 as the color information of VXL12. In this way, the three-dimensional data encoding device 1300 can control the reduction of the data volume by changing the depth of the octree.

[0755] The three-dimensional data encoding device 1300 can also set the depth information of the octree using any one of the world space unit, space unit, and volume unit. And at this time, the three-dimensional data encoding device 1300 can also attach the depth information to the header information of the world space, the header information of the space, or the header information of the volume. Also, the same value can be used as the depth information in all world spaces, spaces, and volumes at different times. In this case, the three-dimensional data encoding device 1300 can also attach the depth information to the header information for managing the world space for all times.

[0756] When the voxel contains color information, the transformation unit 1303 applies a frequency transformation such as an orthogonal transformation to the prediction residual of the color information of the voxel in the volume. For example, the transformation unit 1303 scans the prediction residual in a certain scanning order to create a one-dimensional arrangement. After that, the transformation unit 1303 transforms the created one-dimensional arrangement into the frequency domain by applying a one-dimensional orthogonal transformation. Accordingly, when the values of the prediction residuals in the volume are close, the values of the frequency components in the low-frequency band become larger, and the values of the frequency components in the high-frequency band become smaller. Therefore, the quantization unit 1304 can more effectively reduce the encoding amount.

[0757] Also, the transformation unit 1303 can use an orthogonal transformation of two or more dimensions instead of using a one-dimensional orthogonal transformation. For example, the transformation unit 1303 maps the prediction residual into a two-dimensional arrangement in a certain scanning order and applies a two-dimensional orthogonal transformation to the obtained two-dimensional arrangement. Also, the transformation unit 1303 can select the orthogonal transformation method to be used from multiple orthogonal transformation methods. In this case, the three-dimensional data encoding device 1300 attaches information indicating which orthogonal transformation method is used to the bitstream. And it can be that the transformation unit 1303 selects the orthogonal transformation method to be used from multiple orthogonal transformation methods with different dimensions. In this case, the three-dimensional data encoding device 1300 attaches information indicating which dimension of the orthogonal transformation method is used to the bitstream.

[0758] For example, the transformation unit 1303 aligns the scan order of the prediction residual with the scan order (such as breadth-first or depth-first) in the octree within the volume. Accordingly, since there is no need to attach information indicating the scan order of the prediction residual to the bitstream, the overhead can be reduced. Also, the transformation unit 1303 may apply a scan order different from the scan order of the octree. In this case, the 3D data encoding device 1300 attaches information indicating the scan order of the prediction residual to the bitstream. Accordingly, the 3D data encoding device 1300 can efficiently encode the prediction residual. It may also be that the 3D data encoding device 1300 attaches information (such as a flag) indicating whether the scan order of the octree is applied to the bitstream, and in the case where the scan order of the octree is not applied, attaches information indicating the scan order of the prediction residual to the bitstream.

[0759] The transformation unit 1303 can transform not only the prediction residual of color information but also other attribute information possessed by the voxels. For example, it may be that the transformation unit 1303 transforms and encodes information such as reflectance obtained when acquiring point clouds through LiDAR or the like.

[0760] When the space does not have attribute information such as color information, the transformation unit 1303 can skip the processing. Also, the 3D data encoding device 1300 can attach information (a flag) indicating whether to skip the processing of the transformation unit 1303 to the bitstream.

[0761] The quantization unit 1304 quantizes the frequency components of the prediction residual generated in the transformation unit 1303 using quantization control parameters to generate quantization coefficients. Thereby, the amount of information is reduced. The generated quantization coefficients are output to the entropy encoding unit 1313. The quantization unit 1304 can control the quantization control parameters in world space units, space units, or volume units. At this time, the 3D data encoding device 1300 attaches the quantization control parameters to respective header information and the like. Also, the quantization unit 1304 can change the weights for quantization control according to the frequency components of each prediction residual. For example, the quantization unit 1304 can perform fine quantization on low-frequency components and rough quantization on high-frequency components. In this case, the 3D data encoding device 1300 can attach parameters indicating the weights of the respective frequency components to the header.

[0762] When the space does not have attribute information such as color information, the quantization unit 1304 can skip the processing. Also, the 3D data encoding device 1300 can attach information (a flag) indicating whether to skip the processing of the quantization unit 1304 to the bitstream.

[0763] The inverse quantization unit 1305 uses the quantization control parameter to perform inverse quantization on the quantization coefficients generated by the quantization unit 1304. Accordingly, the inverse quantization coefficients of the prediction residual are generated, and the generated inverse quantization coefficients are output to the inverse transform unit 1306.

[0764] The inverse transform unit 1306 applies an inverse transform to the inverse quantization coefficients generated in the inverse quantization unit 1305, thereby generating the prediction residual after the inverse transform is applied. Since the prediction residual after the inverse transform is applied is the prediction residual generated after quantization, it may not be exactly the same as the prediction residual output by the transform unit 1303.

[0765] The addition unit 1307 adds the prediction volume generated by the inverse transform unit 1306 after the inverse transform is applied and the prediction volume used in the generation of the prediction residual before quantization and generated by the intra prediction or inter prediction described later to generate the reconstructed volume. The reconstructed volume is stored in the reference volume memory 1308 or the reference space memory 1310.

[0766] The intra prediction unit 1309 uses the attribute information of the adjacent volume stored in the reference volume memory 1308 to generate the prediction volume of the volume to be encoded. The attribute information includes the color information or reflectivity of the voxel. The intra prediction unit 1309 generates the predicted value of the color information or reflectivity of the volume to be encoded.

[0767] Figure 61 It is a diagram for explaining the operation of the intra prediction unit 1309. For example, Figure 61 As shown, the intra prediction unit 1309 generates the prediction volume of the volume to be encoded (volume idx = 3) based on the adjacent volume (volume idx = 0). Here, the volume idx is the identifier information attached to the volume in the space, and different values are assigned to each volume. The order of assignment of the volume idx may be the same as the encoding order or different from the encoding order. For example, as Figure 61 As the predicted value of the color information of the volume to be encoded shown, the intra prediction unit 1309 uses the average value of the color information of the voxels included in the adjacent volume with volume idx = 0. In this case, by subtracting the predicted value of the color information from the color information of each voxel included in the volume to be encoded, the prediction residual is generated. The processing after the transform unit 1303 is performed on the prediction residual. And, in this case, the three-dimensional data encoding device 1300 attaches the adjacent volume information and the prediction mode information to the bitstream. Here, the adjacent volume information is the information showing the adjacent volume used in the prediction, for example, showing the volume idx of the adjacent volume used in the prediction. And, the prediction mode information shows the mode used in the generation of the prediction volume. The mode is, for example, the average value mode that generates the predicted value based on the average value of the voxels in the adjacent volume, or the median value mode that generates the predicted value based on the median value of the voxels in the adjacent volume, etc.

[0768] The intra prediction unit 1309 can also generate a prediction volume based on multiple adjacent volumes. For example, in the Figure 61 configuration shown, the intra prediction unit 1309 generates prediction volume 0 based on the volume with volume idx = 0, and generates prediction volume 1 based on the volume with volume idx = 1. Then, the intra prediction unit 1309 generates the average of prediction volume 0 and prediction volume 1 as the final prediction volume. In this case, the three-dimensional data encoding device 1300 can also attach the multiple volume idxs of the multiple volumes used in the generation of the prediction volume to the bitstream.

[0769] Figure 62 FIG. is a diagram schematically showing the inter prediction process according to the present embodiment. The inter prediction unit 1311 performs encoding (inter prediction) on the space (SPC) at a certain time T_Cur using the encoded spaces at different times T_LX. In this case, the inter prediction unit 1311 applies rotation and translation processing to the encoded spaces at different times T_LX to perform the encoding process.

[0770] Furthermore, the three-dimensional data encoding device 1300 attaches RT information related to the rotation and translation processing of the spaces at different times T_LX to the bitstream. The different times T_LX are, for example, the time T_L0 before the certain time T_Cur. At this time, the three-dimensional data encoding device 1300 can also attach the RT information RT_L0 related to the rotation and translation processing of the space at time T_L0 to the bitstream.

[0771] Alternatively, the different times T_LX are, for example, the time T_L1 after the certain time T_Cur. At this time, the three-dimensional data encoding device 1300 can attach the RT information RT_L1 related to the rotation and translation processing of the space at time T_L1 to the bitstream.

[0772] Alternatively, the inter prediction unit 1311 performs encoding (dual prediction) by referring to the spaces at both different times T_L0 and time T_L1. In this case, the three-dimensional data encoding device 1300 can attach both the RT information RT_L0 and RT_L1 related to the rotation and translation applied to the spaces respectively to the bitstream.

[0773] In addition, although T_L0 is set as the time before T_Cur and T_L1 is set as the time after T_Cur above, it is not limited thereto. For example, both T_L0 and T_L1 can be the times before T_Cur. Or, both T_L0 and T_L1 can be the times after T_Cur.

[0774] Alternatively, when the three-dimensional data encoding device 1300 performs encoding with reference to spaces at multiple different times, RT information related to the rotation and translation applicable to each space is appended to the bitstream. For example, the three-dimensional data encoding device 1300 manages the multiple encoded spaces to be referred to through two reference lists (L0 list and L1 list). When the first reference space in the L0 list is set as L0R0, the second reference space in the L0 list is set as L0R1, the first reference space in the L1 list is set as L1R0, and the second reference space in the L1 list is set as L1R1, the three-dimensional data encoding device 1300 appends the RT information RT_L0R0 of L0R0, the RT information RT_L0R1 of L0R1, the RT information RT_L1R0 of L1R0, and the RT information RT_L1R1 of L1R1 to the bitstream. For example, the three-dimensional data encoding device 1300 appends this RT information to the header of the bitstream or the like.

[0775] Alternatively, when the three-dimensional data encoding device 1300 performs encoding with reference to reference spaces at multiple different times, it determines whether rotation and translation are applicable for each reference space. At this time, the three-dimensional data encoding device 1300 may append information (such as an RT application flag) indicating whether rotation and translation are applicable for each reference space to the header information of the bitstream or the like. For example, the three-dimensional data encoding device 1300 calculates the RT information and the ICP error value for each reference space to be referred to according to the encoding target space using the ICP (Interactive Closest Point) algorithm. When the ICP error value is equal to or less than a predetermined fixed value, the three-dimensional data encoding device 1300 determines that rotation and translation are not required and sets the RT application flag to OFF (invalid). In addition, when the ICP error value is greater than the above-mentioned fixed value, the three-dimensional data encoding device 1300 sets the RT application flag to ON (valid) and appends the RT information to the bitstream.

[0776] Figure 63 FIG. shows a syntax example of appending RT information and an RT application flag to the header. In addition, the number of bits allocated to each syntax can be determined according to the range that the syntax can take. For example, when the number of reference spaces included in the reference list L0 is 8, 3 bits can be allocated to MaxRefSpc_l0. The number of allocated bits can be changed according to the values that each syntax can take, or the number of allocated bits can be fixed regardless of the values that can be taken. When the number of allocated bits is fixed, the three-dimensional data encoding device 1300 can append this fixed number of bits to other header information.

[0777] Here, Figure 63The shown MaxRefSpc_l0 indicates the number of reference spaces included in the reference list L0. RT_flag_l0[i] is the RT application flag for the reference space i in the reference list L0. When RT_flag_l0[i] is 1, rotation and translation are applied to the reference space i. When RT_flag_l0[i] is 0, rotation and translation are not applied to the reference space i.

[0778] R_l0[i] and T_l0[i] are the RT information of the reference space i in the reference list L0. R_l0[i] is the rotation information of the reference space i in the reference list L0. The rotation information indicates the content of the applied rotation process, such as a rotation matrix or a quaternion, etc. T_l0[i] is the translation information of the reference space i in the reference list L0. The translation information indicates the content of the applied translation process, such as a translation vector, etc.

[0779] MaxRefSpc_l1 indicates the number of reference spaces included in the reference list L1. RT_flag_l1[i] is the RT application flag for the reference space i in the reference list L1. When RT_flag_l1[i] is 1, rotation and translation are applied to the reference space i. When RT_flag_l1[i] is 0, rotation and translation are not applied to the reference space i.

[0780] R_l1[i] and T_l1[i] are the RT information of the reference space i in the reference list L1. R_l1[i] is the rotation information of the reference space i in the reference list L1. The rotation information indicates the content of the applied rotation process, such as a rotation matrix or a quaternion, etc. T_l1[i] is the translation information of the reference space i in the reference list L1. The translation information indicates the content of the applied translation process, such as a translation vector, etc.

[0781] The inter-frame prediction unit 1311 generates a predicted volume of the coding target volume by using the information of the encoded reference spaces stored in the reference space memory 1310. As described above, before generating the predicted volume of the coding target volume, the inter-frame prediction unit 1311 uses the ICP (Interactive Closest Point) algorithm in the coding target space and the reference spaces to find the RT information in order to make the positional relationship between the coding target space and the entire reference space closer. Then, the inter-frame prediction unit 1311 applies rotation and translation processing to the reference space by using the obtained RT information, thereby obtaining the reference space B. After that, the inter-frame prediction unit 1311 generates a predicted volume of the coding target volume in the coding target space by using the information in the reference space B. Here, the three-dimensional data coding device 1300 attaches the RT information used to obtain the reference space B to the header information, etc., of the coding target space.

[0782] In this way, the inter-frame prediction unit 1311 applies rotation and translation processing to the reference space, so that after making the positional relationship between the coding target space and the overall reference space closer, the prediction volume is generated using the information of the reference space. In this way, the accuracy of the prediction volume can be improved. Also, since the prediction residual can be suppressed, the coding amount can be reduced. Additionally, although this is an example showing ICP using the coding target space and the reference space, it is not limited thereto. For example, in order to reduce the processing amount, the inter-frame prediction unit 1311 can also perform ICP using at least one of the coding target space with the voxel or point cloud number extracted and the reference space with the voxel or point cloud number extracted, so as to obtain the RT information.

[0783] Also, when the ICP error value obtained from the result of ICP is smaller than a predetermined first threshold, that is, for example, when the positional relationship between the coding target space and the reference space is close, the inter-frame prediction unit 1311 can determine that rotation and translation processing are not required and does not perform rotation and translation. In this case, the three-dimensional data coding device 1300 can suppress the overhead by not attaching the RT information to the bitstream.

[0784] Also, when the ICP error value is larger than a predetermined second threshold, it is determined that the shape change in space is large, and intra-frame prediction can be applied to all volumes of the coding target space. Hereinafter, the space to which intra-frame prediction is applied is referred to as the intra-frame space. Also, the second threshold is a value larger than the above-mentioned first threshold. Also, it is not limited to ICP, and any method can be applied as long as it is a method for obtaining the RT information from two voxel sets or two point cloud sets.

[0785] Also, when attribute information such as shape or color is included in the three-dimensional data, as the prediction volume of the coding target volume in the coding target space, the inter-frame prediction unit 1311 searches, for example, for the volume in the reference space that is closest to the shape or color attribute information of the coding target volume. Also, the reference space is, for example, the reference space after the above-mentioned rotation and translation processing. The inter-frame prediction unit 1311 generates the prediction volume based on the volume (reference volume) obtained through the search. Figure 64 It is a diagram for explaining the generation operation of the prediction volume. The inter-frame prediction unit 1311 is for Figure 64In the case of encoding the shown encoded object volume (volume idx = 0) using inter-frame prediction, while sequentially scanning the reference volumes in the reference space, the volume with the smallest difference, i.e., the prediction residual, between the encoded object volume and the reference volume is searched for. The inter-frame prediction unit 1311 selects the volume with the smallest prediction residual as the prediction volume. The prediction residual between the encoded object volume and the prediction volume is encoded by the processing after the transform unit 1303. Here, the prediction residual refers to the difference between the attribute information of the encoded object volume and the attribute information of the prediction volume. Further, the three-dimensional data encoding device 1300 attaches the volume idx of the reference volume in the reference space used as the prediction volume to the head of the bitstream or the like.

[0786] In Figure 64 In the example shown, the reference volume with volume idx = 4 in the reference space L0R0 is selected as the prediction volume of the encoded object volume. Then, the prediction residual between the encoded object volume and the reference volume and the reference volume idx = 4 are encoded and attached to the bitstream.

[0787] In addition, although an example of generating a prediction volume for attribute information has been described here, the same processing can be performed for the prediction volume of the position information.

[0788] The prediction control unit 1312 controls which of intra-frame prediction and inter-frame prediction is used to encode the encoded object volume. Here, the mode including intra-frame prediction and inter-frame prediction is referred to as the prediction mode. For example, the prediction control unit 1312 calculates the prediction residual in the case where the encoded object volume is predicted by intra-frame prediction and the prediction residual in the case where it is predicted by inter-frame prediction as evaluation values, and selects the prediction mode with the smaller evaluation value. Alternatively, the prediction control unit 1312 may apply orthogonal transformation, quantization, and entropy coding to the prediction residual of intra-frame prediction and the prediction residual of inter-frame prediction, respectively, to calculate the actual coding amount, and use the calculated coding amount as the evaluation value to select the prediction mode. Further, overhead information other than the prediction residual (such as reference volume idx information) may be added to the evaluation value. Also, the prediction control unit 1312 may usually select intra-frame prediction even when the encoded object space is predetermined to be encoded in the intra-frame space.

[0789] The entropy coding unit 1313 generates an encoded signal (encoded bitstream) by performing variable-length coding on the input from the quantization unit 1304, i.e., the quantization coefficients. Specifically, the entropy coding unit 1313 binarizes the quantization coefficients, for example, and performs arithmetic coding on the obtained binary signal.

[0790] Next, a three-dimensional data decoding device that decodes the encoded signal generated by the three-dimensional data encoding device 1300 will be described. Figure 65It is a block diagram of the 3D data decoding device 1400 according to this embodiment. The 3D data decoding device 1400 includes: an entropy decoding unit 1401, an inverse quantization unit 1402, an inverse transformation unit 1403, an addition unit 1404, a reference volume memory 1405, an intra prediction unit 1406, a reference space memory 1407, an inter prediction unit 1408, and a prediction control unit 1409.

[0791] The entropy decoding unit 1401 performs variable length decoding on the encoded signal (encoded bitstream). For example, the entropy decoding unit 1401 performs arithmetic decoding on the encoded signal to generate a binary signal, and generates quantization coefficients based on the generated binary signal.

[0792] The inverse quantization unit 1402 performs inverse quantization on the quantization coefficients input from the entropy decoding unit 1401 using the quantization parameters attached to the bitstream or the like, thereby generating inverse quantization coefficients.

[0793] The inverse transformation unit 1403 performs inverse transformation on the inverse quantization coefficients input from the inverse quantization unit 1402, thereby generating a prediction residual. For example, the inverse transformation unit 1403 performs inverse orthogonal transformation on the inverse quantization coefficients according to the information attached to the bitstream, thereby generating a prediction residual.

[0794] The addition unit 1404 adds the prediction residual generated by the inverse transformation unit 1403 and the prediction volume generated by intra prediction or inter prediction to generate a reconstructed volume. This reconstructed volume is output as decoded 3D data and stored in the reference volume memory 1405 or the reference space memory 1407.

[0795] The intra prediction unit 1406 generates a prediction volume by intra prediction using the reference volume in the reference volume memory 1405 and the information attached to the bitstream. Specifically, the intra prediction unit 1406 obtains prediction mode information and adjacent volume information (such as volume idx) attached to the bitstream, and uses the adjacent volume indicated by the adjacent volume information to generate a prediction volume in the mode indicated by the prediction mode information. In addition, for the details of these processes, except for using the information attached to the bitstream, it is the same as the process of the intra prediction unit 1309 described above.

[0796] The inter-frame prediction unit 1408 generates a prediction volume through inter-frame prediction by using the reference spaces in the reference space memory 1407 and the information appended to the bitstream. Specifically, the inter-frame prediction unit 1408 applies rotation and translation processing to the reference spaces by using the RT information for each reference space appended to the bitstream, and generates a prediction volume by using the reference spaces after the application. In addition, when the RT application flag for each reference space exists in the bitstream, the inter-frame prediction unit 1408 applies rotation and translation processing to the reference spaces in accordance with the RT application flag. In addition, the details of the above processing are the same as those of the above inter-frame prediction unit 1311 except for using the information appended to the bitstream.

[0797] Regarding whether to decode the volume to be decoded by intra-frame prediction or inter-frame prediction, it will be controlled by the prediction control unit 1409. For example, the prediction control unit 1409 selects intra-frame prediction or inter-frame prediction in accordance with the information appended to the bitstream and indicating the prediction mode to be used. In addition, the prediction control unit 1409 may generally select intra-frame prediction when it is predetermined that the decoding target space is decoded by the intra-frame space.

[0798] A modification example of the present embodiment will be described below. In the present embodiment, although rotation and translation are applied in units of space as an example, rotation and translation may also be applied in smaller units. For example, the three-dimensional data encoding device 1300 may divide the space into sub-spaces and apply rotation and translation in units of sub-spaces. In this case, the three-dimensional data encoding device 1300 generates RT information for each sub-space and appends the generated RT information to the head of the bitstream or the like. And, the three-dimensional data encoding device 1300 may apply rotation and translation in units of volume as the encoding unit. In this case, the three-dimensional data encoding device 1300 generates RT information in units of encoding volume and appends the generated RT information to the head of the bitstream or the like. Moreover, the above may be combined. That is, the three-dimensional data encoding device 1300 may apply rotation and translation in a large unit first, and then apply rotation and translation in a smaller unit. For example, it may be that the three-dimensional data encoding device 1300 applies rotation and translation in units of space, and applies different rotations and translations to each of the multiple volumes included in the obtained space.

[0799] Further, although the rotation and translation are applied to the reference space in this embodiment as an example, it is not limited thereto. For example, the three-dimensional data encoding device 1300 may apply a scaling process to change the size of the three-dimensional data. Further, the three-dimensional data encoding device 1300 may also apply any one or two of rotation, translation, and scaling. Further, as described above, when the processes are applied in different units in multiple stages, the types of processes applied in each unit may be different. For example, rotation and translation may be applied in the space unit, and translation may be applied in the volume unit.

[0800] In addition, regarding these modification examples, the three-dimensional data decoding device 1400 can be similarly applied.

[0801] As described above, the three-dimensional data encoding device 1300 according to this embodiment performs the following processes. Figure 65 It is a flowchart of the inter-frame prediction process performed by the three-dimensional data encoding device 1300.

[0802] First, the three-dimensional data encoding device 1300 generates prediction position information (for example, a prediction volume) by using the position information of the three-dimensional points included in the target three-dimensional data (for example, the encoding target space) and the reference three-dimensional data (for example, the reference space) at different times (S1301). Specifically, the three-dimensional data encoding device 1300 generates prediction position information by applying rotation and translation processes to the position information of the three-dimensional points included in the reference three-dimensional data.

[0803] In addition, the three-dimensional data encoding device 1300 performs rotation and translation processes in the first unit (for example, space), and generates prediction position information in the second unit (for example, volume) that is finer than the first unit. For example, the three-dimensional data encoding device 1300 searches for the volume in which the difference between the encoded target volume included in the encoded target space and the position information is the smallest among the multiple volumes included in the reference space after the rotation and translation processes, and uses the obtained volume as the prediction volume. In addition, the three-dimensional data encoding device 1300 may also perform the rotation and translation processes and the generation of prediction position information in the same unit.

[0804] And it may be that the three-dimensional data encoding device 1300 applies the first rotation and translation process to the position information of the three-dimensional points included in the reference three-dimensional data in the first unit (for example, space), and applies the second rotation and translation process to the position information of the three-dimensional points obtained by the first rotation and translation process in the second unit (for example, volume) that is finer than the first unit, thereby generating prediction position information.

[0805] Here, the position information of the three-dimensional points and the prediction position information are as Figure 58As shown, it is represented in an octree structure. For example, the position information of three-dimensional points and the predicted position information are represented in the order of width-first scanning in the depth and width of the octree structure. Alternatively, the position information of three-dimensional points and the predicted po...

Claims

1. A three-dimensional data encoding method, wherein: Encode the object nodes contained in the point group information represented by the N-ary tree structure, In the encoding, the object node is encoded using an adjacent pattern, wherein the adjacent pattern represents the occupation status of a plurality of adjacent nodes that are spatially adjacent to the object node. The adjacent pattern when N is 2 or 4 is expressed in the same format as the adjacent pattern when N is 8.

2. The three-dimensional data encoding method according to claim 1, wherein: The point group information includes first point group information including a first target node and represented by a binary tree or a quadtree, and second point group information including a second target node and represented by an octree. In the encoding, selecting a combination of contexts for encoding the first object node based on a first neighboring pattern indicating occupation states of a plurality of neighboring nodes spatially adjacent to the first object node; A combination of contexts used for encoding the second target node is selected based on a second neighboring pattern indicating occupation states of a plurality of neighboring nodes spatially adjacent to the second target node.

3. The three-dimensional data encoding method according to claim 2, wherein: The first adjacent pattern includes a 6-bit third bit pattern composed of a first bit pattern and a second bit pattern, the first bit pattern is composed of one or more bits, the one or more bits represent one or more first adjacent nodes that are spatially adjacent to the first object node in a specified direction among multiple directions, and represent that each is not occupied by a point group, the second bit pattern is composed of multiple bits, the multiple bits represent multiple second adjacent nodes that are spatially adjacent to the first object node in a direction other than the specified direction among the multiple directions, The second adjacent pattern includes a 6-bit fourth bit pattern, and the fourth bit pattern is composed of a plurality of bits, wherein the plurality of bits represent a plurality of third adjacent nodes that are spatially adjacent to the second object node in the plurality of directions.

4. The three-dimensional data encoding method according to claim 1, wherein: The point group information includes first point group information including a first target node and represented by a binary tree or a quadtree, and second point group information including a second target node and represented by an octree. In the encoding, selecting a combination of first contexts based on a first neighboring pattern indicating occupation states of a plurality of neighboring nodes spatially adjacent to the first object node, and performing entropy coding on the first object node using the selected combination of first contexts; selecting a combination of second contexts based on a second neighboring pattern indicating occupation states of a plurality of neighboring nodes spatially adjacent to the second object node, and performing entropy coding on the second object node using the selected combination of second contexts, The first adjacent pattern and the second adjacent pattern are expressed in a common format.

5. The three-dimensional data encoding method according to any one of claims 2 to 4, wherein: In the neighboring pattern when N is 2 or 4, neighboring nodes located in a predetermined axis direction from the target node are invalidated or regarded as not including a point.

6. The three-dimensional data encoding method according to any one of claims 2 to 4, wherein: Furthermore, a bit stream including identification information indicating whether the encoding target is the first target node or the second target node is generated.

7. The three-dimensional data encoding method according to any one of claims 2 to 4, wherein: The first point group information is a point group arranged on a plane. The second point group information is a point group arranged around the plane.

8. A three-dimensional data decoding method, wherein: Decode the object nodes contained in the point group information represented by the N-ary tree structure, In the decoding, the object node is decoded using a neighboring pattern, wherein the neighboring pattern represents the occupation status of a plurality of neighboring nodes that are spatially adjacent to the object node. The adjacent pattern when N is 2 or 4 is expressed in the same format as the adjacent pattern when N is 8.

9. The three-dimensional data decoding method according to claim 8, wherein: The point group information includes first point group information including a first target node and represented by a binary tree or a quadtree, and second point group information including a second target node and represented by an octree. In the decoding, selecting a combination of contexts for decoding the first target node based on a first neighboring pattern indicating occupation states of a plurality of neighboring nodes spatially adjacent to the first target node; A combination of contexts used for decoding the second target node is selected based on a second neighboring pattern indicating occupation states of a plurality of neighboring nodes spatially adjacent to the second target node.

10. The three-dimensional data decoding method according to claim 9, wherein: The first adjacent pattern includes a 6-bit third bit pattern composed of a first bit pattern and a second bit pattern, the first bit pattern is composed of one or more bits, the one or more bits represent one or more first adjacent nodes that are spatially adjacent to the first object node in a specified direction among multiple directions, and represent that each is not occupied by a point group, the second bit pattern is composed of multiple bits, the multiple bits represent multiple second adjacent nodes that are spatially adjacent to the first object node in a direction other than the specified direction among the multiple directions, The second adjacent pattern includes a 6-bit fourth bit pattern, and the fourth bit pattern is composed of a plurality of bits, wherein the plurality of bits represent a plurality of third adjacent nodes that are spatially adjacent to the second object node in the plurality of directions.

11. The three-dimensional data decoding method according to claim 8, wherein: The point group information includes first point group information including a first target node and represented by a binary tree or a quadtree, and second point group information including a second target node and represented by an octree. In the decoding, selecting a combination of first contexts based on a first neighboring pattern indicating occupation states of a plurality of neighboring nodes spatially adjacent to the first object node, and performing entropy decoding on the first object node using the selected combination of first contexts, selecting a combination of second contexts based on a second neighboring pattern indicating occupation states of a plurality of neighboring nodes spatially adjacent to the second object node, and performing entropy decoding on the second object node using the selected combination of second contexts, The first adjacent pattern and the second adjacent pattern are expressed in a common format.

12. The three-dimensional data decoding method according to any one of claims 9 to 11, wherein: In the neighboring pattern when N is 2 or 4, neighboring nodes located in a predetermined axis direction from the target node are invalidated or regarded as not including a point.

13. The three-dimensional data decoding method according to claim 12, wherein: In the decoding, obtaining a bit stream including identification information indicating whether the encoding object is the first object node or the second object node, When the identification information indicates that it is the first object node, information corresponding to the first object node in the bit stream is decoded.

14. The three-dimensional data decoding method according to any one of claims 9 to 11, wherein: The first point group information is a point group arranged on a plane. The second point group information is a point group arranged around the plane.

15. A three-dimensional data encoding device, wherein: have: Processor; and Memory, The processor uses the memory, Encode the object nodes contained in the point group information represented by the N-ary tree structure, In the encoding, the object node is encoded using an adjacent pattern, wherein the adjacent pattern represents the occupation status of a plurality of adjacent nodes that are spatially adjacent to the object node. The adjacent pattern when N is 2 or 4 is expressed in the same format as the adjacent pattern when N is 8.

16. A three-dimensional data decoding device, wherein: have: Processor; and Memory, The processor uses the memory, Decode the object nodes contained in the point group information represented by the N-ary tree structure, In the decoding, the object node is decoded using a neighboring pattern, wherein the neighboring pattern represents the occupation status of a plurality of neighboring nodes that are spatially adjacent to the object node. The adjacent pattern when N is 2 or 4 is expressed in the same format as the adjacent pattern when N is 8.

Citation Information

Patent Citations

  • Map display device

    WO2014020663A1

  • Three-dimensional shape data processing method and three-dimensional shape data processor

    JP2013077165A