Method executed by client device, and client device
The proposed method addresses the challenge of reducing data transmission and simplifying configurations in three-dimensional data creation by using a client-server system where the client transmits sensor information only when significant changes occur, allowing the server to create updated three-dimensional data.
Patent Information
- Application Number
- JP2025044180
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2017-09-29
- Filing Date
- 2025-03-19
- Publication Date
- 2025-06-12
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Existing three-dimensional data creation methods face challenges in reducing data transmission volume and simplifying device configurations, particularly in encoding and decoding three-dimensional data for applications like autonomous vehicles and robots.
A method involving a client device with multiple sensors that acquires sensor data, creates three-dimensional data, and transmits sensor information to a server only when the difference between the three-dimensional data and a three-dimensional map exceeds a predetermined threshold. The server then creates three-dimensional data from the received sensor information.
This approach reduces the amount of data transmitted and simplifies device configurations by only sending necessary data updates, thereby enhancing efficiency in data processing and transmission.
Smart Images

Figure 2025089362000001_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to a three-dimensional data creation method, a client device, and a server.
Background Art
[0002] In the future, the spread of devices or services that utilize three-dimensional data is expected in a wide range of fields such as computer vision, map information, monitoring, infrastructure inspection, or video distribution for autonomous operation of automobiles or robots. Three-dimensional data is acquired by various methods such as a distance sensor such as a range finder, a stereo camera, or a combination of a plurality of monocular cameras.
[0003] As one of the methods for expressing three-dimensional data, there is a representation method called point cloud that represents the shape of a three-dimensional structure by a point group in a three-dimensional space. In a point cloud, the positions and colors of the point group are stored. Although point cloud is expected to become the mainstream as a method for expressing three-dimensional data, the point group has a very large data volume. Therefore, in the accumulation or transmission of three-dimensional data, compression of the data volume by encoding is essential, similar to two-dimensional moving images (for example, MPEG-4 AVC or HEVC standardized by MPEG).
[0004] Also, regarding the compression of point cloud, it is partially supported by a public library (Point Cloud Library) that performs point cloud-related processing.
[0005] Also, a technique is known for searching and displaying facilities located around a vehicle using three-dimensional map data (see, for example, Patent Document 1).
Prior Art Documents
Patent Documents
[0006]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0007] In such a three-dimensional data creation method for creating three-dimensional data, it is desired to reduce the amount of data to be transmitted and simplify the configuration of the device.
[0008] An object of the present disclosure is to provide a three-dimensional data creation method, a client device, or a server that can reduce the amount of data to be transmitted or simplify the configuration of the device.
Means for Solving the Problems
[0009] A method according to an aspect of the present disclosure is a method performed by a client device including a plurality of sensors, the method including: acquiring sensor data using the plurality of sensors; creating three-dimensional data including three-dimensional coordinate values of a plurality of points using the sensor data; and transmitting sensor information including the sensor data to a server when a difference between the three-dimensional data and a three-dimensional map is greater than a predetermined threshold value.
[0010] A three-dimensional data creation method according to an aspect of the present disclosure is a three-dimensional data creation method in a server communicable with a client device mounted on a moving body, the method including: receiving, from the client device, sensor information indicating a surrounding situation of the moving body obtained by a sensor mounted on the moving body; and creating three-dimensional data around the moving body from the received sensor information.
Advantages of the Invention
[0011] The present disclosure can provide a three-dimensional data creation method, a client device, or a server that can reduce the amount of data to be transmitted or simplify the configuration of the device.
Brief Description of the Drawings
[0012]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
Figure 9
Figure 10
Figure 11
Figure 12
Figure 13
Figure 14
Figure 15
Figure 16
Figure 17
Figure 18
Figure 19
Figure 20
Figure 21
Figure 22
Figure 23
Figure 24
Figure 25
Figure 26
Figure 27
Figure 28
Figure 29
Figure 30
Figure 31
Figure 32
Figure 33
Figure 34
Figure 35
Figure 36
Figure 37
Figure 38
Figure 39
Figure 40
Figure 41
Figure 42
Figure 43
Figure 44
Figure 45
Figure 46
Figure 47
Figure 48
Figure 49
Figure 50
Figure 51
Figure 52
Figure 53
[0013] When using encoded data such as point clouds in an actual device or service, random access to a desired spatial position or object is essential. However, until now, random access in three-dimensional encoded data has not existed as a function, and there has also been no encoding method therefor.
[0014] In the present disclosure, a three-dimensional data encoding method, a three-dimensional data decoding method, a three-dimensional data encoding device, or a three-dimensional data decoding device that can provide a random access function in encoded three-dimensional data will be described.
[0015] A three-dimensional data encoding method according to an aspect of the present disclosure is a three-dimensional data encoding method for encoding three-dimensional data, the method including: a dividing step of dividing the three-dimensional data into first processing units that are random access units and each of which is associated with three-dimensional coordinates; and an encoding step of generating encoded data by encoding each of the plurality of first processing units.
[0016] According to this, random access for each first processing unit becomes possible. Thus, the three-dimensional data encoding method can provide a random access function in the encoded three-dimensional data.
[0017] For example, the three-dimensional data encoding method includes a generating step of generating first information indicating the plurality of first processing units and three-dimensional coordinates associated with each of the plurality of first processing units, and the encoded data may include the first information.
[0018] For example, the first information may further indicate at least one of an object, a time, and a data storage destination, which is associated with each of the plurality of first processing units.
[0019] For example, in the dividing step, the first processing unit may be further divided into a plurality of second processing units, and in the encoding step, each of the plurality of second processing units may be encoded.
[0020] For example, in the encoding step, a second processing unit to be processed included in the first processing unit to be processed may be encoded with reference to another second processing unit included in the first processing unit to be processed.
[0021] According to this, the encoding efficiency can be improved by referring to another second processing unit.
[0022] For example, in the encoding step, as the type of the second processing unit to be processed, any one of a first type that does not refer to another second processing unit, a second type that refers to one other second processing unit, and a third type that refers to two other second processing units may be selected, and the second processing unit to be processed may be encoded according to the selected type.
[0023] For example, in the encoding step, the frequency of selecting the first type may be changed according to the number or density of objects included in the three-dimensional data.
[0024] According to this, the random accessibility and the encoding efficiency, which are in a trade-off relationship, can be appropriately set.
[0025] For example, in the encoding step, the size of the first processing unit may be determined according to the number or density of objects or dynamic objects included in the three-dimensional data.
[0026] According to this, the random accessibility and the encoding efficiency, which are in a trade-off relationship, can be appropriately set.
[0027] For example, the first processing unit includes a plurality of layers that are spatially divided in a predetermined direction and each include one or more of the second processing units. In the encoding step, the second processing unit may be encoded with reference to the second processing unit included in the same layer as the second processing unit or a layer lower than the second processing unit.
[0028] According to this, for example, the random accessibility of an important layer in the system can be improved, and a decrease in the encoding efficiency can be suppressed.
[0029] For example, in the dividing step, a second processing unit including only static objects and a second processing unit including only dynamic objects may be assigned to different first processing units.
[0030] According to this, control of dynamic objects and static objects can be easily performed.
[0031] For example, in the encoding step, a plurality of dynamic objects may be individually encoded, and the encoded data of the plurality of dynamic objects may be associated with a second processing unit including only static objects.
[0032] According to this, control of dynamic objects and static objects can be easily performed.
[0033] For example, in the dividing step, the second processing unit may be further divided into a plurality of third processing units, and in the encoding step, each of the plurality of third processing units may be encoded.
[0034] For example, the third processing unit may include one or more voxels that are the minimum units to which position information is associated.
[0035] For example, the second processing unit may include a feature point group derived from information obtained by a sensor.
[0036] For example, the encoded data may include information indicating the encoding order of the plurality of first processing units.
[0037] For example, the encoded data may include information indicating the sizes of the plurality of first processing units.
[0038] For example, in the encoding step, the plurality of first processing units may be encoded in parallel.
[0039] Further, a three-dimensional data decoding method according to an aspect of the present disclosure is a three-dimensional data decoding method for decoding three-dimensional data, including a decoding step of generating three-dimensional data of the first processing unit by decoding each of the encoded data of the first processing units, which are random access units and each associated with three-dimensional coordinates.
[0040] According to this, random access for each first processing unit becomes possible. Thus, the three-dimensional data decoding method can provide a random access function in the encoded three-dimensional data.
[0041] Further, a three-dimensional data encoding apparatus according to an aspect of the present disclosure is a three-dimensional data encoding apparatus for encoding three-dimensional data, including a dividing unit that divides the three-dimensional data into first processing units, which are random access units and each associated with three-dimensional coordinates, and an encoding unit that generates encoded data by encoding each of the plurality of first processing units.
[0042] According to this, random access for each first processing unit becomes possible. Thus, the three-dimensional data encoding apparatus can provide a random access function in the encoded three-dimensional data.
[0043] Also, a three-dimensional data decoding device according to an aspect of the present disclosure is a three-dimensional data decoding device that decodes three-dimensional data, and may include a decoding unit that generates three-dimensional data of the first processing unit by decoding each of the encoded data of the first processing units, which are random access units and each associated with three-dimensional coordinates.
[0044] According to this, random access for each first processing unit becomes possible. Thus, the three-dimensional data decoding device can provide a random access function in the encoded three-dimensional data.
[0045] Note that the present disclosure is effective even when it does not necessarily perform random access, by virtue of a configuration that divides and encodes space, enabling quantization, prediction, etc. of space.
[0046] Also, a three-dimensional data encoding method according to an aspect of the present disclosure includes an extraction step of extracting second three-dimensional data having a feature amount equal to or greater than a threshold from first three-dimensional data, and a first encoding step of generating first encoded three-dimensional data by encoding the second three-dimensional data.
[0047] According to this, the three-dimensional data encoding method generates first encoded three-dimensional data obtained by encoding data having a feature amount equal to or greater than a threshold. Thereby, the data amount of the encoded three-dimensional data can be reduced as compared with the case of directly encoding the first three-dimensional data. Therefore, the three-dimensional data encoding method can reduce the data amount to be transmitted.
[0048] For example, the three-dimensional data encoding method may further include a second encoding step of generating second encoded three-dimensional data by encoding the first three-dimensional data.
[0049] According to this, the three-dimensional data encoding method can selectively transmit the first encoded three-dimensional data and the second encoded three-dimensional data, for example, according to the usage purpose or the like.
[0050] For example, the second three-dimensional data may be encoded by a first encoding method, and the first three-dimensional data may be encoded by a second encoding method different from the first encoding method.
[0051] According to this, the three-dimensional data encoding method can use encoding methods suitable for the first three-dimensional data and the second three-dimensional data, respectively.
[0052] For example, in the first encoding method, inter prediction may be prioritized over intra prediction and inter prediction compared to the second encoding method.
[0053] According to this, the three-dimensional data encoding method can increase the priority of inter prediction for the second three-dimensional data in which the correlation between adjacent data is likely to be low.
[0054] For example, the first encoding method and the second encoding method may differ in the expression method of three-dimensional positions.
[0055] According to this, the three-dimensional data encoding method can use a more suitable expression method of three-dimensional positions for three-dimensional data with different numbers of data.
[0056] For example, at least one of the first encoded three-dimensional data and the second encoded three-dimensional data may include an identifier indicating whether the encoded three-dimensional data is the encoded three-dimensional data obtained by encoding the first three-dimensional data or the encoded three-dimensional data obtained by encoding a part of the first three-dimensional data.
[0057] According to this, the decoding device can easily determine whether the acquired encoded three-dimensional data is the first encoded three-dimensional data or the second encoded three-dimensional data.
[0058] For example, in the first encoding step, the second three-dimensional data may be encoded such that the data amount of the first encoded three-dimensional data is smaller than the data amount of the second encoded three-dimensional data.
[0059] According to this, the three-dimensional data encoding method can make the data amount of the first encoded three-dimensional data smaller than the data amount of the second encoded three-dimensional data.
[0060] For example, in the extraction step, further, data corresponding to an object having a predetermined attribute may be extracted from the first three-dimensional data as the second three-dimensional data.
[0061] According to this, the three-dimensional data encoding method can generate the first encoded three-dimensional data including the data required by the decoding device.
[0062] For example, the three-dimensional data encoding method may further include a transmission step of transmitting one of the first encoded three-dimensional data and the second encoded three-dimensional data to the client according to the state of the client.
[0063] According to this, the three-dimensional data encoding method can transmit appropriate data according to the state of the client.
[0064] For example, the state of the client may include the communication status of the client or the moving speed of the client.
[0065] For example, the three-dimensional data encoding method may further include a transmission step of transmitting one of the first encoded three-dimensional data and the second encoded three-dimensional data to the client according to the request of the client.
[0066] According to this, the three-dimensional data encoding method can transmit appropriate data according to the request of the client.
[0067] In addition, a three-dimensional data decoding method according to one aspect of the present disclosure includes a first decoding step of decoding first encoded three-dimensional data obtained by encoding second three-dimensional data whose feature amount extracted from first three-dimensional data is equal to or greater than a threshold value by a first decoding method, and a second decoding step of decoding second encoded three-dimensional data obtained by encoding the first three-dimensional data by a second decoding method different from the first decoding method.
[0068] According to this, the three-dimensional data decoding method can selectively receive the first encoded three-dimensional data obtained by encoding data whose feature amount is equal to or greater than the threshold value and the second encoded three-dimensional data according to, for example, the usage purpose. Thereby, the three-dimensional data decoding method can reduce the amount of data to be transmitted. Further, the three-dimensional data decoding method can use a decoding method suitable for the first three-dimensional data and the second three-dimensional data, respectively.
[0069] For example, in the first decoding method, inter prediction may be prioritized over intra prediction among the intra prediction and the inter prediction compared to the second decoding method.
[0070] According to this, the three-dimensional data decoding method can increase the priority of inter prediction for the second three-dimensional data in which the correlation between adjacent data is likely to be low.
[0071] For example, the first decoding method and the second decoding method may differ in the expression method of the three-dimensional position.
[0072] According to this, the three-dimensional data decoding method can use a more suitable expression method of the three-dimensional position for three-dimensional data with different numbers of data.
[0073] For example, at least one of the first encoded three-dimensional data and the second encoded three-dimensional data includes an identifier indicating whether the encoded three-dimensional data is the encoded three-dimensional data obtained by encoding the first three-dimensional data or the encoded three-dimensional data obtained by encoding a part of the first three-dimensional data. With reference to the identifier, the first encoded three-dimensional data and the second encoded three-dimensional data may be identified.
[0074] According to this, the three-dimensional data decoding method can easily determine whether the acquired encoded three-dimensional data is the first encoded three-dimensional data or the second encoded three-dimensional data.
[0075] For example, the three-dimensional data decoding method may further include a notification step of notifying the server of the state of the client, and a reception step of receiving one of the first encoded three-dimensional data and the second encoded three-dimensional data transmitted from the server according to the state of the client.
[0076] According to this, the three-dimensional data decoding method can receive appropriate data according to the state of the client.
[0077] For example, the state of the client may include the communication status of the client or the moving speed of the client.
[0078] For example, the three-dimensional data decoding method may further include a request step of requesting one of the first encoded three-dimensional data and the second encoded three-dimensional data from the server, and a reception step of receiving one of the first encoded three-dimensional data and the second encoded three-dimensional data transmitted from the server according to the request.
[0079] According to this, the three-dimensional data decoding method can receive appropriate data according to the application.
[0080] In addition, a three-dimensional data encoding device according to one aspect of the present disclosure includes an extraction unit that extracts second three-dimensional data having a feature amount equal to or greater than a threshold value from first three-dimensional data, and a first encoding unit that generates first encoded three-dimensional data by encoding the second three-dimensional data.
[0081] According to this, the three-dimensional data encoding device generates first encoded three-dimensional data obtained by encoding data having a feature amount equal to or greater than a threshold value. As a result, the data amount can be reduced as compared with the case where the first three-dimensional data is encoded as it is. Therefore, the three-dimensional data encoding device can reduce the data amount to be transmitted.
[0082] In addition, a three-dimensional data decoding device according to one aspect of the present disclosure includes a first decoding unit that decodes first encoded three-dimensional data obtained by encoding second three-dimensional data having a feature amount equal to or greater than a threshold value extracted from first three-dimensional data by a first decoding method, and a second decoding unit that decodes second encoded three-dimensional data obtained by encoding the first three-dimensional data by a second decoding method different from the first decoding method.
[0083] According to this, the three-dimensional data decoding device can selectively receive the first encoded three-dimensional data obtained by encoding data having a feature amount equal to or greater than a threshold value and the second encoded three-dimensional data, for example, according to the usage purpose or the like. As a result, the three-dimensional data decoding device can reduce the data amount to be transmitted. Further, the three-dimensional data decoding device can use decoding methods suitable for the first three-dimensional data and the second three-dimensional data, respectively.
[0084] In addition, a three-dimensional data creation method according to one aspect of the present disclosure includes a creation step of creating first three-dimensional data from information detected by a sensor, a reception step of receiving encoded three-dimensional data obtained by encoding second three-dimensional data, a decoding step of obtaining the second three-dimensional data by decoding the received encoded three-dimensional data, and a synthesis step of creating third three-dimensional data by synthesizing the first three-dimensional data and the second three-dimensional data.
[0085] According to this, the three-dimensional data creation method can create detailed third three-dimensional data using the created first three-dimensional data and the received second three-dimensional data.
[0086] For example, in the synthesis step, by synthesizing the first three-dimensional data and the second three-dimensional data, third three-dimensional data having a higher density than the first three-dimensional data and the second three-dimensional data may be created.
[0087] For example, the second three-dimensional data may be three-dimensional data generated by extracting data having a feature amount equal to or greater than a threshold value from fourth three-dimensional data.
[0088] According to this, the three-dimensional data creation method can reduce the data amount of the three-dimensional data to be transmitted.
[0089] For example, the three-dimensional data creation method may further include a search step of searching for a transmission device that is a transmission source of the encoded three-dimensional data, and in the reception step, receiving the encoded three-dimensional data from the searched transmission device.
[0090] According to this, the three-dimensional data creation method can identify, for example, a transmission device that owns necessary three-dimensional data by searching.
[0091] For example, the three-dimensional data creation method may further include a determination step of determining a request range that is a range of a three-dimensional space that requests three-dimensional data, and a transmission step of transmitting information indicating the request range to the transmission device, and the second three-dimensional data may include three-dimensional data within the request range.
[0092] According to this, the three-dimensional data creation method can receive necessary three-dimensional data and reduce the data amount of the three-dimensional data to be transmitted.
[0093] For example, in the determination step, a space range including an occlusion area that cannot be detected by the sensor may be determined as the request range.
[0094] A three-dimensional data transmission method according to an aspect of the present disclosure includes a creation step of creating fifth three-dimensional data from information detected by a sensor, an extraction step of creating sixth three-dimensional data by extracting a part of the fifth three-dimensional data, an encoding step of generating encoded three-dimensional data by encoding the sixth three-dimensional data, and a transmission step of transmitting the encoded three-dimensional data.
[0095] According to this, the three-dimensional data transmission method can transmit the three-dimensional data created by itself to another device and can reduce the data amount of the transmitted three-dimensional data.
[0096] For example, in the creation step, seventh three-dimensional data may be created from the information detected by the sensor, and the fifth three-dimensional data may be created by extracting data having a feature amount equal to or greater than a threshold value from the seventh three-dimensional data.
[0097] According to this, the three-dimensional data transmission method can reduce the data amount of the transmitted three-dimensional data.
[0098] For example, the three-dimensional data transmission method may further include a reception step of receiving, from a receiving device, information indicating a request range that is a range of a three-dimensional space for requesting three-dimensional data. In the extraction step, the sixth three-dimensional data may be created by extracting the three-dimensional data within the request range from the fifth three-dimensional data. In the transmission step, the encoded three-dimensional data may be transmitted to the receiving device.
[0099] According to this, the three-dimensional data transmission method can reduce the data amount of the transmitted three-dimensional data.
[0100] In addition, a three-dimensional data creation device according to one aspect of the present disclosure includes a creation unit that creates first three-dimensional data from information detected by a sensor, a reception unit that receives encoded three-dimensional data in which second three-dimensional data is encoded, a decoding unit that acquires the second three-dimensional data by decoding the received encoded three-dimensional data, and a synthesis unit that creates third three-dimensional data by synthesizing the first three-dimensional data and the second three-dimensional data.
[0101] According to this, the three-dimensional data creation device can create detailed third three-dimensional data using the created first three-dimensional data and the received second three-dimensional data.
[0102] In addition, a three-dimensional data transmission device according to one aspect of the present disclosure includes a creation unit that creates fifth three-dimensional data from information detected by a sensor, an extraction unit that creates sixth three-dimensional data by extracting a part of the fifth three-dimensional data, an encoding unit that generates encoded three-dimensional data by encoding the sixth three-dimensional data, and a transmission unit that transmits the encoded three-dimensional data.
[0103] According to this, the three-dimensional data transmission device can transmit the three-dimensional data created by itself to another device and can reduce the data amount of the transmitted three-dimensional data.
[0104] In addition, a three-dimensional information processing method according to one aspect of the present disclosure includes an acquisition step of acquiring map data including first three-dimensional position information via a communication path, a generation step of generating second three-dimensional position information from information detected by a sensor, a determination step of determining whether the first three-dimensional position information or the second three-dimensional position information is abnormal by performing an abnormality determination process on the first three-dimensional position information or the second three-dimensional position information, a determination step of determining a coping operation for the abnormality when it is determined that the first three-dimensional position information or the second three-dimensional position information is abnormal, and an operation control step of performing control necessary for the implementation of the coping operation.
[0105] According to this, the three-dimensional information processing method can detect an abnormality in the first three-dimensional position information or the second three-dimensional position information and perform a countermeasure operation.
[0106] For example, the first three-dimensional position information may be encoded in units of partial spaces having three-dimensional coordinate information, each being an aggregate of one or more partial spaces, and may include a plurality of random access units that can be independently decoded.
[0107] According to this, the three-dimensional information processing method can reduce the data amount of the first three-dimensional position information to be acquired.
[0108] For example, the first three-dimensional position information may be data in which feature points where three-dimensional feature amounts are equal to or greater than a predetermined threshold are encoded.
[0109] According to this, the three-dimensional information processing method can reduce the data amount of the first three-dimensional position information to be acquired.
[0110] For example, in the determination step, it may be determined whether the first three-dimensional position information can be acquired via the communication path, and if the first three-dimensional position information cannot be acquired via the communication path, it may be determined that the first three-dimensional position information is abnormal.
[0111] According to this, the three-dimensional information processing method can perform an appropriate countermeasure operation when the first three-dimensional position information cannot be acquired according to the communication situation or the like.
[0112] For example, the three-dimensional information processing method further includes a self-position estimation step of performing self-position estimation of the moving body having the sensor using the first three-dimensional position information and the second three-dimensional position information. In the determination step, it is predicted whether the moving body enters a region with a poor communication state, and in the operation control step, if it is predicted that the moving body enters a region with a poor communication state, before the moving body enters the region, the moving body may acquire the first three-dimensional information.
[0113] According to this, when there is a possibility that the first three-dimensional position information cannot be obtained, the three-dimensional information processing method can obtain the first three-dimensional position information in advance.
[0114] For example, in the operation control step, when the first three-dimensional position information cannot be obtained via the communication path, third three-dimensional position information in a range narrower than the first three-dimensional position information may be obtained via the communication path.
[0115] According to this, since the three-dimensional information processing method can reduce the data amount of data obtained via the communication path, three-dimensional position information can be obtained even when the communication situation is poor.
[0116] For example, the three-dimensional information processing method further includes a self-position estimation step of estimating the self-position of the moving body having the sensor using the first three-dimensional position information and the second three-dimensional position information. In the operation control step, when the first three-dimensional position information cannot be obtained via the communication path, map data including two-dimensional position information is obtained via the communication path, and the self-position of the moving body having the sensor may be estimated using the two-dimensional position information and the second three-dimensional position information.
[0117] According to this, since the three-dimensional information processing method can reduce the data amount of data obtained via the communication path, three-dimensional position information can be obtained even when the communication situation is poor.
[0118] For example, the three-dimensional information processing method further includes an automatic driving step of automatically driving the moving body using the result of the self-position estimation. In the determination step, further, based on the moving environment of the moving body, it may be determined whether to perform the automatic driving of the moving body using the result of the self-position estimation of the moving body using the two-dimensional position information and the second three-dimensional position information.
[0119] According to this, the three-dimensional information processing method can determine whether to continue the automatic driving according to the moving environment of the moving body.
[0120] For example, the three-dimensional information processing method may further include an automatic driving step of automatically driving the moving body using the result of the self-position estimation, and in the operation control step, the mode of the automatic driving may be switched based on the moving environment of the moving body.
[0121] According to this, the three-dimensional information processing method can set an appropriate automatic driving mode according to the moving environment of the moving body.
[0122] For example, in the determination step, it may be determined whether the data of the first three-dimensional position information is complete, and if the data of the first three-dimensional position information is not complete, it may be determined that the first three-dimensional position information is abnormal.
[0123] According to this, the three-dimensional information processing method can perform an appropriate countermeasure operation, for example, when the first three-dimensional position information is damaged.
[0124] For example, in the determination step, it may be determined whether the generation accuracy of the data of the second three-dimensional position information is equal to or higher than a reference value, and if the generation accuracy of the data of the second three-dimensional position information is not equal to or higher than the reference value, it may be determined that the second three-dimensional position information is abnormal.
[0125] According to this, the three-dimensional information processing method can perform an appropriate countermeasure operation when the accuracy of the second three-dimensional position information is low.
[0126] For example, in the operation control step, when the generation accuracy of the data of the second three-dimensional position information is not equal to or higher than the reference value, the fourth three-dimensional position information may be generated from the information detected by an alternative sensor different from the sensor.
[0127] According to this, the three-dimensional information processing method can acquire three-dimensional position information using an alternative sensor, for example, when the sensor is malfunctioning.
[0128] For example, the three-dimensional information processing method further includes a self-position estimation step of estimating the self-position of the moving body having the sensor using the first three-dimensional position information and the second three-dimensional position information, and an automatic driving step of automatically driving the moving body using the result of the self-position estimation. In the operation control step, when the generation accuracy of the data of the second three-dimensional position information is not equal to or higher than the reference value, the automatic driving mode may be switched.
[0129] According to this, when the generation accuracy of the second three-dimensional position information is low, the three-dimensional information processing method can perform an appropriate countermeasure operation.
[0130] For example, in the operation control step, when the generation accuracy of the data of the second three-dimensional position information is not equal to or higher than the reference value, operation correction of the sensor may be performed.
[0131] According to this, when the generation accuracy of the second three-dimensional position information is low, the three-dimensional information processing method can increase the accuracy of the second three-dimensional position information.
[0132] In addition, a three-dimensional information processing apparatus according to an aspect of the present disclosure includes an acquisition unit that acquires map data including first three-dimensional position information via a communication path, a generation unit that generates second three-dimensional position information from information detected by a sensor, a determination unit that determines whether the first three-dimensional position information or the second three-dimensional position information is abnormal by performing an abnormality determination process on the first three-dimensional position information or the second three-dimensional position information, a determination unit that determines whether the first three-dimensional position information or the second three-dimensional position information is abnormal, and a determination unit that determines a countermeasure operation for the abnormality when it is determined that the first three-dimensional position information or the second three-dimensional position information is abnormal, and an operation control unit that performs control necessary for implementing the countermeasure operation.
[0133] According to this, the three-dimensional information processing apparatus can detect an abnormality in the first three-dimensional position information or the second three-dimensional position information and perform a countermeasure operation.
[0134] Moreover, a three-dimensional data creation method according to an aspect of the present disclosure is a three-dimensional data creation method in a mobile body including a sensor and a communication unit that transmits and receives three-dimensional data to and from the outside, the method including: a three-dimensional data creation step of creating second three-dimensional data based on information detected by the sensor and first three-dimensional data received by the communication unit; and a transmission step of transmitting third three-dimensional data, which is a part of the second three-dimensional data, to the outside.
[0135] According to this, the three-dimensional data creation method can generate three-dimensional data in a range that cannot be detected by the mobile body. That is, the three-dimensional data creation method can generate detailed three-dimensional data. Further, the three-dimensional data creation method can transmit three-dimensional data in a range that cannot be detected by other mobile bodies or the like to the other mobile bodies or the like.
[0136] For example, the three-dimensional data creation step and the transmission step may be repeatedly performed, and the third three-dimensional data may be three-dimensional data of a small space having a predetermined size located at a predetermined distance in front of the moving direction of the mobile body from the current position of the mobile body.
[0137] According to this, the data amount of the third three-dimensional data to be transmitted can be reduced.
[0138] For example, the predetermined distance may change according to the moving speed of the mobile body.
[0139] According to this, the three-dimensional data creation method can set an appropriate small space according to the moving speed of the mobile body and transmit the three-dimensional data of the small space to other mobile bodies or the like.
[0140] For example, the predetermined size may change according to the moving speed of the mobile body.
[0141] According to this, the three-dimensional data creation method can set an appropriate small space according to the moving speed of the mobile body and transmit the three-dimensional data of the small space to other mobile bodies or the like.
[0142] For example, the three-dimensional data creation method may further determine whether there is a change in the second three-dimensional data of the small space corresponding to the transmitted third three-dimensional data, and if there is a change, transmit at least a part of the fourth three-dimensional data, which is a part of the second three-dimensional data with the change, to the outside.
[0143] According to this, the three-dimensional data creation method can transmit the fourth three-dimensional data of the space with the change to other moving bodies or the like.
[0144] For example, the fourth three-dimensional data may be transmitted with higher priority than the third three-dimensional data.
[0145] According to this, the three-dimensional data creation method preferentially transmits the fourth three-dimensional data of the space with the change to other moving bodies or the like, so that other moving bodies or the like can quickly make judgments based on the three-dimensional data, for example.
[0146] For example, if there is a change, the fourth three-dimensional data may be transmitted before the transmission of the third three-dimensional data.
[0147] For example, the fourth three-dimensional data may indicate the difference between the second three-dimensional data of the small space corresponding to the transmitted third three-dimensional data and the second three-dimensional data after the change.
[0148] According to this, the three-dimensional data creation method can reduce the data volume of the three-dimensional data to be transmitted.
[0149] For example, in the transmission step, if there is no difference between the third three-dimensional data of the small space and the first three-dimensional data of the small space, the third three-dimensional data may not be transmitted.
[0150] According to this, the data volume of the transmitted third three-dimensional data can be reduced.
[0151] For example, in the transmission step, when there is no difference between the third three-dimensional data of the small space and the first three-dimensional data of the small space, information indicating that there is no difference between the third three-dimensional data of the small space and the first three-dimensional data of the small space may be transmitted to the outside.
[0152] For example, the information detected by the sensor may be three-dimensional data.
[0153] Moreover, a three-dimensional data creation device according to an aspect of the present disclosure is a three-dimensional data creation device mounted on a moving body, including a sensor, a receiving unit that receives first three-dimensional data from the outside, a creation unit that creates second three-dimensional data based on the information detected by the sensor and the first three-dimensional data, and a transmission unit that transmits third three-dimensional data, which is a part of the second three-dimensional data, to the outside.
[0154] According to this, the three-dimensional data creation device can generate three-dimensional data in a range that cannot be detected by the moving body. That is, the three-dimensional data creation device can generate detailed three-dimensional data. Further, the three-dimensional data creation device can transmit three-dimensional data in a range that cannot be detected by other moving bodies or the like to the other moving bodies or the like.
[0155] Moreover, a display method according to an aspect of the present disclosure is a display method in a display device that cooperates with and operates with a moving body, including a determination step of determining whether to display first peripheral information that is information indicating the peripheral situation of the moving body and is generated using two-dimensional data or second peripheral information that is information indicating the peripheral situation of the moving body and is generated using three-dimensional data based on the driving situation of the moving body, and a display step of displaying the determined first peripheral information or the second peripheral information.
[0156] According to this, the display method can switch between displaying either the first peripheral information generated using two-dimensional data and the second peripheral information generated using three-dimensional data based on the driving situation of the moving body. For example, when detailed information is required, the display method displays the second peripheral information with a large amount of information, and when not, it displays the first peripheral information with a small amount of data and processing amount, etc. Thereby, the display method can display appropriate information according to the situation and can reduce the communication data amount, processing amount, etc.
[0157] For example, the driving situation is whether the moving body is in autonomous movement or in manual driving. In the determination step, when the moving body is in autonomous movement, it may be determined to display the first peripheral information, and when the moving body is in manual driving, it may be determined to display the second peripheral information.
[0158] According to this, detailed information can be displayed during manual driving, and the processing amount, etc. during automatic driving can be reduced.
[0159] For example, the driving situation may be the area where the moving body is located.
[0160] According to this, appropriate information can be displayed according to the position of the moving body.
[0161] For example, the three-dimensional data may be data obtained by extracting point groups with a feature amount equal to or greater than a threshold value from three-dimensional point cloud data.
[0162] According to this, the communication data amount or the data amount to be stored can be reduced.
[0163] For example, the three-dimensional data may be data having a mesh structure generated from three-dimensional point cloud data.
[0164] According to this, the communication data amount or the data amount to be stored can be reduced.
[0165] For example, the three-dimensional data may be data obtained by extracting, from three-dimensional point cloud data, point clouds whose feature amounts are equal to or greater than a threshold value and that are necessary for self-position estimation, driving assistance, or autonomous driving.
[0166] According to this, the communication data amount or the data amount to be stored can be reduced.
[0167] For example, the three-dimensional data may be three-dimensional point cloud data.
[0168] According to this, the accuracy of the second peripheral information can be improved.
[0169] For example, in the display step, the second peripheral information is displayed on a head-up display, and the display method may further adjust the display position of the second peripheral information according to the posture, body shape, or eye position of the user boarding the moving body.
[0170] According to this, information can be displayed at an appropriate position according to the posture, body shape, or eye position of the user.
[0171] Moreover, a display device according to an aspect of the present disclosure is a display device that cooperates with and operates with a moving body, and includes a determination unit that determines whether to display first peripheral information that is an image showing the peripheral situation of the moving body and is generated using two-dimensional data or second peripheral information that is an image showing the peripheral situation of the moving body and is generated using three-dimensional data based on the driving situation of the moving body, and a display unit that displays the first peripheral information or the second peripheral information determined to be displayed.
[0172] According to this, the display device can switch between displaying either the first peripheral information generated using two-dimensional data or the second peripheral information generated using three-dimensional data based on the driving situation of the moving body. For example, when detailed information is required, the display device displays the second peripheral information with a large amount of information, and when not, it displays the first peripheral information with a small amount of data and processing amount, etc. Thereby, the display device can display appropriate information according to the situation and can reduce the communication data amount, processing amount, etc.
[0173] Moreover, a three-dimensional data creation method according to an aspect of the present disclosure is a three-dimensional data creation method in a client device mounted on a moving body, which creates three-dimensional data of the periphery of the moving body from sensor information indicating the peripheral situation of the moving body obtained by a sensor mounted on the moving body, estimates the self-position of the moving body using the created three-dimensional data, and transmits the acquired sensor information to a server or another moving body.
[0174] According to this, the three-dimensional data creation method transmits sensor information to a server or the like. Thereby, there is a possibility that the data amount of the transmitted data can be reduced compared to the case of transmitting three-dimensional data. Also, since it is not necessary to perform processing such as compression or encoding of three-dimensional data in the client device, the processing amount of the client device can be reduced. Therefore, the three-dimensional data creation method can achieve reduction of the data amount to be transmitted or simplification of the device configuration.
[0175] For example, the three-dimensional data creation method may further transmit a transmission request for a three-dimensional map to the server, receive the three-dimensional map from the server, and estimate the self-position using the three-dimensional data and the three-dimensional map in the estimation of the self-position.
[0176] For example, the sensor information may include at least one of information obtained by a laser sensor, a luminance image, an infrared image, a depth image, sensor position information, and sensor speed information.
[0177] For example, the sensor information may include at least one of information obtained by a laser sensor, a luminance image, an infrared image, a depth image, the position information of the sensor, and the speed information of the sensor.
[0178] For example, the sensor information may be encoded or compressed, and in the transmission of the sensor information, the encoded or compressed sensor information may be transmitted to the server or the other mobile body.
[0179] According to this, the three-dimensional data creation method can reduce the amount of data to be transmitted.
[0180] Further, a three-dimensional data creation method according to an aspect of the present disclosure is a three-dimensional data creation method in a server capable of communicating with a client device mounted on a mobile body, and is obtained by a sensor mounted on the mobile body. Sensor information indicating the surrounding situation of the mobile body is received from the client device, and three-dimensional data around the mobile body is created from the received sensor information.
[0181] According to this, the three-dimensional data creation method creates three-dimensional data using the sensor information transmitted from the client device. Thereby, there is a possibility that the amount of transmitted data can be reduced as compared with the case where the client device transmits three-dimensional data. In addition, since it is not necessary for the client device to perform processing such as compression or encoding of three-dimensional data, the processing amount of the client device can be reduced. Therefore, the three-dimensional data creation method can realize reduction of the amount of data to be transmitted or simplification of the device configuration.
[0182] For example, the three-dimensional data creation method may further transmit a transmission request for the sensor information to the client device.
[0183] For example, the three-dimensional data creation method may further update a three-dimensional map using the created three-dimensional data, and transmit the three-dimensional map to the client device in response to a transmission request for the three-dimensional map from the client device.
[0184] For example, the sensor information may include at least one of information obtained by a laser sensor, a luminance image, an infrared image, a depth image, the position information of the sensor, and the speed information of the sensor.
[0185] For example, the sensor information may include information indicating the performance of the sensor.
[0186] For example, the three-dimensional data creation method may further correct the three-dimensional data according to the performance of the sensor.
[0187] According to this, the three-dimensional data creation method can improve the quality of the three-dimensional data.
[0188] For example, in receiving the sensor information, a plurality of sensor information may be received from a plurality of client devices, and based on the plurality of information indicating the performance of the sensor included in the plurality of sensor information, the sensor information to be used for creating the three-dimensional data may be selected.
[0189] According to this, the three-dimensional data creation method can improve the quality of the three-dimensional data.
[0190] For example, the received sensor information may be decoded or decompressed, and the three-dimensional data may be created from the decoded or decompressed sensor information.
[0191] According to this, the three-dimensional data creation method can reduce the amount of data to be transmitted.
[0192] In addition, a client device according to an aspect of the present disclosure is a client device mounted on a moving body, including a processor and a memory. The processor uses the memory to create three-dimensional data of the periphery of the moving body from sensor information indicating the peripheral situation of the moving body obtained by a sensor mounted on the moving body, estimates the self-position of the moving body using the created three-dimensional data, and transmits the acquired sensor information to a server or another moving body.
[0193] According to this, the client device transmits sensor information to a server or the like. Thereby, compared with the case of transmitting three-dimensional data, there is a possibility of reducing the amount of data to be transmitted. In addition, since it is not necessary for the client device to perform processing such as compression or encoding of three-dimensional data, the processing amount of the client device can be reduced. Therefore, the client device can reduce the amount of data to be transmitted or simplify the configuration of the device.
[0194] In addition, a server according to an aspect of the present disclosure is a server capable of communicating with a client device mounted on a moving body, and includes a processor and a memory. The processor uses the memory to receive, from the client device, sensor information indicating the surrounding situation of the moving body obtained by a sensor mounted on the moving body, and creates three-dimensional data around the moving body from the received sensor information.
[0195] According to this, the server creates three-dimensional data using the sensor information transmitted from the client device. Thereby, compared with the case where the client device transmits three-dimensional data, there is a possibility of reducing the amount of data to be transmitted. In addition, since it is not necessary for the client device to perform processing such as compression or encoding of three-dimensional data, the processing amount of the client device can be reduced. Therefore, the server can reduce the amount of data to be transmitted or simplify the configuration of the device.
[0196] Note that these general or specific aspects may be implemented by a system, a method, an integrated circuit, a computer program, or a recording medium such as a computer-readable CD-ROM, or may be implemented by any combination of a system, a method, an integrated circuit, a computer program, and a recording medium.
[0197] Hereinafter, embodiments will be specifically described with reference to the drawings. Note that each of the embodiments described below shows a specific example of the present disclosure. The numerical values, shapes, materials, components, arrangement positions and connection forms of the components, steps, order of steps, etc. shown in the following embodiments are merely examples and are not intended to limit the present disclosure. In addition, among the components in the following embodiments, components not described in the independent claims indicating the most general concept are described as optional components.
[0198] (Embodiment 1) First, the data structure of the encoded three-dimensional data (hereinafter also referred to as encoded data) according to the present embodiment will be described. FIG. 1 is a diagram showing the configuration of the encoded three-dimensional data according to the present embodiment.
[0199] In the present embodiment, the three-dimensional space is divided into spaces (SPC) corresponding to pictures in the encoding of moving images, and three-dimensional data is encoded in units of spaces. The space is further divided into volumes (VLM) corresponding to macroblocks or the like in moving image encoding, and prediction and conversion are performed in units of VLM. A volume includes a plurality of voxels (VXL) which are the smallest units to which position coordinates are associated. Note that prediction is, similar to the prediction performed on a two-dimensional image, to generate predicted three-dimensional data similar to the processing unit to be processed by referring to other processing units, and to encode the difference between the predicted three-dimensional data and the processing unit to be processed. Further, this prediction includes not only spatial prediction that refers to other prediction units at the same time but also temporal prediction that refers to prediction units at different times.
[0200] For example, when encoding a three-dimensional space represented by point cloud data such as point cloud, a three-dimensional data encoding device (hereinafter also referred to as an encoding device) encodes each point of the point cloud or a plurality of points included in a voxel together according to the size of the voxel. If the voxel is subdivided, the three-dimensional shape of the point cloud can be expressed with high precision, and if the size of the voxel is increased, the three-dimensional shape of the point cloud can be expressed roughly.
[0201] In the following, the case where the three-dimensional data is a point cloud will be described as an example. However, the three-dimensional data is not limited to a point cloud and may be three-dimensional data in any format.
[0202] Also, hierarchical voxels may be used. In this case, in the n-th hierarchy, it may be sequentially indicated whether there are sample points in the hierarchies of n-1 or lower (the lower layers of the n-th hierarchy). For example, when decoding only the n-th hierarchy, if there are sample points in the hierarchies of n-1 or lower, it can be decoded assuming that there are sample points at the center of the voxels in the n-th hierarchy.
[0203] Also, the encoding device acquires point cloud data using a distance sensor, a stereo camera, a monocular camera, a gyro, or an inertial sensor, etc.
[0204] Space is classified into any of at least three prediction structures including an intra-space (I-SPC) that can be decoded independently, a predictive space (P-SPC) that allows only unidirectional reference, and a bidirectional space (B-SPC) that allows bidirectional reference, similar to the encoding of moving images. Also, space has two types of time information, the decoding time and the display time.
[0205] Also, as shown in FIG. 1, there is a GOS (Group Of Space) which is a random access unit as a processing unit including a plurality of spaces. Further, there is a world (WLD) as a processing unit including a plurality of GOSs.
[0206] The space region occupied by the world is associated with the absolute position on the earth by GPS or latitude and longitude information, etc. This position information is stored as meta information. Note that the meta information may be included in the encoded data or may be transmitted separately from the encoded data.
[0207] Also, within the GOS, all SPCs may be three-dimensionally adjacent, or there may be SPCs that are not three-dimensionally adjacent to other SPCs.
[0208] In the following, the processes such as encoding, decoding, or referencing for three-dimensional data included in processing units such as GOS, SPC, or VLM are also simply described as encoding, decoding, or referencing the processing unit, etc. Further, the three-dimensional data included in the processing unit includes, for example, at least one set of a spatial position such as three-dimensional coordinates and a characteristic value such as color information.
[0209] Next, the prediction structure of SPC in GOS will be described. A plurality of SPCs within the same GOS or a plurality of VLMs within the same SPC occupy different spaces from each other but have the same time information (decoding time and display time).
[0210] Also, the SPC that is the first in the decoding order within the GOS is the I-SPC. Further, there are two types of GOS: closed GOS and open GOS. The closed GOS is a GOS in which all SPCs within the GOS can be decoded when starting decoding from the first I-SPC. In the open GOS, some SPCs whose display time is earlier than that of the first I-SPC within the GOS refer to different GOSs, and decoding cannot be performed only with this GOS.
[0211] Note that in encoded data such as map information, the WLD may be decoded from the direction opposite to the encoding order, and reverse playback is difficult if there is a dependency between GOSs. Therefore, in such a case, basically, a closed GOS is used.
[0212] Also, the GOS has a layer structure in the height direction, and encoding or decoding is performed in order from the SPCs in the lower layer.
[0213] FIG. 2 is a diagram showing an example of the prediction structure between SPCs belonging to the bottom layer of the GOS. FIG. 3 is a diagram showing an example of the prediction structure between layers.
[0214] There is one or more I-SPCs in the GOS. In the three-dimensional space, there are objects such as humans, animals, cars, bicycles, signals, or buildings serving as landmarks. In particular, objects with a relatively small size are effectively encoded as I-SPCs. For example, when a three-dimensional data decoder (hereinafter also referred to as a decoder) decodes the GOS with a low processing volume or at high speed, only the I-SPCs in the GOS are decoded.
[0215] Also, the encoding device may switch the encoding interval or the appearance frequency of the I-SPC according to the coarseness of the objects in the WLD.
[0216] Also, in the configuration shown in FIG. 3, the encoding device or the decoding device encodes or decodes a plurality of layers in order from the lower layer (layer 1). Thereby, for example, for an autonomous vehicle, the priority of data near the ground with a larger amount of information can be increased.
[0217] Note that in the encoded data used in a drone or the like, encoding or decoding may be performed in order from the SPC of the upper layer in the height direction within the GOS.
[0218] Also, the encoding device or the decoding device may encode or decode a plurality of layers so that the decoding device can roughly grasp the GOS and gradually increase the resolution. For example, the encoding device or the decoding device may encode or decode in the order of layer 3, 8, 1, 9,....
[0219] Next, how to handle static objects and dynamic objects will be described.
[0220] In the three-dimensional space, there are static objects or scenes such as buildings or roads (hereinafter collectively referred to as static objects) and dynamic objects such as cars or humans (hereinafter referred to as dynamic objects). Detection of objects is separately performed by extracting feature points from point cloud data, camera images such as stereo cameras, and the like. Here, an example of an encoding method for dynamic objects will be described.
[0221] The first method is a method of encoding without distinguishing between static objects and dynamic objects. The second method is a method of distinguishing between static objects and dynamic objects using identification information.
[0222] For example, GOS is used as an identification unit. In this case, a GOS including SPCs constituting a static object and a GOS including SPCs constituting a dynamic object are distinguished by identification information stored in the encoded data or separately from the encoded data.
[0223] Alternatively, SPC may be used as an identification unit. In this case, an SPC including VLMs constituting a static object and an SPC including VLMs constituting a dynamic object are distinguished by the above identification information.
[0224] Alternatively, VLM or VXL may be used as an identification unit. In this case, a VLM or VXL including a static object and a VLM or VXL including a dynamic object are distinguished by the above identification information.
[0225] In addition, the encoding device may encode a dynamic object as one or more VLMs or SPCs, and encode an SPC or VLM including a static object and an SPC including a dynamic object as different GOSs. Further, when the size of the GOS is variable according to the size of the dynamic object, the encoding device stores the size of the GOS separately as meta information.
[0226] In addition, the encoding device may encode a static object and a dynamic object independently of each other, and may superimpose the dynamic object on a world composed of static objects. At this time, the dynamic object is composed of one or more SPCs, and each SPC is associated with one or more SPCs constituting the static object on which the SPC is superimposed. Note that the dynamic object may be represented by one or more VLMs or VXLs instead of SPCs.
[0227] Further, the encoding device may encode static objects and dynamic objects as different streams from each other.
[0228] Also, the encoding device may generate a GOS including one or more SPCs constituting a dynamic object. Further, the encoding device may set a GOS (GOS_M) including a dynamic object and a GOS of a static object corresponding to the spatial region of GOS_M to the same size (occupying the same spatial region). Thereby, the overlapping process can be performed in units of GOS.
[0229] The P-SPC or B-SPC constituting the dynamic object may refer to SPCs included in different encoded GOSs. In a case where the position of the dynamic object changes temporally and the same dynamic object is encoded as GOSs at different times, the cross-GOS reference is effective from the viewpoint of the compression ratio.
[0230] Also, depending on the use of the encoded data, the above first method and second method may be switched. For example, when using the encoded three-dimensional data as a map, it is desirable to be able to separate the dynamic object, so the encoding device uses the second method. On the other hand, when encoding three-dimensional data of an event such as a concert or a sport, the encoding device uses the first method if there is no need to separate the dynamic object.
[0231] Also, the decoding time and display time of GOS or SPC can be stored in the encoded data or as meta information. Also, all the time information of static objects may be the same. At this time, the actual decoding time and display time may be determined by the decoding device. Alternatively, different values may be assigned for each GOS or SPC as the decoding time, and the same value may be assigned for all as the display time. Further, a model may be introduced that guarantees that the decoder has a buffer of a predetermined size and can decode without breakdown by reading the bit stream at a predetermined bit rate according to the decoding time, such as the decoder model in video coding such as HEVC's HRD (Hypothetical Reference Decoder).
[0232] Next, the arrangement of GOS in the world will be described. The coordinates of the three-dimensional space in the world are represented by three mutually orthogonal coordinate axes (x-axis, y-axis, z-axis). By providing a predetermined rule in the encoding order of GOS, encoding can be performed so that spatially adjacent GOS are continuous in the encoded data. For example, in the example shown in FIG. 4, the GOS in the xz plane is encoded continuously. After encoding all the GOS in a certain xz plane, the value of the y-axis is updated. That is, as the encoding progresses, the world extends in the y-axis direction. Also, the index number of GOS is set in the encoding order.
[0233] Here, the three-dimensional space of the world is associated one-to-one with geographical absolute coordinates such as GPS or latitude and longitude. Alternatively, the three-dimensional space may be represented by the relative position from a preset reference position. The directions of the x-axis, y-axis, and z-axis of the three-dimensional space are represented as direction vectors determined based on latitude and longitude, etc., and the direction vectors are stored as meta information together with the encoded data.
[0234] Also, the size of the GOS is fixed, and the encoding device stores the size as meta information. Also, the size of the GOS may be switched according to, for example, whether it is an urban area or whether it is indoor or outdoor. That is, the size of the GOS may be switched according to the amount or nature of the object that is valuable as information. Alternatively, the encoding device may adaptively switch the size of the GOS or the interval of I-SPCs in the GOS according to, for example, the density of objects in the same world. For example, the higher the density of the objects, the smaller the size of the GOS and the shorter the interval of I-SPCs in the GOS.
[0235] In the example of FIG. 5, in the regions of the 3rd to 10th GOSs, since the density of objects is high, the GOS is subdivided in order to realize random access with a fine granularity. Note that the 7th to 10th GOSs are respectively located behind the 3rd to 6th GOSs.
[0236] Next, the configuration and operation flow of the three-dimensional data encoding device according to the present embodiment will be described. FIG. 6 is a block diagram of a three-dimensional data encoding device 100 according to the present embodiment. FIG. 7 is a flowchart showing an operation example of the three-dimensional data encoding device 100.
[0237] The three-dimensional data encoding device 100 shown in FIG. 6 generates encoded three-dimensional data 112 by encoding three-dimensional data 111. This three-dimensional data encoding device 100 includes an acquisition unit 101, an encoding region determination unit 102, a division unit 103, and an encoding unit 104.
[0238] As shown in FIG. 7, first, the acquisition unit 101 acquires three-dimensional data 111 which is point cloud data (S101).
[0239] Next, the encoding region determination unit 102 determines a region to be encoded among the spatial regions corresponding to the acquired point cloud data (S102). For example, the encoding region determination unit 102 determines the spatial region around the position as the region to be encoded according to the position of the user or the vehicle.
[0240] Next, the division unit 103 divides the point cloud data included in the area to be encoded into each processing unit. Here, the processing unit is the above-mentioned GOS, SPC, etc. Also, this area to be encoded corresponds to, for example, the above-mentioned world. Specifically, the division unit 103 divides the point cloud data into processing units based on a preset GOS size, or the presence or size of a dynamic object (S103). Also, the division unit 103 determines the start position of the SPC that is the head in the encoding order in each GOS.
[0241] Next, the encoding unit 104 generates encoded three-dimensional data 112 by sequentially encoding a plurality of SPCs within each GOS (S104).
[0242] Here, an example in which the area to be encoded is divided into GOS and SPC and then each GOS is encoded is shown, but the processing procedure is not limited to the above. For example, after determining the configuration of one GOS, that GOS may be encoded, and then procedures such as determining the configuration of the next GOS may be used.
[0243] In this way, the three-dimensional data encoding device 100 generates encoded three-dimensional data 112 by encoding the three-dimensional data 111. Specifically, the three-dimensional data encoding device 100 divides the three-dimensional data into first processing units (GOS) that are random access units and each of which is associated with three-dimensional coordinates, divides the first processing units (GOS) into a plurality of second processing units (SPC), and divides the second processing units (SPC) into a plurality of third processing units (VLM). Also, the third processing unit (VLM) includes one or more voxels (VXL) that are the minimum units to which position information is associated.
[0244] Next, the three-dimensional data encoding device 100 generates encoded three-dimensional data 112 by encoding each of a plurality of first processing units (GOS). Specifically, the three-dimensional data encoding device 100 encodes each of a plurality of second processing units (SPC) in each first processing unit (GOS). Further, the three-dimensional data encoding device 100 encodes each of a plurality of third processing units (VLM) in each second processing unit (SPC).
[0245] For example, when the first processing unit (GOS) to be processed is a closed GOS, the three-dimensional data encoding device 100 encodes the second processing unit (SPC) to be processed included in the first processing unit (GOS) to be processed with reference to other second processing units (SPC) included in the first processing unit (GOS) to be processed. That is, the three-dimensional data encoding device 100 does not refer to the second processing units (SPC) included in a first processing unit (GOS) different from the first processing unit (GOS) to be processed.
[0246] On the other hand, when the first processing unit (GOS) to be processed is an open GOS, the three-dimensional data encoding device 100 encodes the second processing unit (SPC) to be processed included in the first processing unit (GOS) to be processed with reference to other second processing units (SPC) included in the first processing unit (GOS) to be processed or the second processing units (SPC) included in a first processing unit (GOS) different from the first processing unit (GOS) to be processed.
[0247] In addition, the three-dimensional data encoding device 100 selects any one of a first type (I-SPC) that does not refer to other second processing units (SPC), a second type (P-SPC) that refers to one other second processing unit (SPC), and a third type that refers to two other second processing units (SPC) as the type of the second processing unit (SPC) to be processed, and encodes the second processing unit (SPC) to be processed according to the selected type.
[0248] Next, the configuration and operation flow of the three-dimensional data decoding apparatus according to the present embodiment will be described. FIG. 8 is a block diagram of the blocks of the three-dimensional data decoding apparatus 200 according to the present embodiment. FIG. 9 is a flowchart showing an operation example of the three-dimensional data decoding apparatus 200.
[0249] The three-dimensional data decoding apparatus 200 shown in FIG. 8 generates decoded three-dimensional data 212 by decoding the encoded three-dimensional data 211. Here, the encoded three-dimensional data 211 is, for example, the encoded three-dimensional data 112 generated by the three-dimensional data encoding apparatus 100. This three-dimensional data decoding apparatus 200 includes an acquisition unit 201, a decoding start GOS determination unit 202, a decoding SPC determination unit 203, and a decoding unit 204.
[0250] First, the acquisition unit 201 acquires the encoded three-dimensional data 211 (S201). Next, the decoding start GOS determination unit 202 determines the GOS to be decoded (S202). Specifically, the decoding start GOS determination unit 202 refers to the meta information stored in the encoded three-dimensional data 211 or separately from the encoded three-dimensional data, and determines the GOS including the SPC corresponding to the spatial position, object, or time at which decoding is to be started as the GOS to be decoded.
[0251] Next, the decoding SPC determination unit 203 determines the type (I, P, B) of the SPC to be decoded within the GOS (S203). For example, the decoding SPC determination unit 203 determines whether to decode only the I-SPC, (2) whether to decode the I-SPC and the P-SPC, or (3) whether to decode all types. Note that if the type of the SPC to be decoded is determined in advance, such as decoding all SPCs, this step may not be performed.
[0252] Next, the decoding unit 204 acquires the address position at which the first SPC in the decoding order (the same as the encoding order) within the GOS starts in the encoded three-dimensional data 211, acquires the encoded data of the first SPC from the address position, and sequentially decodes each SPC in order from the first SPC (S204). Note that the above address position is stored in meta information or the like.
[0253] In this way, the three-dimensional data decoder 200 decodes the decoded three-dimensional data 212. Specifically, the three-dimensional data decoder 200 decodes each of the encoded three-dimensional data 211 of the first processing units (GOS), which are random access units and each associated with three-dimensional coordinates, to generate the decoded three-dimensional data 212 of the first processing units (GOS). More specifically, the three-dimensional data decoder 200 decodes each of the plurality of second processing units (SPC) in each first processing unit (GOS). Further, the three-dimensional data decoder 200 decodes each of the plurality of third processing units (VLM) in each second processing unit (SPC).
[0254] Next, the meta information for random access will be described. This meta information is generated by the three-dimensional data encoder 100 and included in the encoded three-dimensional data 112 (211).
[0255] In the random access in conventional two-dimensional moving images, decoding starts from the first frame of the random access unit near the specified time. On the other hand, in the world, random access to space (coordinates or objects, etc.) in addition to time is assumed.
[0256] Therefore, in order to realize random access to at least three elements of coordinates, objects, and time, a table is prepared that associates each element with the index number of the GOS. Further, the index number of the GOS is associated with the address of the first I-SPC of the GOS. FIG. 10 is a diagram showing an example of the table included in the meta information. Note that it is not necessary to use all the tables shown in FIG. 10, and at least one table may be used.
[0257] Hereinafter, as an example, random access starting from coordinates will be described. When accessing the coordinates (x2, y2, z2), first, by referring to the coordinate-GOS table, it can be known that the point with the coordinates (x2, y2, z2) is included in the second GOS. Next, by referring to the GOS address table, since it is found that the address of the first I-SPC in the second GOS is addr(2), the decoding unit 204 acquires data from this address and starts decoding.
[0258] Note that the address may be an address in the logical format or a physical address of the HDD or memory. Also, information specifying a file segment may be used instead of the address. For example, a file segment is a unit obtained by segmenting one or more GOSs, etc.
[0259] Also, when an object spans multiple GOSs, in the object-GOS table, multiple GOSs to which the object belongs may be indicated. If the multiple GOSs are closed GOSs, the encoding device and the decoding device can perform encoding or decoding in parallel. On the other hand, if the multiple GOSs are open GOSs, the compression efficiency can be improved by the multiple GOSs referring to each other.
[0260] Examples of objects include humans, animals, cars, bicycles, signals, or landmark buildings. For example, the three-dimensional data encoding device 100 can extract feature points unique to an object from three-dimensional point clouds or the like during world encoding, detect the object based on the feature points, and set the detected object as a random access point.
[0261] In this way, the three-dimensional data encoding device 100 generates first information indicating a plurality of first processing units (GOS) and three-dimensional coordinates associated with each of the plurality of first processing units (GOS). Further, the encoded three-dimensional data 112(211) includes this first information. The first information further indicates at least one of an object, a time, and a data storage destination associated with each of the plurality of first processing units (GOS).
[0262] The three-dimensional data decoding device 200 acquires the first information from the encoded three-dimensional data 211, uses the first information to identify the encoded three-dimensional data 211 of the first processing unit corresponding to the specified three-dimensional coordinates, object, or time, and decodes the encoded three-dimensional data 211.
[0263] Hereinafter, examples of other meta-information will be described. In addition to the meta-information for random access, the three-dimensional data encoding device 100 may generate and store the following meta-information. Further, the three-dimensional data decoding device 200 may use this meta-information at the time of decoding.
[0264] When using three-dimensional data as map information, etc., a profile is defined according to the application, and information indicating the profile may be included in the meta-information. For example, profiles for urban areas or suburbs, or for flying objects are defined, and the maximum or minimum sizes of the world, SPC, or VLM are defined in each case. For example, for urban areas, more detailed information is required than for suburbs, so the minimum size of the VLM is set small.
[0265] The meta-information may include a tag value indicating the type of object. This tag value is associated with the VLM, SPC, or GOS that constitutes the object. For example, the tag value "0" indicates "person", the tag value "1" indicates "car", the tag value "2" indicates "traffic signal", etc., and tag values may be set for each type of object. Alternatively, when the type of object is difficult to determine or there is no need to determine it, a tag value indicating a property such as size or whether it is a dynamic object or a static object may be used.
[0266] Further, the meta information may include information indicating the range of the spatial area occupied by the world.
[0267] Further, the meta information may store the size of the SPC or VXL as the entire stream of encoded data or as header information common to a plurality of SPCs such as SPCs within the GOS, e.g., the SPCs within the GOS.
[0268] Further, the meta information may include identification information of a distance sensor or a camera used for generating the point cloud, or information indicating the positional accuracy of the point group within the point cloud.
[0269] Further, the meta information may include information indicating whether the world is composed only of static objects or includes dynamic objects.
[0270] Hereinafter, modifications of the present embodiment will be described.
[0271] The encoding device or the decoding device may encode or decode two or more different SPCs or GOSs in parallel. The GOSs to be encoded or decoded in parallel can be determined based on meta information indicating the spatial position of the GOSs.
[0272] In a case where the three-dimensional data is used as a spatial map when a vehicle or a flying object moves, or such a spatial map is generated, the encoding device or the decoding device may encode or decode the GOS or SPC included in the space specified based on GPS, route information, or zoom ratio.
[0273] Further, the decoding device may perform decoding in order from the space close to its own position or the traveling route. The encoding device or the decoding device may encode or decode the space far from its own position or the traveling route with a lower priority than the space close to it. Here, lowering the priority means lowering the processing order, lowering the resolution (subsampling for processing), or lowering the image quality (increasing the encoding efficiency. For example, increasing the quantization step).
[0274] Also, when decoding encoded data hierarchically encoded in space, the decoding device may decode only the lower layer.
[0275] Also, the decoding device may preferentially decode from the lower layer according to the zoom magnification or usage of the map.
[0276] Also, in applications such as self-position estimation or object recognition performed during autonomous driving of a vehicle or a robot, the encoding device or the decoding device may reduce the resolution and perform encoding or decoding outside the area within a specific height from the road surface (the area where recognition is performed).
[0277] Also, the encoding device may separately encode point clouds representing the spatial shapes of the indoor and outdoor spaces. For example, by separating the GOS representing the indoor area (indoor GOS) and the GOS representing the outdoor area (outdoor GOS), the decoding device can select the GOS to be decoded according to the viewpoint position when using the encoded data.
[0278] Also, the encoding device may encode the indoor GOS and the outdoor GOS with close coordinates so that they are adjacent in the encoding stream. For example, the encoding device associates the identifiers of both and stores information indicating the associated identifiers in the encoding stream or separately stored meta information. Thereby, the decoding device can identify the indoor GOS and the outdoor GOS with close coordinates by referring to the information in the meta information.
[0279] Also, the encoding device may switch the size of the GOS or SPC between the indoor GOS and the outdoor GOS. For example, the encoding device sets the size of the GOS smaller indoors than outdoors. Also, the encoding device may change the accuracy when extracting feature points from the point cloud or the accuracy of object detection, etc. between the indoor GOS and the outdoor GOS.
[0280] Further, the encoding device may add information for the decoding device to distinguish and display dynamic objects from static objects to the encoded data. Thereby, the decoding device can display a dynamic object together with a red frame or explanatory characters. Note that the decoding device may display only a red frame or explanatory characters instead of the dynamic object. Further, the decoding device may display more detailed object types. For example, a red frame may be used for a vehicle, and a yellow frame may be used for a human.
[0281] Further, the encoding device or the decoding device may determine whether to encode or decode a dynamic object and a static object as different SPCs or GOs according to the appearance frequency of the dynamic object or the ratio between the static object and the dynamic object. For example, when the appearance frequency or ratio of the dynamic object exceeds a threshold, an SPC or GO in which the dynamic object and the static object are mixed is allowed, and when the appearance frequency or ratio of the dynamic object does not exceed the threshold, an SPC or GO in which the dynamic object and the static object are mixed is not allowed.
[0282] When detecting a dynamic object from two-dimensional image information of a camera instead of a point cloud, the encoding device may separately obtain information (such as a frame or characters) for identifying the detection result and the object position, and encode these information as part of the three-dimensional encoded data. In this case, the decoding device superimposes and displays auxiliary information (a frame or characters) indicating the dynamic object on the decoding result of the static object.
[0283] Further, the encoding device may change the coarseness of VXL or VLM in the SPC according to the complexity of the shape of the static object. For example, the encoding device sets VXL or VLM more densely as the shape of the static object is more complex. Further, the encoding device may determine quantization steps when quantizing spatial position or color information according to the coarseness of VXL or VLM. For example, the encoding device sets a smaller quantization step as VXL or VLM is denser.
[0284] As described above, the encoding device or decoding device according to the present embodiment performs spatial encoding or decoding in space units having coordinate information.
[0285] In addition, the encoding device and the decoding device perform encoding or decoding in volume units within the space. A volume includes voxels which are the minimum units to which position information is associated.
[0286] In addition, the encoding device and the decoding device perform encoding or decoding by associating arbitrary elements with each other using a table that associates each element of spatial information including coordinates, objects, time, etc. with a GOP, or a table that associates between each element. Further, the decoding device determines coordinates using the value of the selected element, specifies a volume, voxel or space from the coordinates, and decodes the space including the volume or voxel, or the specified space.
[0287] In addition, the encoding device determines a volume, voxel or space that can be selected by an element by feature point extraction or object recognition, and encodes it as a volume, voxel or space that can be randomly accessed.
[0288] Spaces are classified into three types: I-SPC that can be encoded or decoded as the space unit itself, P-SPC that is encoded or decoded with reference to any one processed space, and B-SPC that is encoded or decoded with reference to any two processed spaces.
[0289] One or more volumes correspond to static objects or dynamic objects. The space including static objects and the space including dynamic objects are encoded or decoded as different GOSs. That is, the SPC including static objects and the SPC including dynamic objects are assigned to different GOSs.
[0290] Dynamic objects are encoded or decoded on a per-object basis and associated with one or more spaces that include static objects. That is, multiple dynamic objects are encoded individually, and the encoded data of the resulting multiple dynamic objects is associated with an SPC that includes static objects.
[0291] The encoding device and the decoding device perform encoding or decoding by increasing the priority of the I-SPC in the GOS. For example, the encoding device performs encoding so that the degradation of the I-SPC is reduced (so that the original three-dimensional data is more faithfully reproduced after decoding). Also, the decoding device decodes, for example, only the I-SPC.
[0292] The encoding device may perform encoding by changing the frequency of using the I-SPC according to the density or number (quantity) of objects in the world. That is, the encoding device changes the frequency of selecting the I-SPC according to the number or density of objects included in the three-dimensional data. For example, the encoding device increases the frequency of using the I-space as the objects in the world are denser.
[0293] Also, the encoding device sets random access points in units of GOS and stores information indicating the spatial region corresponding to the GOS in the header information.
[0294] The encoding device uses, for example, a default value as the spatial size of the GOS. Note that the encoding device may change the size of the GOS according to the number (quantity) or density of objects or dynamic objects. For example, the encoding device reduces the spatial size of the GOS as the objects or dynamic objects are denser or the number is larger.
[0295] Also, the space or volume includes a set of feature points derived using information obtained by sensors such as a depth sensor, a gyro, or a camera. The coordinates of the feature points are set at the center position of the voxel. Also, high-precision position information can be realized by subdividing the voxel.
[0296] The feature point group is derived using a plurality of pictures. The plurality of pictures have at least two types of time information, namely actual time information and the same time information (e.g., the encoding time used for rate control, etc.) in a plurality of pictures associated with space.
[0297] Also, encoding or decoding is performed in GOS units including one or more spaces.
[0298] The encoding device and the decoding device predict the P space or the B space in the GOS to be processed by referring to the spaces in the processed GOS.
[0299] Alternatively, the encoding device and the decoding device predict the P space or the B space in the GOS to be processed by using the processed spaces in the GOS to be processed without referring to different GOSs.
[0300] Also, the encoding device and the decoding device transmit or receive an encoding stream in world units including one or more GOSs.
[0301] Also, the GOS has a layer structure in at least one direction within the world, and the encoding device and the decoding device perform encoding or decoding from the lower layer. For example, a randomly accessible GOS belongs to the lowest layer. A GOS belonging to a higher layer refers to a GOS belonging to the same layer or lower layers. That is, the GOS is spatially divided in a predetermined direction and includes a plurality of layers each containing one or more SPCs. The encoding device and the decoding device encode or decode each SPC by referring to the SPCs included in the same layer or lower layers than the SPC.
[0302] Also, the encoding device and the decoding device continuously encode or decode GOSs within a world unit including a plurality of GOSs. The encoding device and the decoding device write or read information indicating the order (direction) of encoding or decoding as metadata. That is, the encoded data includes information indicating the encoding order of a plurality of GOSs.
[0303] In addition, the encoding device and the decoding device encode or decode two or more different spaces or GOSs in parallel.
[0304] In addition, the encoding device and the decoding device encode or decode the spatial information (coordinates, size, etc.) of the space or GOS.
[0305] In addition, the encoding device and the decoding device encode or decode the space or GOS included in a specific space specified based on external information such as GPS, route information, or magnification regarding their own position or / and area size.
[0306] The encoding device or the decoding device encodes or decodes a space far from its own position with a lower priority than a nearby space.
[0307] The encoding device sets one direction of the world according to the magnification or application, and encodes a GOS having a layer structure in that direction. In addition, the decoding device preferentially decodes a GOS having a layer structure in one direction of the world set according to the magnification or application from the lower layer.
[0308] The encoding device changes the feature point extraction, object recognition accuracy, or spatial area size included in the space between indoors and outdoors. However, the encoding device and the decoding device encode or decode an indoor GOS and an outdoor GOS with adjacent coordinates in the world adjacent to each other, and also encode or decode their identifiers in association with each other.
[0309] (Embodiment 2) When using the encoded data of the point cloud in an actual device or service, it is desirable to transmit and receive necessary information according to the application in order to suppress the network bandwidth. However, until now, such a function has not existed in the encoding structure of three-dimensional data, and there has also been no encoding method therefor.
[0310] In this embodiment, a three-dimensional data encoding method, a three-dimensional data encoding apparatus for providing a function of transmitting and receiving only necessary information according to the application in the encoded data of three-dimensional point clouds, a three-dimensional decoding method for decoding the encoded data, and a three-dimensional data apparatus will be described.
[0311] A voxel (VXL) having a feature amount equal to or greater than a certain value is defined as a feature voxel (FVXL), and a world (WLD) composed of FVXLs is defined as a sparse world (SWLD). FIG. 11 is a diagram showing a configuration example of a sparse world and a world. The SWLD includes a FGOS that is a GOS composed of FVXLs, an FSPC that is an SPC composed of FVXLs, and an FVLM that is a VLM composed of FVXLs. The data structure and prediction structure of the FGOS, FSPS, and FVLM may be the same as those of the GOS, SPS, and VLM.
[0312] The feature amount is a feature amount representing the three-dimensional position information of the VXL or the visible light information of the VXL position, and is a feature amount that is particularly frequently detected at corners and edges of three-dimensional objects. Specifically, this feature amount is a three-dimensional feature amount or a visible light feature amount as described below, but any feature amount that represents the position, luminance, or color information of the VXL may be used.
[0313] As the three-dimensional feature amount, a SHOT feature amount (Signature of Histograms of OrienTations), a PFH feature amount (Point Feature Histograms), or a PPF feature amount (Point Pair Feature) is used.
[0314] The SHOT feature amount is obtained by dividing the periphery of the VXL and calculating the inner product of the normal vector of the reference point and the divided region and then histogramming. This SHOT feature amount has the characteristics of a high number of dimensions and high feature representation ability.
[0315] The PFH feature quantity is obtained by selecting a large number of point pairs in the vicinity of the VXL, calculating the normal vector and the like from the two points, and creating a histogram. Since this PFH feature quantity is a histogram feature, it has robustness against some disturbances and has the feature of high feature expressiveness.
[0316] The PPF feature quantity is a feature quantity calculated using the normal vector and the like for each pair of VXLs. Since all VXLs are used for this PPF feature quantity, it has robustness against occlusion.
[0317] Also, as a feature quantity of visible light, SIFT (Scale-Invariant Feature Transform), SURF (Speeded Up Robust Features), or HOG (Histogram of Oriented Gradients) using information such as the luminance gradient information of the image can be used.
[0318] The SWLD is generated by calculating the above feature quantity from each VXL of the WLD and extracting the FVXL. Here, the SWLD may be updated every time the WLD is updated, or may be updated periodically after a certain period of time regardless of the update timing of the WLD.
[0319] The SWLD may be generated for each feature quantity. For example, separate SWLDs may be generated for each feature quantity, such as SWLD1 based on the SHOT feature quantity and SWLD2 based on the SIFT feature quantity, and the SWLDs may be used appropriately according to the application. Also, the feature quantity of each calculated FVXL may be held in each FVXL as feature quantity information.
[0320] Next, the usage method of the sparse world (SWLD) will be described. Since the SWLD includes only the feature voxels (FVXL), the data size is generally smaller than that of the WLD including all VXLs.
[0321] In an application that achieves some purpose using feature amounts, by using the information of SWLD instead of WLD, it is possible to suppress the read time from the hard disk and the bandwidth and transfer time during network transfer. For example, as map information, both WDL and SWLD are held in the server, and by switching the map information to be transmitted to WLD or SWLD according to the request from the client, the network bandwidth and transfer time can be suppressed. Hereinafter, specific examples will be shown.
[0322] FIGS. 12 and 13 are diagrams showing usage examples of SWLD and WLD. As shown in FIG. 12, when the in-vehicle device client 1 needs map information for its own position determination use, the client 1 sends a request for acquiring map data for its own position estimation to the server (S301). The server transmits the SWLD to the client 1 in response to the acquisition request (S302). The client 1 performs its own position determination using the received SWLD (S303). At this time, the client 1 acquires VXL information around the client 1 by various methods such as a distance sensor such as a range finder, a stereo camera, or a combination of a plurality of monocular cameras, and estimates its own position information from the obtained VXL information and the SWLD. Here, the own position information includes the three-dimensional position information and orientation of the client 1.
[0323] As shown in FIG. 13, when the in-vehicle device client 2 needs map information for map drawing uses such as a three-dimensional map, the client 2 sends a request for acquiring map data for map drawing to the server (S311). The server transmits the WLD to the client 2 in response to the acquisition request (S312). The client 2 performs map drawing using the received WLD (S313). At this time, the client 2 creates a rendering image using, for example, an image captured by its own visible light camera or the like and the WLD acquired from the server, and draws the created image on a screen such as a car navigation.
[0324] As described above, the server transmits the SWLD to the client for applications that mainly require feature amounts of each VXL such as self-position estimation, and transmits the WLD to the client when detailed VXL information such as map drawing is required. Thereby, it becomes possible to efficiently transmit and receive map data.
[0325] Note that the client may determine which of the SWLD and the WLD is necessary by itself and request the server to transmit the SWLD or the WLD. Also, the server may determine which of the SWLD and the WLD should be transmitted according to the client or network situation.
[0326] Next, a method for switching the transmission and reception of the sparse world (SWLD) and the world (WLD) will be described.
[0327] It may be possible to switch whether to receive the WLD or the SWLD according to the network bandwidth. FIG. 14 is a diagram showing an operation example in this case. For example, when a low-speed network with a limited available network bandwidth such as in an LTE (Long Term Evolution) environment is used, the client accesses the server via the low-speed network (S321) and acquires the SWLD as map information from the server (S322). On the other hand, when a high-speed network with a sufficient network bandwidth such as in a Wi-Fi (registered trademark) environment is used, the client accesses the server via the high-speed network (S323) and acquires the WLD from the server (S324). Thereby, the client can acquire appropriate map information according to the network bandwidth of the client.
[0328] Specifically, the client receives the SWLD via LTE outdoors, and acquires the WLD via Wi-Fi (registered trademark) when entering indoors such as a facility. Thereby, the client can acquire more detailed map information indoors.
[0329] In this way, the client may request the WLD or SWLD from the server according to the bandwidth of the network it uses. Or, the client may send information indicating the bandwidth of the network it uses to the server, and the server may send data (WLD or SWLD) suitable for the client according to the information. Or, the server may determine the network bandwidth of the client and send data (WLD or SWLD) suitable for the client.
[0330] Also, it may be possible to switch whether to receive the WLD or SWLD according to the moving speed. FIG. 15 is a diagram showing an operation example in this case. For example, when the client is moving at high speed (S331), the client receives the SWLD from the server (S332). On the other hand, when the client is moving at low speed (S333), the client receives the WLD from the server (S334). Thereby, the client can obtain map information suitable for the speed while suppressing the network bandwidth. Specifically, while driving on a highway, the client can update the general map information at an appropriate speed by receiving the SWLD with a small amount of data. On the other hand, while driving on a general road, the client can obtain more detailed map information by receiving the WLD.
[0331] In this way, the client may request the WLD or SWLD from the server according to its own moving speed. Or, the client may send information indicating its own moving speed to the server, and the server may send data (WLD or SWLD) suitable for the client according to the information. Or, the server may determine the moving speed of the client and send data (WLD or SWLD) suitable for the client.
[0332] Also, the client may first obtain the SWLD from the server and then obtain the WLD of important areas therein. For example, when obtaining map data, the client first obtains rough map information using the SWLD, then narrows down the areas where features such as buildings, signs, or people appear frequently, and later obtains the WLD of the narrowed-down areas. This enables the client to obtain detailed information on the necessary areas while suppressing the amount of received data from the server.
[0333] In addition, the server may create separate SWLDs for each object from the WLD, and the client may receive each of them according to the application. This can suppress the network bandwidth. For example, the server recognizes people or vehicles in advance from the WLD and creates the SWLD for people and the SWLD for vehicles. When the client wants to obtain information about the surrounding people, it receives the SWLD for people, and when it wants to obtain information about vehicles, it receives the SWLD for vehicles. Also, such types of SWLDs may be distinguished by information (such as flags or types) added to the header or the like.
[0334] Next, the configuration and operation flow of the three-dimensional data encoding device (such as a server) according to this embodiment will be described. FIG. 16 is a block diagram of the three-dimensional data encoding device 400 according to this embodiment. FIG. 17 is a flowchart of the three-dimensional data encoding process by the three-dimensional data encoding device 400.
[0335] The three-dimensional data encoding device 400 shown in FIG. 16 generates encoded three-dimensional data 413 and 414, which are encoded streams, by encoding the input three-dimensional data 411. Here, the encoded three-dimensional data 413 is the encoded three-dimensional data corresponding to the WLD, and the encoded three-dimensional data 414 is the encoded three-dimensional data corresponding to the SWLD. This three-dimensional data encoding device 400 includes an acquisition unit 401, an encoding area determination unit 402, an SWLD extraction unit 403, a WLD encoding unit 404, and an SWLD encoding unit 405.
[0336] As shown in FIG. 17, first, the acquisition unit 401 acquires input three-dimensional data 411, which is point cloud data in a three-dimensional space (S401).
[0337] Next, the encoding region determination unit 402 determines a spatial region to be encoded based on the spatial region where the point cloud data exists (S402).
[0338] Next, the SWLD extraction unit 403 defines the spatial region to be encoded as the WLD, and calculates feature amounts from each VXL included in the WLD. Then, the SWLD extraction unit 403 extracts VXLs whose feature amounts are equal to or greater than a predetermined threshold value, defines the extracted VXLs as FVXLs, and generates extracted three-dimensional data 412 by adding the FVXLs to the SWLD. That is, the extracted three-dimensional data 412 whose feature amounts are equal to or greater than the threshold value is extracted from the input three-dimensional data 411.
[0339] Next, the WLD encoding unit 404 generates encoded three-dimensional data 413 corresponding to the WLD by encoding the input three-dimensional data 411 corresponding to the WLD (S404). At this time, the WLD encoding unit 404 adds information for distinguishing that the encoded three-dimensional data 413 is a stream including the WLD to the header of the encoded three-dimensional data 413.
[0340] Also, the SWLD encoding unit 405 generates encoded three-dimensional data 414 corresponding to the SWLD by encoding the extracted three-dimensional data 412 corresponding to the SWLD (S405). At this time, the SWLD encoding unit 405 adds information for distinguishing that the encoded three-dimensional data 414 is a stream including the SWLD to the header of the encoded three-dimensional data 414.
[0341] Note that the processing order of the process of generating the encoded three-dimensional data 413 and the process of generating the encoded three-dimensional data 414 may be reversed from the above. Also, a part or all of these processes may be performed in parallel.
[0342] As information attached to the headers of the encoded three-dimensional data 413 and 414, for example, a parameter named "world_type" is defined. When world_type = 0, it indicates that the stream includes WLD, and when world_type = 1, it indicates that the stream includes SWLD. When defining a number of other types, the numerical values assigned can be increased, such as world_type = 2. Also, a specific flag may be included in one of the encoded three-dimensional data 413 and 414. For example, a flag indicating that the stream includes SWLD may be attached to the encoded three-dimensional data 414. In this case, the decoding device can determine whether the stream includes WLD or SWLD based on the presence or absence of the flag.
[0343] Also, the encoding method used when the WLD encoding unit 404 encodes WLD may be different from the encoding method used when the SWLD encoding unit 405 encodes SWLD.
[0344] For example, since data is decimated in SWLD, the correlation with surrounding data may be lower than that in WLD. Therefore, in the encoding method used for SWLD, inter prediction may be prioritized over intra prediction and inter prediction in the encoding method used for WLD.
[0345] Also, the method of expressing the three-dimensional position may be different between the encoding method used for SWLD and the encoding method used for WLD. For example, in SWLD, the three-dimensional position of SVXL may be expressed by three-dimensional coordinates, and in WLD, the three-dimensional position may be expressed by an octree described later, or vice versa.
[0346] Also, the SWLD encoding unit 405 performs encoding so that the data size of the encoded three-dimensional data 414 of SWLD is smaller than the data size of the encoded three-dimensional data 413 of WLD. For example, as described above, SWLD may have lower correlation between data compared to WLD. As a result, the encoding efficiency decreases, and the data size of the encoded three-dimensional data 414 may become larger than the data size of the encoded three-dimensional data 413 of WLD. Therefore, when the data size of the obtained encoded three-dimensional data 414 is larger than the data size of the encoded three-dimensional data 413 of WLD, the SWLD encoding unit 405 performs re-encoding to regenerate the encoded three-dimensional data 414 with a reduced data size.
[0347] For example, the SWLD extraction unit 403 regenerates the extracted three-dimensional data 412 with a reduced number of feature points to be extracted, and the SWLD encoding unit 405 encodes the extracted three-dimensional data 412. Alternatively, the degree of quantization in the SWLD encoding unit 405 may be made coarser. For example, in the octree structure described later, the degree of quantization can be made coarser by rounding the data in the bottom layer.
[0348] Also, when the SWLD encoding unit 405 cannot make the data size of the encoded three-dimensional data 414 of SWLD smaller than the data size of the encoded three-dimensional data 413 of WLD, it may not be necessary to generate the encoded three-dimensional data 414 of SWLD. Alternatively, the encoded three-dimensional data 413 of WLD may be copied to the encoded three-dimensional data 414 of SWLD. That is, the encoded three-dimensional data 413 of WLD may be used as it is as the encoded three-dimensional data 414 of SWLD.
[0349] Next, the configuration and operation flow of the three-dimensional data decoding device (for example, a client) according to the present embodiment will be described. FIG. 18 is a block diagram of the three-dimensional data decoding device 500 according to the present embodiment. FIG. 19 is a flowchart of the three-dimensional data decoding process by the three-dimensional data decoding device 500.
[0350] The three-dimensional data decoding device 500 shown in FIG. 18 generates decoded three-dimensional data 512 or 513 by decoding the encoded three-dimensional data 511. Here, the encoded three-dimensional data 511 is, for example, the encoded three-dimensional data 413 or 414 generated by the three-dimensional data encoding device 400.
[0351] This three-dimensional data decoding device 500 includes an acquisition unit 501, a header analysis unit 502, a WLD decoding unit 503, and a SWLD decoding unit 504.
[0352] As shown in FIG. 19, first, the acquisition unit 501 acquires the encoded three-dimensional data 511 (S501). Next, the header analysis unit 502 analyzes the header of the encoded three-dimensional data 511 to determine whether the encoded three-dimensional data 511 is a stream including WLD or a stream including SWLD (S502). For example, the above-described parameter of world_type is referred to for determination.
[0353] When the encoded three-dimensional data 511 is a stream including WLD (Yes in S503), the WLD decoding unit 503 generates the decoded three-dimensional data 512 of WLD by decoding the encoded three-dimensional data 511 (S504). On the other hand, when the encoded three-dimensional data 511 is a stream including SWLD (No in S503), the SWLD decoding unit 504 generates the decoded three-dimensional data 513 of SWLD by decoding the encoded three-dimensional data 511 (S505).
[0354] Also, similar to the encoding device, the decoding method used when the WLD decoding unit 503 decodes WLD and the decoding method used when the SWLD decoding unit 504 decodes SWLD may be different. For example, in the decoding method used for SWLD, inter prediction may be prioritized over intra prediction and inter prediction in the decoding method used for WLD.
[0355] Also, the decoding method used for SWLD and the decoding method used for WLD may differ in the method of expressing the three-dimensional position. For example, in SWLD, the three-dimensional position of SVXL may be expressed by three-dimensional coordinates, and in WLD, the three-dimensional position may be expressed by an octree described later, or vice versa.
[0356] Next, the octree representation, which is a method of expressing the three-dimensional position, will be described. The VXL data included in the three-dimensional data is encoded after being converted into an octree structure. FIG. 20 is a diagram showing an example of the VXL of WLD. FIG. 21 is a diagram showing the octree structure of the WLD shown in FIG. 20. In the example shown in FIG. 20, there are three VXLs 1 to 3, which are VXLs (hereinafter referred to as valid VXLs) including a point cloud. As shown in FIG. 21, the octree structure is composed of nodes and leaves. Each node has up to eight nodes or leaves. Each leaf has VXL information. Here, among the leaves shown in FIG. 21, leaves 1, 2, and 3 represent VXLs 1, 2, and 3 shown in FIG. 20, respectively.
[0357] Specifically, each node and leaf corresponds to a three-dimensional position. Node 1 corresponds to the entire block shown in FIG. 20. The block corresponding to Node 1 is divided into eight blocks. Among the eight blocks, the blocks containing valid VXLs are set as nodes, and the other blocks are set as leaves. The block corresponding to the node is further divided into eight nodes or leaves, and this process is repeated for the hierarchical levels of the tree structure. Also, all the blocks in the bottom layer are set as leaves.
[0358] Further, FIG. 22 is a diagram showing an example of the SWLD generated from the WLD shown in FIG. 20. VXL1 and VXL2 shown in FIG. 20 are determined as FVXL1 and FVXL2 as a result of feature extraction and are added to the SWLD. On the other hand, VXL3 is not determined as FVXL and is not included in the SWLD. FIG. 23 is a diagram showing the octree structure of the SWLD shown in FIG. 22. In the octree structure shown in FIG. 23, leaf 3 corresponding to VXL3 shown in FIG. 21 is deleted. As a result, node 3 shown in FIG. 21 no longer has valid VXLs and is changed to a leaf. In general, the number of leaves of the SWLD is thus smaller than the number of leaves of the WLD, and the encoded three-dimensional data of the SWLD is also smaller than the encoded three-dimensional data of the WLD.
[0359] Hereinafter, a modification example of the present embodiment will be described.
[0360] For example, when a client such as an in-vehicle device performs self-position estimation, it receives the SWLD from the server and performs self-position estimation using the SWLD. When performing obstacle detection, the client may perform obstacle detection based on three-dimensional information of the surroundings obtained by itself using various methods such as a distance sensor such as a range finder, a stereo camera, or a combination of a plurality of monocular cameras.
[0361] Also, generally, the SWLD is unlikely to include VXL data of a flat area. Therefore, the server may hold a sub-sampled world (subWLD) obtained by sub-sampling the WLD for detecting static obstacles, and transmit the SWLD and the subWLD to the client. Thereby, while suppressing the network bandwidth, self-position estimation and obstacle detection can be performed on the client side.
[0362] Also, when the client rapidly renders three-dimensional map data, it may be more convenient if the map information has a mesh structure. Therefore, the server may generate a mesh from the WLD and pre-hold it as a Mesh World Map (MWLD). For example, when the client requires rough three-dimensional rendering, it receives the MWLD, and when it requires detailed three-dimensional rendering, it receives the WLD. This can suppress the network bandwidth.
[0363] Also, among each VXL, the server has set the VXL whose feature amount is equal to or greater than the threshold as the FVXL, but the FVXL may be calculated by different methods. For example, the server may determine that the VXL, VLM, SPS, or GOS that constitutes a signal or an intersection, etc. is necessary for self-position estimation, driving assistance, or autonomous driving, etc., and include them in the SWLD as the FVXL, FVLM, FSPS, FGOS. Also, the above determination may be made manually. In addition, the FVXL, etc. obtained by the above method may be added to the FVXL, etc. set based on the feature amount. That is, the SWLD extraction unit 403 may further extract, as the extracted three-dimensional data 412, data corresponding to an object having a predetermined attribute from the input three-dimensional data 411.
[0364] Also, you may label the necessity for those applications separately from the feature amount. Further, the server may separately hold the FVXL necessary for self-position estimation, driving assistance, or autonomous driving, etc. of signals or intersections, etc. as the upper layer (for example, the lane world) of the SWLD.
[0365] Also, the server may add an attribute to the VXL in the WLD for each random access unit or a predetermined unit. The attribute includes, for example, information indicating whether it is necessary or unnecessary for self-position estimation, or information indicating whether it is important as traffic information such as a signal or an intersection. Also, the attribute may include the correspondence relationship with a Feature (such as an intersection or a road) in the lane information (GDF: Geographic Data Files, etc.).
[0366] Also, the following methods may be used as the method for updating the WLD or SWLD.
[0367] Update information indicating changes in people, construction, or trees (for trucks), etc. is uploaded to the server as point clouds or metadata. Based on the upload, the server updates the WLD, and then updates the SWLD using the updated WLD.
[0368] Also, when the client detects an inconsistency between the three-dimensional information generated by itself during self-position estimation and the three-dimensional information received from the server, the client may send the three-dimensional information generated by itself to the server together with an update notification. In this case, the server updates the SWLD using the WLD. If the SWLD is not updated, the server determines that the WLD itself is old.
[0369] Also, although information distinguishing between the WLD and the SWLD is added as the header information of the encoded stream, for example, when there are multiple types of worlds such as a mesh world or a lane world, information for distinguishing them may be added to the header information. Also, when there are a large number of SWLDs with different feature amounts, information for distinguishing each of them may be added to the header information.
[0370] Also, although the SWLD is assumed to be composed of FVXLs, it may include VXLs that are not determined to be FVXLs. For example, the SWLD may include adjacent VXLs used when calculating the feature amounts of the FVXLs. Thereby, even when feature amount information is not added to each FVXL of the SWLD, the client can calculate the feature amounts of the FVXLs when receiving the SWLD. At that time, the SWLD may include information for distinguishing whether each VXL is an FVXL or a VXL.
[0371] As described above, the three-dimensional data encoding device 400 extracts the extracted three-dimensional data 412 (second three-dimensional data) with a feature amount equal to or greater than the threshold value from the input three-dimensional data 411 (first three-dimensional data), and generates the encoded three-dimensional data 414 (first encoded three-dimensional data) by encoding the extracted three-dimensional data 412.
[0372] According to this, the three-dimensional data encoding device 400 generates the encoded three-dimensional data 414 obtained by encoding the data with a feature amount equal to or greater than the threshold value. As a result, the data amount can be reduced as compared with the case where the input three-dimensional data 411 is directly encoded. Therefore, the three-dimensional data encoding device 400 can reduce the data amount to be transmitted.
[0373] In addition, the three-dimensional data encoding device 400 further generates the encoded three-dimensional data 413 (second encoded three-dimensional data) by encoding the input three-dimensional data 411.
[0374] According to this, the three-dimensional data encoding device 400 can selectively transmit the encoded three-dimensional data 413 and the encoded three-dimensional data 414 according to, for example, the usage purpose.
[0375] In addition, the extracted three-dimensional data 412 is encoded by the first encoding method, and the input three-dimensional data 411 is encoded by the second encoding method different from the first encoding method.
[0376] According to this, the three-dimensional data encoding device 400 can use the encoding methods suitable for the input three-dimensional data 411 and the extracted three-dimensional data 412, respectively.
[0377] In addition, in the first encoding method, inter prediction is prioritized over intra prediction among the intra prediction and the inter prediction as compared with the second encoding method.
[0378] According to this, the three-dimensional data encoding device 400 can increase the priority of inter prediction for the extracted three-dimensional data 412 in which the correlation between adjacent data is likely to be low.
[0379] Also, in the first encoding method and the second encoding method, the expression methods of three-dimensional positions are different. For example, in the second encoding method, the three-dimensional position is expressed by an octree, and in the first encoding method, the three-dimensional position is expressed by three-dimensional coordinates.
[0380] According to this, the three-dimensional data encoding device 400 can use a more suitable expression method for three-dimensional positions for three-dimensional data with different numbers of data (the number of VXLs or SVXLs).
[0381] Also, at least one of the encoded three-dimensional data 413 and 414 includes an identifier indicating whether the encoded three-dimensional data is the encoded three-dimensional data obtained by encoding the input three-dimensional data 411 or the encoded three-dimensional data obtained by encoding a part of the input three-dimensional data 411. That is, the identifier indicates whether the encoded three-dimensional data is the encoded three-dimensional data 413 of the WLD or the encoded three-dimensional data 414 of the SWLD.
[0382] According to this, the decoding device can easily determine whether the acquired encoded three-dimensional data is the encoded three-dimensional data 413 or the encoded three-dimensional data 414.
[0383] Also, the three-dimensional data encoding device 400 encodes the extracted three-dimensional data 412 so that the data amount of the encoded three-dimensional data 414 is smaller than the data amount of the encoded three-dimensional data 413.
[0384] According to this, the three-dimensional data encoding device 400 can make the data amount of the encoded three-dimensional data 414 smaller than the data amount of the encoded three-dimensional data 413.
[0385] Also, the three-dimensional data encoding device 400 further extracts, as the extracted three-dimensional data 412, data corresponding to an object having a predetermined attribute from the input three-dimensional data 411. For example, an object having a predetermined attribute is an object necessary for self-position estimation, driving assistance, or autonomous driving, such as a signal or an intersection.
[0386] According to this, the three-dimensional data encoding device 400 can generate encoded three-dimensional data 414 including data required by the decoding device.
[0387] Further, the three-dimensional data encoding device 400 (server) further transmits one of the encoded three-dimensional data 413 and 414 to the client according to the state of the client.
[0388] According to this, the three-dimensional data encoding device 400 can transmit appropriate data according to the state of the client.
[0389] Further, the state of the client includes the communication status of the client (for example, network bandwidth) or the moving speed of the client.
[0390] Further, the three-dimensional data encoding device 400 further transmits one of the encoded three-dimensional data 413 and 414 to the client according to the request of the client.
[0391] According to this, the three-dimensional data encoding device 400 can transmit appropriate data according to the request of the client.
[0392] Further, the three-dimensional data decoding device 500 according to the present embodiment decodes the encoded three-dimensional data 413 or 414 generated by the three-dimensional data encoding device 400.
[0393] That is, the three-dimensional data decoding device 500 decodes the encoded three-dimensional data 414 obtained by encoding the extracted three-dimensional data 412 whose feature amount extracted from the input three-dimensional data 411 is equal to or greater than the threshold value by the first decoding method. Further, the three-dimensional data decoding device 500 decodes the encoded three-dimensional data 413 obtained by encoding the input three-dimensional data 411 by a second decoding method different from the first decoding method.
[0394] According to this, the three-dimensional data decoding device 500 can selectively receive the encoded three-dimensional data 414 obtained by encoding data with a feature amount equal to or greater than the threshold value and the encoded three-dimensional data 413, for example, according to the usage purpose or the like. Thereby, the three-dimensional data decoding device 500 can reduce the amount of data to be transmitted. Further, the three-dimensional data decoding device 500 can use decoding methods suitable for the input three-dimensional data 411 and the extracted three-dimensional data 412, respectively.
[0395] Also, in the first decoding method, inter prediction among intra prediction and inter prediction is prioritized over the second decoding method.
[0396] According to this, the three-dimensional data decoding device 500 can increase the priority of inter prediction for the extracted three-dimensional data in which the correlation between adjacent data is likely to be low.
[0397] Also, in the first decoding method and the second decoding method, the expression method of the three-dimensional position is different. For example, in the second decoding method, the three-dimensional position is expressed by an octree, and in the first decoding method, the three-dimensional position is expressed by three-dimensional coordinates.
[0398] According to this, the three-dimensional data decoding device 500 can use a more suitable expression method of the three-dimensional position for three-dimensional data with different numbers of data (the number of VXLs or SVXLs).
[0399] Also, at least one of the encoded three-dimensional data 413 and 414 includes an identifier indicating whether the encoded three-dimensional data is the encoded three-dimensional data obtained by encoding the input three-dimensional data 411 or the encoded three-dimensional data obtained by encoding a part of the input three-dimensional data 411. The three-dimensional data decoding device 500 identifies the encoded three-dimensional data 413 and 414 with reference to the identifier.
[0400] According to this, the three-dimensional data decoding device 500 can easily determine whether the acquired encoded three-dimensional data is the encoded three-dimensional data 413 or the encoded three-dimensional data 414.
[0401] Further, the three-dimensional data decoding device 500 further notifies the server of the state of the client (the three-dimensional data decoding device 500). The three-dimensional data decoding device 500 receives one of the encoded three-dimensional data 413 and 414 transmitted from the server according to the state of the client.
[0402] According to this, the three-dimensional data decoding device 500 can receive appropriate data according to the state of the client.
[0403] Also, the state of the client includes the communication status of the client (for example, network bandwidth) or the moving speed of the client.
[0404] Further, the three-dimensional data decoding device 500 further requests one of the encoded three-dimensional data 413 and 414 from the server, and receives one of the encoded three-dimensional data 413 and 414 transmitted from the server in response to the request.
[0405] According to this, the three-dimensional data decoding device 500 can receive appropriate data according to the application.
[0406] (Embodiment 3) In this embodiment, a method for transmitting and receiving three-dimensional data between vehicles will be described.
[0407] FIG. 24 is a schematic diagram showing a state of transmission and reception of three-dimensional data 607 between the host vehicle 600 and surrounding vehicles 601.
[0408] When acquiring three-dimensional data by a sensor (such as a distance sensor such as a range finder, a stereo camera, or a combination of a plurality of monocular cameras) mounted on the host vehicle 600, an area (hereinafter referred to as an occlusion area 604) that cannot create three-dimensional data although it is within the sensor detection range 602 of the host vehicle 600 due to obstacles such as surrounding vehicles 601 occurs. Further, as the space for acquiring three-dimensional data increases, the accuracy of autonomous operation increases, but the sensor detection range of only the host vehicle 600 is limited.
[0409] The sensor detection range 602 of the host vehicle 600 includes a region 603 capable of acquiring three-dimensional data and an occlusion region 604. The range in which the host vehicle 600 wants to acquire three-dimensional data includes the sensor detection range 602 of the host vehicle 600 and other regions. Further, the sensor detection range 605 of the surrounding vehicle 601 includes the occlusion region 604 and a region 606 not included in the sensor detection range 602 of the host vehicle 600.
[0410] The surrounding vehicle 601 transmits the information detected by the surrounding vehicle 601 to the host vehicle 600. By acquiring the information detected by the surrounding vehicle 601 such as the preceding vehicle, the host vehicle 600 can acquire the three-dimensional data 607 of the occlusion region 604 and the region 606 outside the sensor detection range 602 of the host vehicle 600. The host vehicle 600 uses the information acquired by the surrounding vehicle 601 to complement the three-dimensional data of the occlusion region 604 and the region 606 outside the sensor detection range.
[0411] The use of three-dimensional data in the autonomous operation of a vehicle or a robot is for self-position estimation, detection of the surrounding situation, or both. For example, for self-position estimation, three-dimensional data generated by the host vehicle 600 based on the sensor information of the host vehicle 600 is used. For detection of the surrounding situation, in addition to the three-dimensional data generated by the host vehicle 600, the three-dimensional data acquired from the surrounding vehicle 601 is also used.
[0412] The surrounding vehicle 601 that transmits the three-dimensional data 607 to the host vehicle 600 may be determined according to the state of the host vehicle 600. For example, this surrounding vehicle 601 is the preceding vehicle when the host vehicle 600 is going straight, the oncoming vehicle when the host vehicle 600 is turning right, and the following vehicle when the host vehicle 600 is reversing. Further, the driver of the host vehicle 600 may directly specify the surrounding vehicle 601 that transmits the three-dimensional data 607 to the host vehicle 600.
[0413] In addition, the host vehicle 600 may search for a surrounding vehicle 601 that owns three-dimensional data of an area that cannot be acquired by the host vehicle 600 and is included in the space for which the three-dimensional data 607 is to be acquired. The area that cannot be acquired by the host vehicle 600 is, for example, an occlusion area 604 or an area 606 outside the sensor detection range 602.
[0414] Further, the host vehicle 600 may specify the occlusion area 604 based on the sensor information of the host vehicle 600. For example, the host vehicle 600 specifies, as the occlusion area 604, an area within the sensor detection range 602 of the host vehicle 600 where three-dimensional data cannot be created.
[0415] Hereinafter, an operation example in the case where the leading vehicle transmits the three-dimensional data 607 will be described. FIG. 25 is a diagram showing an example of the three-dimensional data transmitted in this case.
[0416] As shown in FIG. 25, the three-dimensional data 607 transmitted from the leading vehicle is, for example, a sparse world (SWLD) of a point cloud. That is, the leading vehicle creates three-dimensional data (point cloud) of the WLD from the information detected by the sensors of the leading vehicle, and creates three-dimensional data (point cloud) of the SWLD by extracting data with a feature amount equal to or greater than a threshold value from the three-dimensional data of the WLD. Then, the leading vehicle transmits the created three-dimensional data of the SWLD to the host vehicle 600.
[0417] The host vehicle 600 receives the SWLD and merges the received SWLD with the point cloud created by the host vehicle 600.
[0418] The transmitted SWLD has information on absolute coordinates (the position of the SWLD in the coordinate system of the three-dimensional map). The host vehicle 600 can realize the merge process by overwriting the point cloud generated by the host vehicle 600 based on this absolute coordinate.
[0419] The SWLD transmitted from the surrounding vehicle 601 may be the SWLD of the area 606 outside the sensor detection range 602 of the host vehicle 600 and within the sensor detection range 605 of the surrounding vehicle 601, or the SWLD of the occlusion area 604 for the host vehicle 600, or both. Also, the transmitted SWLD may be the SWLD of the area used by the surrounding vehicle 601 for detecting the surrounding situation among the above SWLDs.
[0420] Also, the surrounding vehicle 601 may change the density of the transmitted point cloud according to the communication available time based on the speed difference between the host vehicle 600 and the surrounding vehicle 601. For example, when the speed difference is large and the communication available time is short, the surrounding vehicle 601 may reduce the density (data amount) of the point cloud by extracting three-dimensional points with large feature amounts from the SWLD.
[0421] Also, detecting the surrounding situation is to determine the presence or absence of people, vehicles, equipment for road construction, etc., identify their types, and detect their positions, moving directions, moving speeds, etc.
[0422] Also, the host vehicle 600 may obtain the braking information of the surrounding vehicle 601 instead of, or in addition to, the three-dimensional data 607 generated by the surrounding vehicle 601. Here, the braking information of the surrounding vehicle 601 is, for example, information indicating that the accelerator or brake of the surrounding vehicle 601 has been depressed or the degree thereof.
[0423] Also, in the point cloud generated by each vehicle, the three-dimensional space is subdivided into random access units in consideration of low-latency communication between vehicles. On the other hand, three-dimensional maps and the like, which are map data downloaded from the server, have the three-dimensional space divided into larger random access units compared to the case of vehicle-to-vehicle communication.
[0424] Data of areas that are likely to be occlusion areas, such as the area in front of the leading vehicle or the area behind the trailing vehicle, is divided into fine random access units as low-latency-oriented data.
[0425] Since the importance of the front increases during high-speed driving, each vehicle creates SWLD in a narrow field of view range in fine random access units during high-speed driving.
[0426] When an area where the host vehicle 600 can acquire point clouds is included in the SWLD created by the leading vehicle for transmission, the leading vehicle may reduce the transmission amount by removing the point clouds in that area.
[0427] Next, the configuration and operation of the three-dimensional data creation device 620, which is a three-dimensional data reception device according to the present embodiment, will be described.
[0428] FIG. 26 is a block diagram of the three-dimensional data creation device 620 according to the present embodiment. This three-dimensional data creation device 620 is included in, for example, the host vehicle 600 described above, and creates denser third three-dimensional data 636 by synthesizing the received second three-dimensional data 635 with the first three-dimensional data 632 created by the three-dimensional data creation device 620.
[0429] This three-dimensional data creation device 620 includes a three-dimensional data creation unit 621, a required range determination unit 622, a search unit 623, a reception unit 624, a decoding unit 625, and a synthesis unit 626. FIG. 27 is a flowchart showing the operation of the three-dimensional data creation device 620.
[0430] First, the three-dimensional data creation unit 621 creates first three-dimensional data 632 using sensor information 631 detected by sensors included in the host vehicle 600 (S621). Next, the required range determination unit 622 determines a required range, which is a three-dimensional space range lacking data, in the created first three-dimensional data 632 (S622).
[0431] Next, the search unit 623 searches for surrounding vehicles 601 that possess the three-dimensional data within the requested range, and transmits request range information 633 indicating the requested range to the surrounding vehicles 601 identified through the search (S623). Next, the reception unit 624 receives encoded three-dimensional data 634, which is the encoded stream of the requested range, from the surrounding vehicles 601 (S624). Note that the search unit 623 may issue requests indiscriminately to all vehicles existing within the specific range, and receive the encoded three-dimensional data 634 from the party that responded. Also, the search unit 623 may issue requests not only to vehicles but also to objects such as traffic signals or signs, and receive the encoded three-dimensional data 634 from the object.
[0432] Next, the decoding unit 625 obtains the second three-dimensional data 635 by decoding the received encoded three-dimensional data 634 (S625). Next, the combining unit 626 creates denser third three-dimensional data 636 by combining the first three-dimensional data 632 and the second three-dimensional data 635 (S626).
[0433] Next, the configuration and operation of the three-dimensional data transmission device 640 according to the present embodiment will be described. FIG. 28 is a block diagram of the three-dimensional data transmission device 640.
[0434] The three-dimensional data transmission device 640 is included in, for example, the above-described surrounding vehicle 601, processes the fifth three-dimensional data 652 created by the surrounding vehicle 601 into sixth three-dimensional data 654 requested by the host vehicle 600, generates encoded three-dimensional data 634 by encoding the sixth three-dimensional data 654, and transmits the encoded three-dimensional data 634 to the host vehicle 600.
[0435] The three-dimensional data transmission device 640 includes a three-dimensional data creation unit 641, a reception unit 642, an extraction unit 643, an encoding unit 644, and a transmission unit 645. FIG. 29 is a flowchart showing the operation of the three-dimensional data transmission device 640.
[0436] First, the three-dimensional data creation unit 641 creates fifth three-dimensional data 652 using the sensor information 651 detected by the sensors provided in the surrounding vehicle 601 (S641). Next, the reception unit 642 receives the requested range information 633 transmitted from the host vehicle 600 (S642).
[0437] Next, the extraction unit 643 processes the fifth three-dimensional data 652 into sixth three-dimensional data 654 by extracting the three-dimensional data within the requested range indicated by the requested range information 633 from the fifth three-dimensional data 652 (S643). Next, the encoding unit 644 encodes the sixth three-dimensional data 654 to generate encoded three-dimensional data 634, which is an encoded stream (S644). Then, the transmission unit 645 transmits the encoded three-dimensional data 634 to the host vehicle 600 (S645).
[0438] Here, an example in which the host vehicle 600 includes the three-dimensional data creation device 620 and the surrounding vehicle 601 includes the three-dimensional data transmission device 640 is described. However, each vehicle may have the functions of the three-dimensional data creation device 620 and the three-dimensional data transmission device 640.
[0439] Hereinafter, the configuration and operation of the three-dimensional data creation device 620 when it is a surrounding situation detection device that realizes the detection process of the surrounding situation of the host vehicle 600 will be described. FIG. 30 is a block diagram showing the configuration of the three-dimensional data creation device 620A in this case. The three-dimensional data creation device 620A shown in FIG. 30 further includes a detection area determination unit 627, a surrounding situation detection unit 628, and an autonomous operation control unit 629 in addition to the configuration of the three-dimensional data creation device 620 shown in FIG. 26. The three-dimensional data creation device 620A is included in the host vehicle 600.
[0440] FIG. 31 is a flowchart of the surrounding situation detection process of the host vehicle 600 by the three-dimensional data creation device 620A.
[0441] First, the three-dimensional data creation unit 621 creates first three-dimensional data 632, which is a point cloud, using sensor information 631 within the detection range of the host vehicle 600 detected by sensors included in the host vehicle 600 (S661). Note that the three-dimensional data creation device 620A may further perform self-position estimation using the sensor information 631.
[0442] Next, the detection area determination unit 627 determines a detection target range, which is a spatial area for which the surrounding situation is to be detected (S662). For example, the detection area determination unit 627 calculates an area necessary for detecting the surrounding situation for safe autonomous operation (automatic driving) according to the situation of autonomous operation such as the traveling direction and speed of the host vehicle 600, and determines the area as the detection target range.
[0443] Next, the required range determination unit 622 determines an occlusion area 604 and a spatial area outside the detection range of the sensors of the host vehicle 600 but necessary for detecting the surrounding situation as a required range (S663).
[0444] When the required range determined in step S663 exists (Yes in S664), the search unit 623 searches for surrounding vehicles that possess information regarding the required range. For example, the search unit 623 may inquire of surrounding vehicles whether they possess information regarding the required range, or may determine whether a surrounding vehicle possesses information regarding the required range based on the position of the required range and the surrounding vehicle. Next, the search unit 623 transmits a request signal 637 requesting transmission of three-dimensional data to the surrounding vehicle 601 identified by the search. Then, after receiving a permission signal indicating that the request of the request signal 637 has been received from the surrounding vehicle 601, the search unit 623 transmits required range information 633 indicating the required range to the surrounding vehicle 601 (S665).
[0445] Next, the reception unit 624 detects a transmission notification of transmission data 638, which is information regarding the required range, and receives the transmission data 638 (S666).
[0446] Note that the three-dimensional data creation device 620A may send requests to all vehicles existing in a specific range without searching for the recipient of the request, and may receive the transmission data 638 from the recipient that responds with information regarding the request range. Further, the search unit 623 may send requests not only to vehicles but also to objects such as traffic lights or signs, and may receive the transmission data 638 from the object.
[0447] In addition, the transmission data 638 includes at least one of the encoded three-dimensional data 634 in which the three-dimensional data of the request range generated by the surrounding vehicle 601 is encoded and the surrounding situation detection result 639 of the request range. The surrounding situation detection result 639 indicates the positions, moving directions, moving speeds, etc. of people and vehicles detected by the surrounding vehicle 601. Further, the transmission data 638 may include information indicating the position and movement, etc. of the surrounding vehicle 601. For example, the transmission data 638 may include the braking information of the surrounding vehicle 601.
[0448] When the encoded three-dimensional data 634 is included in the received transmission data 638 (Yes in S667), the decoding unit 625 obtains the second three-dimensional data 635 of the SWLD by decoding the encoded three-dimensional data 634 (S668). That is, the second three-dimensional data 635 is three-dimensional data (SWLD) generated by extracting data having a feature amount equal to or greater than the threshold value from the fourth three-dimensional data (WLD).
[0449] Next, the combining unit 626 generates the third three-dimensional data 636 by combining the first three-dimensional data 632 and the second three-dimensional data 635 (S669).
[0450] Next, the surrounding situation detection unit 628 detects the surrounding situation of the host vehicle 600 using the third three-dimensional data 636, which is a point cloud of the spatial region necessary for surrounding situation detection (S670). When the received transmission data 638 includes the surrounding situation detection result 639, the surrounding situation detection unit 628 detects the surrounding situation of the host vehicle 600 using the surrounding situation detection result 639 in addition to the third three-dimensional data 636. When the received transmission data 638 includes the braking information of the surrounding vehicle 601, the surrounding situation detection unit 628 detects the surrounding situation of the host vehicle 600 using the braking information in addition to the third three-dimensional data 636.
[0451] Next, the autonomous operation control unit 629 controls the autonomous operation (automatic driving) of the host vehicle 600 based on the surrounding situation detection result by the surrounding situation detection unit 628 (S671). Note that the surrounding situation detection result may be presented to the driver through a UI (user interface) or the like.
[0452] On the other hand, when there is no required range in step S663 (No in S664), that is, when the information of all the spatial regions necessary for surrounding situation detection can be created based on the sensor information 631, the surrounding situation detection unit 628 detects the surrounding situation of the host vehicle 600 using the first three-dimensional data 632, which is a point cloud of the spatial region necessary for surrounding situation detection (S672). Then, the autonomous operation control unit 629 controls the autonomous operation (automatic driving) of the host vehicle 600 based on the surrounding situation detection result by the surrounding situation detection unit 628 (S671).
[0453] When the received transmission data 638 does not include the encoded three-dimensional data 634 (No in S667), that is, when the transmission data 638 includes only the surrounding situation detection result 639 or the braking information of the surrounding vehicle 601, the surrounding situation detection unit 628 detects the surrounding situation of the host vehicle 600 using the first three-dimensional data 632 and the surrounding situation detection result 639 or the braking information (S673). Then, the autonomous operation control unit 629 controls the autonomous operation (automatic driving) of the host vehicle 600 based on the surrounding situation detection result by the surrounding situation detection unit 628 (S671).
[0454] Next, a three-dimensional data transmission device 640A that transmits transmission data 638 to the three-dimensional data creation device 620A will be described. FIG. 32 is a block diagram of this three-dimensional data transmission device 640A.
[0455] The three-dimensional data transmission device 640A shown in FIG. 32 includes a transmission permission determination unit 646 in addition to the configuration of the three-dimensional data transmission device 640 shown in FIG. 28. Further, the three-dimensional data transmission device 640A is included in the surrounding vehicle 601.
[0456] FIG. 33 is a flowchart showing an operation example of the three-dimensional data transmission device 640A. First, the three-dimensional data creation unit 641 creates fifth three-dimensional data 652 using sensor information 651 detected by sensors provided in the surrounding vehicle 601 (S681).
[0457] Next, the receiving unit 642 receives a request signal 637 for requesting a transmission request of three-dimensional data from the host vehicle 600 (S682). Next, the transmission permission determination unit 646 determines whether to respond to the request indicated by the request signal 637 (S683). For example, the transmission permission determination unit 646 determines whether to respond to the request based on the content preset by the user. Note that the receiving unit 642 may receive a request from the other party such as a request range first, and the transmission permission determination unit 646 may determine whether to respond to the request according to the content. For example, the transmission permission determination unit 646 may determine to respond to the request when it has the three-dimensional data within the request range, and may determine not to respond to the request when it does not have the three-dimensional data within the request range.
[0458] When responding to the request (Yes in S683), the three-dimensional data transmission device 640A transmits a permission signal to the host vehicle 600, and the receiving unit 642 receives request range information 633 indicating the request range (S684). Next, the extraction unit 643 cuts out the point cloud within the request range from the fifth three-dimensional data 652 that is a point cloud, and creates transmission data 638 including sixth three-dimensional data 654 that is the SWLD of the cut-out point cloud (S685).
[0459] That is, the three-dimensional data transmission device 640A creates the seventh three-dimensional data (WLD) from the sensor information 651, and creates the fifth three-dimensional data 652 (SWLD) by extracting data with a feature amount equal to or greater than the threshold value from the seventh three-dimensional data (WLD). Note that the three-dimensional data creation unit 641 may create the three-dimensional data of SWLD in advance, and the extraction unit 643 may extract the three-dimensional data of SWLD within the required range from the three-dimensional data of SWLD, or the extraction unit 643 may generate the three-dimensional data of SWLD within the required range from the three-dimensional data of WLD within the required range.
[0460] In addition, the transmission data 638 may include the surrounding situation detection result 639 within the required range by the surrounding vehicle 601 and the braking information of the surrounding vehicle 601. Further, the transmission data 638 may not include the sixth three-dimensional data 654, and may include only at least one of the surrounding situation detection result 639 within the required range by the surrounding vehicle 601 and the braking information of the surrounding vehicle 601.
[0461] When the transmission data 638 includes the sixth three-dimensional data 654 (Yes in S686), the encoding unit 644 generates the encoded three-dimensional data 634 by encoding the sixth three-dimensional data 654 (S687).
[0462] Then, the transmission unit 645 transmits the transmission data 638 including the encoded three-dimensional data 634 to the host vehicle 600 (S688).
[0463] On the other hand, when the transmission data 638 does not include the sixth three-dimensional data 654 (No in S686), the transmission unit 645 transmits the transmission data 638 including at least one of the surrounding situation detection result 639 within the required range by the surrounding vehicle 601 and the braking information of the surrounding vehicle 601 to the host vehicle 600 (S688).
[0464] Hereinafter, a modification example of the present embodiment will be described.
[0465] For example, the information transmitted from the surrounding vehicle 601 may not be the three-dimensional data created by the surrounding vehicle or the surrounding situation detection result, but may be the accurate feature point information of the surrounding vehicle 601 itself. The host vehicle 600 corrects the feature point information of the leading vehicle in the point cloud acquired by the host vehicle 600 using the feature point information of this surrounding vehicle 601. Thereby, the host vehicle 600 can improve the matching accuracy during self-position estimation.
[0466] Also, the feature point information of the leading vehicle is, for example, three-dimensional point information composed of color information and coordinate information. Thereby, even if the sensor of the host vehicle 600 is a laser sensor or a stereo camera, the feature point information of the leading vehicle can be used without depending on its type.
[0467] Note that the host vehicle 600 may use the point cloud of the SWLD not only during transmission but also when calculating the accuracy during self-position estimation. For example, when the sensor of the host vehicle 600 is an imaging device such as a stereo camera, the host vehicle 600 detects two-dimensional points on the image captured by the camera and estimates its own position using the two-dimensional points. Also, the host vehicle 600 creates a point cloud of surrounding objects at the same time as its own position. The host vehicle 600 reprojects the three-dimensional points of the SWLD therein onto the two-dimensional image and evaluates the accuracy of self-position estimation based on the error between the detected points and the reprojected points on the two-dimensional image.
[0468] Also, when the sensor of the host vehicle 600 is a laser sensor such as LIDAR, the host vehicle 600 evaluates the accuracy of self-position estimation based on the error calculated by Iterative Closest Point using the SWLD of the created point cloud and the SWLD of the three-dimensional map.
[0469] Also, when the communication state via a base station or a server such as 5G is poor, the host vehicle 600 may acquire a three-dimensional map from the surrounding vehicle 601.
[0470] Further, information from a distance that cannot be obtained from vehicles around the host vehicle 600 may be obtained by vehicle-to-vehicle communication. For example, the host vehicle 600 may obtain information such as traffic accident information immediately after occurrence several hundred meters or several kilometers ahead through oncoming vehicle passing communication or a relay method in which the information is sequentially transmitted to surrounding vehicles. At this time, the data format of the transmitted data is transmitted as meta information of the upper layer of the dynamic three-dimensional map.
[0471] Further, the detection result of the surrounding situation and the information detected by the host vehicle 600 may be presented to the user through the user interface. For example, these pieces of information are presented by superimposing them on the screen of the car navigation or the front window.
[0472] In addition, a vehicle that does not support autonomous driving but has cruise control may detect surrounding vehicles traveling in the autonomous driving mode and follow the surrounding vehicles.
[0473] Further, when the host vehicle 600 cannot obtain a three-dimensional map or cannot estimate its own position because there are too many occlusion areas, etc., the operation mode may be switched from the autonomous driving mode to the surrounding vehicle following mode.
[0474] In addition, the vehicle being followed may be equipped with a user interface that warns the user that the vehicle is being followed and allows the user to specify whether to permit the following. At that time, an advertisement may be displayed on the following vehicle, and a mechanism such as paying an incentive to the followed side may be provided.
[0475] Further, the transmitted information is based on SWLD which is three-dimensional data, but may be information according to the requirement settings set in the host vehicle 600 or the public settings of the preceding vehicle. For example, the transmitted information may be WLD which is dense point cloud, the detection result of the surrounding situation by the preceding vehicle, or the braking information of the preceding vehicle.
[0476] In addition, the host vehicle 600 may receive the WLD, visualize the three-dimensional data of the WLD, and present the visualized three-dimensional data to the driver using the GUI. At this time, the host vehicle 600 may present the information with color coding or the like so that the user can distinguish between the point cloud created by the host vehicle 600 and the received point cloud.
[0477] In addition, when the host vehicle 600 presents the information detected by the host vehicle 600 and the detection result of the surrounding vehicle 601 to the driver using the GUI, the host vehicle 600 may present the information with color coding or the like so that the user can distinguish between the information detected by the host vehicle 600 and the received detection result.
[0478] As described above, in the three-dimensional data creation device 620 according to the present embodiment, the three-dimensional data creation unit 621 creates the first three-dimensional data 632 from the sensor information 631 detected by the sensor. The reception unit 624 receives the encoded three-dimensional data 634 in which the second three-dimensional data 635 is encoded. The decoding unit 625 acquires the second three-dimensional data 635 by decoding the received encoded three-dimensional data 634. The combining unit 626 creates the third three-dimensional data 636 by combining the first three-dimensional data 632 and the second three-dimensional data 635.
[0479] According to this, the three-dimensional data creation device 620 can create detailed third three-dimensional data 636 using the created first three-dimensional data 632 and the received second three-dimensional data 635.
[0480] In addition, the combining unit 626 creates the third three-dimensional data 636 having a higher density than the first three-dimensional data 632 and the second three-dimensional data 635 by combining the first three-dimensional data 632 and the second three-dimensional data 635.
[0481] In addition, the second three-dimensional data 635 (for example, SWLD) is three-dimensional data generated by extracting data having a feature amount equal to or greater than a threshold value from the fourth three-dimensional data (for example, WLD).
[0482] According to this, the three-dimensional data creation device 620 can reduce the data amount of the transmitted three-dimensional data.
[0483] Further, the three-dimensional data creation device 620 further includes a search unit 623 that searches for a transmission device that is the transmission source of the encoded three-dimensional data 634. The reception unit 624 receives the encoded three-dimensional data 634 from the searched transmission device.
[0484] According to this, the three-dimensional data creation device 620 can identify, for example, a transmission device that owns necessary three-dimensional data by searching.
[0485] Further, the three-dimensional data creation device further includes a request range determination unit 622 that determines a request range that is the range of the three-dimensional space for which three-dimensional data is requested. The search unit 623 transmits request range information 633 indicating the request range to the transmission device. The second three-dimensional data 635 includes the three-dimensional data within the request range.
[0486] According to this, the three-dimensional data creation device 620 can receive necessary three-dimensional data and reduce the data amount of the transmitted three-dimensional data.
[0487] Further, the request range determination unit 622 determines a space range including the occlusion area 604 that cannot be detected by the sensor as the request range.
[0488] Further, in the three-dimensional data transmission device 640 according to the present embodiment, the three-dimensional data creation unit 641 creates fifth three-dimensional data 652 from sensor information 651 detected by the sensor. The extraction unit 643 creates sixth three-dimensional data 654 by extracting a part of the fifth three-dimensional data 652. The encoding unit 644 generates encoded three-dimensional data 634 by encoding the sixth three-dimensional data 654. The transmission unit 645 transmits the encoded three-dimensional data 634.
[0489] According to this, the three-dimensional data transmission device 640 can transmit the three-dimensional data created by itself to other devices and reduce the data amount of the transmitted three-dimensional data.
[0490] In addition, the three-dimensional data creation unit 641 creates seventh three-dimensional data (e.g., WLD) from the sensor information 651 detected by the sensor, and creates fifth three-dimensional data 652 (e.g., SWLD) by extracting data with a feature amount equal to or greater than a threshold value from the seventh three-dimensional data.
[0491] According to this, the three-dimensional data transmission device 640 can reduce the data amount of the three-dimensional data to be transmitted.
[0492] In addition, the three-dimensional data transmission device 640 further includes a receiving unit 642 that receives request range information 633 indicating a request range, which is the range of the three-dimensional space for requesting three-dimensional data, from the receiving device. The extraction unit 643 creates sixth three-dimensional data 654 by extracting the three-dimensional data within the request range from the fifth three-dimensional data 652. The transmission unit 645 transmits the encoded three-dimensional data 634 to the receiving device.
[0493] According to this, the three-dimensional data transmission device 640 can reduce the data amount of the three-dimensional data to be transmitted.
[0494] (Embodiment 4) In this embodiment, an abnormal operation in self-position estimation based on a three-dimensional map will be described.
[0495] It is expected that applications such as automatic driving of vehicles, or autonomous movement of moving bodies such as robots or flying bodies such as drones will expand in the future. As an example of a means for realizing such autonomous movement, there is a method in which a moving body travels according to a map while estimating its own position (self-position estimation) within a three-dimensional map.
[0496] Self-position estimation can be realized by matching a three-dimensional map with three-dimensional information around the host vehicle (hereinafter, host vehicle detection three-dimensional data) acquired by sensors such as a range finder (such as LiDAR) or a stereo camera mounted on the host vehicle, and estimating the position of the host vehicle within the three-dimensional map.
[0497] The three-dimensional map may include not only three-dimensional point clouds, but also two-dimensional map data such as the shape information of roads and intersections, or information that changes in real time such as traffic jams and accidents, like the HD map proposed by HERE. The three-dimensional map is composed of multiple layers such as three-dimensional data, two-dimensional data, and metadata that changes in real time, and the device can also acquire or refer to only the necessary data.
[0498] The data of the point cloud may be the above-mentioned SWLD, or may include point cloud data that is not feature points. In addition, the transmission and reception of the data of the point cloud are performed based on one or more random access units.
[0499] The following method can be used as a method for matching the three-dimensional map and the three-dimensional data of the host vehicle detection. For example, the device compares the shapes of the point clouds in the respective point clouds, and determines that the part with a high similarity between the feature points is the same position. In addition, when the three-dimensional map is composed of SWLD, the device performs matching by comparing the feature points constituting the SWLD with the three-dimensional feature points extracted from the three-dimensional data of the host vehicle detection.
[0500] Here, in order to perform highly accurate self-position estimation, it is necessary that (A) the three-dimensional map and the three-dimensional data of the host vehicle detection can be acquired, and (B) their accuracy meets a predetermined standard. However, in the following abnormal cases, (A) or (B) cannot be satisfied.
[0501] (1) The three-dimensional map cannot be acquired via communication.
[0502] (2) The three-dimensional map does not exist, or the three-dimensional map has been acquired but is damaged.
[0503] (3) The sensor of the host vehicle is malfunctioning, or due to bad weather, the generation accuracy of the three-dimensional data of the host vehicle detection is not sufficient.
[0504] The operations for handling these abnormal cases will be described below. In the following, the operations will be described by taking a vehicle as an example, but the following method can be applied to all moving objects that move autonomously, such as robots or drones.
[0505] Next, the configuration and operation of the three-dimensional information processing apparatus according to the present embodiment for dealing with abnormal cases in the three-dimensional map or the ego-vehicle detection three-dimensional data will be described. FIG. 34 is a block diagram showing a configuration example of the three-dimensional information processing apparatus 700 according to the present embodiment. FIG. 35 is a flowchart of a three-dimensional information processing method by the three-dimensional information processing apparatus 700.
[0506] The three-dimensional information processing apparatus 700 is mounted on a moving object such as an automobile, for example. As shown in FIG. 34, the three-dimensional information processing apparatus 700 includes a three-dimensional map acquisition unit 701, an ego-vehicle detection data acquisition unit 702, an abnormal case determination unit 703, a countermeasure operation determination unit 704, and an operation control unit 705.
[0507] Note that the three-dimensional information processing apparatus 700 may include a two-dimensional or one-dimensional sensor (not shown) for detecting structures or moving objects around the ego-vehicle, such as a camera for acquiring two-dimensional images or a sensor for one-dimensional data using ultrasonic waves or lasers. Further, the three-dimensional information processing apparatus 700 may include a communication unit (not shown) for acquiring the three-dimensional map via a mobile communication network such as 4G or 5G, or vehicle-to-vehicle communication or road-to-vehicle communication.
[0508] As shown in FIG. 35, the three-dimensional map acquisition unit 701 acquires a three-dimensional map 711 in the vicinity of the travel route (S701). For example, the three-dimensional map acquisition unit 701 acquires the three-dimensional map 711 via a mobile communication network, or vehicle-to-vehicle communication or road-to-vehicle communication.
[0509] Next, the ego-vehicle detection data acquisition unit 702 acquires ego-vehicle detection three-dimensional data 712 based on the sensor information (S702). For example, the ego-vehicle detection data acquisition unit 702 generates the ego-vehicle detection three-dimensional data 712 based on the sensor information acquired by the sensors provided in the ego-vehicle.
[0510] Next, the abnormal case determination unit 703 detects an abnormal case by performing a predetermined check on at least one of the acquired three-dimensional map 711 and the own vehicle detection three-dimensional data 712 (S703). That is, the abnormal case determination unit 703 determines whether at least one of the acquired three-dimensional map 711 and the own vehicle detection three-dimensional data 712 is abnormal.
[0511] In step S703, when an abnormal case is detected (Yes in S704), the countermeasure operation determination unit 704 determines a countermeasure operation for the abnormal case (S705). Next, the operation control unit 705 controls the operations of each processing unit required for implementing the countermeasure operation, such as the three-dimensional map acquisition unit 701 (S706).
[0512] On the other hand, in step S703, when no abnormal case is detected (No in S704), the three-dimensional information processing device 700 ends the process.
[0513] Also, the three-dimensional information processing device 700 estimates the self-position of the vehicle having the three-dimensional information processing device 700 using the three-dimensional map 711 and the own vehicle detection three-dimensional data 712. Next, the three-dimensional information processing device 700 automatically drives the vehicle using the result of the self-position estimation.
[0514] In this way, the three-dimensional information processing device 700 acquires map data (three-dimensional map 711) including the first three-dimensional position information via a communication path. For example, the first three-dimensional position information is encoded in units of partial spaces having three-dimensional coordinate information, each being an aggregate of one or more partial spaces, and includes a plurality of random access units that can be decoded independently. For example, the first three-dimensional position information is data (SWLD) in which feature points where three-dimensional feature amounts are equal to or greater than a predetermined threshold are encoded.
[0515] In addition, the three-dimensional information processing device 700 generates second three-dimensional position information (own vehicle detection three-dimensional data 712) from the information detected by the sensor. Next, the three-dimensional information processing device 700 determines whether the first three-dimensional position information or the second three-dimensional position information is abnormal by performing an abnormality determination process on the first three-dimensional position information or the second three-dimensional position information.
[0516] When it is determined that the first three-dimensional position information or the second three-dimensional position information is abnormal, the three-dimensional information processing device 700 determines a coping operation for the abnormality. Next, the three-dimensional information processing device 700 performs control necessary for implementing the coping operation.
[0517] Thereby, the three-dimensional information processing device 700 can detect an abnormality in the first three-dimensional position information or the second three-dimensional position information and perform a coping operation.
[0518] Hereinafter, a coping operation for the case where the three-dimensional map 711, which is abnormal case 1, cannot be acquired via communication will be described.
[0519] The three-dimensional map 711 is necessary for self-position estimation. However, when the vehicle has not acquired the three-dimensional map 711 corresponding to the route to the destination in advance, it is necessary to acquire the three-dimensional map 711 by communication. However, due to congestion of the communication path or deterioration of the radio wave reception state, etc., the vehicle may not be able to acquire the three-dimensional map 711 on the driving route.
[0520] The abnormality case determination unit 703 checks whether the three-dimensional map 711 in all sections on the route to the destination or in a section within a predetermined range from the current position has been acquired. If not, it is determined as abnormal case 1. That is, the abnormality case determination unit 703 determines whether the three-dimensional map 711 (first three-dimensional position information) can be acquired via the communication path. If the three-dimensional map 711 cannot be acquired via the communication path, it is determined that the three-dimensional map 711 is abnormal.
[0521] When it is determined as abnormal case 1, the countermeasure operation determination unit 704 selects one of two types of countermeasure operations: (1) continuing the self-position estimation and (2) stopping the self-position estimation.
[0522] First, a specific example of the countermeasure operation when (1) continuing the self-position estimation will be described. When continuing the self-position estimation, a three-dimensional map 711 on the route to the destination is required.
[0523] For example, the vehicle determines a location where the communication path is available within the range where the three-dimensional map 711 has been acquired, moves to that location, and acquires the three-dimensional map 711. At this time, the vehicle may acquire all the three-dimensional maps 711 to the destination, or may acquire the three-dimensional map 711 for each random access unit within the upper limit size that can be held in a recording unit such as the vehicle's memory or HDD.
[0524] In addition, the vehicle separately acquires the communication state on the route. If it can be predicted that the communication state on the route will be poor, the vehicle acquires the three-dimensional map 711 of the section in advance before reaching the section where the communication state is poor, or operates to acquire the three-dimensional map 711 within the maximum acquirable range. That is, the three-dimensional information processing device 700 predicts whether the vehicle will enter an area with a poor communication state. When it is predicted that the vehicle will enter an area with a poor communication state, the three-dimensional information processing device 700 acquires the three-dimensional map 711 before the vehicle enters the area.
[0525] Further, the vehicle may specify a random access unit that constitutes the minimum three-dimensional map 711 required for self-position estimation on the route, which is a narrower range than normal, and receive the specified random access unit. That is, when the three-dimensional information processing device 700 cannot acquire the three-dimensional map 711 (first three-dimensional position information) via the communication path, it may acquire the third three-dimensional position information in a range narrower than the first three-dimensional position information via the communication path.
[0526] Also, when the vehicle cannot access the distribution server of the three-dimensional map 711, the vehicle may obtain the three-dimensional map 711 on the route to the destination, such as other vehicles traveling around the host vehicle, from a moving object that has already obtained the three-dimensional map 711 and can communicate with the host vehicle.
[0527] Next, a specific example of the coping operation when (2) stopping the self-position estimation will be described. In this case, the three-dimensional map 711 on the route to the destination is not necessary.
[0528] For example, the vehicle notifies the driver that functions such as automatic driving based on self-position estimation cannot be continued, and shifts the operation mode to a manual mode in which the driver drives the vehicle.
[0529] Normally, when performing self-position estimation, although there are differences in levels according to the degree of human intervention, automatic driving is performed. On the other hand, the result of self-position estimation can also be used as navigation or the like when a person drives. Therefore, the result of self-position estimation does not necessarily have to be used for automatic driving.
[0530] Also, when a communication path normally used, such as a mobile communication network such as 4G or 5G, cannot be used, the vehicle checks whether it can obtain the three-dimensional map 711 via another communication path, such as vehicle-to-roadside Wi-Fi (registered trademark) or millimeter-wave communication, or vehicle-to-vehicle communication, and may switch the communication path used to a communication path capable of obtaining the three-dimensional map 711.
[0531] Also, when the vehicle cannot obtain the three-dimensional map 711, the vehicle may obtain a two-dimensional map and continue automatic driving using the two-dimensional map and the self-vehicle detection three-dimensional data 712. That is, when the three-dimensional information processing device 700 cannot obtain the three-dimensional map 711 via the communication path, the map data (two-dimensional map) including two-dimensional position information may be obtained via the communication path, and the self-position estimation of the vehicle may be performed using the two-dimensional position information and the self-vehicle detection three-dimensional data 712.
[0532] Specifically, for self-position estimation, the vehicle uses a two-dimensional map and self-vehicle detection three-dimensional data 712, and for detecting surrounding vehicles, pedestrians, obstacles, etc., the self-vehicle detection three-dimensional data 712 is used.
[0533] Here, the map data such as the HD map includes, together with the three-dimensional map 711 composed of three-dimensional point clouds, etc., two-dimensional map data (two-dimensional map), a simplified version of the map data obtained by extracting characteristic information such as road shape or intersection from the two-dimensional map data, and metadata representing real-time information such as traffic jams, accidents, or construction. For example, the map data has a layer structure in which three-dimensional data (three-dimensional map 711), two-dimensional data (two-dimensional map), and metadata are arranged in order from the lower layer.
[0534] Here, the two-dimensional data has a smaller data size than the three-dimensional data. Therefore, the vehicle may be able to acquire the two-dimensional map even when the communication state is poor. Or, the vehicle can acquire a wide range of two-dimensional maps collectively in a section with a good communication state. Therefore, when the communication path state is poor and it is difficult to acquire the three-dimensional map 711, the vehicle may receive the layer including the two-dimensional map without receiving the three-dimensional map 711. Since the metadata has a small data size, for example, the vehicle always receives the metadata regardless of the communication state.
[0535] There are, for example, the following two methods for the self-position estimation method using the two-dimensional map and the self-vehicle detection three-dimensional data 712.
[0536] The first method is a method of performing two-dimensional feature quantity matching. Specifically, the vehicle extracts two-dimensional feature quantities from the self-vehicle detection three-dimensional data 712 and performs matching between the extracted two-dimensional feature quantities and the two-dimensional map.
[0537] For example, the vehicle projects the self-vehicle detection three-dimensional data 712 onto the same plane as the two-dimensional map and matches the obtained two-dimensional data with the two-dimensional map. The matching is performed using two-dimensional image feature quantities extracted from both.
[0538] When the three - dimensional map 711 includes SWLD, the three - dimensional map 711 may store two - dimensional feature amounts in the same plane as the two - dimensional map, together with the three - dimensional feature amounts at the feature points in the three - dimensional space. For example, identification information is attached to the two - dimensional feature amounts. Or, the two - dimensional feature amounts are stored in a layer separate from the three - dimensional data and the two - dimensional map, and the vehicle acquires the data of the two - dimensional feature amounts together with the two - dimensional map.
[0539] When the two - dimensional map shows information on positions with different heights (not in the same plane), such as white lines, guardrails, and buildings on the road, within the same map, the vehicle extracts feature amounts from the data of multiple heights in the ego - vehicle detection three - dimensional data 712.
[0540] Also, information indicating the correspondence relationship between the feature points in the two - dimensional map and the feature points in the three - dimensional map 711 may be stored as meta - information of the map data.
[0541] The second method is a method of performing three - dimensional feature amount matching. Specifically, the vehicle acquires the three - dimensional feature amounts corresponding to the feature points in the two - dimensional map, and matches the acquired three - dimensional feature amounts with the three - dimensional feature amounts of the ego - vehicle detection three - dimensional data 712.
[0542] Specifically, the three - dimensional feature amounts corresponding to the feature points in the two - dimensional map are stored in the map data. The vehicle acquires these three - dimensional feature amounts together when acquiring the two - dimensional map. When the three - dimensional map 711 includes SWLD, by attaching information for identifying the feature points corresponding to the feature points of the two - dimensional map among the feature points in SWLD, the vehicle can determine the three - dimensional feature amounts acquired together with the two - dimensional map based on the identification information. In this case, since it is only necessary to represent the two - dimensional position, the data amount can be reduced compared to the case of representing the three - dimensional position.
[0543] In addition, when estimating the self-position using a two-dimensional map, the accuracy of self-position estimation is lower than that of the three-dimensional map 711. Therefore, the vehicle determines whether it can continue the automatic driving even if the estimation accuracy decreases, and may continue the automatic driving only when it is determined that it can continue.
[0544] Whether the automatic driving can continue also depends on whether the road on which the vehicle is traveling is an urban area or a road with few other vehicles or pedestrians, such as a highway, and is also affected by the driving environment such as the road width or the traffic congestion (density of vehicles or pedestrians). Furthermore, it is also possible to arrange markers for recognition by sensors such as cameras in the premises of a business, a street, or a building. In these specific areas, since markers can be recognized with high accuracy by two-dimensional sensors, for example, by including the position information of the markers in the two-dimensional map, self-position estimation can be performed with high accuracy.
[0545] In addition, by including identification information indicating whether each area in the map is a specific area, the vehicle can determine whether the vehicle exists within the specific area. When the vehicle exists within the specific area, the vehicle determines to continue the automatic driving. In this way, the vehicle may determine whether to continue the automatic driving based on the accuracy of self-position estimation when using the two-dimensional map or the driving environment of the vehicle.
[0546] In this way, the three-dimensional information processing device 700 determines whether to perform the automatic driving of the vehicle using the result of self-position estimation of the vehicle using the two-dimensional map and the self-vehicle detection three-dimensional data 712 based on the driving environment (moving environment of the moving body) of the vehicle.
[0547] In addition, the vehicle may switch the level (mode) of autonomous driving according to the accuracy of self-position estimation or the driving environment of the vehicle, rather than whether autonomous driving can continue. Here, switching the level (mode) of autonomous driving means, for example, limiting the speed, increasing the driver's operation amount (lowering the automatic level of autonomous driving), switching to a mode of driving with reference to the driving information of the vehicle running ahead, switching to a mode of autonomous driving using the driving information of the vehicle with the same destination set, and so on.
[0548] Further, the map may include information indicating the recommended level of autonomous driving when self-position estimation is performed using a two-dimensional map associated with position information. The recommended level may be metadata that dynamically changes according to traffic volume and the like. Thereby, the vehicle can determine the level simply by acquiring the information in the map without sequentially determining the level according to the surrounding environment and the like. Also, by multiple vehicles referring to the same map, the level of autonomous driving of each vehicle can be kept constant. Note that the recommended level may be a level that is not a recommendation but must be complied with.
[0549] In addition, the vehicle may switch the level of autonomous driving according to the presence or absence of a driver (whether it is manned or unmanned). For example, the vehicle lowers the level of autonomous driving if it is manned and stops when it is unmanned. The vehicle determines a position where it can stop safely by recognizing surrounding pedestrians, vehicles, and traffic signs. Or, the map may include position information indicating a position where the vehicle can stop safely, and the vehicle may determine a position where it can stop safely by referring to the position information.
[0550] Next, the coping operation for the abnormal case 2 where the three-dimensional map 711 does not exist or the three-dimensional map 711 has been acquired but is damaged will be described.
[0551] The abnormal case determination unit 703 checks whether either (1) the three-dimensional map 711 in some or all of the sections on the route to the destination does not exist in the access destination distribution server or the like and cannot be acquired, or (2) some or all of the acquired three-dimensional map 711 is damaged. If it applies, it determines that it is an abnormal case 2. That is, the abnormal case determination unit 703 determines whether the data of the three-dimensional map 711 is complete. If the data of the three-dimensional map 711 is not complete, it determines that the three-dimensional map 711 is abnormal.
[0552] When it is determined to be abnormal case 2, the following coping operations are performed. First, (1) an example of the coping operation when the three-dimensional map 711 cannot be acquired will be described.
[0553] For example, the vehicle sets a route that does not pass through the section where the three-dimensional map 711 does not exist.
[0554] In addition, when there is no alternative route or the alternative route cannot be set because the distance of the alternative route increases significantly, etc., the vehicle sets a route including the section where the three-dimensional map 711 does not exist. Also, the vehicle notifies the driver to switch the driving mode in that section and switches the driving mode to the manual mode.
[0555] (2) When some or all of the acquired three-dimensional map 711 is damaged, the following coping operations are performed.
[0556] The vehicle identifies the damaged part in the three-dimensional map 711, requests the data of the damaged part by communication, acquires the data of the damaged part, and updates the three-dimensional map 711 using the acquired data. At this time, the vehicle may specify the damaged part by position information such as absolute coordinates or relative coordinates in the three-dimensional map 711, or may specify the damaged part by the index number of the random access unit constituting the damaged part. In this case, the vehicle replaces the random access unit including the damaged part with the acquired random access unit.
[0557] Next, the coping operation for the case where the on-vehicle sensor malfunctions or due to bad weather in abnormal case 3 and the on-vehicle detection three-dimensional data 712 cannot be generated will be described.
[0558] The abnormal case determination unit 703 checks whether the generation error of the on-vehicle detection three-dimensional data 712 is within the allowable range, and determines it as abnormal case 3 if it is not within the allowable range. That is, the abnormal case determination unit 703 determines whether the generation accuracy of the data of the on-vehicle detection three-dimensional data 712 is equal to or higher than the reference value. If the generation accuracy of the data of the on-vehicle detection three-dimensional data 712 is not equal to or higher than the reference value, it is determined that the on-vehicle detection three-dimensional data 712 is abnormal.
[0559] As a method for checking whether the generation error of the on-vehicle detection three-dimensional data 712 is within the allowable range, the following method can be used.
[0560] Based on the resolution in the depth direction and the scan direction in the three-dimensional sensor of the vehicle such as a range finder or a stereo camera, or the density of the point cloud that can be generated, the spatial resolution of the on-vehicle detection three-dimensional data 712 during normal operation is determined in advance. In addition, the vehicle acquires the spatial resolution of the three-dimensional map 711 from the meta-information included in the three-dimensional map 711 and the like.
[0561] The vehicle uses the spatial resolutions of both to estimate the reference value of the matching error when matching the on-vehicle detection three-dimensional data 712 and the three-dimensional map 711 based on three-dimensional feature amounts and the like. As the matching error, statistical quantities such as the error of the three-dimensional feature amount for each feature point, the average value of the errors of the three-dimensional feature amounts between a plurality of feature points, or the error of the spatial distance between a plurality of feature points can be used. The allowable range of deviation from the reference value is set in advance.
[0562] If the matching error between the on-vehicle detection three-dimensional data 712 generated before or during driving and the three-dimensional map 711 is not within the allowable range, the vehicle determines it as abnormal case 3.
[0563] Alternatively, the vehicle may use a test pattern having a known three-dimensional shape for accuracy checking to obtain the vehicle detection three-dimensional data 712 for the test pattern before starting to run or the like, and determine whether it is an abnormal case 3 based on whether the shape error is within the allowable range.
[0564] For example, the vehicle performs the above determination every time before starting to run. Alternatively, the vehicle obtains the time-series change of the matching error by performing the above determination at regular time intervals or the like during running. When the matching error has an increasing tendency, the vehicle may determine that it is an abnormal case 3 even if the error is within the allowable range. Also, when it can be predicted that the vehicle will become abnormal based on the time-series change, the vehicle may notify the user that it is predicted to become abnormal, such as by displaying a message prompting inspection or repair. Further, the vehicle may determine an abnormality based on a transient factor such as bad weather and an abnormality based on a sensor failure by the time-series change, and notify the user only of the abnormality based on the sensor failure.
[0565] Also, when it is determined that it is an abnormal case 3, the vehicle performs any one of or selectively performs three types of countermeasures: (1) operating an emergency alternative sensor (rescue mode), (2) switching the driving mode, and (3) performing operation correction of the three-dimensional sensor.
[0566] First, the case of (1) operating an emergency alternative sensor will be described. The vehicle operates an emergency alternative sensor different from the three-dimensional sensor used during normal driving. That is, when the data generation accuracy of the vehicle detection three-dimensional data 712 is not equal to or higher than the reference value, the three-dimensional information processing device 700 generates the vehicle detection three-dimensional data 712 (fourth three-dimensional position information) from the information detected by an alternative sensor different from the normal sensor.
[0567] Specifically, when the vehicle acquires the ego-vehicle detection three-dimensional data 712 by using a plurality of cameras or LiDAR in combination, the vehicle identifies a malfunctioning sensor based on, for example, the direction in which the matching error of the ego-vehicle detection three-dimensional data 712 exceeds the allowable range. Then, the vehicle activates an alternative sensor corresponding to the malfunctioning sensor.
[0568] The alternative sensor may be a three-dimensional sensor, a camera capable of acquiring two-dimensional images, or a one-dimensional sensor such as an ultrasonic wave. When the alternative sensor is a sensor other than a three-dimensional sensor, the accuracy of self-position estimation may decrease or self-position estimation may not be possible. Therefore, the vehicle may switch the driving mode according to the type of the alternative sensor.
[0569] For example, when the alternative sensor is a three-dimensional sensor, the vehicle continues the autonomous driving mode. Also, when the alternative sensor is a two-dimensional sensor, the vehicle changes the driving mode to a semi-autonomous driving mode assuming human driving operation from the fully autonomous driving mode. Further, when the alternative sensor is a one-dimensional sensor, the vehicle switches the driving mode to a manual mode in which automatic braking control is not performed.
[0570] In addition, the vehicle may switch the autonomous driving mode based on the driving environment. For example, when the alternative sensor is a two-dimensional sensor, if the vehicle is driving on a highway, it continues the fully autonomous driving mode, and if it is driving in an urban area, it switches the driving mode to the semi-autonomous driving mode.
[0571] Also, even when there is no alternative sensor, if the vehicle can acquire a sufficient number of feature points with only the normally operating sensors, the vehicle may continue self-position estimation. However, since detection in a specific direction becomes impossible, the vehicle switches the driving mode to the semi-autonomous or manual mode.
[0572] Next, the coping operation for switching the (2) driving mode will be described. The vehicle switches the driving mode from the automatic driving mode to the manual mode. Alternatively, the vehicle may continue automatic driving to a place where it can stop safely, such as the road shoulder, and then stop. Further, the vehicle may switch the driving mode to the manual mode after stopping. In this way, when the generation accuracy of the own vehicle detection three-dimensional data 712 is not equal to or higher than the reference value, the three-dimensional information processing device 700 switches the automatic driving mode.
[0573] Next, the coping operation for performing (3) operation correction of the three-dimensional sensor will be described. The vehicle identifies a malfunctioning three-dimensional sensor from the direction in which a matching error occurs, etc., and calibrates the identified sensor. Specifically, when a plurality of LiDARs or cameras are used as sensors, a part of the three-dimensional space reconstructed by each sensor overlaps. That is, the data in the overlapping part is acquired by a plurality of sensors. The three-dimensional point cloud data acquired for the overlapping part is different between a normal sensor and a malfunctioning sensor. Therefore, the vehicle performs origin correction of the LiDAR or adjusts the operation of a predetermined part such as the exposure or focus of the camera so that the malfunctioning sensor can acquire three-dimensional point cloud data equivalent to that of a normal sensor.
[0574] After the adjustment, if the matching error is within the allowable range, the vehicle continues the previous driving mode. On the other hand, if the matching accuracy does not fall within the allowable range even after the adjustment, the vehicle performs the coping operation of activating the (1) emergency alternative sensor or the coping operation of switching the (2) driving mode.
[0575] In this way, when the generation accuracy of the data of the own vehicle detection three-dimensional data 712 is not equal to or higher than the reference value, the three-dimensional information processing device 700 performs operation correction of the sensor.
[0576] Hereinafter, the method for selecting the coping operation will be described. The coping operation may be selected by a user such as a driver, or may be automatically selected by the vehicle without going through the user.
[0577] In addition, the vehicle may switch control according to whether the driver is on board. For example, when the driver is on board, the vehicle gives priority to switching to the manual mode. On the other hand, when the driver is not on board, the vehicle gives priority to the mode of moving to a safe place and stopping.
[0578] The information indicating the stopping place may be included as meta information in the three-dimensional map 711. Alternatively, the vehicle may issue a response request for the stopping place to the service that manages the driving information of the autonomous driver and acquire the information indicating the stopping place.
[0579] In addition, when the vehicle runs on a predetermined route, etc., the driving mode of the vehicle may shift to the mode in which the operator manages the operation of the vehicle via the communication path. In particular, an abnormality in the self-position estimation function of a vehicle running in the fully autonomous driving mode is highly dangerous. Therefore, when an abnormal case is detected or when the detected abnormality cannot be corrected, the vehicle notifies the service that manages the driving information of the occurrence of the abnormality via the communication path. The said service may notify vehicles running around the said vehicle, etc. of the presence of the vehicle in which the abnormality has occurred, or may instruct to vacate a nearby stopping place.
[0580] In addition, when an abnormal case is detected, the vehicle may reduce its running speed compared to normal times.
[0581] When the vehicle is an autonomous vehicle that provides a vehicle dispatching service such as a taxi and an abnormal case occurs in the vehicle, the vehicle contacts the operation management center and stops at a safe place. In addition, the vehicle dispatching service dispatches a substitute vehicle. Alternatively, the user of the vehicle dispatching service may drive the vehicle. In these cases, a discount on the fare or the granting of privilege points may also be used in combination.
[0582] In addition, in the method for dealing with abnormal case 1, although the method of performing self-position estimation based on the two-dimensional map has been described, self-position estimation may also be performed using the two-dimensional map even during normal times. FIG. 36 is a flowchart of the self-position estimation process in this case.
[0583] First, the vehicle acquires a three-dimensional map 711 near the driving route (S711). Next, the vehicle acquires self-vehicle detection three-dimensional data 712 based on sensor information (S712).
[0584] Next, the vehicle determines whether the three-dimensional map 711 is necessary for self-position estimation (S713). Specifically, the vehicle determines the necessity of the three-dimensional map 711 based on the accuracy of self-position estimation when using a two-dimensional map and the driving environment. For example, the same method as the coping method for the above-described abnormal case 1 is used.
[0585] When it is determined that the three-dimensional map 711 is not necessary (No in S714), the vehicle acquires a two-dimensional map (S715). At this time, the vehicle may also acquire additional information as described in the coping method for abnormal case 1. Further, the vehicle may generate a two-dimensional map from the three-dimensional map 711. For example, the vehicle may generate a two-dimensional map by cutting out an arbitrary plane from the three-dimensional map 711.
[0586] Next, the vehicle performs self-position estimation using the self-vehicle detection three-dimensional data 712 and the two-dimensional map (S716). Note that the method of self-position estimation using a two-dimensional map is the same as, for example, the method described in the coping method for the above-described abnormal case 1.
[0587] On the other hand, when it is determined that the three-dimensional map 711 is necessary (Yes in S714), the vehicle acquires the three-dimensional map 711 (S717). Next, the vehicle performs self-position estimation using the self-vehicle detection three-dimensional data 712 and the three-dimensional map 711 (S718).
[0588] Note that the vehicle may switch between using a two-dimensional map as a basis and using a three-dimensional map 711 as a basis according to the corresponding speed of the communication device of the vehicle or the status of the communication path. For example, a communication speed required when traveling while receiving the three-dimensional map 711 is preset, and the vehicle may use a two-dimensional map as a basis when the communication speed during traveling is equal to or lower than the set value, and use the three-dimensional map 711 as a basis when the communication speed during traveling is greater than the set value. Note that the vehicle may use a two-dimensional map as a basis without making a determination on whether to adopt either the two-dimensional map or the three-dimensional map.
[0589] (Embodiment 5) In this embodiment, a method for transmitting three-dimensional data to a following vehicle and the like will be described. FIG. 37 is a diagram showing an example of a target space of three-dimensional data to be transmitted to a following vehicle or the like.
[0590] The vehicle 801 transmits three-dimensional data such as a point cloud (point group) included in a rectangular parallelepiped space 802 with a width W, a height H, and a depth D at a distance L from the vehicle 801 in front of the vehicle 801 to a traffic monitoring cloud or a following vehicle that monitors the road condition at time intervals of Δt.
[0591] When a change occurs in the three-dimensional data included in the space 802 that has been transmitted in the past, such as a vehicle or a person entering the space 802 from the outside, the vehicle 801 also transmits the three-dimensional data of the space in which the change has occurred.
[0592] Note that in FIG. 37, an example in which the shape of the space 802 is a rectangular parallelepiped is shown, but the space 802 only needs to include a space on the forward road that is a blind spot for the following vehicle, and does not necessarily have to be a rectangular parallelepiped.
[0593] The distance L is preferably set to a distance at which the following vehicle can safely stop after receiving the three-dimensional data. For example, the distance L is the sum of the distance that the following vehicle moves during the reception of the three-dimensional data, the distance that the following vehicle moves until it starts to decelerate according to the received data, and the distance that the following vehicle requires to safely stop after starting to decelerate. Since these distances change according to the speed, the distance L may change according to the speed V of the vehicle, such as L = a×V + b (a and b are constants).
[0594] The width W is set to a value that is at least larger than the width of the lane in which the vehicle 801 is traveling. More preferably, the width W is set to a size that includes adjacent spaces such as the left and right lanes or the roadside strip.
[0595] The depth D may be a fixed value, or may change according to the speed V of the vehicle, such as D = c×V + d (c and d are constants). Also, by setting D so that D > V×Δt, the transmitted space can be made to overlap with the space that has already been transmitted in the past. Thereby, the vehicle 801 can more reliably transmit the space on the road without omission to the following vehicle or the like.
[0596] In this way, by limiting the three-dimensional data transmitted by the vehicle 801 to a useful space for the following vehicle, the capacity of the three-dimensional data to be transmitted can be effectively reduced, and low latency and low cost of communication can be achieved.
[0597] Next, the configuration of the three-dimensional data creation device 810 according to the present embodiment will be described. FIG. 38 is a block diagram showing a configuration example of the three-dimensional data creation device 810 according to the present embodiment. This three-dimensional data creation device 810 is mounted on, for example, the vehicle 801. The three-dimensional data creation device 810 transmits and receives three-dimensional data to and from an external traffic monitoring cloud, a preceding vehicle, or a following vehicle, and creates and stores the three-dimensional data.
[0598] The three-dimensional data creation device 810 includes a data reception unit 811, a communication unit 812, a reception control unit 813, a format conversion unit 814, a plurality of sensors 815, a three-dimensional data creation unit 816, a three-dimensional data synthesis unit 817, a three-dimensional data storage unit 818, a communication unit 819, a transmission control unit 820, a format conversion unit 821, and a data transmission unit 822.
[0599] The data reception unit 811 receives three-dimensional data 831 from a traffic monitoring cloud or a preceding vehicle. The three-dimensional data 831 includes information such as point cloud, visible light video, depth information, sensor position information, or speed information, for example, including areas that cannot be detected by the sensors 815 of the host vehicle.
[0600] The communication unit 812 communicates with a traffic monitoring cloud or a preceding vehicle and transmits a data transmission request or the like to the traffic monitoring cloud or the preceding vehicle.
[0601] The reception control unit 813 exchanges information such as the corresponding format with the communication destination via the communication unit 812 and establishes communication with the communication destination.
[0602] The format conversion unit 814 generates three-dimensional data 832 by performing format conversion or the like on the three-dimensional data 831 received by the data reception unit 811. Further, when the three-dimensional data 831 is compressed or encoded, the format conversion unit 814 performs decompression or decoding processing.
[0603] The plurality of sensors 815 is a group of sensors that acquire information outside the vehicle 801, such as LIDAR, a visible light camera, or an infrared camera, and generates sensor information 833. For example, when the sensor 815 is a laser sensor such as LIDAR, the sensor information 833 is three-dimensional data such as point cloud (point group data). Note that the number of sensors 815 does not have to be plural.
[0604] The three-dimensional data creation unit 816 generates three-dimensional data 834 from the sensor information 833. The three-dimensional data 834 includes information such as, for example, point cloud, visible light video, depth information, sensor position information, or speed information.
[0605] The three-dimensional data synthesis unit 817 synthesizes the three-dimensional data 832 created by the traffic monitoring cloud or the vehicle ahead, etc., with the three-dimensional data 834 created based on the sensor information 833 of the host vehicle, thereby constructing three-dimensional data 835 that includes the space in front of the vehicle ahead that cannot be detected by the sensors 815 of the host vehicle.
[0606] The three-dimensional data storage unit 818 stores the generated three-dimensional data 835, etc.
[0607] The communication unit 819 communicates with the traffic monitoring cloud or the following vehicle, and transmits a data transmission request, etc., to the traffic monitoring cloud or the following vehicle.
[0608] The transmission control unit 820 exchanges information such as the corresponding format with the communication destination via the communication unit 819, and establishes communication with the communication destination. Further, the transmission control unit 820 determines a transmission area, which is the space of the three-dimensional data to be transmitted, based on the three-dimensional data construction information of the three-dimensional data 832 generated by the three-dimensional data synthesis unit 817 and the data transmission request from the communication destination.
[0609] Specifically, the transmission control unit 820 determines a transmission area that includes the space in front of the host vehicle that cannot be detected by the sensors of the following vehicle in response to a data transmission request from the traffic monitoring cloud or the following vehicle. Further, the transmission control unit 820 determines the transmission area by judging the availability of updating the space that can be transmitted or the transmitted space based on the three-dimensional data construction information. For example, the transmission control unit 820 determines the area specified in the data transmission request and where the corresponding three-dimensional data 835 exists as the transmission area. Then, the transmission control unit 820 notifies the format conversion unit 821 of the corresponding format of the communication destination and the transmission area.
[0610] The format conversion unit 821 generates three-dimensional data 837 by converting the three-dimensional data 836 in the transmission area among the three-dimensional data 835 stored in the three-dimensional data storage unit 818 into a format supported by the receiving side. Note that the format conversion unit 821 may reduce the data amount by compressing or encoding the three-dimensional data 837.
[0611] The data transmission unit 822 transmits the three-dimensional data 837 to the traffic monitoring cloud or the following vehicle. This three-dimensional data 837 includes information such as point cloud, visible light video, depth information, or sensor position information in front of the host vehicle, for example, including areas that are blind spots for the following vehicle.
[0612] Here, an example in which format conversion and the like are performed by the format conversion units 814 and 821 has been described, but the format conversion may not be performed.
[0613] With such a configuration, the three-dimensional data creation device 810 acquires three-dimensional data 831 of an area that cannot be detected by the sensor 815 of the host vehicle from the outside, and generates three-dimensional data 835 by synthesizing the three-dimensional data 831 and three-dimensional data 834 based on the sensor information 833 detected by the sensor 815 of the host vehicle. Thereby, the three-dimensional data creation device 810 can generate three-dimensional data in a range that cannot be detected by the sensor 815 of the host vehicle.
[0614] In addition, the three-dimensional data creation device 810 can transmit three-dimensional data including the space in front of the host vehicle that cannot be detected by the sensors of the following vehicle to the traffic monitoring cloud or the following vehicle in response to a data transmission request from the traffic monitoring cloud or the following vehicle.
[0615] Next, the transmission procedure of the three-dimensional data to the following vehicle in the three-dimensional data creation device 810 will be described. FIG. 39 is a flowchart showing an example of the procedure for transmitting three-dimensional data to the traffic monitoring cloud or the following vehicle by the three-dimensional data creation device 810.
[0616] First, the three-dimensional data creation device 810 generates and updates three-dimensional data 835 of a space including a space 802 on the road ahead of the host vehicle 801 (S801). Specifically, the three-dimensional data creation device 810 synthesizes the three-dimensional data 831 created by the traffic monitoring cloud or the preceding vehicle or the like with the three-dimensional data 834 created based on the sensor information 833 of the host vehicle 801, etc., to construct the three-dimensional data 835 including the space in front of the preceding vehicle that cannot be detected by the sensor 815 of the host vehicle.
[0617] Next, the three-dimensional data creation device 810 determines whether the three-dimensional data 835 included in the transmitted space has changed (S802).
[0618] If a vehicle or a person enters the transmitted space from the outside and the three-dimensional data 835 included in the space changes (Yes in S802), the three-dimensional data creation device 810 transmits the three-dimensional data including the three-dimensional data 835 of the changed space to the traffic monitoring cloud or the following vehicle (S803).
[0619] Note that the three-dimensional data creation device 810 may transmit the three-dimensional data of the changed space in accordance with the transmission timing of the three-dimensional data transmitted at a predetermined interval, or may transmit it immediately after detecting the change. That is, the three-dimensional data creation device 810 may transmit the three-dimensional data of the changed space with higher priority than the three-dimensional data transmitted at a predetermined interval.
[0620] Further, the three-dimensional data creation device 810 may transmit all of the three-dimensional data of the changed space as the three-dimensional data of the changed space, or may transmit only the difference of the three-dimensional data (for example, information on three-dimensional points that have appeared or disappeared, or displacement information of three-dimensional points, etc.).
[0621] In addition, the three-dimensional data creation device 810 may transmit metadata related to the risk avoidance operation of the host vehicle, such as an emergency braking warning, to the following vehicle prior to the three-dimensional data of the space in which the change has occurred. According to this, the following vehicle can recognize the emergency braking of the preceding vehicle earlier and can start a risk avoidance operation such as deceleration earlier.
[0622] When there is no change in the three-dimensional data 835 included in the transmitted space (No in S802), or after step S803, the three-dimensional data creation device 810 transmits the three-dimensional data included in the space of a predetermined shape at the forward distance L of the host vehicle 801 to the traffic monitoring cloud or the following vehicle (S804).
[0623] In addition, for example, the processes of steps S801 to S804 are repeatedly performed at a predetermined time interval.
[0624] In addition, when there is no difference between the three-dimensional data 835 of the current transmission target space 802 and the three-dimensional map, the three-dimensional data creation device 810 may not transmit the three-dimensional data 837 of the space 802.
[0625] FIG. 40 is a flowchart showing the operation of the three-dimensional data creation device 810 in this case.
[0626] First, the three-dimensional data creation device 810 generates and updates the three-dimensional data 835 of the space including the space 802 on the forward road of the host vehicle 801 (S811).
[0627] Next, the three-dimensional data creation device 810 determines whether there is an update from the three-dimensional map in the generated three-dimensional data 835 of the space 802 (S812). That is, the three-dimensional data creation device 810 determines whether there is a difference between the generated three-dimensional data 835 of the space 802 and the three-dimensional map. Here, the three-dimensional map is three-dimensional map information managed by an infrastructure-side device such as a traffic monitoring cloud. For example, this three-dimensional map is acquired as the three-dimensional data 831.
[0628] When there is an update (Yes in S812), the three-dimensional data creation device 810 transmits the three-dimensional data included in the space 802 to the traffic monitoring cloud or the following vehicle in the same manner as described above (S813).
[0629] On the other hand, when there is no update (No in S812), the three-dimensional data creation device 810 does not transmit the three-dimensional data included in the space 802 to the traffic monitoring cloud and the following vehicle (S814). Note that the three-dimensional data creation device 810 may control so that the three-dimensional data of the space 802 is not transmitted by setting the volume of the space 802 to zero. Further, the three-dimensional data creation device 810 may transmit information indicating that there is no update in the space 802 to the traffic monitoring cloud or the following vehicle.
[0630] As described above, for example, when there is no obstacle on the road, there is no difference between the generated three-dimensional data 835 and the three-dimensional map on the infrastructure side, and data transmission is not performed. In this way, transmission of unnecessary data can be suppressed.
[0631] In the above description, an example in which the three-dimensional data creation device 810 is mounted on a vehicle has been described. However, the three-dimensional data creation device 810 is not limited to a vehicle and may be mounted on any moving body.
[0632] As described above, the three-dimensional data creation device 810 according to the present embodiment is mounted on a moving body including a sensor 815 and a communication unit (such as a data receiving unit 811 or a data transmitting unit 822) that transmits and receives three-dimensional data to and from the outside. The three-dimensional data creation device 810 creates three-dimensional data 835 (second three-dimensional data) based on the sensor information 833 detected by the sensor 815 and the three-dimensional data 831 (first three-dimensional data) received by the data receiving unit 811. The three-dimensional data creation device 810 transmits three-dimensional data 837, which is a part of the three-dimensional data 835, to the outside.
[0633] As a result, the three-dimensional data creation device 810 can generate three-dimensional data in a range that cannot be detected by the host vehicle. Also, the three-dimensional data creation device 810 can transmit the three-dimensional data in a range that cannot be detected by other vehicles or the like to the other vehicles.
[0634] Also, the three-dimensional data creation device 810 repeatedly creates the three-dimensional data 835 and transmits the three-dimensional data 837 at a predetermined interval. The three-dimensional data 837 is the three-dimensional data of a small space 802 having a predetermined size located at a predetermined distance L in the forward direction of the moving direction of the vehicle 801 from the position of the current vehicle 801.
[0635] As a result, since the range of the transmitted three-dimensional data 837 is limited, the data amount of the transmitted three-dimensional data 837 can be reduced.
[0636] Also, the predetermined distance L changes according to the moving speed V of the vehicle 801. For example, as the moving speed V increases, the predetermined distance L becomes longer. As a result, the vehicle 801 can set an appropriate small space 802 according to the moving speed V of the vehicle 801 and transmit the three-dimensional data 837 of the small space 802 to a following vehicle or the like.
[0637] Also, the predetermined size changes according to the moving speed V of the vehicle 801. For example, as the moving speed V increases, the predetermined size becomes larger. For example, as the moving speed V increases, the depth D, which is the length of the small space 802 in the moving direction of the vehicle, becomes larger. As a result, the vehicle 801 can set an appropriate small space 802 according to the moving speed V of the vehicle 801 and transmit the three-dimensional data 837 of the small space 802 to a following vehicle or the like.
[0638] Also, the three-dimensional data creation device 810 determines whether there is a change in the three-dimensional data 835 of the small space 802 corresponding to the transmitted three-dimensional data 837. When the three-dimensional data creation device 810 determines that there is a change, it transmits the three-dimensional data 837 (fourth three-dimensional data), which is at least a part of the changed three-dimensional data 835, to an external following vehicle or the like.
[0639] As a result, the vehicle 801 can transmit the three-dimensional data 837 of the changed space to a following vehicle or the like.
[0640] In addition, the three-dimensional data creation device 810 transmits the changed three-dimensional data 837 (fourth three-dimensional data) with priority over the normal three-dimensional data 837 (third three-dimensional data) that is transmitted periodically. Specifically, the three-dimensional data creation device 810 transmits the changed three-dimensional data 837 (fourth three-dimensional data) before transmitting the normal three-dimensional data 837 (third three-dimensional data) that is transmitted periodically. That is, the three-dimensional data creation device 810 transmits the changed three-dimensional data 837 (fourth three-dimensional data) non-periodically without waiting for the transmission of the normal three-dimensional data 837 that is transmitted periodically.
[0641] As a result, the vehicle 801 can preferentially transmit the three-dimensional data 837 of the changed space to a following vehicle or the like, so that the following vehicle or the like can quickly make a determination based on the three-dimensional data.
[0642] In addition, the changed three-dimensional data 837 (fourth three-dimensional data) indicates the difference between the three-dimensional data 835 of the small space 802 corresponding to the transmitted three-dimensional data 837 and the three-dimensional data 835 after the change. Thereby, the data amount of the three-dimensional data 837 to be transmitted can be reduced.
[0643] In addition, when there is no difference between the three-dimensional data 837 of the small space 802 and the three-dimensional data 831 of the small space 802, the three-dimensional data creation device 810 does not transmit the three-dimensional data 837 of the small space 802. Further, the three-dimensional data creation device 810 may transmit information indicating that there is no difference between the three-dimensional data 837 of the small space 802 and the three-dimensional data 831 of the small space 802 to the outside.
[0644] Thereby, it is possible to suppress the transmission of unnecessary three-dimensional data 837, so that the data amount of the three-dimensional data 837 to be transmitted can be reduced.
[0645] (Embodiment 6) In the present embodiment, a display device and a display method for displaying information obtained from a three-dimensional map or the like, and a storage device and a storage method for a three-dimensional map or the like will be described.
[0646] For a moving body such as a vehicle or a robot, for the purpose of automatic driving of the vehicle or autonomous movement of the robot, a three-dimensional map obtained by communication with a server or another vehicle and two-dimensional video or vehicle detection three-dimensional data obtained from sensors mounted on the own vehicle are utilized. It is considered that the data that the user wants to view or save among these data varies depending on the situation. Hereinafter, a display device that switches the display according to the situation will be described.
[0647] FIG. 41 is a flowchart showing an outline of a display method by a display device. The display device is mounted on a moving body such as a vehicle or a robot. Hereinafter, an example in which the moving body is a vehicle (automobile) will be described.
[0648] First, the display device determines whether to display two-dimensional surrounding information or three-dimensional surrounding information according to the driving situation of the vehicle (S901). Here, the two-dimensional surrounding information corresponds to the first surrounding information in the claims, and the three-dimensional surrounding information corresponds to the second surrounding information in the claims. Here, the surrounding information is information indicating the surroundings of the moving body. For example, it is a video seen from a predetermined direction from the vehicle or a map of the surroundings of the vehicle.
[0649] The two-dimensional surrounding information is information generated using two-dimensional data. Here, the two-dimensional data is two-dimensional map information or video. For example, the two-dimensional surrounding information is a map of the surroundings of the vehicle obtained from a two-dimensional map or a video obtained by a camera mounted on the vehicle. Also, the two-dimensional surrounding information does not include three-dimensional information, for example. That is, when the two-dimensional surrounding information is a map of the surroundings of the vehicle, the map does not include information in the height direction. Also, when the two-dimensional surrounding information is a video obtained by a camera, the video does not include information in the depth direction.
[0650] Furthermore, the three-dimensional surrounding information is information generated using three-dimensional data. Here, the three-dimensional data is, for example, a three-dimensional map. Note that the three-dimensional data may be information indicating the three-dimensional position or three-dimensional shape of an object around the vehicle, obtained from another vehicle or a server, or detected by the host vehicle. For example, the three-dimensional surrounding information is a two-dimensional or three-dimensional video or map around the vehicle generated using a three-dimensional map. Also, the three-dimensional surrounding information includes, for example, three-dimensional information. For example, when the three-dimensional surrounding information is a video in front of the vehicle, the video includes information indicating the distance to an object in the video. Or, in the video, for example, a pedestrian or the like existing behind a vehicle in front is displayed. Also, the three-dimensional surrounding information may be a video obtained from a sensor mounted on the vehicle, with information indicating these distances or pedestrians or the like superimposed thereon. Also, the three-dimensional surrounding information may be a two-dimensional map with height-direction information superimposed thereon.
[0651] Also, the three-dimensional data may be three-dimensionally displayed, or a two-dimensional video or two-dimensional map obtained from the three-dimensional data may be displayed on a two-dimensional display or the like.
[0652] When it is determined in step S901 to display the three-dimensional surrounding information (Yes in S902), the display device displays the three-dimensional surrounding information (S903). On the other hand, when it is determined in step S901 to display the two-dimensional surrounding information (No in S902), the display device displays the two-dimensional surrounding information (S904). In this way, the display device displays the three-dimensional surrounding information or the two-dimensional surrounding information determined to be displayed in step S901.
[0653] Hereinafter, specific examples will be described. In the first example, the display device switches the surrounding information to be displayed according to whether the vehicle is being automatically driven or manually driven. Specifically, during automatic driving, since the driver does not need to know in detail the detailed surrounding road information, the display device displays two-dimensional surrounding information (for example, a two-dimensional map). On the other hand, during manual driving, three-dimensional surrounding information (for example, a three-dimensional map) is displayed so that the driver can know the details of the surrounding road information for safe driving.
[0654] Also, during autonomous driving, in order to show the user based on what information the host vehicle is driving, the display device may display information that has affected the driving operation (for example, SWLD used for self-position estimation, lanes, road signs, and surrounding situation detection results, etc.). For example, the display device may display these information in addition to a two-dimensional map.
[0655] Note that the surrounding information displayed during the above-mentioned autonomous driving and manual driving is just an example. The display device may display three-dimensional surrounding information during autonomous driving and two-dimensional surrounding information during manual driving. Also, the display device may display metadata or surrounding situation detection results in addition to a two-dimensional or three-dimensional map or video during at least one of autonomous driving and manual driving, or may display metadata or surrounding situation detection results instead of a two-dimensional or three-dimensional map or video. Here, the metadata is information indicating the three-dimensional position or three-dimensional shape of an object obtained from a server or another vehicle. Also, the surrounding situation detection result is information indicating the three-dimensional position or three-dimensional shape of an object detected by the host vehicle.
[0656] In the second example, the display device switches the surrounding information to be displayed according to the driving environment. For example, the display device switches the surrounding information to be displayed according to the brightness of the external environment. Specifically, when the surroundings of the host vehicle are bright, the display device displays a two-dimensional video obtained by a camera mounted on the host vehicle, or three-dimensional surrounding information created using the two-dimensional video. On the other hand, when the surroundings of the host vehicle are dark, since the two-dimensional video obtained from the camera mounted on the host vehicle is dark and difficult to view, the display device displays three-dimensional surrounding information created using a lidar or a millimeter-wave radar.
[0657] In addition, the display device may switch the surrounding information to be displayed according to the driving area, which is the area where the current host vehicle is located. For example, in a tourist destination, the city center, or near the destination, the display device may display three-dimensional surrounding information so as to provide the user with information about surrounding buildings. On the other hand, in mountainous areas or the suburbs, since there are many cases where detailed surrounding information is not required, the display device may display two-dimensional surrounding information.
[0658] In addition, the display device may switch the surrounding information to be displayed based on the weather condition. For example, in sunny weather, the display device may display three-dimensional surrounding information created using a camera or a lidar. On the other hand, in rainy or foggy weather, since noise is likely to be included in the three-dimensional surrounding information obtained by a camera or a lidar, the display device may display three-dimensional surrounding information created using a millimeter-wave radar.
[0659] In addition, these display switches may be automatically performed by the system or manually performed by the user.
[0660] In addition, the three-dimensional surrounding information is generated from any one or more of the following data: dense point cloud data generated based on the WLD, mesh data generated based on the MWLD, sparse data generated based on the SWLD, lane data generated based on the lane world, two-dimensional map data including three-dimensional shape information of roads and intersections, and metadata or host vehicle detection results including three-dimensional position or three-dimensional shape information that changes in real time.
[0661] Note that as described above, the WLD is three-dimensional point cloud data, the SWLD is data obtained by extracting point clouds with a feature amount equal to or greater than a threshold from the WLD. In addition, the MWLD is data having a mesh structure generated from the WLD. The lane world is data obtained by extracting point clouds with a feature amount equal to or greater than a threshold from the WLD and that are necessary for self-position estimation, driving assistance, or autonomous driving.
[0662] Here, MWLD and SWLD have less data volume compared to WLD. Therefore, when more detailed data is required, WLD can be used, and when this is not the case, MWLD or SWLD can be used, thereby appropriately reducing the communication data volume and the processing volume. Also, the lane world has less data volume compared to SWLD. Therefore, by using the lane world, the communication data volume and the processing volume can be further reduced.
[0663] Also, in the above, an example of switching between two-dimensional peripheral information and three-dimensional peripheral information was described, but the display device may switch the type of data (such as WLD, SWLD, etc.) used for generating the three-dimensional peripheral information based on the above conditions. That is, in the above description, in the case where the display device displays three-dimensional peripheral information, it displays three-dimensional peripheral information generated from first data (for example, WLD or SWLD) with a larger data volume, and in the case where it displays two-dimensional peripheral information, instead of the two-dimensional peripheral information, it may display three-dimensional peripheral information generated from second data (for example, SWLD or lane world) with a smaller data volume than the first data.
[0664] Also, the display device displays two-dimensional peripheral information or three-dimensional peripheral information on, for example, a two-dimensional display, a head-up display, or a head-mounted display mounted on the own vehicle. Alternatively, the display device may transfer and display two-dimensional peripheral information or three-dimensional peripheral information to a mobile terminal such as a smartphone by wireless communication. That is, the display device is not limited to being mounted on a moving body, and may be any device that cooperates with and operates with a moving body. For example, when a user having a display device such as a smartphone boards a moving body or drives a moving body, information of the moving body such as the position of the moving body based on the self-position estimation of the moving body is displayed on the display device, or these information are displayed on the display device together with the peripheral information.
[0665] Also, when the display device displays a three-dimensional map, it may render the three-dimensional map and display it as two-dimensional data, or display it as three-dimensional data using a three-dimensional display or a three-dimensional hologram.
[0666] Next, a method for storing a three-dimensional map will be described. A moving object such as a vehicle or a robot utilizes a three-dimensional map obtained through communication with a server or another vehicle, a two-dimensional video obtained from a sensor mounted on the own vehicle, or own vehicle detection three-dimensional data for automatic driving of the vehicle or autonomous movement of the robot. It is considered that the data that the user wants to view or store among these data varies depending on the situation. Below, a method for storing data according to the situation will be described.
[0667] The storage device is mounted on a moving object such as a vehicle or a robot. Hereinafter, an example in which the moving object is a vehicle (automobile) will be described. Also, the storage device may be included in the above-described display device.
[0668] In the first example, the storage device determines whether to store a three-dimensional map based on the region. Here, by storing the three-dimensional map in the storage medium of the own vehicle, automatic driving becomes possible without communication with the server within the stored space. However, since there is a limit to the storage capacity, only limited data can be stored. Therefore, the storage device limits the region to be stored as shown below.
[0669] For example, the storage device preferentially stores a three-dimensional map of a region frequently passed through such as a commuting route or the vicinity of the home. Thereby, it is not necessary to acquire the data of the region frequently used each time, so that the communication data volume can be effectively reduced. Note that preferentially storing means storing data with a higher priority within a predetermined storage capacity. For example, when new data cannot be stored within the storage capacity, data with a lower priority than the new data is deleted.
[0670] Alternatively, the storage device preferentially stores a three-dimensional map of a region with a poor communication environment. Thereby, in a region with a poor communication environment, it is not necessary to acquire data through communication, so that the occurrence of a case where a three-dimensional map cannot be acquired due to poor communication can be suppressed.
[0671] Alternatively, the storage device preferentially stores the three-dimensional map of areas with high traffic volume. Thereby, the three-dimensional map of areas with a high accident rate can be preferentially stored. Thus, in such areas, it is possible to suppress the reduction in the accuracy of autonomous driving or driving support due to poor communication preventing the acquisition of the three-dimensional map.
[0672] Alternatively, the storage device preferentially stores the three-dimensional map of areas with low traffic volume. Here, in areas with low traffic volume, the likelihood of not being able to use the autonomous driving mode of automatically following the vehicle ahead increases. Thereby, there are cases where more detailed surrounding information is required. Thus, by preferentially storing the three-dimensional map of areas with low traffic volume, it is possible to improve the accuracy of autonomous driving or driving support in such areas.
[0673] Note that a combination of the above-described multiple storage methods may be used. Also, the areas where these three-dimensional maps are preferentially stored may be automatically determined by the system or specified by the user.
[0674] Further, the storage device may delete the three-dimensional map after a predetermined period has passed since storage, or may update it with the latest data. Thereby, it is possible to prevent old map data from being used. Also, when updating the map data, the storage device detects a difference area, which is a spatial area with a difference, by comparing the old map and the new map, and adds the data of the difference area of the new map to the old map, or removes the data of the difference area from the old map, so that only the data of the changed area may be updated.
[0675] Also, in this example, the stored three-dimensional map is used for autonomous driving. Thus, by using SWLD as this three-dimensional map, it is possible to reduce the communication data volume and the like. Note that the three-dimensional map is not limited to SWLD and may be other types of data such as WLD.
[0676] In the second example, the storage device stores the three-dimensional map based on an event.
[0677] For example, the storage device stores a special event encountered during driving as a three-dimensional map. This enables the user to view the details of the event later, etc. Examples of events stored as three-dimensional maps are shown below. Note that the storage device may store three-dimensional surrounding information generated from the three-dimensional map.
[0678] For example, the storage device stores a three-dimensional map before and after a collision accident or when danger is detected.
[0679] Alternatively, the storage device stores a three-dimensional map of a characteristic scene such as a beautiful scenery, a crowded place, or a tourist destination.
[0680] These events to be stored may be automatically determined by the system or specified by the user in advance. For example, machine learning may be used as a method for determining these events.
[0681] Also, in this example, the stored three-dimensional map is used for viewing. Therefore, by using the WLD as this three-dimensional map, high-quality video can be provided. Note that the three-dimensional map is not limited to the WLD and may be other types of data such as the SWLD.
[0682] Hereinafter, a method for the display device to control the display according to the user will be described. When the display device superimposes and displays the surrounding situation detection result obtained by vehicle-to-vehicle communication on the map, the surrounding vehicles are represented by wireframes or transparency is given to the surrounding vehicles so that the detection objects behind the surrounding vehicles can be seen. Alternatively, the display device may display a video from an overhead viewpoint so that the host vehicle, surrounding vehicles, and the surrounding situation detection result can be seen from above.
[0683] When using a head-up display to superimpose the surrounding situation detection results or point cloud data on the surrounding environment seen through the windshield as shown in Fig. 42, the position at which the information is superimposed may shift depending on the user's posture, body type, or eye position. Fig. 43 is a diagram showing an example of the head-up display when the position is shifted.
[0684] To correct such misalignment, the display device detects the user's posture, body type, or eye position using information from an in-car camera or a sensor mounted on the seat. The display device adjusts the position at which information is superimposed according to the detected posture, body type, or eye position of the user. Figure 44 is a diagram showing an example of the display of the head-up display after adjustment.
[0685] The user may manually adjust the superimposition position using a control device mounted on the vehicle.
[0686] The display device may display safe locations on a map in the event of a disaster and present the location to the user. Alternatively, the vehicle may inform the user of the details of the disaster and that the user is heading to a safe location, and automatically drive to the safe location.
[0687] For example, the vehicle may set its destination to an area with a high altitude above sea level so as not to be swept away by a tsunami when an earthquake occurs. At this time, the vehicle may acquire information on roads that have become impassable due to the earthquake by communicating with a server, and may perform processing according to the nature of the disaster, such as taking a route that avoids those roads.
[0688] In addition, the autonomous driving may include multiple modes, such as a travel mode and a drive mode.
[0689] In the travel mode, the vehicle determines a route to the destination, taking into consideration factors such as fast arrival time, low fare, short driving distance, and low energy consumption, and then drives automatically along the determined route.
[0690] In the drive mode, the vehicle automatically determines a route to reach the destination at the time specified by the user. For example, when the user sets the destination and the arrival time, the vehicle determines a route that can reach the destination at the set time by going around the surrounding tourist attractions and performs automatic driving according to the determined route.
[0691] (Embodiment 7) In Embodiment 5, an example in which a client device such as a vehicle transmits three-dimensional data to a server such as another vehicle or a traffic monitoring cloud has been described. In this embodiment, the client device transmits sensor information obtained by a sensor to a server or another client device.
[0692] First, the configuration of the system according to this embodiment will be described. FIG. 45 is a diagram showing the configuration of a three-dimensional map and a sensor information transmission / reception system according to this embodiment. This system includes a server 901 and client devices 902A and 902B. When the client devices 902A and 902B are not particularly distinguished, they are also referred to as the client device 902.
[0693] The client device 902 is, for example, in-vehicle equipment mounted on a moving body such as a vehicle. The server 901 is, for example, a traffic monitoring cloud or the like and can communicate with a plurality of client devices 902.
[0694] The server 901 transmits a three-dimensional map composed of point clouds to the client device 902. Note that the configuration of the three-dimensional map is not limited to point clouds and may represent other three-dimensional data such as a mesh structure.
[0695] The client device 902 transmits the sensor information acquired by the client device 902 to the server 901. The sensor information includes, for example, at least one of LIDAR acquisition information, visible light image, infrared image, depth image, sensor position information, and speed information.
[0696] Data transmitted and received between the server 901 and the client device 902 may be compressed for data reduction, or may remain uncompressed to maintain data accuracy. When compressing data, for example, a three-dimensional compression method based on an octree structure can be used for the point cloud. Also, a two-dimensional image compression method can be used for visible light images, infrared images, and depth images. Examples of the two-dimensional image compression method include MPEG-4 AVC or HEVC standardized by MPEG.
[0697] In addition, the server 901 transmits the three-dimensional map managed by the server 901 to the client device 902 in response to a transmission request for the three-dimensional map from the client device 902. Note that the server 901 may transmit the three-dimensional map without waiting for a transmission request for the three-dimensional map from the client device 902. For example, the server 901 may broadcast the three-dimensional map to one or more client devices 902 located in a predetermined space. Also, the server 901 may transmit a three-dimensional map suitable for the position of the client device 902 to the client device 902 that has received a transmission request at regular intervals. Further, the server 901 may transmit the three-dimensional map to the client device 902 each time the three-dimensional map managed by the server 901 is updated.
[0698] The client device 902 issues a transmission request for the three-dimensional map to the server 901. For example, when the client device 902 wants to estimate its own position during driving, the client device 902 transmits a transmission request for the three-dimensional map to the server 901.
[0699] Note that in the following cases, the client device 902 may issue a transmission request for the three-dimensional map to the server 901. When the three-dimensional map held by the client device 902 is old, the client device 902 may issue a transmission request for the three-dimensional map to the server 901. For example, when a certain period has elapsed since the client device 902 acquired the three-dimensional map, the client device 902 may issue a transmission request for the three-dimensional map to the server 901.
[0700] Before a certain time when the client device 902 exits from the space shown by the three-dimensional map held by the client device 902, the client device 902 may send a transmission request for the three-dimensional map to the server 901. For example, when the client device 902 exists within a predetermined distance from the boundary of the space shown by the three-dimensional map held by the client device 902, the client device 902 may send a transmission request for the three-dimensional map to the server 901. Also, when the movement path and movement speed of the client device 902 can be grasped, based on these, the time when the client device 902 exits from the space shown by the three-dimensional map held by the client device 902 may be predicted.
[0701] When the error at the time of alignment between the three-dimensional data created by the client device 902 from the sensor information and the three-dimensional map is a certain value or more, the client device 902 may send a transmission request for the three-dimensional map to the server 901.
[0702] The client device 902 transmits the sensor information to the server 901 in response to the transmission request for the sensor information sent from the server 901. Note that the client device 902 may send the sensor information to the server 901 without waiting for the transmission request for the sensor information from the server 901. For example, when the client device 902 once obtains a transmission request for the sensor information from the server 901, the client device 902 may periodically transmit the sensor information to the server 901 for a certain period. Also, when the error at the time of alignment between the three-dimensional data created by the client device 902 based on the sensor information and the three-dimensional map obtained from the server 901 is a certain value or more, the client device 902 determines that there may be a change in the three-dimensional map around the client device 902, and may send that fact and the sensor information to the server 901.
[0703] Server 901 sends a request to the client device 902 to transmit sensor information. For example, server 901 receives the location information of client device 902 such as GPS from client device 902. If server 901 determines based on the location information of client device 902 that client device 902 is approaching a space with less information in the three-dimensional map managed by server 901, server 901 sends a request to client device 902 to transmit sensor information in order to generate a new three-dimensional map. Also, when server 901 wants to update the three-dimensional map, when it wants to check the road conditions such as during snow accumulation or disasters, when it wants to check the traffic congestion situation, or the accident situation, etc., server 901 may send a request to transmit sensor information.
[0704] Also, client device 902 may set the data volume of the sensor information to be transmitted to server 901 according to the communication state or bandwidth at the time of receiving the request to transmit sensor information received from server 901. Setting the data volume of the sensor information to be transmitted to server 901 means, for example, increasing or decreasing the data itself, or appropriately selecting a compression method.
[0705] FIG. 46 is a block diagram showing a configuration example of client device 902. Client device 902 receives a three-dimensional map composed of a point cloud or the like from server 901, and estimates its own position from the three-dimensional data created based on the sensor information of client device 902. Also, client device 902 transmits the acquired sensor information to server 901.
[0706] Client device 902 includes a data receiving unit 1011, a communication unit 1012, a reception control unit 1013, a format conversion unit 1014, a plurality of sensors 1015, a three-dimensional data creation unit 1016, a three-dimensional image processing unit 1017, a three-dimensional data storage unit 1018, a format conversion unit 1019, a communication unit 1020, a transmission control unit 1021, and a data transmission unit 1022.
[0707] The data reception unit 1011 receives the three-dimensional map 1031 from the server 901. The three-dimensional map 1031 is data including point clouds such as WLD or SWLD. The three-dimensional map 1031 may include either compressed data or uncompressed data.
[0708] The communication unit 1012 communicates with the server 901 and transmits a data transmission request (for example, a request to transmit a three-dimensional map) to the server 901.
[0709] The reception control unit 1013 exchanges information such as a corresponding format with the communication destination via the communication unit 1012 and establishes communication with the communication destination.
[0710] The format conversion unit 1014 generates a three-dimensional map 1032 by performing format conversion or the like on the three-dimensional map 1031 received by the data reception unit 1011. Further, when the three-dimensional map 1031 is compressed or encoded, the format conversion unit 1014 performs decompression or decoding processing. Note that when the three-dimensional map 1031 is uncompressed data, the format conversion unit 1014 does not perform decompression or decoding processing.
[0711] The plurality of sensors 1015 is a group of sensors that acquire information outside the vehicle on which the client device 902 is mounted, such as a LIDAR, a visible light camera, an infrared camera, or a depth sensor, and generates sensor information 1033. For example, when the sensor 1015 is a laser sensor such as a LIDAR, the sensor information 1033 is three-dimensional data such as a point cloud (point group data). Note that the number of sensors 1015 may not be plural.
[0712] The three-dimensional data creation unit 1016 creates three-dimensional data 1034 around the host vehicle based on the sensor information 1033. For example, the three-dimensional data creation unit 1016 creates point cloud data with color information around the host vehicle using the information acquired by the LIDAR and the visible light video obtained by the visible light camera.
[0713] The three-dimensional image processing unit 1017 performs self-vehicle position estimation processing and the like using the received three-dimensional map 1032 such as point cloud and the three-dimensional data 1034 of the surroundings of the self-vehicle generated from the sensor information 1033. Note that the three-dimensional image processing unit 1017 may create the three-dimensional data 1035 of the surroundings of the self-vehicle by synthesizing the three-dimensional map 1032 and the three-dimensional data 1034, and perform self-vehicle position estimation processing using the created three-dimensional data 1035.
[0714] The three-dimensional data storage unit 1018 stores the three-dimensional map 1032, the three-dimensional data 1034, the three-dimensional data 1035, and the like.
[0715] The format conversion unit 1019 generates the sensor information 1037 by converting the sensor information 1033 into a format that the receiving side can handle. Note that the format conversion unit 1019 may reduce the data amount by compressing or encoding the sensor information 1037. Also, the format conversion unit 1019 may omit the process when format conversion is not necessary. Further, the format conversion unit 1019 may control the data amount to be transmitted according to the specified transmission range.
[0716] The communication unit 1020 communicates with the server 901 and receives a data transmission request (a transmission request for sensor information) and the like from the server 901.
[0717] The transmission control unit 1021 exchanges information such as the corresponding format with the communication destination via the communication unit 1020 and establishes communication.
[0718] The data transmission unit 1022 transmits the sensor information 1037 to the server 901. The sensor information 1037 includes information acquired by a plurality of sensors 1015 such as information acquired by LIDAR, a luminance image acquired by a visible light camera, an infrared image acquired by an infrared camera, a depth image acquired by a depth sensor, sensor position information, and speed information.
[0719] Next, the configuration of the server 901 will be described. FIG. 47 is a block diagram showing a configuration example of the server 901. The server 901 receives sensor information transmitted from the client device 902 and creates three-dimensional data based on the received sensor information. The server 901 updates the three-dimensional map managed by the server 901 using the created three-dimensional data. In addition, the server 901 transmits the updated three-dimensional map to the client device 902 in response to a transmission request for the three-dimensional map from the client device 902.
[0720] The server 901 includes a data reception unit 1111, a communication unit 1112, a reception control unit 1113, a format conversion unit 1114, a three-dimensional data creation unit 1116, a three-dimensional data synthesis unit 1117, a three-dimensional data storage unit 1118, a format conversion unit 1119, a communication unit 1120, a transmission control unit 1121, and a data transmission unit 1122.
[0721] The data reception unit 1111 receives the sensor information 1037 from the client device 902. The sensor information 1037 includes, for example, information acquired by LIDAR, a luminance image acquired by a visible light camera, an infrared image acquired by an infrared camera, a depth image acquired by a depth sensor, sensor position information, and speed information.
[0722] The communication unit 1112 communicates with the client device 902 and transmits a data transmission request (for example, a transmission request for sensor information) to the client device 902.
[0723] The reception control unit 1113 exchanges information such as the corresponding format with the communication destination via the communication unit 1112 and establishes communication.
[0724] When the received sensor information 1037 is compressed or encoded, the format conversion unit 1114 generates the sensor information 1132 by performing decompression or decoding processing. Note that the format conversion unit 1114 does not perform decompression or decoding processing if the sensor information 1037 is uncompressed data.
[0725] The three-dimensional data creation unit 1116 creates three-dimensional data 1134 around the client device 902 based on the sensor information 1132. For example, the three-dimensional data creation unit 1116 creates point cloud data with color information around the client device 902 using the information acquired by the LIDAR and the visible light video obtained by the visible light camera.
[0726] The three-dimensional data synthesis unit 1117 updates the three-dimensional map 1135 by synthesizing the three-dimensional data 1134 created based on the sensor information 1132 into the three-dimensional map 1135 managed by the server 901.
[0727] The three-dimensional data storage unit 1118 stores the three-dimensional map 1135 and the like.
[0728] The format conversion unit 1119 generates the three-dimensional map 1031 by converting the three-dimensional map 1135 into a format supported by the receiving side. Note that the format conversion unit 1119 may reduce the data volume by compressing or encoding the three-dimensional map 1135. Also, the format conversion unit 1119 may omit the process when format conversion is not necessary. Further, the format conversion unit 1119 may control the data volume to be transmitted according to the specified transmission range.
[0729] The communication unit 1120 communicates with the client device 902 and receives a data transmission request (a transmission request for the three-dimensional map) and the like from the client device 902.
[0730] The transmission control unit 1121 exchanges information such as the corresponding format with the communication destination via the communication unit 1120 to establish communication.
[0731] The data transmission unit 1122 transmits the three-dimensional map 1031 to the client device 902. The three-dimensional map 1031 is data including a point cloud such as WLD or SWLD. The three-dimensional map 1031 may include either compressed data or uncompressed data.
[0732] Next, the operation flow of the client device 902 will be described. FIG. 48 is a flowchart showing the operation when the client device 902 acquires a three-dimensional map.
[0733] First, the client device 902 requests the server 901 to transmit a three-dimensional map (such as a point cloud) (S1001). At this time, the client device 902 may request the server 901 to transmit a three-dimensional map related to the position information by transmitting the position information of the client device 902 obtained by GPS or the like together.
[0734] Next, the client device 902 receives a three-dimensional map from the server 901 (S1002). If the received three-dimensional map is compressed data, the client device 902 decrypts the received three-dimensional map to generate an uncompressed three-dimensional map (S1003).
[0735] Next, the client device 902 creates three-dimensional data 1034 around the client device 902 from the sensor information 1033 obtained by the plurality of sensors 1015 (S1004). Next, the client device 902 estimates its own position using the three-dimensional map 1032 received from the server 901 and the three-dimensional data 1034 created from the sensor information 1033 (S1005).
[0736] FIG. 49 is a flowchart showing the operation when the client device 902 transmits sensor information. First, the client device 902 receives a sensor information transmission request from the server 901 (S1011). The client device 902 that has received the transmission request transmits the sensor information 1037 to the server 901 (S1012). Note that when the sensor information 1033 includes a plurality of pieces of information obtained by the plurality of sensors 1015, the client device 902 may generate the sensor information 1037 by compressing each piece of information using a compression method suitable for each piece of information.
[0737] Next, the operation flow of the server 901 will be described. FIG. 50 is a flowchart showing the operation when the server 901 acquires sensor information. First, the server 901 requests the client device 902 to transmit sensor information (S1021). Next, the server 901 receives the sensor information 1037 transmitted from the client device 902 in response to the request (S1022). Next, the server 901 creates three-dimensional data 1134 using the received sensor information 1037 (S1023). Next, the server 901 reflects the created three-dimensional data 1134 in the three-dimensional map 1135 (S1024).
[0738] FIG. 51 is a flowchart showing the operation when the server 901 transmits a three-dimensional map. First, the server 901 receives a request to transmit a three-dimensional map from the client device 902 (S1031). The server 901 that has received the request to transmit the three-dimensional map transmits the three-dimensional map 1031 to the client device 902 (S1032). At this time, the server 901 may extract the three-dimensional map in the vicinity according to the position information of the client device 902 and transmit the extracted three-dimensional map. Further, the server 901 may compress the three-dimensional map composed of the point cloud using, for example, a compression method based on an octree structure, and transmit the compressed three-dimensional map.
[0739] Hereinafter, a modification example of the present embodiment will be described.
[0740] Server 901 creates three-dimensional data 1134 near the location of client device 902 using the sensor information 1037 received from client device 902. Next, server 901 calculates the difference between the three-dimensional data 1134 and the three-dimensional map 1135 of the same area managed by server 901 by performing matching between the created three-dimensional data 1134 and the three-dimensional map 1135. When the difference is equal to or greater than a predetermined threshold, server 901 determines that some abnormality has occurred around client device 902. For example, a large difference may occur between the three-dimensional map 1135 managed by server 901 and the three-dimensional data 1134 created based on the sensor information 1037 when ground subsidence or the like occurs due to a natural disaster such as an earthquake.
[0741] The sensor information 1037 may include information indicating at least one of the type of the sensor, the performance of the sensor, and the model number of the sensor. Also, a class ID or the like corresponding to the performance of the sensor may be added to the sensor information 1037. For example, when the sensor information 1037 is information acquired by LIDAR, it is conceivable to assign an identifier to the performance of the sensor such that a sensor capable of acquiring information with an accuracy of several millimeters is class 1, a sensor capable of acquiring information with an accuracy of several centimeters is class 2, and a sensor capable of acquiring information with an accuracy of several meters is class 3. Also, the server 901 may estimate the performance information of the sensor, etc., from the model number of the client device 902. For example, when the client device 902 is mounted on a vehicle, the server 901 may determine the sensor specification information from the vehicle type of the vehicle. In this case, the server 901 may have acquired the information on the vehicle type of the vehicle in advance, or the information may be included in the sensor information. Also, the server 901 may use the acquired sensor information 1037 to switch the degree of correction for the three-dimensional data 1134 created using the sensor information 1037. For example, when the sensor performance is high accuracy (class 1), the server 901 does not correct the three-dimensional data 1134. When the sensor performance is low accuracy (class 3), the server 901 applies correction corresponding to the accuracy of the sensor to the three-dimensional data 1134. For example, the server 901 increases the degree (intensity) of correction as the accuracy of the sensor decreases.
[0742] The server 901 may simultaneously send a request to transmit sensor information to a plurality of client devices 902 in a certain space. When the server 901 receives a plurality of sensor information from the plurality of client devices 902, it is not necessary to use all the sensor information for creating the three-dimensional data 1134. For example, the server 901 may select the sensor information to be used according to the performance of the sensor. For example, when updating the three-dimensional map 1135, the server 901 may select high-accuracy sensor information (class 1) from among the plurality of received sensor information and create the three-dimensional data 1134 using the selected sensor information.
[0743] The server 901 is not limited to only servers such as traffic monitoring clouds, and may be other client devices (in-vehicle). FIG. 52 is a diagram of the system configuration in this case.
[0744] For example, the client device 902C sends a transmission request for sensor information to the client device 902A nearby and acquires the sensor information from the client device 902A. Then, the client device 902C creates three-dimensional data using the acquired sensor information of the client device 902A and updates the three-dimensional map of the client device 902C. As a result, the client device 902C can generate a three-dimensional map of the space that can be acquired from the client device 902A by making use of the performance of the client device 902C. For example, such a case is considered to occur when the performance of the client device 902C is high.
[0745] Also, in this case, the client device 902A that provided the sensor information is given the right to acquire the high-precision three-dimensional map generated by the client device 902C. The client device 902A receives the high-precision three-dimensional map from the client device 902C according to that right.
[0746] Also, the client device 902C may send a transmission request for sensor information to a plurality of nearby client devices 902 (client device 902A and client device 902B). When the sensors of the client device 902A or the client device 902B are high-performance, the client device 902C can create three-dimensional data using the sensor information obtained by this high-performance sensor.
[0747] FIG. 53 is a block diagram showing the functional configuration of the server 901 and the client device 902. The server 901 includes, for example, a three-dimensional map compression / decompression processing unit 1201 that compresses and decompresses a three-dimensional map, and a sensor information compression / decompression processing unit 1202 that compresses and decompresses sensor information.
[0748] The client device 902 includes a three-dimensional map decoding processing unit 1211 and a sensor information compression processing unit 1212. The three-dimensional map decoding processing unit 1211 receives the encoded data of the compressed three-dimensional map, decodes the encoded data, and acquires the three-dimensional map. The sensor information compression processing unit 1212 compresses the sensor information itself instead of the three-dimensional data created from the acquired sensor information, and transmits the encoded data of the compressed sensor information to the server 901. With this configuration, the client device 902 only needs to hold internally a processing unit (device or LSI) that performs the process of decoding a three-dimensional map (such as a point cloud), and does not need to hold internally a processing unit that performs the process of compressing the three-dimensional data of the three-dimensional map (such as a point cloud). Thereby, the cost and power consumption of the client device 902 can be suppressed.
[0749] As described above, the client device 902 according to the present embodiment is mounted on a moving body, and creates three-dimensional data 1034 around the moving body from sensor information 1033 indicating the surrounding situation of the moving body obtained by the sensor 1015 mounted on the moving body. The client device 902 estimates its own position of the moving body using the created three-dimensional data 1034. The client device 902 transmits the acquired sensor information 1033 to the server 901 or another moving body 902.
[0750] According to this, the client device 902 transmits the sensor information 1033 to the server 901 or the like. Thereby, there is a possibility that the amount of transmitted data can be reduced compared with the case of transmitting three-dimensional data. Further, since it is not necessary for the client device 902 to perform processes such as compression or encoding of three-dimensional data, the processing amount of the client device 902 can be reduced. Therefore, the client device 902 can achieve reduction of the amount of transmitted data or simplification of the device configuration.
[0751] Further, the client device 902 further sends a transmission request for the three-dimensional map to the server 901 and receives the three-dimensional map 1031 from the server 901. The client device 902 estimates its own position using the three-dimensional data 1034 and the three-dimensional map 1032 in the estimation of its own position.
[0752] Also, the sensor information 1033 includes at least one of information obtained by a laser sensor, a luminance image, an infrared image, a depth image, the position information of the sensor, and the speed information of the sensor.
[0753] Also, the sensor information 1033 includes information indicating the performance of the sensor.
[0754] Also, the client device 902 encodes or compresses the sensor information 1033, and in the transmission of the sensor information, transmits the encoded or compressed sensor information 1037 to the server 901 or another mobile body 902. According to this, the client device 902 can reduce the amount of data to be transmitted.
[0755] For example, the client device 902 includes a processor and a memory, and the processor performs the above processing using the memory.
[0756] Also, the server 901 according to the present embodiment is communicable with the client device 902 mounted on the mobile body, and receives the sensor information 1037 indicating the surrounding situation of the mobile body obtained by the sensor 1015 mounted on the mobile body from the client device 902. The server 901 creates three-dimensional data 1134 around the mobile body from the received sensor information 1037.
[0757] According to this, the server 901 creates three-dimensional data 1134 using the sensor information 1037 transmitted from the client device 902. Thereby, there is a possibility that the data volume of the transmitted data can be reduced as compared with the case where the client device 902 transmits three-dimensional data. Further, since it is not necessary for the client device 902 to perform processes such as compression or encoding of the three-dimensional data, the processing amount of the client device 902 can be reduced. Therefore, the server 901 can achieve reduction of the data volume to be transmitted or simplification of the device configuration.
[0758] Further, the server 901 further transmits a transmission request for sensor information to the client device 902.
[0759] Further, the server 901 further updates the three-dimensional map 1135 using the created three-dimensional data 1134, and transmits the three-dimensional map 1135 to the client device 902 in response to a transmission request for the three-dimensional map 1135 from the client device 902.
[0760] Further, the sensor information 1037 includes at least one of information obtained by a laser sensor, a luminance image, an infrared image, a depth image, sensor position information, and sensor speed information.
[0761] Further, the sensor information 1037 includes information indicating the performance of the sensor.
[0762] Further, the server 901 further corrects the three-dimensional data according to the performance of the sensor. According to this, the three-dimensional data creation method can improve the quality of the three-dimensional data.
[0763] Further, when receiving the sensor information, the server 901 receives a plurality of sensor information 1037 from a plurality of client devices 902, and selects the sensor information 1037 to be used for creating the three-dimensional data 1134 based on a plurality of information indicating the performance of the sensors included in the plurality of sensor information 1037. According to this, the server 901 can improve the quality of the three-dimensional data 1134.
[0764] In addition, the server 901 decrypts or decompresses the received sensor information 1037, and creates three-dimensional data 1134 from the decrypted or decompressed sensor information 1132. According to this, the server 901 can reduce the amount of data to be transmitted.
[0765] For example, the server 901 includes a processor and a memory, and the processor performs the above processing using the memory.
[0766] As described above, the server and the client device according to the embodiment of the present disclosure have been described. However, the present disclosure is not limited to this embodiment.
[0767] In addition, each processing unit included in the server and the client device according to the above embodiment is typically realized as an LSI which is an integrated circuit. These may be individually formed into one chip, or may be formed into one chip so as to include some or all of them.
[0768] In addition, the integration into an integrated circuit is not limited to an LSI, and may be realized by a dedicated circuit or a general-purpose processor. An FPGA (Field Programmable Gate Array) that can be programmed after manufacturing the LSI, or a reconfigurable processor that can reconfigure the connection and setting of circuit cells inside the LSI may be used.
[0769] In addition, in each of the above embodiments, each component may be configured by dedicated hardware, or may be realized by executing a software program suitable for each component. Each component may be realized by a program execution unit such as a CPU or a processor reading and executing a software program recorded on a recording medium such as a hard disk or a semiconductor memory.
[0770] In addition, the present disclosure may be realized as a three-dimensional data creation method or the like executed by a server and a client device or the like.
[0771] Also, the division of functional blocks in the block diagram is just an example. It is possible to implement multiple functional blocks as one functional block, divide one functional block into multiple ones, or transfer some functions to other functional blocks. Further, the functions of multiple functional blocks having similar functions may be processed by a single piece of hardware or software in parallel or in a time-sharing manner.
[0772] Also, the order in which each step in the flowchart is executed is for illustration purposes to specifically describe the present disclosure, and it may be in an order other than the above. Further, some of the above steps may be executed simultaneously (in parallel) with other steps.
[0773] As described above, the server and client device etc. according to one or more aspects have been described based on the embodiments. However, the present disclosure is not limited to these embodiments. As long as it does not depart from the spirit of the present disclosure, various modifications conceived by those skilled in the art applied to these embodiments or forms constructed by combining components in different embodiments may also be included within the scope of one or more aspects.
Industrial Applicability
[0774] The present disclosure can be applied to client devices and servers that create three-dimensional data.
Explanation of Reference Numerals
[0775] 100, 400 Three-dimensional data encoding device 101, 201, 401, 501 Acquisition unit 102, 402 Encoding area determination unit 103 Division unit 104, 644 Encoding unit 111, 607 Three-dimensional data 112, 211, 413, 414, 511, 634 Encoded three-dimensional data 200, 500 Three-dimensional data decoding device 202 Decoding start GOS determination unit 203 Decoding SPC determination unit 204, 625 Decryption Unit 212, 512, 513 Decrypted Three-Dimensional Data 403 SWLD Extraction Unit 404 WLD Encoding Unit 405 SWLD Encoding Unit 411 Input Three-Dimensional Data 412 Extracted Three-Dimensional Data 502 Header Analysis Unit 503 WLD Decryption Unit 504 SWLD Decryption Unit 600 Own Vehicle 601 Surrounding Vehicles 602, 605 Sensor Detection Range 603, 606 Regions 604 Occlusion Region 620, 620A Three-Dimensional Data Creation Device 621, 641 Three-Dimensional Data Creation Unit 622 Required Range Determination Unit 623 Search Unit 624, 642 Receiver 626 Synthesis Unit 627 Detection Region Determination Unit 628 Surrounding Situation Detection Unit 629 Autonomous Operation Control Unit 631, 651 Sensor Information 632 First Three-Dimensional Data 633 Required Range Information 635 Second Three-Dimensional Data 636 Third Three-Dimensional Data 637 Request Signal 638 Transmission Data 639 Surrounding Situation Detection Result 640, 640A Three-Dimensional Data Transmission Device 643 Extraction Unit 645 Transmitter 646 Trans...
Claims
1. 1. A method implemented by a client device equipped with a plurality of sensors, comprising: acquiring sensor data using the plurality of sensors; creating three-dimensional data including three-dimensional coordinate values of a plurality of points using the sensor data; transmitting sensor information including the sensor data to a server when a difference between the three-dimensional data and the three-dimensional map is greater than a predetermined threshold; method.
2. 2. The method of claim 1, further comprising: estimating a location of the client device using the three-dimensional data and the three-dimensional map; method.
3. 2. The method of claim 1 , The sensor information includes information indicating positions of the plurality of sensors. method.
4. 2. The method of claim 1 , The sensor information further includes information indicating an estimation accuracy of the positions of the plurality of points. method.
5. A client device, A plurality of sensors for acquiring sensor data; a processor for generating three-dimensional data including three-dimensional coordinate values of a plurality of points; transmitting sensor information including the sensor data to a server when a difference between the three-dimensional data and the three-dimensional map is greater than a predetermined threshold; Client device.
Citation Information
Patent Citations
Travel control system and travel control method
JP2013153280A
Information management system, vehicle and information management method
JP2016045609A
Map display device
WO2014020663A1