Information provision system, client device, and method for generating and updating three-dimensional maps.

The system efficiently transmits and processes three-dimensional data by dividing spatial areas into geographical units, addressing the challenge of high data volume in existing technologies and ensuring accurate data transmission for autonomous applications.

JP7897399B2Active Publication Date: 2026-07-29PANASONIC INTELLECTUAL PROPERTY CORP OF AMERICA
View PDF 5 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
PANASONIC INTELLECTUAL PROPERTY CORP OF AMERICA
Filing Date
2025-07-30
Publication Date
2026-07-29

AI Technical Summary

Technical Problem

Existing technologies face challenges in efficiently transmitting and processing large amounts of three-dimensional data, such as point clouds, for applications like autonomous vehicles and robotics, due to the high data volume and the need for effective compression and transmission methods.

Method used

An information transmission system that hierarchically divides spatial areas into geographical units, generating and transmitting reduced-size three-dimensional maps based on feature quantities, allowing for efficient data transmission and processing.

Benefits of technology

Enables effective transmission and processing of three-dimensional data, reducing data size while maintaining accuracy for applications like autonomous vehicle navigation and robotics.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007897399000001
    Figure 0007897399000001
  • Figure 0007897399000002
    Figure 0007897399000002
  • Figure 0007897399000003
    Figure 0007897399000003
Patent Text Reader

Abstract

To achieve further improvement.SOLUTION: An information providing system that provides a client device mounted on a movable body with a three-dimensional map indicating a situation in a three-dimensional space, includes: a map database that stores a first three-dimensional map in which a space is hierarchically divided on a geographical area basis; a data creator that creates a second three-dimensional map composed of space elements including feature quantities more than or equal to a predetermined threshold value among a plurality of space elements included in the first three-dimensional map and having a data size smaller than that of the first three-dimensional map; and a data transmitter that receives a request from the client device, specifies either the first three-dimensional map or the second three-dimensional map based on request information included in the request, and transmits the specified first three-dimensional map or second three-dimensional map as a three-dimensional map to the client device.SELECTED DRAWING: Figure 12
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to an information transmission method and a client device.

Background Art

[0002] In the future, the spread of devices or services that utilize three-dimensional data is expected in a wide range of fields such as computer vision, map information, monitoring, infrastructure inspection, or video distribution for autonomous operation of automobiles or robots. Three-dimensional data is acquired by various methods such as distance sensors such as lidar, stereo cameras, or combinations of multiple monocular cameras.

[0003] As one of the methods for representing three-dimensional data, there is a representation method called point cloud that represents the shape of a three-dimensional structure by a point group in a three-dimensional space. In a point cloud, the positions and colors of the point group are stored. Although the point cloud is expected to become mainstream as a method for representing three-dimensional data, the amount of data of the point group is very large. Therefore, in the accumulation or transmission of three-dimensional data, similar to two-dimensional moving images (for example, MPEG-4 AVC or HEVC standardized by MPEG), compression of the data amount by encoding is essential.

[0004] Also, regarding the compression of point clouds, it is partially supported by public libraries (Point Cloud Library) that perform point cloud-related processing.

[0005] Also, a technique for searching and displaying facilities located around a vehicle using three-dimensional map data is known (for example, see Patent Document 1).

Prior Art Documents

Patent Documents

[0006]

Patent Document 1

Summary of the Invention

[0007] It is desirable to properly transmit the information necessary to create such three-dimensional data.

[0008] This disclosure aims to provide an information transmission method or client device that can appropriately transmit information for creating three-dimensional data. [Means for solving the problem]

[0009] An information provision system according to one aspect of the present disclosure is an information provision system that provides a three-dimensional map showing the situation in a three-dimensional space to a client device mounted on a mobile body, comprising: a map database that holds a first three-dimensional map in which space is hierarchically divided into geographical area units; a data generation unit that generates a second three-dimensional map which is composed of spatial elements from among a plurality of spatial elements included in the first three-dimensional map that include feature quantities of a predetermined threshold or higher, and which has a smaller data size than the first three-dimensional map; and a data transmission unit that receives a request from the client device, identifies either the first three-dimensional map or the second three-dimensional map based on the request information included in the request, and transmits the identified first three-dimensional map or second three-dimensional map to the client device as the three-dimensional map. [Effects of the Invention]

[0010] This disclosure provides an information transmission method or client device that can appropriately transmit information for creating three-dimensional data. [Brief explanation of the drawing]

[0011] [Figure 1] Figure 1 is a diagram showing the structure of encoded three-dimensional data according to Embodiment 1. [Figure 2] Figure 2 shows an example of a prediction structure between SPCs belonging to the lowest layer of GOS according to Embodiment 1. [Figure 3] FIG. 3 is a diagram showing an example of a prediction structure between layers according to Embodiment 1. [Figure 4] FIG. 4 is a diagram showing an example of the encoding order of GOS according to Embodiment 1. [Figure 5] FIG. 5 is a diagram showing an example of the encoding order of GOS according to Embodiment 1. [Figure 6] FIG. 6 is a block diagram of a three-dimensional data encoding device according to Embodiment 1. [Figure 7] FIG. 7 is a flowchart of the encoding process according to Embodiment 1. [Figure 8] FIG. 8 is a block diagram of a three-dimensional data decoding device according to Embodiment 1. [Figure 9] FIG. 9 is a flowchart of the decoding process according to Embodiment 1. [Figure 10] FIG. 10 is a diagram showing an example of meta information according to Embodiment 1. [Figure 11] FIG. 11 is a diagram showing a configuration example of SWLD according to Embodiment 2. [Figure 12] FIG. 12 is a diagram showing an operation example of a server and a client according to Embodiment 2. [Figure 13] FIG. 13 is a diagram showing an operation example of a server and a client according to Embodiment 2. [Figure 14] FIG. 14 is a diagram showing an operation example of a server and a client according to Embodiment 2. [Figure 15] FIG. 15 is a diagram showing an operation example of a server and a client according to Embodiment 2. [Figure 16] FIG. 16 is a block diagram of a three-dimensional data encoding device according to Embodiment 2. [Figure 17] FIG. 17 is a flowchart of the encoding process according to Embodiment 2. [Figure 18] FIG. 18 is a block diagram of a three-dimensional data decoding device according to Embodiment 2. [Figure 19] FIG. 19 is a flowchart of the decoding process according to Embodiment 2. [Figure 20] FIG. 20 is a diagram showing a configuration example of a WLD according to Embodiment 2. [Figure 21] FIG. 21 is a diagram showing an example of an octree structure of a WLD according to Embodiment 2. [Figure 22] FIG. 22 is a diagram showing a configuration example of a SWLD according to Embodiment 2. [Figure 23] FIG. 23 is a diagram showing an example of an octree structure of a SWLD according to Embodiment 2. [Figure 24] FIG. 24 is a schematic diagram showing the state of transmission and reception of three-dimensional data between vehicles according to Embodiment 3. [Figure 25] FIG. 25 is a diagram showing an example of three-dimensional data transmitted between vehicles according to Embodiment 3. [Figure 26] FIG. 26 is a block diagram of a three-dimensional data creation device according to Embodiment 3. [Figure 27] FIG. 27 is a flowchart of a three-dimensional data creation process according to Embodiment 3. [Figure 28] FIG. 28 is a block diagram of a three-dimensional data transmission device according to Embodiment 3. [Figure 29] FIG. 29 is a flowchart of a three-dimensional data transmission process according to Embodiment 3. [Figure 30] FIG. 30 is a block diagram of a three-dimensional data creation device according to Embodiment 3. [Figure 31] FIG. 31 is a flowchart of a three-dimensional data creation process according to Embodiment 3. [Figure 32] FIG. 32 is a block diagram of a three-dimensional data transmission device according to Embodiment 3. [Figure 33] FIG. 33 is a flowchart of a three-dimensional data transmission process according to Embodiment 3. [Figure 34] FIG. 34 is a block diagram of a three-dimensional information processing device according to Embodiment 4. [Figure 35] FIG. 35 is a flowchart of a three-dimensional information processing method according to Embodiment 4. [Figure 36]Figure 36 is a flowchart of the three-dimensional information processing method according to Embodiment 4. [Figure 37] Figure 37 is a diagram illustrating the transmission process of three-dimensional data according to Embodiment 5. [Figure 38] Figure 38 is a block diagram of a three-dimensional data creation device according to Embodiment 5. [Figure 39] Figure 39 is a flowchart of the three-dimensional data creation method according to Embodiment 5. [Figure 40] Figure 40 is a flowchart of the three-dimensional data creation method according to Embodiment 5. [Figure 41] Figure 41 is a flowchart of the display method according to Embodiment 6. [Figure 42] Figure 42 is a diagram showing an example of the surrounding environment as seen through the windshield according to Embodiment 6. [Figure 43] Figure 43 is a diagram showing an example of the display of the head-up display according to Embodiment 6. [Figure 44] Figure 44 shows an example of the display of the adjusted head-up display according to Embodiment 6. [Figure 45] Figure 45 is a diagram showing the configuration of the system according to Embodiment 7. [Figure 46] Figure 46 is a block diagram of the client device according to Embodiment 7. [Figure 47] Figure 47 is a block diagram of the server according to Embodiment 7. [Figure 48] Figure 48 is a flowchart of the three-dimensional data creation process by the client device according to Embodiment 7. [Figure 49] Figure 49 is a flowchart of the sensor information transmission process by the client device according to Embodiment 7. [Figure 50] Figure 50 is a flowchart of the three-dimensional data creation process performed by the server according to Embodiment 7. [Figure 51] Figure 51 is a flowchart of the three-dimensional map transmission process by the server according to Embodiment 7. [Figure 52] Figure 52 shows a modified configuration of the system according to Embodiment 7. [Figure 53] Figure 53 is a diagram showing the configuration of the server and client device according to Embodiment 7. [Figure 54] Figure 54 is a diagram showing the configuration of the server and client device according to Embodiment 8. [Figure 55] Figure 55 is a flowchart of the processing performed by the client device according to Embodiment 8. [Figure 56] Figure 56 is a diagram showing the configuration of the sensor information collection system according to Embodiment 8. [Modes for carrying out the invention]

[0012] An information transmission method according to one aspect of the present disclosure is an information transmission method for a client device mounted on a mobile body, comprising: acquiring sensor information indicating the surrounding conditions of the mobile body obtained by a sensor mounted on the mobile body; storing the sensor information in a storage unit; determining whether the mobile body is in an environment where it can transmit the sensor information to a server; and, if it is determined that the mobile body is in an environment where it can transmit the sensor information to a server, transmitting the sensor information to the server.

[0013] According to this, the information transmission method can appropriately transmit information for creating three-dimensional data.

[0014] For example, the information transmission method may further create three-dimensional data of the surroundings of the moving object from the sensor information, and estimate the self-position of the moving object using the created three-dimensional data.

[0015] For example, the information transmission method may further send a request to the server to transmit a three-dimensional map, receive the three-dimensional map from the server, and estimate the self-position using the three-dimensional data and the three-dimensional map.

[0016] According to this, the information transmission method can improve the accuracy of self-localization.

[0017] For example, the sensor information may include at least one of the following: information obtained from a laser sensor, brightness images, infrared images, depth images, sensor position information, and sensor velocity information.

[0018] For example, the sensor information may include acquisition location information indicating the position of the moving object or the sensor at the time the sensor information was acquired by the sensor.

[0019] For example, the sensor information may include acquisition time information indicating the time when the sensor information was acquired by the sensor.

[0020] For example, the information transmission method may further obtain time information from the server and generate the obtained time information using the obtained time information.

[0021] According to this, the acquisition time information of sensor data transmitted from multiple client devices can be synchronized.

[0022] For example, the information transmission method may further receive a sensor information transmission request from the server that includes designation information specifying a location and time, and if the storage unit has stored sensor information obtained at the location and time indicated by the designation information, and determines that the mobile body is in an environment where it can transmit the sensor information to the server, it may transmit the sensor information obtained at the location and time indicated by the designation information to the server.

[0023] For example, the information transmission method may further delete the sensor information that has already been transmitted to the server from the storage unit.

[0024] According to this, the storage capacity can be reduced.

[0025] For example, the information transmission method may further delete the sensor information from the storage unit if the difference between the position of the moving body or the sensor when the sensor information was acquired by the sensor and the current position of the moving body or the sensor exceeds a predetermined distance.

[0026] According to this, the storage capacity can be reduced.

[0027] For example, the information transmission method may further delete the sensor information from the storage unit if the difference between the time the sensor information was acquired by the sensor and the current time exceeds a predetermined time.

[0028] According to this, the storage capacity can be reduced.

[0029] Furthermore, a client device according to one aspect of the present disclosure is a client device mounted on a mobile body, comprising a processor and a memory, wherein the processor uses the memory to acquire sensor information indicating the surrounding conditions of the mobile body obtained by a sensor mounted on the mobile body, stores the sensor information in a storage unit, determines whether the mobile body is in an environment where it can transmit the sensor information to a server, and if it determines that the mobile body is in an environment where it can transmit the sensor information to a server, transmits the sensor information to the server.

[0030] According to this, the client device can properly transmit information for creating three-dimensional data.

[0031] These comprehensive or specific embodiments may be implemented as a system, method, integrated circuit, computer program, or recording medium such as a computer-readable CD-ROM, or as any combination of a system, method, integrated circuit, computer program, and recording medium.

[0032] The embodiments will be described in detail below with reference to the drawings. Note that the embodiments described below are all specific examples of this disclosure. The numerical values, shapes, materials, components, arrangement and connection configurations of components, steps, and the order of steps shown in the following embodiments are examples only and are not intended to limit this disclosure. Furthermore, among the components in the following embodiments, those not described in the independent claim representing the highest-level concept will be described as optional components.

[0033] (Embodiment 1) First, the data structure of the encoded three-dimensional data (hereinafter also referred to as encoded data) according to this embodiment will be described. Figure 1 is a diagram showing the configuration of the encoded three-dimensional data according to this embodiment.

[0034] In this embodiment, the three-dimensional space is divided into spaces (SPCs) corresponding to pictures in video encoding, and three-dimensional data is encoded using these spaces as units. The spaces are further divided into volumes (VLMs) corresponding to macroblocks in video encoding, and prediction and transformation are performed using the VLMs as units. Each volume contains multiple voxels (VXLs), which are the smallest units to which position coordinates are associated. Prediction, similar to prediction performed on two-dimensional images, involves referencing other processing units to generate predicted three-dimensional data similar to the processing unit being processed, and then encoding the difference between this predicted three-dimensional data and the processing unit being processed. Furthermore, this prediction includes not only spatial prediction that references other prediction units at the same time, but also temporal prediction that references prediction units at different times.

[0035] For example, a three-dimensional data encoding device (hereinafter also referred to as the encoding device) encodes a three-dimensional space represented by point cloud data, such as a point cloud, by encoding each point in the point cloud, or multiple points contained within a voxel, depending on the size of the voxel. Subdividing the voxel allows for a highly accurate representation of the three-dimensional shape of the point cloud, while increasing the voxel size allows for a rougher representation of the three-dimensional shape of the point cloud.

[0036] In the following explanation, we will use the example of a point cloud as the 3D data, but the 3D data is not limited to a point cloud; any format of 3D data is acceptable.

[0037] Alternatively, a hierarchical structure of voxels may be used. In this case, for the nth-order hierarchy, it may be indicated sequentially whether or not sample points exist in the (n-1)th-order hierarchy and below (the lower layers of the nth-order hierarchy). For example, when decoding only the nth-order hierarchy, if sample points exist in the (n-1)th-order hierarchy and below, the sample points can be assumed to be at the center of the voxel of the nth-order hierarchy and decoded accordingly.

[0038] Furthermore, the encoding device acquires point cloud data using distance sensors, stereo cameras, monocular cameras, gyroscopes, or inertial sensors.

[0039] Spaces, like video encodings, are classified into at least three predictive structures, including intra-spaces (I-SPCs) that can be decoded independently, predictive spaces (P-SPCs) that allow only unidirectional referencing, and bidirectional spaces (B-SPCs) that allow bidirectional referencing. Furthermore, spaces contain two types of time information: the decoding time and the display time.

[0040] Furthermore, as shown in Figure 1, there is a processing unit called GOS (Group of Space), which is a random access unit, that contains multiple spaces. In addition, there is a processing unit called WLD (World), which contains multiple GOS.

[0041] The spatial area occupied by a world is associated with an absolute location on Earth using GPS or latitude and longitude information. This location information is stored as metadata. This metadata may be included in the encoded data or transmitted separately from the encoded data.

[0042] Furthermore, within a GOS, all SPCs may be adjacent in three dimensions, or there may be SPCs that are not adjacent in three dimensions to other SPCs.

[0043] In the following, the processing of three-dimensional data contained in processing units such as GOS, SPC, or VLM, including encoding, decoding, or referencing, will also be simply referred to as encoding, decoding, or referencing the processing unit. Furthermore, the three-dimensional data contained in the processing unit includes, for example, at least one pair of spatial position such as three-dimensional coordinates and characteristic values ​​such as color information.

[0044] Next, we will explain the prediction structure of SPCs in GOS. Multiple SPCs within the same GOS, or multiple VLMs within the same SPC, occupy different spaces from each other, but they have the same time information (decoded time and display time).

[0045] Furthermore, the SPC that is first in the decryption order within a GOS is the I-SPC. There are also two types of GOSs: closed GOS and open GOS. A closed GOS is one in which all SPCs within the GOS can be decrypted when decryption starts from the first I-SPC. In an open GOS, some SPCs whose displayed time is earlier than the first I-SPC refer to a different GOS, and decryption cannot be performed using only that GOS.

[0046] Furthermore, with encoded data such as map information, the WLD may be decoded in the reverse direction of the encoding order, and if there are dependencies between GOSs, reverse playback becomes difficult. Therefore, in such cases, a closed GOS is generally used.

[0047] Furthermore, GOS has a layered structure in the height direction, and encoding or decoding is performed sequentially from the SPC of the lower layer.

[0048] Figure 2 shows an example of the prediction structure between SPCs belonging to the lowest layer of GOS. Figure 3 shows an example of the prediction structure between layers.

[0049] One or more I-SPCs exist within a GOS. While objects such as people, animals, cars, bicycles, traffic lights, or landmark buildings exist in three-dimensional space, it is particularly effective to encode small objects as I-SPCs. For example, a three-dimensional data decoding device (hereinafter also referred to as the decoding device) decodes only the I-SPCs within the GOS when decoding a GOS with low processing load or at high speed.

[0050] Furthermore, the encoding device may switch the encoding interval or frequency of I-SPCs according to the density of objects in the WLD.

[0051] Furthermore, in the configuration shown in Figure 3, the encoding or decoding device encodes or decodes multiple layers sequentially from the bottom layer (Layer 1). This allows for prioritizing data near the ground, which contains more information, for applications such as autonomous vehicles.

[0052] Furthermore, in the case of encoded data used in drones and the like, encoding or decoding may be done sequentially within the GOS, starting from the SPC layer at the top in the height direction.

[0053] Furthermore, the encoding or decoding device may encode or decode multiple layers so that the decoding device can grasp the GOS roughly and gradually increase the resolution. For example, the encoding or decoding device may encode or decode layers 3, 8, 1, 9, and so on.

[0054] Next, we will explain how to handle static and dynamic objects.

[0055] In three-dimensional space, there are static objects or scenes such as buildings or roads (hereinafter collectively referred to as static objects) and dynamic objects such as cars or people (hereinafter referred to as dynamic objects). Object detection is performed separately, for example, by extracting feature points from point cloud data or camera images such as stereo cameras. Here, we will explain an example of an encoding method for dynamic objects.

[0056] The first method is to encode static and dynamic objects without distinguishing between them. The second method is to distinguish between static and dynamic objects using identification information.

[0057] For example, GOS is used as the identification unit. In this case, GOS containing SPCs that constitute static objects and GOS containing SPCs that constitute dynamic objects are distinguished by identification information stored within the encoded data or separately from the encoded data.

[0058] Alternatively, an SPC may be used as the identification unit. In this case, an SPC containing a VLM that constitutes a static object and an SPC containing a VLM that constitutes a dynamic object are distinguished by the above identification information.

[0059] Alternatively, VLM or VXL may be used as the identification unit. In this case, VLM or VXL containing static objects and VLM or VXL containing dynamic objects are distinguished by the above identification information.

[0060] Furthermore, the encoding device may encode dynamic objects as one or more VLMs or SPCs, and encode the VLM or SPC containing static objects and the SPC containing dynamic objects as different GOSs. Also, if the size of the GOS is variable depending on the size of the dynamic objects, the encoding device stores the size of the GOS separately as metadata.

[0061] Furthermore, the encoding device may encode static objects and dynamic objects independently of each other and superimpose dynamic objects onto a world composed of static objects. In this case, a dynamic object is composed of one or more SPCs, and each SPC is associated with one or more SPCs that constitute the static object on which it is superimposed. Note that dynamic objects may be represented by one or more VLMs or VXLs instead of SPCs.

[0062] Furthermore, the encoding device may encode static objects and dynamic objects as separate streams.

[0063] Furthermore, the encoding device may generate a GOS containing one or more SPCs that constitute a dynamic object. In addition, the encoding device may set the GOS containing the dynamic object (GOS_M) and the GOS of the static object corresponding to the spatial region of GOS_M to be the same size (occupy the same spatial region). This allows superposition processing to be performed on a GOS-by-GOS basis.

[0064] The P-SPC or B-SPC that constitute a dynamic object may reference SPCs contained in different encoded GOS. In cases where the position of a dynamic object changes over time and the same dynamic object is encoded as a GOS at different times, cross-GOS references are effective from a compression standpoint.

[0065] Furthermore, the first and second methods described above may be switched depending on the intended use of the encoded data. For example, when using encoded three-dimensional data as a map, it is desirable to be able to separate dynamic objects, so the encoding device uses the second method. On the other hand, when encoding three-dimensional data of an event such as a concert or sporting event, if there is no need to separate dynamic objects, the encoding device uses the first method.

[0066] Furthermore, the decoding time and display time of GOS or SPC can be stored within the encoded data or as metadata. The time information for static objects may also be identical. In this case, the actual decoding time and display time may be determined by the decoding device. Alternatively, different values ​​may be assigned to each GOS or SPC as the decoding time, while the same value may be assigned to all as the display time. Furthermore, a decoder model may be introduced, such as the HEVC HRD (Hypothetical Reference Decoder) in video encoding, which guarantees that decoding can be performed without failure if the decoder has a buffer of a predetermined size and reads the bitstream at a predetermined bitrate according to the decoding time.

[0067] Next, we will explain the arrangement of GOS within the world. The coordinates of the three-dimensional space in the world are represented by three mutually orthogonal coordinate axes (x-axis, y-axis, and z-axis). By establishing a predetermined rule for the coding order of GOS, coding can be performed so that spatially adjacent GOS are continuous within the coded data. For example, in the example shown in Figure 4, GOS in the xz plane are coded continuously. The value of the y-axis is updated after coding all GOS in a given xz plane is completed. That is, as coding progresses, the world expands in the y-axis direction. Also, the index numbers of the GOS are set in the coding order.

[0068] Here, the world's three-dimensional space is mapped one-to-one with geographical absolute coordinates such as GPS, latitude, and longitude. Alternatively, the three-dimensional space may be represented by relative positions from a pre-defined reference position. The directions of the x, y, and z axes of the three-dimensional space are represented as direction vectors determined based on latitude and longitude, and these direction vectors are stored as metadata along with encoded data.

[0069] Furthermore, the size of the GOS is fixed, and the encoding device stores this size as metadata. Alternatively, the size of the GOS may be switched depending on, for example, whether it is an urban area or not, or whether it is indoors or outdoors. In other words, the size of the GOS may be switched depending on the quantity or nature of objects that have informational value. Or, the encoding device may adaptively switch the size of the GOS or the spacing of I-SPCs within the GOS depending on the density of objects within the same world. For example, the encoding device may reduce the size of the GOS and shorten the spacing of I-SPCs within the GOS as the density of objects increases.

[0070] In the example in Figure 5, the GOS regions from the 3rd to the 10th are subdivided to enable fine-grained random access due to the high object density. Note that GOS regions 7 through 10 are located behind GOS regions 3 through 6, respectively.

[0071] Next, the configuration and operation flow of the three-dimensional data encoding device according to this embodiment will be described. Figure 6 is a block diagram of the three-dimensional data encoding device 100 according to this embodiment. Figure 7 is a flowchart showing an example of the operation of the three-dimensional data encoding device 100.

[0072] The three-dimensional data encoding device 100 shown in Figure 6 generates encoded three-dimensional data 112 by encoding three-dimensional data 111. This three-dimensional data encoding device 100 comprises an acquisition unit 101, an encoding region determination unit 102, a division unit 103, and an encoding unit 104.

[0073] As shown in Figure 7, first, the acquisition unit 101 acquires three-dimensional data 111, which is point cloud data (S101).

[0074] Next, the encoding region determination unit 102 determines the region to be encoded from among the spatial regions corresponding to the acquired point cloud data (S102). For example, the encoding region determination unit 102 determines the spatial region around the location of the user or vehicle as the region to be encoded.

[0075] Next, the division unit 103 divides the point cloud data included in the region to be encoded into processing units. Here, the processing units are the GOS and SPC mentioned above. The region to be encoded corresponds to, for example, the world mentioned above. Specifically, the division unit 103 divides the point cloud data into processing units based on a pre-set GOS size, or the presence or size of dynamic objects (S103). The division unit 103 also determines the starting position of the SPC that will be the first in the encoding order for each GOS.

[0076] Next, the encoding unit 104 generates encoded three-dimensional data 112 by sequentially encoding multiple SPCs within each GOS (S104).

[0077] Note that while this example shows the region to be encoded being divided into GOS and SPC before encoding each GOS, the processing procedure is not limited to the above. For example, one could determine the structure of one GOS, encode that GOS, and then determine the structure of the next GOS.

[0078] In this way, the three-dimensional data encoding device 100 generates encoded three-dimensional data 112 by encoding the three-dimensional data 111. Specifically, the three-dimensional data encoding device 100 divides the three-dimensional data into first processing units (GOS), which are random access units, each of which is associated with a three-dimensional coordinate. The first processing units (GOS) are then divided into a plurality of second processing units (SPCs), and the second processing units (SPCs) are then divided into a plurality of third processing units (VLMs). The third processing unit (VLM) also contains one or more voxels (VXLs), which are the smallest units to which positional information is associated.

[0079] Next, the three-dimensional data encoding device 100 generates encoded three-dimensional data 112 by encoding each of the multiple first processing units (GOS). Specifically, the three-dimensional data encoding device 100 encodes each of the multiple second processing units (SPC) in each first processing unit (GOS). Furthermore, the three-dimensional data encoding device 100 encodes each of the multiple third processing units (VLM) in each second processing unit (SPC).

[0080] For example, if the first processing unit (GOS) to be processed is a closed GOS, the three-dimensional data encoding device 100 encodes the second processing unit (SPC) included in the first processing unit (GOS) by referring to other second processing units (SPC) included in the first processing unit (GOS). In other words, the three-dimensional data encoding device 100 does not refer to second processing units (SPC) included in a first processing unit (GOS) different from the first processing unit (GOS) to be processed.

[0081] On the other hand, if the first processing unit (GOS) to be processed is an open GOS, the second processing unit (SPC) included in the first processing unit (GOS) to be processed is encoded by referring to another second processing unit (SPC) included in the first processing unit (GOS) to be processed, or to a second processing unit (SPC) included in a first processing unit (GOS) different from the first processing unit (GOS) to be processed.

[0082] Furthermore, the three-dimensional data encoding device 100 selects one of the following types of second processing units (SPCs) to be processed: a first type (I-SPC) that does not refer to any other second processing units (SPCs), a second type (P-SPC) that refers to one other second processing unit (SPC), and a third type that refers to two other second processing units (SPCs). The device then encodes the second processing unit (SPC) to be processed according to the selected type.

[0083] Next, the configuration and operation flow of the three-dimensional data decoding device according to this embodiment will be described. Figure 8 is a block diagram of the three-dimensional data decoding device 200 according to this embodiment. Figure 9 is a flowchart showing an example of the operation of the three-dimensional data decoding device 200.

[0084] The three-dimensional data decoding device 200 shown in Figure 8 generates decoded three-dimensional data 212 by decoding encoded three-dimensional data 211. Here, encoded three-dimensional data 211 is, for example, encoded three-dimensional data 112 generated by the three-dimensional data encoding device 100. This three-dimensional data decoding device 200 includes an acquisition unit 201, a decoding start GOS determination unit 202, a decoding SPC determination unit 203, and a decoding unit 204.

[0085] First, the acquisition unit 201 acquires encoded three-dimensional data 211 (S201). Next, the decoding start GOS determination unit 202 determines the GOS to be decoded (S202). Specifically, the decoding start GOS determination unit 202 refers to metadata stored in or separately from the encoded three-dimensional data 211 to determine the GOS to be decoded, which includes an SPC corresponding to the spatial position, object, or time to start decoding.

[0086] Next, the decryption SPC determination unit 203 determines the type of SPC (I, P, B) to be decrypted within the GOS (S203). For example, the decryption SPC determination unit 203 determines whether to (1) decrypt only I-SPCs, (2) decrypt I-SPCs and P-SPCs, or (3) decrypt all types. Note that if the type of SPC to be decrypted has been determined in advance, such as decrypting all SPCs, this step may not be performed.

[0087] Next, the decoding unit 204 obtains the address position where the first SPC in the decoding order (same as the encoding order) within the GOS starts in the encoded three-dimensional data 211, obtains the encoded data of the first SPC from that address position, and decodes each SPC sequentially starting from that first SPC (S204). Note that the above address position is stored in metadata, etc.

[0088] In this way, the three-dimensional data decoding device 200 decodes the decoded three-dimensional data 212. Specifically, the three-dimensional data decoding device 200 generates decoded three-dimensional data 212 of the first processing unit (GOS) by decoding each of the encoded three-dimensional data 211 of the first processing unit (GOS), which is a random access unit, and each of which is associated with three-dimensional coordinates. More specifically, the three-dimensional data decoding device 200 decodes each of the multiple second processing units (SPC) in each first processing unit (GOS). Furthermore, the three-dimensional data decoding device 200 decodes each of the multiple third processing units (VLM) in each second processing unit (SPC).

[0089] The metadata for random access is described below. This metadata is generated by the three-dimensional data encoding device 100 and is included in the encoded three-dimensional data 112(211).

[0090] In conventional random access to two-dimensional moving images, decoding began from the first frame of a random access unit that was near the specified time. In contrast, in the world, random access is expected not only to time but also to space (coordinates or objects, etc.).

[0091] Therefore, in order to achieve random access to at least three elements—coordinates, objects, and time—a table is prepared that associates each element with the GOS index number. Furthermore, the GOS index number is associated with the address of the I-SPC that is the starting point of the GOS. Figure 10 shows an example of a table included in the metadata. Note that it is not necessary to use all the tables shown in Figure 10; it is sufficient to use at least one table.

[0092] The following describes random access starting from coordinates as an example. When accessing coordinates (x2, y2, z2), first, the coordinate-GOS table is consulted to find that the location with coordinates (x2, y2, z2) is included in the second GOS. Next, the GOS address table is consulted to find that the address of the first I-SPC in the second GOS is addr(2). Therefore, the decoding unit 204 retrieves data from this address and begins decoding.

[0093] The address may be a logical format address or a physical address of the HDD or memory. Alternatively, information identifying a file segment may be used instead of an address. For example, a file segment is a unit formed by segmenting one or more GOSs (Global Operating Systems).

[0094] Furthermore, if an object spans multiple GOSs, the object-GOS table may indicate multiple GOSs to which the object belongs. If these multiple GOSs are closed GOSs, the encoding and decoding devices can perform encoding or decoding in parallel. On the other hand, if these multiple GOSs are open GOSs, the compression efficiency can be further improved by allowing the multiple GOSs to reference each other.

[0095] Examples of objects include people, animals, cars, bicycles, traffic lights, or landmark buildings. For example, the three-dimensional data encoding device 100 can extract feature points specific to objects from a three-dimensional point cloud or the like when encoding a world, detect objects based on these feature points, and set the detected objects as random access points.

[0096] Thus, the three-dimensional data encoding device 100 generates first information indicating a plurality of first processing units (GOS) and the three-dimensional coordinates associated with each of the plurality of first processing units (GOS). The encoded three-dimensional data 112(211) also includes this first information. Furthermore, the first information indicates at least one of the following: an object, a time, and a data storage location, associated with each of the plurality of first processing units (GOS).

[0097] The three-dimensional data decoding device 200 acquires first information from the encoded three-dimensional data 211, uses the first information to identify the encoded three-dimensional data 211 of the first processing unit corresponding to the specified three-dimensional coordinates, object, or time, and decodes the encoded three-dimensional data 211.

[0098] The following describes examples of other metadata. In addition to metadata for random access, the three-dimensional data encoding device 100 may generate and store the following metadata. The three-dimensional data decoding device 200 may also use this metadata during decoding.

[0099] When using three-dimensional data as map information, profiles may be defined according to the intended use, and information indicating the profile may be included in the metadata. For example, profiles may be defined for urban areas, suburbs, or for flying objects, and the maximum or minimum size of the world, SPC, or VLM may be defined for each. For example, for urban areas, more detailed information is required than for suburbs, so the minimum size of the VLM is set to be smaller.

[0100] Metadata may include tag values ​​indicating the object type. These tag values ​​are associated with the VLM, SPC, or GOS that constitute the object. For example, tag value "0" may indicate "person," tag value "1" may indicate "car," tag value "2" may indicate "traffic light," and so on, with different tag values ​​assigned to each object type. Alternatively, if it is difficult or unnecessary to determine the object type, tag values ​​indicating properties such as size or whether it is a dynamic or static object may be used.

[0101] Furthermore, the metadata may include information indicating the extent of the spatial region occupied by the world.

[0102] Furthermore, the metadata may include the size of the SPC or VXL as header information common to multiple SPCs, such as the entire stream of encoded data or an SPC within a GOS.

[0103] Furthermore, the metadata may include identification information for distance sensors or cameras used to generate the point cloud, or information indicating the positional accuracy of the point cloud within the point cloud.

[0104] Furthermore, the metadata may include information indicating whether the world consists solely of static objects or includes dynamic objects.

[0105] Modifications of this embodiment will be described below.

[0106] The encoding or decoding device may encode or decode two or more different SPCs or GOSs in parallel. The GOSs to be encoded or decoded in parallel can be determined based on metadata indicating the spatial location of the GOS.

[0107] In cases where three-dimensional data is used as a spatial map when a vehicle or flying object moves, or when such a spatial map is generated, the encoding or decoding device may encode or decode the GOS or SPC contained in the space identified based on GPS, route information, or zoom magnification.

[0108] Furthermore, the decoding device may perform decoding starting from the space closest to its own position or travel path. The encoding or decoding device may encode or decode spaces farther from its own position or travel path with lower priority compared to spaces closer to it. Here, lowering priority means lowering the processing order, lowering the resolution (downsampling), or lowering the image quality (increasing encoding efficiency, for example, by increasing the quantization step).

[0109] Furthermore, when a decoding device decodes encoded data that is hierarchically encoded in space, it may decode only the lower layers.

[0110] Furthermore, the decoding device may prioritize decoding from lower layers depending on the map's zoom level or intended use.

[0111] Furthermore, for applications such as self-localization or object recognition during autonomous driving of vehicles or robots, the encoding or decoding device may reduce the resolution of the area outside of the area to be recognized (the area within a specific height from the road surface) when encoding or decoding.

[0112] Furthermore, the encoding device may encode the point clouds representing the spatial shapes of the indoor and outdoor areas separately. For example, by separating the GOS representing the indoor area (indoor GOS) and the GOS representing the outdoor area (outdoor GOS), the decoding device can select the GOS to decode according to the viewpoint position when using the encoded data.

[0113] Furthermore, the encoding device may encode indoor and outdoor GOS locations with similar coordinates so that they are adjacent within the encoding stream. For example, the encoding device associates the identifiers of both locations and stores information indicating the associated identifiers within the encoding stream or in separately stored metadata. This allows the decoding device to identify indoor and outdoor GOS locations with similar coordinates by referring to the information in the metadata.

[0114] Furthermore, the encoding device may switch the size of the GOS or SPC between indoor and outdoor GOS. For example, the encoding device may set the GOS size smaller indoors than outdoors. The encoding device may also change the accuracy of extracting feature points from the point cloud or the accuracy of object detection between indoor and outdoor GOS.

[0115] Furthermore, the encoding device may add information to the encoded data that allows the decoding device to distinguish and display dynamic objects from static objects. This allows the decoding device to display dynamic objects together with a red frame or explanatory text. Alternatively, the decoding device may display only the red frame or explanatory text instead of the dynamic object. The decoding device may also display more detailed object types. For example, a red frame may be used for cars and a yellow frame for people.

[0116] Furthermore, the encoding or decoding device may decide whether to encode or decode dynamic objects and static objects as different SPCs or GOSs depending on the frequency of occurrence of dynamic objects or the ratio of static objects to dynamic objects. For example, if the frequency or ratio of occurrence of dynamic objects exceeds a threshold, an SPC or GOS containing a mixture of dynamic and static objects is permitted, while if the frequency or ratio of occurrence of dynamic objects does not exceed a threshold, an SPC or GOS containing a mixture of dynamic and static objects is not permitted.

[0117] When detecting dynamic objects from two-dimensional image information from a camera rather than a point cloud, the encoding device may separately acquire information to identify the detection result (such as a frame or text) and the object's position, and encode this information as part of the three-dimensional encoded data. In this case, the decoding device overlays auxiliary information (a frame or text) indicating the dynamic object onto the decoded result of the static object.

[0118] Furthermore, the encoding device may change the density of VXL or VLM in the SPC depending on the complexity of the shape of the static object. For example, the encoding device will set the VXL or VLM density to be denser as the shape of the static object becomes more complex. In addition, the encoding device may determine the quantization step when quantizing spatial position or color information according to the density of VXL or VLM. For example, the encoding device will set the quantization step to be smaller as the VXL or VLM density increases.

[0119] As described above, the encoding or decoding device according to this embodiment performs spatial encoding or decoding on a spatial basis that has coordinate information.

[0120] Furthermore, the encoding and decoding devices perform encoding or decoding in volume units within the space. A volume includes a voxel, which is the smallest unit to which location information is associated.

[0121] Furthermore, the encoding and decoding devices encode or decode arbitrary elements by associating each element of spatial information, including coordinates, objects, and time, with the GOP, or by associating each element with another element using a table. The decoding device determines the coordinates using the values ​​of the selected elements, identifies a volume, voxel, or space from the coordinates, and decodes the space containing the volume or voxel, or the identified space.

[0122] Furthermore, the encoding device determines selectable volumes, voxels, or spaces based on the elements through feature point extraction or object recognition, and encodes them as randomly accessible volumes, voxels, or spaces.

[0123] Spaces are classified into three types: I-SPCs, which can be encoded or decoded on their own; P-SPCs, which are encoded or decoded by referencing any one processed space; and B-SPCs, which are encoded or decoded by referencing any two processed spaces.

[0124] One or more volumes correspond to static or dynamic objects. Spaces containing static objects and spaces containing dynamic objects are encoded or decoded as different GOSs. In other words, SPCs containing static objects and SPCs containing dynamic objects are assigned to different GOSs.

[0125] Dynamic objects are encoded or decoded individually and mapped to one or more spaces containing static objects. In other words, multiple dynamic objects are encoded individually, and the resulting encoded data of multiple dynamic objects is mapped to an SPC containing static objects.

[0126] The encoding and decoding devices prioritize the I-SPCs within the GOS when encoding or decoding. For example, the encoding device encodes in a way that minimizes I-SPC degradation (so that the original 3D data is reproduced more faithfully after decoding). The decoding device, on the other hand, decodes only the I-SPCs.

[0127] The encoding device may perform encoding by changing the frequency of using I-SPC depending on the density or number (quantity) of objects in the world. In other words, the encoding device changes the frequency of selecting I-SPC depending on the number or density of objects included in the three-dimensional data. For example, the encoding device will increase the frequency of using I-space as the density of objects in the world increases.

[0128] Furthermore, the encoding device sets random access points in GOS units and stores information indicating the spatial region corresponding to each GOS in the header information.

[0129] The encoding device uses a default value as the spatial size of the GOS. However, the encoding device may change the size of the GOS depending on the number (quantity) or density of objects or dynamic objects. For example, the encoding device will reduce the spatial size of the GOS as the density or number of objects or dynamic objects increases.

[0130] Furthermore, the space or volume includes a set of feature points derived using information obtained from sensors such as depth sensors, gyroscopes, or cameras. The coordinates of the feature points are set to the center position of the voxel. In addition, the accuracy of the positional information can be improved by subdividing the voxels.

[0131] The feature point cloud is derived using multiple pictures. Each of the multiple pictures has at least two types of time information: actual time information and the same time information across multiple pictures mapped to space (for example, the encoded time used for rate control, etc.).

[0132] Furthermore, encoding or decoding is performed in GOS units that contain one or more spaces.

[0133] The encoding and decoding devices refer to the spaces within the processed GOS to predict the P-space or B-space within the GOS to be processed.

[0134] Alternatively, the encoding and decoding devices do not refer to different GOSs, but instead use the processed space within the GOS to be processed to predict the P space or B space within the GOS to be processed.

[0135] Furthermore, the encoding and decoding devices transmit or receive encoded streams in world units containing one or more GOSs.

[0136] Furthermore, the GOS has a layered structure in at least one direction within the world, and the encoding and decoding devices encode or decode from the lower layers. For example, a randomly accessible GOS belongs to the lowest layer. A GOS belonging to a higher layer refers to a GOS belonging to the same layer or lower. In other words, the GOS is spatially divided in a predetermined direction, and each contains multiple layers, each containing one or more SPCs. The encoding and decoding devices encode or decode each SPC by referring to an SPC included in the same layer as that SPC or in a lower layer than that SPC.

[0137] Furthermore, the encoding and decoding devices sequentially encode or decode GOS within a world unit containing multiple GOS. The encoding and decoding devices write or read information indicating the encoding or decoding order (direction) as metadata. In other words, the encoded data includes information indicating the encoding order of multiple GOS.

[0138] Furthermore, the encoding device and the decoding device encode or decode two or more different spaces or GOS in parallel.

[0139] Furthermore, the encoding and decoding devices encode or decode spatial information (coordinates, size, etc.) of space or GOS.

[0140] Furthermore, the encoding and decoding devices encode or decode spaces or GOS contained within a specific space identified based on external information relating to their own position and / or area size, such as GPS, route information, or magnification.

[0141] The encoding or decoding device encodes or decodes spaces farther away from its own position with lower priority compared to spaces closer to it.

[0142] The encoding device sets one direction in the world according to the magnification or application, and encodes a GOS with a layered structure in that direction. The decoding device then decodes the GOS with a layered structure in the one direction in the world set according to the magnification or application, prioritizing from the lower layers.

[0143] The encoding device changes the accuracy of feature point extraction, object recognition, or spatial domain size between indoor and outdoor spaces. However, the encoding and decoding devices encode or decode indoor and outdoor GOS (Geoscopy) points that are close in coordinates adjacent to each other within the world, and encode or decode their identifiers in association with each other.

[0144] (Embodiment 2) When using encoded point cloud data in actual devices or services, it is desirable to send and receive necessary information depending on the application in order to reduce network bandwidth. However, until now, such functionality has not existed in the encoded structure of three-dimensional data, nor has there been an encoding method for that purpose.

[0145] This embodiment describes a three-dimensional data encoding method and a three-dimensional data encoding apparatus for providing a function to transmit and receive only the necessary information in encoded data of a three-dimensional point cloud according to its application, as well as a three-dimensional data decoding method and a three-dimensional data decoding apparatus for decoding said encoded data.

[0146] A voxel (VXL) with a certain number of features is defined as a feature voxel (FVXL), and a world (WLD) composed of FVXLs is defined as a sparse world (SWLD). Figure 11 shows examples of the configuration of sparse worlds and worlds. SWLDs include FGOS, which is a GOS composed of FVXLs; FSPC, which is a SPC composed of FVXLs; and FVLM, which is a VLM composed of FVXLs. The data structure and prediction structure of FGOS, FSPC, and FVLM may be the same as those of GOS, SPC, and VLM.

[0147] A feature is a feature that represents the three-dimensional position information of a VXL, or the visible light information of the VXL's position, and is particularly frequently detected at corners and edges of three-dimensional objects. Specifically, this feature is a three-dimensional feature or a visible light feature as shown below, but any other feature that represents the position, brightness, or color information of the VXL is acceptable.

[0148] Three-dimensional features can be obtained using SHOT features (Signature of Histograms of OrienTations), PFH features (Point Feature Histograms), or PPF features (Point Pair Feature).

[0149] SHOT features are obtained by dividing the area around VXL, calculating the dot product of the reference point and the normal vector of the divided region, and then generating a histogram. These SHOT features have the characteristics of high dimensionality and high feature representation power.

[0150] PFH features are obtained by selecting a large number of pairs of points in the vicinity of the VXL, calculating normal vectors and other parameters from these two points, and then creating a histogram. Because these PFH features are histogram features, they are robust to some disturbances and have high feature representation power.

[0151] PPF features are features calculated using normal vectors and other methods for every two VXLs. Because all VXLs are used in these PPF features, they are robust to occlusion.

[0152] Furthermore, as visible light features, SIFT (Scale-Invariant Feature Transform), SURF (Speeded Up Robust Features), or HOG (Histogram of Oriented Gradients), which use information such as the brightness gradient of the image, can be used.

[0153] SWLD is generated by calculating the above features from each VXL of WLD and extracting FVXL. Here, SWLD can be updated every time WLD is updated, or it can be updated periodically after a certain period of time regardless of when WLD is updated.

[0154] SWLDs can be generated for each feature. For example, separate SWLDs can be generated for each feature, such as SWLD1 based on SHOT features and SWLD2 based on SIFT features, and the appropriate SWLD can be used depending on the application. Alternatively, the features of each calculated FVXL can be stored as feature information within each FVXL.

[0155] Next, we will explain how to use sparse worlds (SWLDs). Because SWLDs contain only feature voxels (FVXLs), they generally have a smaller data size compared to WLDs, which contain all VXLs.

[0156] In applications that utilize features to achieve a specific objective, using SWLD information instead of WLD information can reduce read time from the hard disk, as well as bandwidth and transfer time during network transmission. For example, by storing both WLD and SWLD as map information on a server and switching the map information transmitted to either WLD or SWLD according to client requests, network bandwidth and transfer time can be reduced. A specific example is shown below.

[0157] Figures 12 and 13 illustrate examples of SWLD and WLD usage. As shown in Figure 12, when client 1, an in-vehicle device, requires map information for self-position determination, client 1 sends a request to the server to acquire map data for self-position estimation (S301). The server sends an SWLD to client 1 in response to the acquisition request (S302). Client 1 uses the received SWLD to determine its own position (S303). At this time, client 1 acquires VXL information around client 1 using various methods such as distance sensors like rangefinders, stereo cameras, or combinations of multiple monocular cameras, and estimates its own position information from the obtained VXL information and the SWLD. Here, the self-position information includes the three-dimensional position information and orientation of client 1.

[0158] As shown in Figure 13, when client 2, an in-vehicle device, needs map information for purposes such as drawing three-dimensional maps, client 2 sends a request to the server to acquire map data for map drawing (S311). The server sends a WLD to client 2 in response to the acquisition request (S312). Client 2 uses the received WLD to perform map drawing (S313). In this case, client 2 creates a rendered image using, for example, an image taken by itself with a visible light camera and the WLD acquired from the server, and then draws the created image on the screen of a car navigation system or the like.

[0159] As described above, the server sends SWLDs to the client when primarily needing individual VXL features, such as for self-localization, and sends WLDs to the client when detailed VXL information is required, such as for map plotting. This enables efficient transmission and reception of map data.

[0160] Furthermore, the client may decide for itself whether it needs an SWLD or a WLD and request the server to send either one. The server may also decide whether to send an SWLD or a WLD based on the client or network conditions.

[0161] Next, we will explain how to switch between sending and receiving data in Sparse World (SWLD) and World (WLD) modes.

[0162] The system may switch between receiving a WLD or SWLD depending on the network bandwidth. Figure 14 shows an example of this operation. For example, when a low-speed network with limited usable network bandwidth, such as in an LTE (Long Term Evolution) environment, is used, the client accesses the server via the low-speed network (S321) and obtains an SWLD from the server as map information (S322). On the other hand, when a high-speed network with ample network bandwidth, such as in a Wi-Fi (registered trademark) environment, is used, the client accesses the server via the high-speed network (S323) and obtains a WLD from the server (S324). This allows the client to obtain appropriate map information according to the client's network bandwidth.

[0163] Specifically, the client receives the SWLD via LTE outdoors and acquires the WLD via Wi-Fi (registered trademark) when it enters an indoor facility. This allows the client to obtain more detailed indoor map information.

[0164] Thus, a client may request a WLD or SWLD from the server depending on the bandwidth of the network it is using. Alternatively, the client may send information indicating the bandwidth of the network it is using to the server, and the server may send data (WLD or SWLD) appropriate for that client based on that information. Alternatively, the server may determine the client's network bandwidth and send data (WLD or SWLD) appropriate for that client.

[0165] Furthermore, the system may switch between receiving a WLD or SWLD depending on the travel speed. Figure 15 shows an example of this operation. For example, when the client is traveling at high speed (S331), the client receives an SWLD from the server (S332). On the other hand, when the client is traveling at low speed (S333), the client receives a WLD from the server (S334). This allows the client to acquire map information appropriate to its speed while suppressing network bandwidth. Specifically, when the client is traveling on a highway, it can receive a SWLD with a small amount of data, allowing it to update rough map information at an appropriate speed. On the other hand, when the client is traveling on a general road, it can receive a WLD, allowing it to acquire more detailed map information.

[0166] Thus, the client may request a WLD or SWLD from the server according to its own movement speed. Alternatively, the client may send information indicating its movement speed to the server, and the server may send data (WLD or SWLD) appropriate to the client according to that information. Alternatively, the server may determine the client's movement speed and send data (WLD or SWLD) appropriate to the client.

[0167] Alternatively, the client may first obtain the SWLD from the server and then obtain the WLD for important areas within it. For example, when acquiring map data, the client can first obtain general map information using the SWLD, then narrow down the areas where features such as buildings, signs, or people appear frequently, and then obtain the WLD for those narrowed-down areas later. This allows the client to obtain detailed information for the necessary areas while suppressing the amount of data received from the server.

[0168] Alternatively, the server may create separate SWLDs for each object from the WLD, and the client may receive them according to its purpose. This can reduce network bandwidth usage. For example, the server may recognize people or cars in advance from the WLD and create SWLDs for people and cars. The client receives the SWLD for people if it wants to obtain information about people in the vicinity, or the SWLD for cars if it wants to obtain information about cars. Furthermore, the types of SWLDs may be distinguished by information (flags or types, etc.) added to the header.

[0169] Next, the configuration and operation flow of the three-dimensional data encoding device (e.g., a server) according to this embodiment will be described. Figure 16 is a block diagram of the three-dimensional data encoding device 400 according to this embodiment. Figure 17 is a flowchart of the three-dimensional data encoding process by the three-dimensional data encoding device 400.

[0170] The three-dimensional data encoding device 400 shown in Figure 16 generates encoded streams, encoded three-dimensional data 413 and 414, by encoding the input three-dimensional data 411. Here, encoded three-dimensional data 413 is encoded three-dimensional data corresponding to WLD, and encoded three-dimensional data 414 is encoded three-dimensional data corresponding to SWLD. This three-dimensional data encoding device 400 comprises an acquisition unit 401, an encoding region determination unit 402, an SWLD extraction unit 403, a WLD encoding unit 404, and an SWLD encoding unit 405.

[0171] As shown in Figure 17, first, the acquisition unit 401 acquires input three-dimensional data 411, which is point cloud data in three-dimensional space (S401).

[0172] Next, the encoding region determination unit 402 determines the spatial region to be encoded based on the spatial region where the point cloud data exists (S402).

[0173] Next, the SWLD extraction unit 403 defines the spatial region to be encoded as a WLD and calculates features from each VXL contained in the WLD. Then, the SWLD extraction unit 403 extracts VXLs whose features are equal to or greater than a predetermined threshold, defines the extracted VXLs as FVXLs, and adds these FVXLs to the SWLD to generate extracted three-dimensional data 412 (S403). In other words, extracted three-dimensional data 412 with features equal to or greater than the threshold is extracted from the input three-dimensional data 411.

[0174] Next, the WLD encoding unit 404 generates encoded three-dimensional data 413 corresponding to the WLD by encoding the input three-dimensional data 411 corresponding to the WLD (S404). At this time, the WLD encoding unit 404 adds information to the header of the encoded three-dimensional data 413 to distinguish that the encoded three-dimensional data 413 is a stream containing a WLD.

[0175] Furthermore, the SWLD encoding unit 405 generates encoded three-dimensional data 414 corresponding to the SWLD by encoding the extracted three-dimensional data 412 corresponding to the SWLD (S405). At this time, the SWLD encoding unit 405 adds information to the header of the encoded three-dimensional data 414 to distinguish that the encoded three-dimensional data 414 is a stream containing an SWLD.

[0176] Note that the processing order for generating encoded three-dimensional data 413 and the processing order for generating encoded three-dimensional data 414 may be reversed from the above. Also, some or all of these processes may be performed in parallel.

[0177] A parameter called "world_type" is defined as information to be added to the headers of the encoded three-dimensional data 413 and 414. If world_type=0, it indicates that the stream contains a WLD, and if world_type=1, it indicates that the stream contains an SWLD. If many other types are to be defined, the assigned number can be increased, such as world_type=2. In addition, one of the encoded three-dimensional data 413 or 414 may contain a specific flag. For example, the encoded three-dimensional data 414 may have a flag indicating that the stream contains an SWLD. In this case, the decoder can determine whether the stream contains a WLD or an SWLD based on the presence or absence of the flag.

[0178] Furthermore, the encoding method used by the WLD encoding unit 404 when encoding the WLD and the encoding method used by the SWLD encoding unit 405 when encoding the SWLD may be different.

[0179] For example, because SWLD thins out the data, it may have lower correlation with surrounding data compared to WLD. Therefore, in the encoding method used for SWLD, interpretation may be preferred over intraprediction over interprediction.

[0180] Furthermore, the encoding method used for SWLD and the encoding method used for WLD may differ in their representation of three-dimensional positions. For example, SWLD may represent the three-dimensional position of FVXL using three-dimensional coordinates, while WLD may represent the three-dimensional position using an octree, as described later, or vice versa.

[0181] Furthermore, the SWLD encoding unit 405 encodes the data such that the data size of the SWLD encoded three-dimensional data 414 is smaller than the data size of the WLD encoded three-dimensional data 413. For example, as mentioned above, SWLD may have lower correlation between data compared to WLD. This can reduce encoding efficiency, potentially causing the data size of the encoded three-dimensional data 414 to be larger than the data size of the WLD encoded three-dimensional data 413. Therefore, if the obtained encoded three-dimensional data 414 is larger than the data size of the WLD encoded three-dimensional data 413, the SWLD encoding unit 405 regenerates the encoded three-dimensional data 414 with a reduced data size by re-encoding.

[0182] For example, the SWLD extraction unit 403 regenerates the extracted three-dimensional data 412 with a reduced number of feature points, and the SWLD encoding unit 405 encodes the extracted three-dimensional data 412. Alternatively, the degree of quantization in the SWLD encoding unit 405 may be made coarser. For example, in the octree structure described later, the degree of quantization can be made coarser by rounding the data at the lowest layer.

[0183] Furthermore, if the SWLD encoding unit 405 cannot make the data size of the SWLD encoded three-dimensional data 414 smaller than the data size of the WLD encoded three-dimensional data 413, it does not need to generate the SWLD encoded three-dimensional data 414. Alternatively, the WLD encoded three-dimensional data 413 may be copied to the SWLD encoded three-dimensional data 414. In other words, the WLD encoded three-dimensional data 413 may be used as the SWLD encoded three-dimensional data 414.

[0184] Next, the configuration and operation flow of the three-dimensional data decoding device (e.g., client) according to this embodiment will be described. Figure 18 is a block diagram of the three-dimensional data decoding device 500 according to this embodiment. Figure 19 is a flowchart of the three-dimensional data decoding process by the three-dimensional data decoding device 500.

[0185] The three-dimensional data decoding device 500 shown in Figure 18 generates decoded three-dimensional data 512 or 513 by decoding encoded three-dimensional data 511. Here, encoded three-dimensional data 511 is, for example, encoded three-dimensional data 413 or 414 generated by the three-dimensional data encoding device 400.

[0186] This three-dimensional data decoding device 500 comprises an acquisition unit 501, a header analysis unit 502, a WLD decoding unit 503, and a SWLD decoding unit 504.

[0187] As shown in Figure 19, first, the acquisition unit 501 acquires the encoded three-dimensional data 511 (S501). Next, the header analysis unit 502 analyzes the header of the encoded three-dimensional data 511 and determines whether the encoded three-dimensional data 511 is a stream containing a WLD or a stream containing an SWLD (S502). For example, the world_type parameter mentioned above is referenced to make this determination.

[0188] If the encoded three-dimensional data 511 is a stream containing a WLD (Yes in S503), the WLD decoding unit 503 generates decoded three-dimensional data 512 of the WLD by decoding the encoded three-dimensional data 511 (S504). On the other hand, if the encoded three-dimensional data 511 is a stream containing a SWLD (No in S503), the SWLD decoding unit 504 generates decoded three-dimensional data 513 of the SWLD by decoding the encoded three-dimensional data 511 (S505).

[0189] Furthermore, similar to the encoding device, the decoding method used by the WLD decoding unit 503 when decoding a WLD may be different from the decoding method used by the SWLD decoding unit 504 when decoding an SWLD. For example, in the decoding method used for an SWLD, the inter-prediction method may be given priority over the intra-prediction method used for an inter-prediction method.

[0190] Furthermore, the decoding method used in SWLD and the decoding method used in WLD may differ in their representation of the three-dimensional position. For example, in SWLD, the three-dimensional position of FVXL may be represented by three-dimensional coordinates, while in WLD, the three-dimensional position may be represented by an octree, as described later, or vice versa.

[0191] Next, we will explain the octree representation, a method for representing three-dimensional positions. The VXL data contained in the three-dimensional data is converted into an octree structure and then encoded. Figure 20 shows an example of a VXL in a WLD. Figure 21 shows the octree structure of the WLD shown in Figure 20. In the example shown in Figure 20, there are three VXLs (hereinafter referred to as valid VXLs) VXL1 to VXL3 that contain point clouds. As shown in Figure 21, the octree structure consists of nodes and leaves. Each node has a maximum of eight nodes or leaves. Each leaf has VXL information. Here, among the leaves shown in Figure 21, leaves 1, 2, and 3 represent VXL1, VXL2, and VXL3 shown in Figure 20, respectively.

[0192] Specifically, each node and leaf corresponds to a three-dimensional position. Node 1 corresponds to the entire block shown in Figure 20. The block corresponding to Node 1 is divided into eight blocks, and of these eight blocks, the block containing the valid VXL is set as a node, while the other blocks are set as leaves. The block corresponding to a node is further divided into eight nodes or leaves, and this process is repeated for each level of the tree structure. In addition, all blocks at the lowest level are set as leaves.

[0193] Figure 22 shows an example of an SWLD generated from the WLD shown in Figure 20. VXL1 and VXL2 shown in Figure 20 were determined to be FVXL1 and FVXL2 as a result of feature extraction and were added to the SWLD. On the other hand, VXL3 was not determined to be FVXL and was not included in the SWLD. Figure 23 shows the octree structure of the SWLD shown in Figure 22. In the octree structure shown in Figure 23, leaf 3, which corresponds to VXL3 shown in Figure 21, has been deleted. As a result, node 3 shown in Figure 21 no longer has a valid VXL and has been changed to a leaf. In this way, the number of leaves in an SWLD is generally less than the number of leaves in a WLD, and the encoded three-dimensional data of the SWLD is also smaller than the encoded three-dimensional data of the WLD.

[0194] Modifications of this embodiment will be described below.

[0195] For example, when a client such as an in-vehicle device performs self-position estimation, it may receive a SWLD from the server and perform self-position estimation using the SWLD. When obstacle detection is performed, it may perform obstacle detection based on three-dimensional information of the surroundings acquired by the device itself using various methods such as distance sensors like rangefinders, stereo cameras, or combinations of multiple monocular cameras.

[0196] Furthermore, SWLDs generally do not contain VXL data for flat areas. Therefore, the server may maintain a subsampled world (subWLD) obtained by subsampled WLDs for static obstacle detection and send both the SWLD and subWLD to the client. This allows the client to perform self-localization and obstacle detection while suppressing network bandwidth.

[0197] Furthermore, when clients are rendering 3D map data at high speed, it can be more convenient if the map information is in a mesh structure. Therefore, the server may generate a mesh from the World Map Data (WLD) and store it in advance as a Mesh World Data (MWLD). For example, a client can receive an MWLD if it requires a coarse 3D rendering, and a WLD if it requires a detailed 3D rendering. This can reduce network bandwidth usage.

[0198] Furthermore, the server sets the VXLs whose feature quantities are above a threshold as FVXLs, but it may calculate FVXLs using a different method. For example, the server may decide that VXLs, VLMs, SPCs, or GOSs that constitute signals or intersections are necessary for self-localization, driving assistance, or autonomous driving, and include them in the SWLD as FVXLs, FVLMs, FSPCs, or FGOSs. The above decision may also be made manually. In addition, the FVXLs obtained by the above method may be added to the FVXLs set based on feature quantities. In other words, the SWLD extraction unit 403 may further extract data corresponding to objects having predetermined attributes from the input three-dimensional data 411 as extracted three-dimensional data 412.

[0199] Furthermore, the features may be labeled separately to indicate their necessity for those applications. Additionally, the server may separately maintain FVXLs as a higher layer of the SWLD (e.g., lane world) necessary for self-localization of signals or intersections, driving assistance, or autonomous driving.

[0200] Furthermore, the server may also add attributes to the VXLs within the WLD for each random access unit or predetermined unit. These attributes may include, for example, information indicating whether they are necessary or unnecessary for self-localization, or information indicating whether they are important as traffic information such as signals or intersections. The attributes may also include correspondences with features (such as intersections or roads) in lane information (such as GDF: Geographic Data Files).

[0201] Additionally, the following methods may be used to update the WLD or SWLD.

[0202] Update information indicating changes such as people, construction work, or tree-lined streets (for trucks) is uploaded to the server as point cloud or metadata. Based on this upload, the server updates the WLD, and then updates the SWLD using the updated WLD.

[0203] Furthermore, if the client detects an inconsistency between the 3D information it generates during self-localization and the 3D information it receives from the server, it may send the 3D information it generates to the server along with an update notification. In this case, the server updates the SWLD using the WLD. If the SWLD is not updated, the server determines that the WLD itself is outdated.

[0204] Furthermore, while it was stated that information distinguishing between WLD and SWLD is added to the header information of the encoded stream, if there are multiple types of worlds, such as mesh worlds or lane worlds, information distinguishing between them may also be added to the header information. Also, if there are many SWLDs with different feature quantities, information distinguishing between each of them may also be added to the header information.

[0205] Furthermore, although SWLD is said to consist of FVXLs, it may also include VXLs that were not determined to be FVXLs. For example, SWLD may include adjacent VXLs used when calculating the features of FVXLs. This allows the client to calculate the features of FVXLs when it receives SWLD, even if feature information is not attached to each FVXL in SWLD. In this case, SWLD may also include information to distinguish whether each VXL is an FVXL or a VXL.

[0206] As described above, the three-dimensional data encoding device 400 extracts extracted three-dimensional data 412 (second three-dimensional data) from the input three-dimensional data 411 (first three-dimensional data) in which the feature quantity is equal to or greater than a threshold, and generates encoded three-dimensional data 414 (first encoded three-dimensional data) by encoding the extracted three-dimensional data 412.

[0207] According to this, the three-dimensional data encoding device 400 generates encoded three-dimensional data 414 by encoding data whose feature quantity is greater than or equal to a threshold. This reduces the amount of data compared to encoding the input three-dimensional data 411 as is. Therefore, the three-dimensional data encoding device 400 can reduce the amount of data transmitted.

[0208] Furthermore, the three-dimensional data encoding device 400 generates encoded three-dimensional data 413 (second encoded three-dimensional data) by encoding the input three-dimensional data 411.

[0209] According to this, the three-dimensional data encoding device 400 can selectively transmit encoded three-dimensional data 413 and encoded three-dimensional data 414, for example, depending on the intended use.

[0210] Furthermore, the extracted three-dimensional data 412 is encoded using a first encoding method, and the input three-dimensional data 411 is encoded using a second encoding method different from the first encoding method.

[0211] According to this, the three-dimensional data encoding device 400 can use encoding methods suitable for the input three-dimensional data 411 and the extracted three-dimensional data 412, respectively.

[0212] Furthermore, in the first coding method, interpretation takes precedence over intraprediction over interprediction in the second coding method.

[0213] According to this, the three-dimensional data encoding device 400 can prioritize interpretation for extracted three-dimensional data 412, where the correlation between adjacent data tends to be low.

[0214] Furthermore, the first and second encoding methods differ in their methods of representing three-dimensional positions. For example, in the second encoding method, three-dimensional positions are represented by an octree, while in the first encoding method, three-dimensional positions are represented by three-dimensional coordinates.

[0215] According to this, the three-dimensional data encoding device 400 can use a more suitable three-dimensional position representation method for three-dimensional data with different numbers of data (number of VXLs or FVXLs).

[0216] Furthermore, at least one of the encoded three-dimensional data 413 and 414 includes an identifier indicating whether the encoded three-dimensional data is obtained by encoding the input three-dimensional data 411 or by encoding a portion of the input three-dimensional data 411. In other words, the identifier indicates whether the encoded three-dimensional data is WLD encoded three-dimensional data 413 or SWLD encoded three-dimensional data 414.

[0217] According to this, the decoding device can easily determine whether the acquired encoded three-dimensional data is encoded three-dimensional data 413 or encoded three-dimensional data 414.

[0218] Furthermore, the three-dimensional data encoding device 400 encodes the extracted three-dimensional data 412 such that the amount of data in the encoded three-dimensional data 414 is smaller than the amount of data in the encoded three-dimensional data 413.

[0219] According to this, the three-dimensional data encoding device 400 can make the amount of encoded three-dimensional data 414 smaller than the amount of encoded three-dimensional data 413.

[0220] Furthermore, the three-dimensional data encoding device 400 extracts data corresponding to objects having predetermined attributes from the input three-dimensional data 411 as extracted three-dimensional data 412. For example, objects having predetermined attributes are objects necessary for self-localization, driving assistance, or autonomous driving, such as traffic lights or intersections.

[0221] According to this, the three-dimensional data encoding device 400 can generate encoded three-dimensional data 414 that includes the data required by the decoding device.

[0222] Furthermore, the three-dimensional data encoding device 400 (server) transmits one of the encoded three-dimensional data 413 and 414 to the client, depending on the client's status.

[0223] According to this, the three-dimensional data encoding device 400 can transmit appropriate data according to the client's status.

[0224] Furthermore, the client's status includes the client's communication status (e.g., network bandwidth) or the client's speed of movement.

[0225] Furthermore, the three-dimensional data encoding device 400 transmits one of the encoded three-dimensional data 413 and 414 to the client upon the client's request.

[0226] According to this, the three-dimensional data encoding device 400 can transmit appropriate data in response to the client's request.

[0227] Furthermore, the three-dimensional data decoding device 500 according to this embodiment decodes the encoded three-dimensional data 413 or 414 generated by the three-dimensional data encoding device 400.

[0228] In other words, the three-dimensional data decoding device 500 decodes the encoded three-dimensional data 414 obtained by encoding the extracted three-dimensional data 412, in which the feature quantities extracted from the input three-dimensional data 411 are equal to or greater than a threshold, using the first decoding method. Furthermore, the three-dimensional data decoding device 500 decodes the encoded three-dimensional data 413 obtained by encoding the input three-dimensional data 411 using a second decoding method different from the first decoding method.

[0229] According to this, the three-dimensional data decoding device 500 can selectively receive encoded three-dimensional data 414, which encodes data with feature quantities above a threshold, and encoded three-dimensional data 413, for example, depending on the intended use. This allows the three-dimensional data decoding device 500 to reduce the amount of data transmitted. Furthermore, the three-dimensional data decoding device 500 can use a decoding method suitable for the input three-dimensional data 411 and the extracted three-dimensional data 412, respectively.

[0230] Furthermore, in the first decoding method, interpretation is given priority over intraprediction in the second decoding method.

[0231] According to this, the three-dimensional data decoding device 500 can prioritize interpretation for extracted three-dimensional data where the correlation between adjacent data tends to be low.

[0232] Furthermore, the first decoding method and the second decoding method use different methods for representing three-dimensional positions. For example, in the second decoding method, the three-dimensional position is represented by an octree, while in the first decoding method, the three-dimensional position is represented by three-dimensional coordinates.

[0233] According to this, the three-dimensional data decoding device 500 can use a more suitable three-dimensional position representation method for three-dimensional data with different numbers of data (number of VXLs or FVXLs).

[0234] Furthermore, at least one of the encoded three-dimensional data 413 and 414 includes an identifier indicating whether the encoded three-dimensional data is obtained by encoding the input three-dimensional data 411 or by encoding a portion of the input three-dimensional data 411. The three-dimensional data decoding device 500 identifies the encoded three-dimensional data 413 and 414 by referring to this identifier.

[0235] According to this, the three-dimensional data decoding device 500 can easily determine whether the acquired encoded three-dimensional data is encoded three-dimensional data 413 or encoded three-dimensional data 414.

[0236] Furthermore, the three-dimensional data decoding device 500 also notifies the server of the status of the client (three-dimensional data decoding device 500). Depending on the status of the client, the three-dimensional data decoding device 500 receives one of the encoded three-dimensional data 413 and 414 transmitted from the server.

[0237] According to this, the three-dimensional data decoding device 500 can receive appropriate data according to the client's status.

[0238] Furthermore, the client's status includes the client's communication status (e.g., network bandwidth) or the client's speed of movement.

[0239] Furthermore, the three-dimensional data decoding device 500 requests one of the encoded three-dimensional data 413 and 414 from the server, and in response to the request, receives one of the encoded three-dimensional data 413 and 414 transmitted from the server.

[0240] According to this, the three-dimensional data decoding device 500 can receive appropriate data according to its intended use.

[0241] (Embodiment 3) This embodiment describes a method for transmitting and receiving three-dimensional data between vehicles.

[0242] Figure 24 is a schematic diagram showing the transmission and reception of three-dimensional data 607 between the vehicle 600 and the surrounding vehicle 601.

[0243] When acquiring three-dimensional data using sensors mounted on the vehicle 600 (such as distance sensors like rangefinders, stereo cameras, or combinations of multiple monocular cameras), an occlusion region (hereinafter referred to as the occlusion region 604) occurs where three-dimensional data cannot be created, even though it is within the sensor detection range 602 of the vehicle 600, due to obstacles such as surrounding vehicles 601. Furthermore, while increasing the space in which three-dimensional data can be acquired improves the accuracy of autonomous operation, the sensor detection range of the vehicle 600 alone is finite.

[0244] The sensor detection range 602 of the vehicle 600 includes the area 603 where three-dimensional data can be acquired and the occlusion area 604. The area in which the vehicle 600 wants to acquire three-dimensional data includes the sensor detection range 602 of the vehicle 600 and the other areas. In addition, the sensor detection range 605 of the surrounding vehicle 601 includes the occlusion area 604 and the area 606 that is not included in the sensor detection range 602 of the vehicle 600.

[0245] The surrounding vehicle 601 transmits the information it detects to its own vehicle 600. By acquiring the information detected by the surrounding vehicle 601, such as the vehicle in front, the own vehicle 600 can acquire three-dimensional data 607 of the occlusion region 604 and the region 606 outside the sensor detection range 602 of the own vehicle 600. The own vehicle 600 uses the information acquired by the surrounding vehicle 601 to supplement the three-dimensional data of the occlusion region 604 and the region 606 outside the sensor detection range.

[0246] The uses of three-dimensional data in the autonomous operation of vehicles or robots include self-localization, detection of surrounding conditions, or both. For example, for self-localization, three-dimensional data generated by the vehicle 600 based on sensor information from the vehicle 600 is used. For detection of surrounding conditions, in addition to the three-dimensional data generated by the vehicle 600, three-dimensional data acquired from surrounding vehicles 601 is also used.

[0247] The surrounding vehicle 601 that transmits the three-dimensional data 607 to the vehicle 600 may be determined according to the state of the vehicle 600. For example, this surrounding vehicle 601 is the vehicle in front when the vehicle 600 is moving straight, the vehicle coming from behind when the vehicle 600 is turning right, and the vehicle behind when the vehicle 600 is reversing. Alternatively, the driver of the vehicle 600 may directly specify the surrounding vehicle 601 that transmits the three-dimensional data 607 to the vehicle 600.

[0248] Furthermore, the vehicle 600 may search for surrounding vehicles 601 that possess three-dimensional data for areas within the space where it wants to acquire three-dimensional data 607, but which it cannot acquire itself. Areas that cannot be acquired by the vehicle 600 include occlusion areas 604 or areas 606 outside the sensor detection range 602.

[0249] Furthermore, the vehicle 600 may identify the occlusion region 604 based on the sensor information of the vehicle 600. For example, the vehicle 600 may identify the occlusion region 604 as an area within the sensor detection range 602 of the vehicle 600 where three-dimensional data cannot be created.

[0250] The following describes an example of operation when the vehicle transmitting the three-dimensional data 607 is the vehicle in front. Figure 25 shows an example of the three-dimensional data transmitted in this case.

[0251] As shown in Figure 25, the three-dimensional data 607 transmitted from the preceding vehicle is, for example, a sparse world (SWLD) of a point cloud. In other words, the preceding vehicle creates three-dimensional data of a WLD (point cloud) from information detected by its own sensors, and then creates three-dimensional data of an SWLD (point cloud) by extracting data from the WLD that have a feature value above a threshold. The preceding vehicle then transmits the created SWLD three-dimensional data to its own vehicle 600.

[0252] Vehicle 600 receives the SWLD and merges it into the point cloud created by vehicle 600.

[0253] The transmitted SWLD contains information about its absolute coordinates (the SWLD's position in the coordinate system of the three-dimensional map). Vehicle 600 can perform the merge process by overwriting the point cloud it generates based on these absolute coordinates.

[0254] The SWLD transmitted from the surrounding vehicle 601 may be the SWLD of the area 606 outside the sensor detection range 602 of the host vehicle 600 and within the sensor detection range 605 of the surrounding vehicle 601, or the SWLD of the occlusion area 604 for the host vehicle 600, or both of them. Also, the transmitted SWLD may be the SWLD of the area used by the surrounding vehicle 601 for detecting the surrounding situation among the above SWLDs.

[0255] Also, the surrounding vehicle 601 may change the density of the transmitted point cloud according to the communication available time based on the speed difference between the host vehicle 600 and the surrounding vehicle 601. For example, when the speed difference is large and the communication available time is short, the surrounding vehicle 601 may lower the density (data volume) of the point cloud by extracting three-dimensional points with large feature amounts from the SWLD.

[0256] Also, detecting the surrounding situation is to determine the presence or absence of people, vehicles, equipment for road construction, etc., identify their types, and detect their positions, moving directions, moving speeds, etc.

[0257] Also, the host vehicle 600 may obtain the braking information of the surrounding vehicle 601 instead of, or in addition to, the three-dimensional data 607 generated by the surrounding vehicle 601. Here, the braking information of the surrounding vehicle 601 is, for example, information indicating that the accelerator or brake of the surrounding vehicle 601 has been depressed or the degree thereof.

[0258] Also, in the point cloud generated by each vehicle, the three-dimensional space is subdivided into random access units in consideration of low-latency communication between vehicles. On the other hand, three-dimensional maps and the like, which are map data downloaded from the server, are divided into larger random access units for the three-dimensional space compared to the case of vehicle-to-vehicle communication.

[0259] Data of areas likely to be occlusion areas, such as the area in front of the leading vehicle or the area behind the trailing vehicle, is divided into fine random access units as low-latency-oriented data.

[0260] At high speeds, the importance of the front view increases, so each vehicle creates SWLDs in a narrowed field of view range using fine random access units.

[0261] If the SWLD created by the preceding vehicle for transmission includes an area where the vehicle 600 can acquire a point cloud, the preceding vehicle may reduce the amount of transmission by removing the point cloud in that area.

[0262] Next, the configuration and operation of the three-dimensional data creation device 620, which is a three-dimensional data receiving device according to this embodiment, will be described.

[0263] Figure 26 is a block diagram of the three-dimensional data creation device 620 according to this embodiment. This three-dimensional data creation device 620 is, for example, included in the aforementioned vehicle 600, and creates a denser third three-dimensional data 636 by combining the received second three-dimensional data 635 with the first three-dimensional data 632 created by the three-dimensional data creation device 620.

[0264] The three-dimensional data creation device 620 comprises a three-dimensional data creation unit 621, a request range determination unit 622, a search unit 623, a receiving unit 624, a decoding unit 625, and a synthesis unit 626. Figure 27 is a flowchart showing the operation of the three-dimensional data creation device 620.

[0265] First, the three-dimensional data creation unit 621 creates first three-dimensional data 632 using sensor information 631 detected by sensors on the vehicle 600 (S621). Next, the request range determination unit 622 determines the request range, which is the three-dimensional spatial range in which data is missing from the created first three-dimensional data 632 (S622).

[0266] Next, the search unit 623 searches for surrounding vehicles 601 that possess three-dimensional data for the requested range, and transmits requested range information 633 indicating the requested range to the surrounding vehicles 601 identified through the search (S623). Next, the receiving unit 624 receives encoded three-dimensional data 634, which is an encoded stream of the requested range, from the surrounding vehicles 601 (S624). The search unit 623 may also indiscriminately send requests to all vehicles in a specific range and receive encoded three-dimensional data 634 from those that respond. Furthermore, the search unit 623 may send requests not only to vehicles but also to objects such as traffic lights or signs and receive encoded three-dimensional data 634 from those objects.

[0267] Next, the decoding unit 625 decodes the received encoded three-dimensional data 634 to obtain the second three-dimensional data 635 (S625). Then, the combining unit 626 combines the first three-dimensional data 632 and the second three-dimensional data 635 to create a denser third three-dimensional data 636 (S626).

[0268] Next, the configuration and operation of the three-dimensional data transmission device 640 according to this embodiment will be described. Figure 28 is a block diagram of the three-dimensional data transmission device 640.

[0269] The three-dimensional data transmission device 640 is, for example, included in the surrounding vehicle 601 described above, processes the fifth three-dimensional data 652 created by the surrounding vehicle 601 into the sixth three-dimensional data 654 requested by the vehicle 600, generates encoded three-dimensional data 634 by encoding the sixth three-dimensional data 654, and transmits the encoded three-dimensional data 634 to the vehicle 600.

[0270] The three-dimensional data transmission device 640 comprises a three-dimensional data creation unit 641, a receiving unit 642, an extraction unit 643, an encoding unit 644, and a transmission unit 645. Figure 29 is a flowchart showing the operation of the three-dimensional data transmission device 640.

[0271] First, the three-dimensional data creation unit 641 creates fifth three-dimensional data 652 using sensor information 651 detected by sensors on the surrounding vehicle 601 (S641). Next, the receiving unit 642 receives the requested range information 633 transmitted from its own vehicle 600 (S642).

[0272] Next, the extraction unit 643 processes the fifth three-dimensional data 652 into the sixth three-dimensional data 654 by extracting the three-dimensional data within the requested range indicated by the requested range information 633 from the fifth three-dimensional data 652 (S643). Next, the encoding unit 644 generates encoded three-dimensional data 634, which is an encoded stream, by encoding the sixth three-dimensional data 654 (S644). Then, the transmission unit 645 transmits the encoded three-dimensional data 634 to the vehicle 600 (S645).

[0273] In this example, vehicle 600 is equipped with a three-dimensional data creation device 620 and the surrounding vehicle 601 is equipped with a three-dimensional data transmission device 640. However, each vehicle may also have the functions of both the three-dimensional data creation device 620 and the three-dimensional data transmission device 640.

[0274] The configuration and operation of the three-dimensional data creation device 620 when it is a surrounding situation detection device that performs detection processing of the surrounding conditions of the vehicle 600 will be described below. Figure 30 is a block diagram showing the configuration of the three-dimensional data creation device 620A in this case. The three-dimensional data creation device 620A shown in Figure 30 includes, in addition to the configuration of the three-dimensional data creation device 620 shown in Figure 26, a detection area determination unit 627, a surrounding situation detection unit 628, and an autonomous operation control unit 629. Furthermore, the three-dimensional data creation device 620A is included in the vehicle 600.

[0275] Figure 31 is a flowchart of the process for detecting the surrounding conditions of the vehicle 600 by the three-dimensional data creation device 620A.

[0276] First, the three-dimensional data creation unit 621 creates first three-dimensional data 632 which is a point cloud, using sensor information 631 within the detection range of the host vehicle 600 detected by sensors included in the host vehicle 600 (S661). Note that the three-dimensional data creation device 620A may further perform self-position estimation using the sensor information 631.

[0277] Next, the detection area determination unit 627 determines a detection target range which is a spatial area for which the surrounding situation is to be detected (S662). For example, the detection area determination unit 627 calculates an area necessary for surrounding situation detection to perform autonomous operation (automatic driving) safely, according to the situation of autonomous operation such as the traveling direction and speed of the host vehicle 600, and determines the area as the detection target range.

[0278] Next, the required range determination unit 622 determines an occlusion area 604 and a spatial area which is outside the detection range of the sensors of the host vehicle 600 but necessary for surrounding situation detection, as a required range (S663).

[0279] If the required range determined in step S663 exists (Yes in S664), the search unit 623 searches for surrounding vehicles that possess information regarding the required range. For example, the search unit 623 may inquire the surrounding vehicles as to whether they possess information regarding the required range, or may determine whether the surrounding vehicles possess information regarding the required range, based on the position of the required range and the surrounding vehicles. Next, the search unit 623 transmits a request signal 637 requesting transmission of three-dimensional data to the surrounding vehicle 601 identified by the search. Then, after receiving a permission signal indicating receipt of the request of the request signal 637 transmitted from the surrounding vehicle 601, the search unit 623 transmits required range information 633 indicating the required range to the surrounding vehicle 601 (S665).

[0280] Next, the reception unit 624 detects a transmission notification of transmission data 638 which is information regarding the required range, and receives the transmission data 638 (S666).

[0281] Furthermore, the three-dimensional data creation device 620A may, without searching for a recipient of the request, indiscriminately send requests to all vehicles within a specific range and receive transmission data 638 from recipients who respond that they have information regarding the requested range. In addition, the search unit 623 may send requests not only to vehicles but also to objects such as traffic lights or signs and receive transmission data 638 from those objects.

[0282] Furthermore, the transmitted data 638 includes at least one of encoded three-dimensional data 634, which is encoded from three-dimensional data of the requested range generated by the surrounding vehicle 601, and the surrounding situation detection result 639 of the requested range. The surrounding situation detection result 639 indicates the position, direction of movement, and speed of movement of people and vehicles detected by the surrounding vehicle 601. The transmitted data 638 may also include information indicating the position and movement of the surrounding vehicle 601. For example, the transmitted data 638 may include braking information of the surrounding vehicle 601.

[0283] If the received transmission data 638 contains encoded three-dimensional data 634 (Yes in S667), the decoding unit 625 decodes the encoded three-dimensional data 634 to obtain the second three-dimensional data 635 of the SWLD (S668). In other words, the second three-dimensional data 635 is three-dimensional data (SWLD) generated by extracting data with features exceeding a threshold from the fourth three-dimensional data (WLD).

[0284] Next, the synthesis unit 626 generates the third three-dimensional data 636 by combining the first three-dimensional data 632 and the second three-dimensional data 635 (S669).

[0285] Next, the surrounding situation detection unit 628 uses the third three-dimensional data 636, which is a point cloud of the spatial area necessary for detecting the surrounding situation, to detect the surrounding situation of the vehicle 600 (S670). If the received transmission data 638 includes the surrounding situation detection result 639, the surrounding situation detection unit 628 uses the surrounding situation detection result 639 in addition to the third three-dimensional data 636 to detect the surrounding situation of the vehicle 600. Furthermore, if the received transmission data 638 includes braking information of a surrounding vehicle 601, the surrounding situation detection unit 628 uses the braking information in addition to the third three-dimensional data 636 to detect the surrounding situation of the vehicle 600.

[0286] Next, the autonomous operation control unit 629 controls the autonomous operation (autonomous driving) of the vehicle 600 based on the surrounding situation detection results from the surrounding situation detection unit 628 (S671). The surrounding situation detection results may also be presented to the driver through a UI (user interface) or the like.

[0287] On the other hand, if no requested range exists in step S663 (No in S664), that is, if information for all spatial areas necessary for detecting the surrounding conditions has been created based on the sensor information 631, the surrounding conditions detection unit 628 uses the first three-dimensional data 632, which is a point cloud of the spatial areas necessary for detecting the surrounding conditions, to detect the surrounding conditions of the vehicle 600 (S672). Then, the autonomous operation control unit 629 controls the autonomous operation (autonomous driving) of the vehicle 600 based on the surrounding conditions detection result by the surrounding conditions detection unit 628 (S671).

[0288] Furthermore, if the received transmission data 638 does not contain encoded three-dimensional data 634 (No in S667), that is, if the transmission data 638 contains only the surrounding situation detection result 639 or braking information of the surrounding vehicle 601, the surrounding situation detection unit 628 uses the first three-dimensional data 632 and the surrounding situation detection result 639 or braking information to detect the surrounding situation of the vehicle 600 (S673). Then, the autonomous operation control unit 629 controls the autonomous operation (automatic driving) of the vehicle 600 based on the surrounding situation detection result by the surrounding situation detection unit 628 (S671).

[0289] Next, we will describe the three-dimensional data transmission device 640A, which transmits the transmission data 638 to the three-dimensional data creation device 620A. Figure 32 is a block diagram of this three-dimensional data transmission device 640A.

[0290] The three-dimensional data transmission device 640A shown in Figure 32 includes a transmission feasibility determination unit 646 in addition to the configuration of the three-dimensional data transmission device 640 shown in Figure 28. Furthermore, the three-dimensional data transmission device 640A is included in the surrounding vehicle 601.

[0291] Figure 33 is a flowchart showing an example of the operation of the three-dimensional data transmission device 640A. First, the three-dimensional data creation unit 641 creates the fifth three-dimensional data 652 using sensor information 651 detected by sensors on the surrounding vehicle 601 (S681).

[0292] Next, the receiving unit 642 receives a request signal 637 from its own vehicle 600 requesting the transmission of three-dimensional data (S682). Next, the transmission feasibility determination unit 646 decides whether to respond to the request indicated by the request signal 637 (S683). For example, the transmission feasibility determination unit 646 decides whether to respond to the request based on the content set in advance by the user. Alternatively, the receiving unit 642 may first receive the other party's request, such as the requested range, and the transmission feasibility determination unit 646 may decide whether to respond to the request based on its content. For example, the transmission feasibility determination unit 646 may decide to respond to the request if it possesses three-dimensional data for the requested range, and decide not to respond to the request if it does not possess three-dimensional data for the requested range.

[0293] If the request is granted (Yes in S683), the three-dimensional data transmission device 640A transmits a permission signal to the vehicle 600, and the receiving unit 642 receives request range information 633 indicating the requested range (S684). Next, the extraction unit 643 extracts the point cloud of the requested range from the fifth three-dimensional data 652, which is a point cloud, and creates transmission data 638 including the sixth three-dimensional data 654, which is the SWLD of the extracted point cloud (S685).

[0294] In other words, the three-dimensional data transmission device 640A creates seventh three-dimensional data (WLD) from sensor information 651, and creates fifth three-dimensional data 652 (SWLD) by extracting data from the seventh three-dimensional data (WLD) whose feature quantities are above a threshold. Alternatively, the three-dimensional data creation unit 641 may have already created the three-dimensional data of the SWLD, and the extraction unit 643 may have extracted the SWLD three-dimensional data within the requested range from the SWLD three-dimensional data, or the extraction unit 643 may have generated the SWLD three-dimensional data within the requested range from the WLD three-dimensional data within the requested range.

[0295] Furthermore, the transmitted data 638 may include the surrounding situation detection result 639 of the requested range by the surrounding vehicle 601, and braking information of the surrounding vehicle 601. Alternatively, the transmitted data 638 may not include the sixth three-dimensional data 654, and may include at least one of the surrounding situation detection result 639 of the requested range by the surrounding vehicle 601, and braking information of the surrounding vehicle 601.

[0296] If the transmitted data 638 includes the sixth three-dimensional data 654 (Yes in S686), the encoding unit 644 generates encoded three-dimensional data 634 by encoding the sixth three-dimensional data 654 (S687).

[0297] Then, the transmitting unit 645 transmits the transmission data 638, which includes the encoded three-dimensional data 634, to the vehicle 600 (S688).

[0298] On the other hand, if the transmission data 638 does not include the sixth three-dimensional data 654 (No in S686), the transmission unit 645 transmits the transmission data 638 to its own vehicle 600, which includes at least one of the surrounding situation detection result 639 of the requested range by the surrounding vehicle 601 and the braking information of the surrounding vehicle 601 (S688).

[0299] Modifications of this embodiment will be described below.

[0300] For example, the information transmitted from the surrounding vehicle 601 does not have to be three-dimensional data or surrounding situation detection results created by the surrounding vehicle; it may be accurate feature point information of the surrounding vehicle 601 itself. The vehicle 600 uses this feature point information of the surrounding vehicle 601 to correct the feature point information of the preceding vehicle in the point cloud acquired by the vehicle 600. This allows the vehicle 600 to improve the matching accuracy when estimating its own position.

[0301] Furthermore, the characteristic point information of the preceding vehicle is, for example, three-dimensional point information consisting of color information and coordinate information. This allows the characteristic point information of the preceding vehicle to be used regardless of whether the vehicle 600's sensors are laser sensors or stereo cameras.

[0302] Furthermore, the vehicle 600 may use the SWLD point cloud not only during transmission but also when calculating the accuracy of self-position estimation. For example, if the sensor of the vehicle 600 is an imaging device such as a stereo camera, the vehicle 600 detects two-dimensional points on the image captured by the camera and estimates its own position using these two-dimensional points. The vehicle 600 also creates a point cloud of surrounding objects at the same time as its own position. The vehicle 600 reprojects the three-dimensional points of the SWLD within this point cloud onto the two-dimensional image and evaluates the accuracy of self-position estimation based on the error between the detected points on the two-dimensional image and the reprojected points.

[0303] Furthermore, if the vehicle 600's sensors are laser sensors such as LiDAR, the vehicle 600 evaluates the accuracy of its self-position estimation based on the error calculated by Iterative Closest Point using the SWLD of the created point cloud and the SWLD of the three-dimensional map.

[0304] Furthermore, if the communication status via base stations or servers, such as 5G, is poor, the vehicle 600 may acquire a three-dimensional map from surrounding vehicles 601.

[0305] Furthermore, information from distant locations that cannot be obtained from surrounding vehicles may be acquired through vehicle-to-vehicle communication. For example, vehicle 600 may acquire information about traffic accidents that have just occurred, several hundred meters or several kilometers away, through passing communication from oncoming vehicles or through a relay system that sequentially transmits the information to surrounding vehicles. In this case, the data format of the transmitted data is transmitted as metadata for the upper layer of the dynamic three-dimensional map.

[0306] Furthermore, the detection results of the surrounding conditions and the information detected by the vehicle 600 may be presented to the user through a user interface. For example, this information can be presented on the car navigation screen or superimposed on the front windshield.

[0307] Furthermore, vehicles that do not support autonomous driving but have cruise control may detect surrounding vehicles that are driving in autonomous driving mode and follow those surrounding vehicles.

[0308] Furthermore, if the vehicle 600 is unable to acquire a three-dimensional map or is unable to estimate its own position due to excessive occlusion, it may switch its operating mode from automatic driving mode to tracking surrounding vehicles mode.

[0309] Furthermore, the vehicle being followed may be equipped with a user interface that warns the user that it is being followed and allows the user to specify whether or not to allow the follow. In this case, a mechanism may be put in place to display advertisements on the following vehicle and pay an incentive to the vehicle being followed.

[0310] Furthermore, the transmitted information is based on SWLD, which is three-dimensional data, but may also be information corresponding to the request settings set on the vehicle 600 or the public settings of the preceding vehicle. For example, the transmitted information may be WLD, which is a dense point cloud, the results of the surrounding situation detection by the preceding vehicle, or braking information of the preceding vehicle.

[0311] Furthermore, the vehicle 600 may receive the WLD, visualize the three-dimensional data of the WLD, and present the visualized three-dimensional data to the driver using a GUI. In this case, the vehicle 600 may present the information with color coding or other means so that the user can distinguish between the point cloud created by the vehicle 600 and the point cloud received.

[0312] Furthermore, when the vehicle 600 presents the information it has detected and the detection results of the surrounding vehicles 601 to the driver via a GUI, the information may be presented with color coding or other methods so that the user can distinguish between the information detected by the vehicle 600 and the received detection results.

[0313] As described above, in the three-dimensional data creation device 620 according to this embodiment, the three-dimensional data creation unit 621 creates first three-dimensional data 632 from sensor information 631 detected by the sensor. The receiving unit 624 receives encoded three-dimensional data 634, which contains the second three-dimensional data 635. The decoding unit 625 obtains the second three-dimensional data 635 by decoding the received encoded three-dimensional data 634. The combining unit 626 creates third three-dimensional data 636 by combining the first three-dimensional data 632 and the second three-dimensional data 635.

[0314] According to this, the three-dimensional data creation device 620 can create detailed third three-dimensional data 636 using the created first three-dimensional data 632 and the received second three-dimensional data 635.

[0315] Furthermore, the synthesis unit 626 synthesizes the first three-dimensional data 632 and the second three-dimensional data 635 to create a third three-dimensional data 636 that has a higher density than the first three-dimensional data 632 and the second three-dimensional data 635.

[0316] Furthermore, the second three-dimensional data 635 (e.g., SWLD) is three-dimensional data generated by extracting data from the fourth three-dimensional data (e.g., WLD) whose features exceed a threshold.

[0317] According to this, the three-dimensional data creation device 620 can reduce the amount of data transmitted in the three-dimensional data.

[0318] Furthermore, the three-dimensional data creation device 620 includes a search unit 623 that searches for a transmitting device that is the source of the encoded three-dimensional data 634. The receiving unit 624 receives the encoded three-dimensional data 634 from the searched transmitting device.

[0319] According to this, the three-dimensional data creation device 620 can, for example, identify a transmitting device that possesses the necessary three-dimensional data by searching for it.

[0320] Furthermore, the three-dimensional data creation device includes a request range determination unit 622 that determines the request range, which is the range of the three-dimensional space for which three-dimensional data is requested. The search unit 623 transmits request range information 633 indicating the request range to the transmission device. The second three-dimensional data 635 includes the three-dimensional data of the request range.

[0321] According to this, the three-dimensional data creation device 620 can receive the necessary three-dimensional data and reduce the amount of data transmitted.

[0322] Furthermore, the requested range determination unit 622 determines the requested range to be a spatial range that includes the occlusion region 604 that cannot be detected by the sensor.

[0323] Furthermore, in the three-dimensional data transmission device 640 according to this embodiment, the three-dimensional data creation unit 641 creates fifth three-dimensional data 652 from sensor information 651 detected by the sensor. The extraction unit 643 creates sixth three-dimensional data 654 by extracting a part of the fifth three-dimensional data 652. The encoding unit 644 generates encoded three-dimensional data 634 by encoding the sixth three-dimensional data 654. The transmission unit 645 transmits the encoded three-dimensional data 634.

[0324] According to this, the three-dimensional data transmission device 640 can transmit the three-dimensional data it has created to other devices, and can also reduce the amount of data transmitted.

[0325] Furthermore, the three-dimensional data creation unit 641 creates seventh three-dimensional data (e.g., WLD) from sensor information 651 detected by the sensor, and creates fifth three-dimensional data 652 (e.g., SWLD) by extracting data from the seventh three-dimensional data whose feature quantities are above a threshold.

[0326] According to this, the three-dimensional data transmission device 640 can reduce the amount of data transmitted in the three-dimensional data.

[0327] Furthermore, the three-dimensional data transmission device 640 includes a receiving unit 642 that receives request range information 633 from the receiving device, which indicates the range of the three-dimensional space in which the three-dimensional data is requested. The extraction unit 643 creates the sixth three-dimensional data 654 by extracting the three-dimensional data of the request range from the fifth three-dimensional data 652. The transmission unit 645 transmits the encoded three-dimensional data 634 to the receiving device.

[0328] According to this, the three-dimensional data transmission device 640 can reduce the amount of data transmitted in the three-dimensional data.

[0329] (Embodiment 4) This embodiment describes the behavior of anomalies in self-localization based on a three-dimensional map.

[0330] Applications such as autonomous driving of cars, or autonomous movement of mobile objects like robots or drones, are expected to expand in the future. One example of a means to achieve such autonomous movement is for a mobile object to estimate its own position within a three-dimensional map (self-localization) and then travel according to the map.

[0331] Self-localization can be achieved by matching a three-dimensional map with three-dimensional information about the vehicle's surroundings (hereinafter referred to as "self-detection three-dimensional data") acquired by sensors such as a rangefinder (LiDAR, etc.) or stereo camera mounted on the vehicle, and estimating the vehicle's position within the three-dimensional map.

[0332] Three-dimensional maps, such as the HD maps proposed by HERE, may include not only three-dimensional point clouds but also two-dimensional map data such as road and intersection shape information, or real-time changing information such as traffic congestion and accidents. A three-dimensional map is composed of multiple layers, including three-dimensional data, two-dimensional data, and real-time changing metadata, and the device can acquire or reference only the necessary data.

[0333] The point cloud data may be SWLD as described above, or it may include point cloud data that does not contain feature points. Furthermore, the transmission and reception of point cloud data is based on one or more random access units.

[0334] The following methods can be used to match a three-dimensional map with three-dimensional vehicle detection data. For example, the device compares the shape of the point clouds in each other's point clouds and determines that areas with high similarity between feature points are in the same location. Also, if the three-dimensional map is composed of SWLDs, the device performs matching by comparing the feature points that make up the SWLD with the three-dimensional feature points extracted from the three-dimensional vehicle detection data.

[0335] Here, in order to perform self-localization with high accuracy, (A) a three-dimensional map and three-dimensional self-detection data must be acquired, and (B) the accuracy of these must meet predetermined standards. However, in the following abnormal cases, (A) or (B) cannot be met.

[0336] (1) The 3D map cannot be obtained via communication.

[0337] (2) The 3D map does not exist, or the 3D map was obtained but is corrupted.

[0338] (3) The vehicle's sensors are malfunctioning, or the accuracy of the generated 3D data for vehicle detection is insufficient due to bad weather.

[0339] The following describes the actions needed to address these abnormal cases. While a car will be used as an example, the following methods can be applied to any autonomously moving animal, such as robots or drones.

[0340] The configuration and operation of the three-dimensional information processing device according to this embodiment, for handling abnormal cases in three-dimensional maps or three-dimensional data of self-detection, will be described below. Figure 34 is a block diagram showing an example configuration of the three-dimensional information processing device 700 according to this embodiment. Figure 35 is a flowchart of the three-dimensional information processing method by the three-dimensional information processing device 700.

[0341] The three-dimensional information processing device 700 is mounted on an animal body, such as an automobile. As shown in Figure 34, the three-dimensional information processing device 700 includes a three-dimensional map acquisition unit 701, a vehicle detection data acquisition unit 702, an abnormal case determination unit 703, a response action determination unit 704, and an action control unit 705.

[0342] The three-dimensional information processing device 700 may also include two-dimensional or one-dimensional sensors (not shown) for detecting structures or animals around the vehicle, such as a camera for acquiring two-dimensional images, or a sensor for acquiring one-dimensional data using ultrasound or a laser. Furthermore, the three-dimensional information processing device 700 may also include a communication unit (not shown) for acquiring a three-dimensional map via a mobile communication network such as 4G or 5G, or via vehicle-to-vehicle communication or vehicle-to-infrastructure communication.

[0343] As shown in Figure 35, the three-dimensional map acquisition unit 701 acquires a three-dimensional map 711 of the vicinity of the travel route (S701). For example, the three-dimensional map acquisition unit 701 acquires the three-dimensional map 711 via a mobile communication network, vehicle-to-vehicle communication, or vehicle-to-infrastructure communication.

[0344] Next, the vehicle detection data acquisition unit 702 acquires vehicle detection three-dimensional data 712 based on the sensor information (S702). For example, the vehicle detection data acquisition unit 702 generates vehicle detection three-dimensional data 712 based on the sensor information acquired by the sensors installed in the vehicle.

[0345] Next, the abnormal case determination unit 703 detects an abnormal case by performing a predetermined check on at least one of the acquired three-dimensional map 711 and the vehicle detection three-dimensional data 712 (S703). In other words, the abnormal case determination unit 703 determines whether at least one of the acquired three-dimensional map 711 and the vehicle detection three-dimensional data 712 is abnormal.

[0346] In step S703, if an abnormal case is detected (Yes in S704), the action determination unit 704 determines the corrective action for the abnormal case (S705). Next, the operation control unit 705 controls the operation of each processing unit necessary for carrying out the corrective action, such as the three-dimensional map acquisition unit 701 (S706).

[0347] On the other hand, if no abnormal case is detected in step S703 (No in S704), the three-dimensional information processing device 700 terminates processing.

[0348] Furthermore, the three-dimensional information processing device 700 uses the three-dimensional map 711 and the vehicle detection three-dimensional data 712 to estimate the self-position of the vehicle equipped with the three-dimensional information processing device 700. Next, the three-dimensional information processing device 700 uses the results of the self-position estimation to automatically drive the vehicle.

[0349] In this way, the three-dimensional information processing device 700 acquires map data (three-dimensional map 711) containing the first three-dimensional location information via a communication channel. For example, the first three-dimensional location information is encoded using subspaces having three-dimensional coordinate information as units, each being a collection of one or more subspaces, and containing multiple random access units, each of which can be decoded independently. For example, the first three-dimensional location information is data (SWLD) in which feature points whose three-dimensional feature quantities are greater than or equal to a predetermined threshold are encoded.

[0350] Furthermore, the three-dimensional information processing device 700 generates second three-dimensional position information (self-detection three-dimensional data 712) from the information detected by the sensor. Next, the three-dimensional information processing device 700 determines whether the first three-dimensional position information or the second three-dimensional position information is abnormal by performing an abnormality determination process on the first three-dimensional position information or the second three-dimensional position information.

[0351] If the three-dimensional information processing device 700 determines that the first three-dimensional position information or the second three-dimensional position information is abnormal, it determines a corrective action for the abnormality. Next, the three-dimensional information processing device 700 performs the necessary controls to carry out the corrective action.

[0352] As a result, the three-dimensional information processing device 700 can detect an anomaly in the first three-dimensional position information or the second three-dimensional position information and take corrective action.

[0353] The following describes the corrective actions to take in case of abnormality 1, when the 3D map 711 cannot be obtained via communication.

[0354] Self-localization requires a three-dimensional map 711, but if the vehicle has not previously acquired a three-dimensional map 711 corresponding to the route to its destination, it needs to acquire the three-dimensional map 711 via communication. However, due to congestion in the communication channel or deterioration of radio wave reception conditions, the vehicle may be unable to acquire a three-dimensional map 711 on the road it is traveling on.

[0355] The abnormal case determination unit 703 checks whether the three-dimensional map 711 has been acquired for all sections of the route to the destination, or for sections within a predetermined range from the current location. If it has not been acquired, it determines that it is abnormal case 1. In other words, the abnormal case determination unit 703 determines whether the three-dimensional map 711 (first three-dimensional location information) can be acquired via the communication channel, and if the three-dimensional map 711 cannot be acquired via the communication channel, it determines that the three-dimensional map 711 is abnormal.

[0356] If an abnormal case 1 is determined, the action determination unit 704 selects one of two types of action: (1) continue self-position estimation, or (2) stop self-position estimation.

[0357] First, (1) I will explain a specific example of the actions taken when self-localization is to be continued. When self-localization is to be continued, a three-dimensional map 711 of the route to the destination is required.

[0358] For example, the vehicle determines a location where a communication path is available within the range where the three-dimensional map 711 has already been acquired, moves to that location, and acquires the three-dimensional map 711. At this time, the vehicle may acquire all three-dimensional maps 711 up to the destination, or it may acquire the three-dimensional maps 711 for each random access unit within the upper limit size that can be stored in the vehicle's memory or a recording unit such as an HDD.

[0359] Furthermore, the vehicle may separately acquire the communication status along the route, and if it is predicted that the communication status along the route will be poor, it may operate in such a way that it acquires a three-dimensional map 711 of the section where the communication status is poor in advance, or acquires a three-dimensional map 711 of the maximum range that can be acquired. In other words, the three-dimensional information processing device 700 predicts whether the vehicle will enter an area with poor communication status. If the three-dimensional information processing device 700 predicts that the vehicle will enter an area with poor communication status, it acquires a three-dimensional map 711 before the vehicle enters that area.

[0360] Furthermore, the vehicle may identify random access units that constitute the minimum three-dimensional map 711 necessary for self-position estimation along the route, which is a narrower range than usual, and receive the identified random access units. In other words, if the three-dimensional information processing device 700 cannot obtain the three-dimensional map 711 (first three-dimensional position information) via the communication channel, it may obtain third three-dimensional position information, which is a narrower range than the first three-dimensional position information, via the communication channel.

[0361] Furthermore, if the vehicle is unable to access the distribution server for the three-dimensional map 711, it may obtain the three-dimensional map 711 from another vehicle traveling in its vicinity, or from a mobile entity that has already obtained the three-dimensional map 711 along the route to the destination and is able to communicate with the vehicle.

[0362] Next, we will explain a specific example of the countermeasures to take when (2) self-localization is stopped. In this case, the three-dimensional map 711 along the route to the destination is not necessary.

[0363] For example, the vehicle notifies the driver that functions such as autonomous driving based on self-position estimation cannot be continued, and switches the operating mode to manual mode, where the driver takes control.

[0364] Normally, when self-localization is performed, autonomous driving is carried out, although the level of human intervention varies. On the other hand, the results of self-localization can also be used for navigation when a human is driving. Therefore, the results of self-localization do not necessarily have to be used for autonomous driving.

[0365] Furthermore, if the vehicle cannot use its normally used communication channels, such as 4G or 5G mobile communication networks, it may check whether it can acquire the 3D map 711 via another communication channel, such as vehicle-to-infrastructure Wi-Fi® or millimeter-wave communication, or vehicle-to-vehicle communication, and switch the communication channel to one that can acquire the 3D map 711.

[0366] Furthermore, if the vehicle is unable to acquire the three-dimensional map 711, it may acquire a two-dimensional map and continue autonomous driving using the two-dimensional map and the three-dimensional self-detection data 712. In other words, if the three-dimensional information processing device 700 is unable to acquire the three-dimensional map 711 via the communication channel, it may acquire map data (two-dimensional map) containing two-dimensional position information via the communication channel and perform self-position estimation of the vehicle using the two-dimensional position information and the three-dimensional self-detection data 712.

[0367] Specifically, the vehicle uses a two-dimensional map and three-dimensional self-detection data 712 for self-localization, and the same three-dimensional self-detection data 712 for detecting surrounding vehicles, pedestrians, and obstacles.

[0368] Here, map data such as HD maps can include a three-dimensional map 711 composed of a three-dimensional point cloud, two-dimensional map data (two-dimensional map), a simplified version of map data extracted from the two-dimensional map data to include characteristic information such as road shapes or intersections, and metadata representing real-time information such as traffic congestion, accidents, or construction. For example, the map data has a layer structure in which three-dimensional data (three-dimensional map 711), two-dimensional data (two-dimensional map), and metadata are arranged in order from the bottom layer.

[0369] Here, two-dimensional data has a smaller data size than three-dimensional data. Therefore, even if the communication conditions are poor, the vehicle may be able to acquire a two-dimensional map. Alternatively, the vehicle can acquire a wide-area two-dimensional map in sections where the communication conditions are good. Therefore, if the communication path conditions are poor and it is difficult to acquire the three-dimensional map 711, the vehicle may receive a layer containing the two-dimensional map instead of receiving the three-dimensional map 711. Note that because metadata has a small data size, for example, the vehicle will always receive metadata regardless of the communication conditions.

[0370] There are two methods for estimating the self-position using a two-dimensional map and three-dimensional self-detection data 712, for example:

[0371] The first method involves matching two-dimensional features. Specifically, the vehicle extracts two-dimensional features from the three-dimensional self-detection data 712 and matches the extracted two-dimensional features with a two-dimensional map.

[0372] For example, the vehicle projects the three-dimensional self-detection data 712 onto the same plane as the two-dimensional map, and matches the resulting two-dimensional data with the two-dimensional map. The matching is performed using two-dimensional image features extracted from both.

[0373] If the three-dimensional map 711 includes SWLD, the three-dimensional map 711 may store two-dimensional features coplanar with the two-dimensional map, along with three-dimensional features at feature points in three-dimensional space. For example, identification information may be attached to the two-dimensional features. Alternatively, the two-dimensional features may be stored in a separate layer from the three-dimensional data and the two-dimensional map, and the vehicle may acquire the two-dimensional feature data along with the two-dimensional map.

[0374] If the two-dimensional map displays information about locations at different heights from the ground (not on the same plane), such as white lines on the road, guardrails, and buildings, the vehicle extracts features from multiple height data in the three-dimensional self-detection data 712.

[0375] Furthermore, information indicating the correspondence between feature points in the two-dimensional map and feature points in the three-dimensional map 711 may be stored as metadata for the map data.

[0376] The second method involves matching three-dimensional features. Specifically, the vehicle acquires three-dimensional features corresponding to feature points in a two-dimensional map, and matches these acquired three-dimensional features with the three-dimensional features of the vehicle detection three-dimensional data 712.

[0377] Specifically, three-dimensional features corresponding to feature points in the two-dimensional map are stored in the map data. When the vehicle acquires the two-dimensional map, it also acquires these three-dimensional features. If the three-dimensional map 711 includes a SWLD, information is added to identify the feature points in the SWLD that correspond to the feature points in the two-dimensional map. Based on this identification information, the vehicle can determine which three-dimensional features to acquire along with the two-dimensional map. In this case, since it is sufficient to represent the two-dimensional position, the amount of data can be reduced compared to when representing the three-dimensional position.

[0378] Furthermore, when estimating the vehicle's own position using a two-dimensional map, the accuracy of the self-position estimation is lower than when using a three-dimensional map 711. Therefore, the vehicle may determine whether it can continue autonomous driving even with reduced estimation accuracy, and may continue autonomous driving only if it determines that it can.

[0379] Whether autonomous driving can be sustained depends on the driving environment, including whether the road the vehicle is traveling on is an urban area or a highway with little other traffic or pedestrians, as well as road width and road congestion (density of vehicles or pedestrians). Furthermore, it is possible to place markers for recognition using sensors such as cameras within business premises, streets, or buildings. In these specific areas, markers can be recognized with high accuracy by two-dimensional sensors, so for example, by including the marker location information in a two-dimensional map, self-localization can be performed with high accuracy.

[0380] Furthermore, by including identification information within the map indicating whether each area is a specific area, the vehicle can determine whether it is located within a specific area. If the vehicle is located within a specific area, it will decide to continue autonomous driving. In this way, the vehicle may also decide whether to continue autonomous driving based on the accuracy of its self-position estimation when using a two-dimensional map, or on the vehicle's driving environment.

[0381] In this way, the three-dimensional information processing device 700 determines whether or not to perform autonomous driving of the vehicle based on the vehicle's driving environment (the environment in which the moving object is moving) and the results of self-position estimation of the vehicle using a two-dimensional map and self-detection three-dimensional data 712.

[0382] Furthermore, the vehicle may switch the level (mode) of autonomous driving not based on whether autonomous driving can be continued, but depending on the accuracy of self-position estimation or the vehicle's driving environment. Switching the level (mode) of autonomous driving here means, for example, limiting the speed, increasing the amount of driver input (reducing the level of autonomous driving), switching to a mode that obtains driving information from the vehicle ahead and uses it as a reference, or switching to a mode that obtains driving information from vehicles set to the same destination and uses it for autonomous driving.

[0383] Furthermore, the map may include information indicating the recommended level of autonomous driving when self-localization is performed using a two-dimensional map associated with location information. The recommended level may be metadata that changes dynamically depending on traffic volume, etc. This allows the vehicle to determine the level simply by acquiring information from the map, without having to sequentially determine the level based on the surrounding environment, etc. Also, by having multiple vehicles refer to the same map, the level of autonomous driving for each vehicle can be kept constant. Note that the recommended level may be a level that is mandatory to comply with, rather than just a recommendation.

[0384] Furthermore, the vehicle may switch levels of autonomous driving depending on whether there is a driver (whether it is present or unmanned). For example, the vehicle may reduce the level of autonomous driving if there is a driver, and stop if it is unmanned. The vehicle determines a safe stopping position by recognizing surrounding pedestrians, vehicles, and traffic signs. Alternatively, the map may include location information indicating a safe stopping position for the vehicle, and the vehicle may refer to this location information to determine a safe stopping position.

[0385] Next, we will explain the corrective actions for abnormal case 2, where the 3D map 711 does not exist, or where the 3D map 711 has been acquired but is corrupted.

[0386] The abnormal case determination unit 703 checks whether either of the following applies: (1) the three-dimensional map 711 for some or all sections of the route to the destination does not exist on the distribution server or the access destination and cannot be obtained, or (2) some or all of the obtained three-dimensional map 711 is corrupted. If either of these applies, it determines it to be abnormal case 2. In other words, the abnormal case determination unit 703 determines whether the data of the three-dimensional map 711 is complete, and if the data of the three-dimensional map 711 is incomplete, it determines that the three-dimensional map 711 is abnormal.

[0387] If an abnormal case 2 is detected, the following corrective actions will be taken. First, we will explain an example of the corrective action taken when (1) the 3D map 711 cannot be obtained.

[0388] For example, the vehicle will set a route that does not pass through sections where the 3D map 711 does not exist.

[0389] Furthermore, if an alternative route cannot be set for reasons such as the absence of an alternative route or the significant increase in distance, the vehicle will set a route that includes sections where the 3D map 711 does not exist. In addition, the vehicle will notify the driver that the driving mode will be switched in that section and switch the driving mode to manual mode.

[0390] (2) If some or all of the acquired 3D map 711 is corrupted, the following corrective actions will be taken.

[0391] The vehicle identifies the damaged area in the three-dimensional map 711, requests data for the damaged area via communication, obtains the data for the damaged area, and updates the three-dimensional map 711 using the obtained data. At this time, the vehicle may specify the damaged area using positional information such as absolute or relative coordinates in the three-dimensional map 711, or it may specify it using the index number of the random access unit that constitutes the damaged area. In this case, the vehicle replaces the random access unit containing the damaged area with the obtained random access unit.

[0392] Next, we will explain the corrective actions for abnormal case 3, where the vehicle's sensors malfunction or the vehicle's three-dimensional detection data 712 cannot be generated due to bad weather.

[0393] The abnormal case determination unit 703 checks whether the generation error of the self-detection three-dimensional data 712 is within an acceptable range, and if it is not within an acceptable range, it determines it to be abnormal case 3. In other words, the abnormal case determination unit 703 determines whether the data generation accuracy of the self-detection three-dimensional data 712 is above a standard value, and if the data generation accuracy of the self-detection three-dimensional data 712 is not above a standard value, it determines that the self-detection three-dimensional data 712 is abnormal.

[0394] The following method can be used to check whether the generation error of the vehicle detection 3D data 712 is within an acceptable range.

[0395] The spatial resolution of the vehicle-detected three-dimensional data 712 during normal operation is predetermined based on the resolution in the depth direction and scanning direction of the vehicle's three-dimensional sensor, such as a rangefinder or stereo camera, or the density of the point cloud that can be generated. The vehicle also obtains the spatial resolution of the three-dimensional map 711 from metadata contained in the three-dimensional map 711.

[0396] The vehicle uses the spatial resolution of both to estimate a baseline value for the matching error when matching the vehicle-detected 3D data 712 and the 3D map 711 based on 3D features. The matching error can be calculated using statistical quantities such as the error in the 3D features for each feature point, the average of the errors in the 3D features between multiple feature points, or the error in the spatial distance between multiple feature points. The acceptable range for deviation from the baseline value is predetermined.

[0397] If the matching error between the vehicle's self-detection 3D data 712 generated before or during driving and the 3D map 711 is not within an acceptable range, the vehicle will be determined to be in abnormal case 3.

[0398] Alternatively, the vehicle may use a test pattern with a known three-dimensional shape for accuracy checks to acquire three-dimensional data 712 of its own vehicle detection relative to the test pattern before starting to drive, and determine whether it is abnormal case 3 based on whether the shape error is within an acceptable range.

[0399] For example, the vehicle performs the above determination each time before starting to drive. Alternatively, the vehicle obtains the time-series change in the matching error by performing the above determination at regular time intervals while driving. If the matching error is increasing, the vehicle may determine it as abnormal case 3 even if the error is within the acceptable range. Furthermore, if the vehicle can predict that an abnormality will occur based on the time-series change, it may notify the user that an abnormality is predicted, such as by displaying a message prompting inspection or repair. In addition, the vehicle may distinguish between abnormalities due to transient factors such as bad weather and abnormalities due to sensor failure based on the time-series change, and notify the user only of abnormalities due to sensor failure.

[0400] Furthermore, if the vehicle is determined to be in abnormal case 3, it will perform one or a selection of three types of corrective actions: (1) activate emergency backup sensors (rescue mode), (2) switch the driving mode, or (3) correct the operation of the three-dimensional sensors.

[0401] First, let's explain (1) the case where an emergency alternative sensor is activated. The vehicle activates an emergency alternative sensor that is different from the three-dimensional sensor used during normal operation. In other words, if the data generation accuracy of the self-detection three-dimensional data 712 is not above a standard value, the three-dimensional information processing device 700 generates self-detection three-dimensional data 712 (fourth three-dimensional position information) from information detected by the alternative sensor, which is different from the normal sensor.

[0402] Specifically, when a vehicle acquires three-dimensional self-detection data 712 using multiple cameras or LiDAR, the vehicle identifies a malfunctioning sensor based on factors such as the direction in which the matching error of the three-dimensional self-detection data 712 exceeds an acceptable range. The vehicle then activates a replacement sensor corresponding to the malfunctioning sensor.

[0403] The alternative sensor may be a three-dimensional sensor, a camera capable of acquiring two-dimensional images, or a one-dimensional sensor such as an ultrasonic sensor. If the alternative sensor is anything other than a three-dimensional sensor, the accuracy of self-position estimation may decrease, or self-position estimation may not be possible at all. Therefore, the vehicle may switch its autonomous driving mode depending on the type of alternative sensor.

[0404] For example, if the replacement sensor is a three-dimensional sensor, the vehicle will continue in automatic driving mode. If the replacement sensor is a two-dimensional sensor, the vehicle will change the driving mode from fully automatic driving to a semi-automatic driving mode that requires human operation. If the replacement sensor is a one-dimensional sensor, the vehicle will switch the driving mode to manual mode, which does not perform automatic braking control.

[0405] Furthermore, the vehicle may switch autonomous driving modes based on the driving environment. For example, if the alternative sensor is a two-dimensional sensor, the vehicle may continue in fully autonomous driving mode when driving on a highway and switch to semi-autonomous driving mode when driving in an urban area.

[0406] Furthermore, even if no alternative sensors are available, the vehicle may continue self-localization if a sufficient number of feature points can be acquired using only the sensors that are functioning correctly. However, since detection in a specific direction becomes impossible, the vehicle will switch the driving mode to semi-autonomous driving or manual mode.

[0407] Next, (2) the countermeasure operation for switching the driving mode will be explained. The vehicle switches the driving mode from automatic driving mode to manual mode. Alternatively, the vehicle may continue automatic driving until it reaches a safe place to stop, such as the shoulder of the road, and then stop. The vehicle may also switch the driving mode back to manual mode after stopping. In this way, the three-dimensional information processing device 700 switches the automatic driving mode if the generation accuracy of the self-detection three-dimensional data 712 is not equal to or greater than a standard value.

[0408] Next, we will explain (3) the corrective actions for correcting the operation of the three-dimensional sensor. The vehicle identifies the malfunctioning three-dimensional sensor based on the direction in which the matching error occurs, and then calibrates the identified sensor. Specifically, when multiple LiDARs or cameras are used as sensors, a portion of the three-dimensional space reconstructed by each sensor overlaps. That is, data for the overlapping portion is acquired by multiple sensors. The three-dimensional point cloud data acquired for the overlapping portion will differ between a normal sensor and a malfunctioning sensor. Therefore, the vehicle corrects the origin of the LiDAR, or adjusts the operation of predetermined parts such as the exposure or focus of the camera, so that the malfunctioning sensor can acquire three-dimensional point cloud data equivalent to that of a normal sensor.

[0409] If the matching error falls within the acceptable range after adjustment, the vehicle will continue in the previous driving mode. On the other hand, if the matching accuracy does not fall within the acceptable range after adjustment, the vehicle will perform either (1) the emergency alternative sensor activation action or (2) the driving mode switching action described above.

[0410] Thus, the three-dimensional information processing device 700 corrects the operation of the sensor if the data generation accuracy of the self-detection three-dimensional data 712 is not equal to or greater than a standard value.

[0411] The following describes how to select a corrective action. The corrective action may be selected by the driver or other user, or it may be selected automatically by the vehicle without user intervention.

[0412] Furthermore, the vehicle may switch its control depending on whether a driver is on board or not. For example, if a driver is on board, the vehicle prioritizes switching to manual mode. On the other hand, if there is no driver on board, the vehicle prioritizes moving to a safe location and stopping.

[0413] Information indicating the stopping location may be included as metadata in the three-dimensional map 711. Alternatively, the vehicle may issue a request for a response regarding the stopping location to a service that manages the autonomous driver's operation information and obtain information indicating the stopping location.

[0414] Furthermore, when a vehicle is operating on a predetermined route, the vehicle's driving mode may be switched to a mode in which an operator manages the vehicle's operation via a communication channel. In particular, an anomaly in the self-position estimation function of a vehicle operating in fully autonomous driving mode poses a high risk. Therefore, when an anomaly is detected, or when the detected anomaly cannot be corrected, the vehicle will notify the service that manages operational information of the occurrence of the anomaly via a communication channel. This service may then notify other vehicles traveling in the vicinity of the vehicle experiencing the anomaly, or instruct them to clear nearby stopping areas.

[0415] Furthermore, when an abnormal case is detected, the vehicle may reduce its driving speed compared to normal.

[0416] If the vehicle is an autonomous vehicle used for ride-hailing services such as taxis, and a malfunction occurs in the vehicle, it will contact the operations management center and stop in a safe location. The ride-hailing service will also dispatch a replacement vehicle. Alternatively, the user of the ride-hailing service may drive the vehicle. In these cases, discounts on fares or the awarding of reward points may also be used.

[0417] Furthermore, while the method for handling abnormal case 1 was described as performing self-localization based on a two-dimensional map, self-localization may also be performed using a two-dimensional map under normal circumstances. Figure 36 is a flowchart of the self-localization process in this case.

[0418] First, the vehicle acquires a three-dimensional map 711 of the area near its travel path (S711). Next, the vehicle acquires three-dimensional self-detection data 712 based on the sensor information (S712).

[0419] Next, the vehicle determines whether a three-dimensional map 711 is necessary for self-localization (S713). Specifically, the vehicle determines the necessity of the three-dimensional map 711 based on the accuracy of self-localization when using a two-dimensional map and the driving environment. For example, the same method as the one used to handle abnormal case 1 described above is used.

[0420] If it is determined that the three-dimensional map 711 is not necessary (No in S714), the vehicle acquires a two-dimensional map (S715). At this time, the vehicle may also acquire additional information as described in the method for handling abnormal case 1. Alternatively, the vehicle may generate a two-dimensional map from the three-dimensional map 711. For example, the vehicle may generate a two-dimensional map by cutting out an arbitrary plane from the three-dimensional map 711.

[0421] Next, the vehicle performs self-position estimation using the three-dimensional self-detection data 712 and the two-dimensional map (S716). The method for self-position estimation using the two-dimensional map is the same as the method described above for handling abnormal case 1.

[0422] On the other hand, if it is determined that a three-dimensional map 711 is necessary (Yes in S714), the vehicle acquires the three-dimensional map 711 (S717). Next, the vehicle performs self-position estimation using the self-detection three-dimensional data 712 and the three-dimensional map 711 (S718).

[0423] The vehicle may switch between using the two-dimensional map as the primary method and the three-dimensional map 711 as the primary method, depending on the speed supported by the vehicle's communication equipment or the conditions of the communication path. For example, if the communication speed required when driving while receiving the three-dimensional map 711 is pre-set, the vehicle may use the two-dimensional map as the primary method when the communication speed during driving is less than or equal to the set value, and use the three-dimensional map 711 as the primary method when the communication speed during driving is greater than the set value. The vehicle may also use the two-dimensional map as the primary method without making a decision on whether to adopt the two-dimensional map or the three-dimensional map.

[0424] (Embodiment 5) This embodiment describes a method for transmitting three-dimensional data to a following vehicle. Figure 37 shows an example of the target space of the three-dimensional data to be transmitted to a following vehicle.

[0425] Vehicle 801 transmits three-dimensional data, such as point clouds, contained in a rectangular space 802 with width W, height H, and depth D located at a distance L from the vehicle 801 in front of it, to a traffic monitoring cloud that monitors road conditions or to a following vehicle at time intervals of Δt.

[0426] If a vehicle or person enters space 802 from the outside, causing a change in the three-dimensional data contained in space 802 that has been previously transmitted, vehicle 801 will also transmit the three-dimensional data of the space that has been changed.

[0427] Note that while Figure 37 shows an example where the shape of space 802 is a rectangular prism, space 802 does not necessarily have to be a rectangular prism; it only needs to include the space on the road ahead that is a blind spot for following vehicles.

[0428] It is desirable that the distance L be set to a distance at which a following vehicle can safely stop after receiving the three-dimensional data. For example, the distance L is set to the sum of the distance the following vehicle travels while it is receiving the three-dimensional data, the distance the following vehicle travels before it begins to decelerate in response to the received data, and the distance required for the following vehicle to safely stop after it begins to decelerate. Since these distances change with speed, the distance L may also change according to the vehicle's speed V, such that L = a × V + b (where a and b are constants).

[0429] The width W is set to a value greater than at least the width of the lane in which the vehicle 801 is traveling. More preferably, the width W is set to a size that includes adjacent spaces such as the left and right lanes or shoulders.

[0430] The depth D can be a fixed value, but it can also change according to the vehicle's speed V, as in D = c × V + d (where c and d are constants). Furthermore, by setting D such that D > V × Δt, the transmitted space can overlap with previously transmitted space. This allows vehicle 801 to more reliably transmit the space on the track without any omissions to following vehicles, etc.

[0431] In this way, by limiting the three-dimensional data transmitted by vehicle 801 to a space useful to following vehicles, the capacity of the transmitted three-dimensional data can be effectively reduced, thereby achieving lower communication latency and lower costs.

[0432] Next, the configuration of the three-dimensional data creation device 810 according to this embodiment will be described. Figure 38 is a block diagram showing an example of the configuration of the three-dimensional data creation device 810 according to this embodiment. This three-dimensional data creation device 810 is mounted, for example, on a vehicle 801. The three-dimensional data creation device 810 transmits and receives three-dimensional data with an external traffic monitoring cloud, a preceding vehicle, or a following vehicle, and also creates and stores three-dimensional data.

[0433] The three-dimensional data creation device 810 includes a data receiving unit 811, a communication unit 812, a reception control unit 813, a format conversion unit 814, a plurality of sensors 815, a three-dimensional data creation unit 816, a three-dimensional data synthesis unit 817, a three-dimensional data storage unit 818, a communication unit 819, a transmission control unit 820, a format conversion unit 821, and a data transmission unit 822.

[0434] The data receiving unit 811 receives three-dimensional data 831 from a traffic monitoring cloud or a preceding vehicle. The three-dimensional data 831 includes information such as a point cloud, visible light images, depth information, sensor position information, or speed information, including areas that cannot be detected by the vehicle's sensors 815.

[0435] The communication unit 812 communicates with the traffic monitoring cloud or the preceding vehicle and sends data transmission requests and other messages to the traffic monitoring cloud or the preceding vehicle.

[0436] The receiving control unit 813 exchanges information such as the supported format with the communication destination via the communication unit 812 and establishes communication with the communication destination.

[0437] The format conversion unit 814 generates three-dimensional data 832 by performing format conversion on the three-dimensional data 831 received by the data reception unit 811. Furthermore, if the three-dimensional data 831 is compressed or encoded, the format conversion unit 814 performs decompression or decoding.

[0438] Multiple sensors 815 are a group of sensors that acquire information from outside the vehicle 801, such as LiDAR, visible light cameras, or infrared cameras, and generate sensor information 833. For example, if sensor 815 is a laser sensor such as LiDAR, the sensor information 833 is three-dimensional data such as a point cloud (point cloud data). Note that there are not necessarily multiple sensors 815.

[0439] The three-dimensional data creation unit 816 generates three-dimensional data 834 from the sensor information 833. The three-dimensional data 834 includes information such as a point cloud, visible light image, depth information, sensor position information, or velocity information.

[0440] The three-dimensional data synthesis unit 817 synthesizes three-dimensional data 835, which includes the space in front of the preceding vehicle that cannot be detected by the vehicle's sensors 815, by combining three-dimensional data 834 created based on the vehicle's sensor information 833 with three-dimensional data 832 created by the traffic monitoring cloud or the preceding vehicle.

[0441] The three-dimensional data storage unit 818 stores the generated three-dimensional data 835, etc.

[0442] The communication unit 819 communicates with the traffic monitoring cloud or following vehicles and sends data transmission requests, etc., to the traffic monitoring cloud or following vehicles.

[0443] The transmission control unit 820 exchanges information such as the supported format with the communication destination via the communication unit 819 and establishes communication with the communication destination. The transmission control unit 820 also determines the transmission area, which is the space of the three-dimensional data to be transmitted, based on the three-dimensional data construction information of the three-dimensional data 832 generated by the three-dimensional data synthesis unit 817 and the data transmission request from the communication destination.

[0444] Specifically, the transmission control unit 820 determines a transmission area that includes the space in front of its own vehicle that cannot be detected by the sensors of the following vehicle, in response to a data transmission request from the traffic monitoring cloud or a following vehicle. The transmission control unit 820 also determines the transmission area by determining whether the transmissionable space or the transmitted space has been updated based on the three-dimensional data construction information. For example, the transmission control unit 820 determines the transmission area to be the area specified in the data transmission request and in which the corresponding three-dimensional data 835 exists. The transmission control unit 820 then notifies the format conversion unit 821 of the format supported by the communication destination and the transmission area.

[0445] The format conversion unit 821 generates three-dimensional data 837 by converting the three-dimensional data 836 in the transmission area from the three-dimensional data 835 stored in the three-dimensional data storage unit 818 to a format supported by the receiving side. The format conversion unit 821 may also reduce the amount of data by compressing or encoding the three-dimensional data 837.

[0446] The data transmission unit 822 transmits three-dimensional data 837 to a traffic monitoring cloud or following vehicles. This three-dimensional data 837 includes, for example, information such as a point cloud in front of the vehicle, including areas that are blind spots for following vehicles, visible light images, depth information, or sensor position information.

[0447] Although this example describes a case where format conversion is performed by the format conversion units 814 and 821, format conversion is not required.

[0448] With this configuration, the three-dimensional data creation device 810 acquires three-dimensional data 831 from an external source for areas that cannot be detected by the vehicle's sensors 815, and generates three-dimensional data 835 by combining the three-dimensional data 831 with three-dimensional data 834 based on sensor information 833 detected by the vehicle's sensors 815. In this way, the three-dimensional data creation device 810 can generate three-dimensional data for areas that cannot be detected by the vehicle's sensors 815.

[0449] Furthermore, the three-dimensional data creation device 810 can transmit three-dimensional data, including the space in front of its own vehicle that cannot be detected by the sensors of the following vehicle, to the traffic monitoring cloud or following vehicle in response to a data transmission request from the traffic monitoring cloud or following vehicle.

[0450] Next, the procedure for transmitting three-dimensional data to a following vehicle using the three-dimensional data creation device 810 will be described. Figure 39 is a flowchart showing an example of the procedure for transmitting three-dimensional data to a traffic monitoring cloud or a following vehicle using the three-dimensional data creation device 810.

[0451] First, the three-dimensional data creation device 810 generates and updates three-dimensional data 835 of the space including the space 802 on the road in front of the vehicle 801 (S801). Specifically, the three-dimensional data creation device 810 constructs three-dimensional data 835 that includes the space in front of the vehicle in front of which cannot be detected by the vehicle's sensors 815, by combining three-dimensional data 834 created based on the sensor information 833 of the vehicle 801 with three-dimensional data 831 created by the traffic monitoring cloud or the vehicle in front of it.

[0452] Next, the three-dimensional data creation device 810 determines whether the three-dimensional data 835 contained in the transmitted space has changed (S802).

[0453] If a vehicle or person enters the transmitted space from the outside, causing a change in the three-dimensional data 835 contained in that space (Yes in S802), the three-dimensional data creation device 810 transmits the three-dimensional data, including the three-dimensional data 835 of the space that has been changed, to the traffic monitoring cloud or the following vehicle (S803).

[0454] The 3D data creation device 810 may transmit the 3D data of the space where the change has occurred in accordance with the transmission timing of the 3D data transmitted at predetermined intervals, or it may transmit it immediately after detecting the change. In other words, the 3D data creation device 810 may transmit the 3D data of the space where the change has occurred with priority over the 3D data transmitted at predetermined intervals.

[0455] Furthermore, the three-dimensional data creation device 810 may transmit all of the three-dimensional data of the space in which the change occurred, or it may transmit only the difference in the three-dimensional data (for example, information on three-dimensional points that have appeared or disappeared, or displacement information of three-dimensional points).

[0456] Furthermore, the three-dimensional data creation device 810 may transmit metadata related to its own vehicle's hazard avoidance actions, such as sudden braking warnings, to following vehicles prior to the three-dimensional data of the space where the change has occurred. This allows following vehicles to recognize sudden braking by the preceding vehicle earlier and initiate hazard avoidance actions such as deceleration earlier.

[0457] If no change has occurred in the three-dimensional data 835 contained in the transmitted space (No in S802), or after step S803, the three-dimensional data creation device 810 transmits the three-dimensional data contained in a space of a predetermined shape located at a distance L in front of its own vehicle 801 to the traffic monitoring cloud or a following vehicle (S804).

[0458] Furthermore, for example, the processes in steps S801 to S804 are repeated at predetermined time intervals.

[0459] Furthermore, if there is no difference between the three-dimensional data 835 of the currently transmitted space 802 and the three-dimensional map, the three-dimensional data creation device 810 does not need to transmit the three-dimensional data 837 of space 802.

[0460] Figure 40 is a flowchart showing the operation of the three-dimensional data creation device 810 in this case.

[0461] First, the three-dimensional data creation device 810 generates and updates three-dimensional data 835 of the space including the space 802 on the road in front of the vehicle 801 (S811).

[0462] Next, the three-dimensional data creation device 810 determines whether there are any updates from the three-dimensional map to the three-dimensional data 835 of the generated space 802 (S812). In other words, the three-dimensional data creation device 810 determines whether there is a difference between the three-dimensional data 835 of the generated space 802 and the three-dimensional map. Here, the three-dimensional map is three-dimensional map information managed by infrastructure-side devices such as a traffic monitoring cloud. For example, this three-dimensional map is acquired as three-dimensional data 831.

[0463] If there is an update (Yes in S812), the three-dimensional data creation device 810 transmits the three-dimensional data contained in space 802 to the traffic monitoring cloud or the following vehicle, in the same manner as described above (S813).

[0464] On the other hand, if there is no update (No in S812), the 3D data creation device 810 does not transmit the 3D data contained in space 802 to the traffic monitoring cloud and following vehicles (S814). The 3D data creation device 810 may also control the transmission of the 3D data of space 802 by setting the volume of space 802 to zero. The 3D data creation device 810 may also transmit information indicating that there is no update in space 802 to the traffic monitoring cloud or following vehicles.

[0465] Therefore, for example, if there are no obstacles on the road, there will be no difference between the generated 3D data 835 and the 3D map on the infrastructure side, and no data will be transmitted. In this way, the transmission of unnecessary data can be suppressed.

[0466] In the above description, an example was given in which the three-dimensional data creation device 810 is mounted on a vehicle, but the three-dimensional data creation device 810 is not limited to a vehicle and may be mounted on any moving object.

[0467] As described above, the three-dimensional data creation device 810 according to this embodiment is mounted on a mobile body that includes a sensor 815 and a communication unit (data receiving unit 811 or data transmitting unit 822, etc.) for sending and receiving three-dimensional data to and from the outside. The three-dimensional data creation device 810 creates three-dimensional data 835 (second three-dimensional data) based on sensor information 833 detected by the sensor 815 and three-dimensional data 831 (first three-dimensional data) received by the data receiving unit 811. The three-dimensional data creation device 810 transmits three-dimensional data 837, which is a part of the three-dimensional data 835, to the outside.

[0468] As a result, the 3D data creation device 810 can generate 3D data in areas that cannot be detected by its own vehicle. Furthermore, the 3D data creation device 810 can transmit 3D data in areas that cannot be detected by other vehicles to those other vehicles.

[0469] Furthermore, the three-dimensional data creation device 810 repeatedly creates three-dimensional data 835 and transmits three-dimensional data 837 at predetermined intervals. Three-dimensional data 837 is three-dimensional data of a small space 802 of a predetermined size located at a predetermined distance L in the direction of movement of the vehicle 801 from the current position of the vehicle 801.

[0470] This limits the range of the transmitted three-dimensional data 837, thereby reducing the amount of data transmitted.

[0471] Furthermore, the predetermined distance L changes according to the vehicle 801's speed V. For example, the greater the speed V, the longer the predetermined distance L becomes. This allows the vehicle 801 to set an appropriate small space 802 according to its speed V and transmit three-dimensional data 837 of the small space 802 to a following vehicle or the like.

[0472] Furthermore, the predetermined size changes according to the vehicle 801's movement speed V. For example, the larger the movement speed V, the larger the predetermined size. For example, the larger the movement speed V, the larger the depth D, which is the length of the small space 802 in the direction of vehicle movement. As a result, the vehicle 801 can set an appropriate small space 802 according to the vehicle 801's movement speed V and transmit the three-dimensional data 837 of the small space 802 to a following vehicle or the like.

[0473] Furthermore, the three-dimensional data creation device 810 determines whether there has been a change in the three-dimensional data 835 of the small space 802 corresponding to the transmitted three-dimensional data 837. If the three-dimensional data creation device 810 determines that there has been a change, it transmits the three-dimensional data 837 (fourth three-dimensional data), which is at least a part of the three-dimensional data 835 that has been changed, to an external following vehicle or the like.

[0474] This allows vehicle 801 to transmit three-dimensional data 837 of the space where the change occurred to following vehicles, etc.

[0475] Furthermore, the 3D data creation device 810 prioritizes the transmission of the changed 3D data 837 (fourth 3D data) over the regular 3D data 837 (third 3D data) that is transmitted periodically. Specifically, the 3D data creation device 810 transmits the changed 3D data 837 (fourth 3D data) before the transmission of the regular 3D data 837 (third 3D data) that is transmitted periodically. In other words, the 3D data creation device 810 transmits the changed 3D data 837 (fourth 3D data) irregularly, without waiting for the transmission of the regular 3D data 837 that is transmitted periodically.

[0476] As a result, vehicle 801 can preferentially transmit the three-dimensional data 837 of the space where the change occurred to following vehicles, allowing them to quickly make decisions based on the three-dimensional data.

[0477] Furthermore, the modified three-dimensional data 837 (fourth three-dimensional data) shows the difference between the three-dimensional data 835 of the small space 802 corresponding to the transmitted three-dimensional data 837 and the modified three-dimensional data 835. This reduces the amount of data transmitted for the three-dimensional data 837.

[0478] Furthermore, if there is no difference between the three-dimensional data 837 of small space 802 and the three-dimensional data 831 of small space 802, the three-dimensional data creation device 810 will not transmit the three-dimensional data 837 of small space 802. In addition, the three-dimensional data creation device 810 may transmit information to the outside indicating that there is no difference between the three-dimensional data 837 of small space 802 and the three-dimensional data 831 of small space 802.

[0479] This suppresses the transmission of unnecessary three-dimensional data 837, thereby reducing the amount of three-dimensional data 837 transmitted.

[0480] (Embodiment 6) This embodiment describes a display device and display method for displaying information obtained from a three-dimensional map, as well as a storage device and storage method for the three-dimensional map.

[0481] Mobile devices such as cars or robots utilize three-dimensional maps obtained through communication with servers or other vehicles, as well as two-dimensional images or three-dimensional self-detection data obtained from sensors mounted on the vehicle, for autonomous driving of cars or autonomous movement of robots. The data that users may want to view or save will vary depending on the situation. The following describes a display device that switches the display according to the situation.

[0482] Figure 41 is a flowchart illustrating the overview of the display method using a display device. The display device is mounted on a mobile body such as a car or robot. In the following explanation, we will describe an example where the mobile body is a vehicle (automobile).

[0483] First, the display device determines whether to display two-dimensional or three-dimensional peripheral information depending on the vehicle's driving conditions (S901). The two-dimensional peripheral information corresponds to the first peripheral information in the claim, and the three-dimensional peripheral information corresponds to the second peripheral information in the claim. Here, peripheral information refers to information indicating the surroundings of a moving object, such as an image of the vehicle viewed in a predetermined direction, or a map of the area around the vehicle.

[0484] Two-dimensional peripheral information refers to information generated using two-dimensional data. Here, two-dimensional data refers to two-dimensional map information or video. For example, two-dimensional peripheral information could be a map of the area around a vehicle obtained from a two-dimensional map, or video footage obtained from a camera mounted on the vehicle. Furthermore, two-dimensional peripheral information does not include, for example, three-dimensional information. In other words, if the two-dimensional peripheral information is a map of the area around a vehicle, the map does not include height information. Similarly, if the two-dimensional peripheral information is video footage obtained from a camera, the video does not include depth information.

[0485] Furthermore, three-dimensional surrounding information refers to information generated using three-dimensional data. Here, three-dimensional data refers to, for example, a three-dimensional map. Note that three-dimensional data may also be information indicating the three-dimensional position or shape of objects around the vehicle, obtained from other vehicles or servers, or detected by the vehicle itself. For example, three-dimensional surrounding information is a two-dimensional or three-dimensional image or map of the area around the vehicle, generated using a three-dimensional map. Furthermore, three-dimensional surrounding information may include, for example, three-dimensional information. For example, if the three-dimensional surrounding information is an image of the area in front of the vehicle, the image may include information indicating the distance to objects within the image. Alternatively, the image may display, for example, a pedestrian located behind the vehicle in front. Furthermore, three-dimensional surrounding information may be an image obtained from sensors mounted on the vehicle with this distance or pedestrian information superimposed. Furthermore, three-dimensional surrounding information may be a two-dimensional map with height information superimposed.

[0486] Furthermore, three-dimensional data may be displayed in three dimensions, or two-dimensional images or maps obtained from three-dimensional data may be displayed on a two-dimensional display or the like.

[0487] If it is decided in step S901 to display three-dimensional peripheral information (Yes in S902), the display device displays the three-dimensional peripheral information (S903). On the other hand, if it is decided in step S901 to display two-dimensional peripheral information (No in S902), the display device displays the two-dimensional peripheral information (S904). In this way, the display device displays either the three-dimensional or two-dimensional peripheral information that was decided to be displayed in step S901.

[0488] The following are specific examples. In the first example, the display device switches the surrounding information it shows depending on whether the vehicle is in autonomous or manual driving mode. Specifically, when the vehicle is in autonomous driving mode, the driver does not need to know detailed surrounding road information, so the display device shows two-dimensional surrounding information (e.g., a two-dimensional map). On the other hand, when the vehicle is in manual driving mode, it displays three-dimensional surrounding information (e.g., a three-dimensional map) so that the driver can see detailed surrounding road information for safe driving.

[0489] Furthermore, during autonomous driving, the display device may show the user what information the vehicle is using to drive, including information that influenced the driving operation (for example, SWLD used for self-position estimation, lane markings, road signs, and surrounding environment detection results). For example, the display device may display this information in addition to a two-dimensional map.

[0490] The surrounding information displayed during autonomous and manual driving is merely an example, and the display device may display three-dimensional surrounding information during autonomous driving and two-dimensional surrounding information during manual driving. Furthermore, the display device may, in addition to two-dimensional or three-dimensional maps or images, display metadata or surrounding situation detection results during at least one of autonomous or manual driving, or it may display metadata or surrounding situation detection results instead of two-dimensional or three-dimensional maps or images. Here, metadata refers to information indicating the three-dimensional position or three-dimensional shape of an object obtained from a server or another vehicle. Surrounding situation detection results refer to information indicating the three-dimensional position or three-dimensional shape of an object detected by the vehicle itself.

[0491] In the second example, the display device switches the displayed surrounding information according to the driving environment. For example, the display device switches the displayed surrounding information according to the brightness of the outside world. Specifically, when it is bright around the vehicle, the display device displays a two-dimensional image obtained from a camera mounted on the vehicle, or three-dimensional surrounding information created using that two-dimensional image. On the other hand, when it is dark around the vehicle, the display device displays three-dimensional surrounding information created using LiDAR or millimeter-wave radar, because the two-dimensional image obtained from the camera mounted on the vehicle is too dark to view.

[0492] Furthermore, the display device may switch the surrounding information it displays depending on the driving area, which is the region in which the vehicle is currently located. For example, in tourist areas, urban areas, or near a destination, the display device may display three-dimensional surrounding information to provide the user with information about surrounding buildings, etc. On the other hand, in mountainous areas or suburbs, where detailed surrounding information is often not necessary, the display device may display two-dimensional surrounding information.

[0493] Furthermore, the display device may switch the displayed surrounding information based on weather conditions. For example, in clear weather, the display device displays three-dimensional surrounding information created using a camera or lidar. On the other hand, in rainy or foggy conditions, the display device displays three-dimensional surrounding information created using millimeter-wave radar, as three-dimensional surrounding information from a camera or lidar is prone to noise.

[0494] Furthermore, these displays may be switched automatically by the system or manually by the user.

[0495] Furthermore, three-dimensional surrounding information is generated from two-dimensional map data including dense point cloud data generated based on WLD, mesh data generated based on MWLD, sparse data generated based on SWLD, lane data generated based on lane world, and three-dimensional shape information of roads and intersections, as well as metadata or vehicle detection results that include three-dimensional position or shape information that changes in real time.

[0496] As mentioned above, WLD is three-dimensional point cloud data, SWLD is data obtained by extracting point clouds from WLD where the feature value is above a threshold, and MWLD is data with a mesh structure generated from WLD. Lane world is data obtained by extracting point clouds from WLD where the feature value is above a threshold and is necessary for self-localization, driving assistance, or autonomous driving.

[0497] Here, MWLD and SWLD use less data than WLD. Therefore, by using WLD when more detailed data is needed, and MWLD or SWLD when necessary, the amount of communication data and processing load can be appropriately reduced. Furthermore, lane world uses less data than SWLD. Therefore, by using lane world, the amount of communication data and processing load can be further reduced.

[0498] Furthermore, although the above describes an example of switching between two-dimensional and three-dimensional peripheral information, the display device may also switch the type of data (WLD, SWLD, etc.) used to generate the three-dimensional peripheral information based on the above conditions. In other words, in the above description, when the display device displays three-dimensional peripheral information, it may display three-dimensional peripheral information generated from a first data set with a larger data volume (e.g., WLD or SWLD), and when the display device displays two-dimensional peripheral information, it may display three-dimensional peripheral information generated from a second data set with a smaller data volume than the first data set (e.g., SWLD or lane world) instead of two-dimensional peripheral information.

[0499] Furthermore, the display device may display two-dimensional or three-dimensional peripheral information on, for example, a two-dimensional display, head-up display, or head-mounted display mounted on the vehicle. Alternatively, the display device may transmit and display two-dimensional or three-dimensional peripheral information to a mobile terminal such as a smartphone via wireless communication. In other words, the display device is not limited to those mounted on a moving vehicle, but can be anything that works in conjunction with a moving vehicle. For example, when a user possessing a display device such as a smartphone boards or drives a moving vehicle, information about the vehicle, such as the vehicle's position based on the vehicle's self-position estimation, is displayed on the display device, or this information is displayed on the display device together with the peripheral information.

[0500] Furthermore, when displaying a three-dimensional map, the display device may either render the three-dimensional map and display it as two-dimensional data, or it may display it as three-dimensional data using a three-dimensional display or a three-dimensional hologram.

[0501] Next, we will explain how to save the 3D map. Mobile objects such as cars or robots utilize 3D maps obtained through communication with a server or other vehicles, as well as 2D images obtained from sensors mounted on the vehicle, or 3D data detected by the vehicle itself, for autonomous driving of cars or autonomous movement of robots. Of this data, the data that users may want to view or save will vary depending on the situation. The following describes how to save data according to the situation.

[0502] The storage device is mounted on a mobile body such as a car or robot. In the following example, the mobile body is described as a vehicle (automobile). Furthermore, the storage device may be included in the display device described above.

[0503] In the first example, the storage device decides whether or not to save the 3D map based on the region. By saving the 3D map to the vehicle's storage medium, autonomous driving becomes possible within the saved space without communication with a server. However, due to limitations in storage capacity, only a limited amount of data can be stored. Therefore, the storage device limits the region to which data is saved, as shown below.

[0504] For example, the storage device prioritizes saving 3D maps of frequently used areas, such as commute routes or the area around one's home. This eliminates the need to retrieve data for frequently used areas each time, effectively reducing the amount of communication data. Prioritizing storage means saving data with higher priority within a predetermined storage capacity. For example, if new data cannot be saved within the storage capacity, data with a lower priority than the new data will be deleted.

[0505] Alternatively, the storage device can prioritize saving 3D maps of areas with poor communication environments. This eliminates the need to acquire data via communication in areas with poor communication environments, thus reducing the occurrence of cases where 3D maps cannot be acquired due to poor communication.

[0506] Alternatively, the storage device prioritizes saving 3D maps of areas with high traffic volume. This allows for the priority saving of 3D maps of areas with a high accident rate. Therefore, it is possible to suppress the reduction in the accuracy of autonomous driving or driver assistance that occurs in such areas due to communication failures preventing the acquisition of 3D maps.

[0507] Alternatively, the storage device prioritizes saving 3D maps of areas with low traffic volume. In areas with low traffic volume, the automatic driving mode that automatically follows the vehicle in front is less likely to be available. This may necessitate more detailed surrounding information. Therefore, prioritizing the saving of 3D maps of areas with low traffic volume can improve the accuracy of automatic driving or driver assistance in such areas.

[0508] Furthermore, the above saving methods may be combined. The regions where these 3D maps are preferentially saved may be automatically determined by the system or specified by the user.

[0509] Furthermore, the storage device may delete 3D maps that have been stored for a predetermined period of time or update them with the latest data. This prevents the use of outdated map data. Additionally, when updating map data, the storage device may compare the old map with the new map to detect the difference region, which is the spatial area where there are differences, and update only the data for the areas that have changed by adding the data of the difference region from the new map to the old map, or by removing the data of the difference region from the old map.

[0510] In this example, the saved 3D map is used for autonomous driving. Therefore, by using SWLD as this 3D map, the amount of communication data can be reduced. Note that the 3D map is not limited to SWLD; other types of data such as WLD may also be used.

[0511] In the second example, the storage device saves a three-dimensional map based on events.

[0512] For example, the storage device saves special events encountered during driving as a three-dimensional map. This allows the user to later view details of the events. An example of events saved as a three-dimensional map is shown below. The storage device may also save three-dimensional surrounding information generated from the three-dimensional map.

[0513] For example, the storage device saves a three-dimensional map before and after a collision, or when a hazard is detected.

[0514] Alternatively, the storage device can save three-dimensional maps of distinctive scenes, such as beautiful scenery, crowded places, or tourist attractions.

[0515] These events to be saved may be determined automatically by the system or specified in advance by the user. For example, machine learning may be used as a method to determine these events.

[0516] In this example, the saved 3D map is used for viewing. Therefore, by using WLD as this 3D map, high-quality video can be provided. Note that the 3D map is not limited to WLD; other types of data such as SWLD may also be used.

[0517] The following describes how the display device controls the display according to the user. When the display device overlays the surrounding situation detection results obtained from vehicle-to-vehicle communication onto a map, it may represent surrounding vehicles as wireframes or add transparency to surrounding vehicles to make detected objects behind surrounding vehicles visible. Alternatively, the display device may display an aerial view to provide an overview of the vehicle itself, surrounding vehicles, and surrounding situation detection results.

[0518] When superimposing surrounding environment detection results or point cloud data using a head-up display onto the surrounding environment visible through the windshield, as shown in Figure 42, the position where the information is superimposed may shift due to differences in the user's posture, body shape, or eye position. Figure 43 shows an example of the head-up display when the position is shifted.

[0519] To correct such discrepancies, the display device uses information from an in-car camera or sensors mounted on the seat to detect the user's posture, body shape, or eye position. The display device adjusts the position where the information is superimposed according to the detected user's posture, body shape, or eye position. Figure 44 shows an example of the head-up display after adjustment.

[0520] Alternatively, the user may manually adjust the superimposed position using a control device installed in the vehicle.

[0521] Furthermore, the display device may show safe locations on a map during a disaster and present them to the user. Alternatively, the vehicle may inform the user of the nature of the disaster and that it is heading to a safe location, and then automatically drive to that location.

[0522] For example, a vehicle might be set to target areas with high elevation to avoid being caught in a tsunami during an earthquake. In this case, the vehicle may also obtain information about roads that have become impassable due to the earthquake through communication with a server and take appropriate actions depending on the nature of the disaster, such as taking a route that avoids those roads.

[0523] Furthermore, autonomous driving may include multiple modes, such as a travel mode and a drive mode.

[0524] In travel mode, the vehicle determines a route to its destination, taking into account factors such as speed of arrival, cost, distance, and energy consumption, and then drives autonomously according to the determined route.

[0525] In Drive Mode, the vehicle automatically determines a route to arrive at the destination at the time specified by the user. For example, if the user sets a destination and arrival time, the vehicle will determine a route that takes them around nearby tourist attractions and arrives at the destination at the set time, and then automatically drives according to the determined route.

[0526] (Embodiment 7) Embodiment 5 describes an example in which a client device such as a vehicle transmits three-dimensional data to another vehicle or a server such as a traffic monitoring cloud. In this embodiment, the client device transmits sensor information obtained from the sensor to the server or another client device.

[0527] First, the system configuration according to this embodiment will be described. Figure 45 is a diagram showing the configuration of the three-dimensional map and sensor information transmission and reception system according to this embodiment. This system includes a server 901 and client devices 902A and 902B. When client devices 902A and 902B are not specifically distinguished, they will also be referred to as client device 902.

[0528] The client device 902 is, for example, an in-vehicle device mounted on a moving object such as a vehicle. The server 901 is, for example, a traffic monitoring cloud and is capable of communicating with multiple client devices 902.

[0529] Server 901 transmits a three-dimensional map composed of point clouds to client device 902. Note that the composition of the three-dimensional map is not limited to point clouds; it may also represent other three-dimensional data, such as a mesh structure.

[0530] The client device 902 transmits sensor information acquired by the client device 902 to the server 901. The sensor information includes, for example, at least one of the following: LiDAR acquisition information, visible light image, infrared image, depth image, sensor position information, and velocity information.

[0531] The data transmitted and received between the server 901 and the client device 902 may be compressed to reduce data size, or it may be left uncompressed to maintain data accuracy. When data is compressed, a three-dimensional compression method based on an octave structure, for example, can be used for point clouds. A two-dimensional image compression method can be used for visible light images, infrared images, and depth images. A two-dimensional image compression method is, for example, MPEG-4 AVC or HEVC, which are standardized by MPEG.

[0532] Furthermore, in response to a request from the client device 902 to send a 3D map, the server 901 sends a 3D map managed by the server 901 to the client device 902. The server 901 may also send a 3D map without waiting for a request from the client device 902. For example, the server 901 may broadcast a 3D map to one or more client devices 902 located in a predetermined space. Alternatively, the server 901 may send a 3D map appropriate to the location of the client device 902 at regular intervals after receiving a transmission request from the client device 902. The server 901 may also send a 3D map to the client device 902 whenever the 3D map managed by the server 901 is updated.

[0533] The client device 902 sends a request to the server 901 to send a three-dimensional map. For example, if the client device 902 wants to perform self-position estimation while driving, the client device 902 sends a request to the server 901 to send a three-dimensional map.

[0534] Furthermore, the client device 902 may request the server 901 to send a 3D map in the following cases: If the 3D map held by the client device 902 is outdated, the client device 902 may request the server 901 to send a 3D map. For example, if a certain period of time has elapsed since the client device 902 acquired the 3D map, the client device 902 may request the server 901 to send a 3D map.

[0535] Client device 902 may request server 901 to send the three-dimensional map to the server 901 a certain time before client device 902 leaves the space represented by the three-dimensional map held by client device 902. For example, client device 902 may request server 901 to send the three-dimensional map to the server 901 if it is within a predetermined distance from the boundary of the space represented by the three-dimensional map held by client device 902. Furthermore, if the movement path and speed of client device 902 are known, the time when client device 902 leaves the space represented by the three-dimensional map held by client device 902 may be predicted based on these.

[0536] If the error in the alignment between the three-dimensional data created by the client device 902 from sensor information and the three-dimensional map exceeds a certain level, the client device 902 may request the server 901 to send the three-dimensional map.

[0537] The client device 902 transmits sensor information to the server 901 in response to a request for transmission of sensor information sent from the server 901. The client device 902 may also send sensor information to the server 901 without waiting for a request for transmission of sensor information from the server 901. For example, once the client device 902 receives a request for transmission of sensor information from the server 901, it may periodically transmit sensor information to the server 901 for a certain period. Furthermore, if the error in the alignment between the three-dimensional data created by the client device 902 based on the sensor information and the three-dimensional map obtained from the server 901 exceeds a certain level, the client device 902 may determine that a change has occurred in the three-dimensional map around the client device 902 and transmit this information, along with the sensor information, to the server 901.

[0538] Server 901 requests client device 902 to transmit sensor information. For example, Server 901 receives location information of client device 902, such as GPS, from client device 902. Based on the location information of client device 902, if Server 901 determines that client device 902 is approaching an area with little information on the three-dimensional map managed by Server 901, it requests client device 902 to transmit sensor information in order to generate a new three-dimensional map. Server 901 may also request sensor information transmission if it wants to update the three-dimensional map, check road conditions during snowfall or disasters, check traffic congestion, or check incidents and accidents.

[0539] Furthermore, the client device 902 may set the amount of sensor information data to send to the server 901 depending on the communication status or bandwidth at the time of receiving the sensor information transmission request from the server 901. Setting the amount of sensor information data to send to the server 901 means, for example, increasing or decreasing the data itself, or selecting an appropriate compression method.

[0540] Figure 46 is a block diagram showing an example configuration of the client device 902. The client device 902 receives a three-dimensional map composed of a point cloud, etc., from the server 901 and estimates its own position from the three-dimensional data created based on the sensor information of the client device 902. The client device 902 also transmits the acquired sensor information to the server 901.

[0541] The client device 902 includes a data receiving unit 1011, a communication unit 1012, a reception control unit 1013, a format conversion unit 1014, a plurality of sensors 1015, a three-dimensional data creation unit 1016, a three-dimensional image processing unit 1017, a three-dimensional data storage unit 1018, a format conversion unit 1019, a communication unit 1020, a transmission control unit 1021, and a data transmission unit 1022.

[0542] The data receiving unit 1011 receives the three-dimensional map 1031 from the server 901. The three-dimensional map 1031 is data that includes point clouds such as WLD or SWLD. The three-dimensional map 1031 may contain either compressed or uncompressed data.

[0543] The communication unit 1012 communicates with the server 901 and sends data transmission requests (for example, a request to transmit a 3D map) to the server 901.

[0544] The receiving control unit 1013 exchanges information such as the supported format with the communication destination via the communication unit 1012 and establishes communication with the communication destination.

[0545] The format conversion unit 1014 generates a three-dimensional map 1032 by performing format conversion on the three-dimensional map 1031 received by the data reception unit 1011. Furthermore, if the three-dimensional map 1031 is compressed or encoded, the format conversion unit 1014 performs decompression or decoding. However, if the three-dimensional map 1031 is uncompressed data, the format conversion unit 1014 does not perform decompression or decoding.

[0546] Multiple sensors 1015 are a group of sensors that acquire external information from the vehicle on which the client device 902 is installed, such as LiDAR, visible light cameras, infrared cameras, or depth sensors, and generate sensor information 1033. For example, if sensor 1015 is a laser sensor such as LiDAR, the sensor information 1033 is three-dimensional data such as a point cloud (point cloud data). Note that there are not necessarily multiple sensors 1015.

[0547] The three-dimensional data creation unit 1016 creates three-dimensional data 1034 of the vehicle's surroundings based on the sensor information 1033. For example, the three-dimensional data creation unit 1016 uses information acquired by LiDAR and visible light images obtained by a visible light camera to create point cloud data with color information of the vehicle's surroundings.

[0548] The three-dimensional image processing unit 1017 uses the received three-dimensional map 1032, such as a point cloud, and the three-dimensional data 1034 of the vehicle's surroundings generated from sensor information 1033 to perform self-position estimation processing for the vehicle. Alternatively, the three-dimensional image processing unit 1017 may create three-dimensional data 1035 of the vehicle's surroundings by combining the three-dimensional map 1032 and the three-dimensional data 1034, and then perform self-position estimation processing using the created three-dimensional data 1035.

[0549] The three-dimensional data storage unit 1018 stores the three-dimensional map 1032, three-dimensional data 1034, and three-dimensional data 1035, etc.

[0550] The format conversion unit 1019 generates sensor information 1037 by converting the sensor information 1033 to a format supported by the receiving side. The format conversion unit 1019 may also reduce the amount of data by compressing or encoding the sensor information 1037. Furthermore, the format conversion unit 1019 may omit processing if format conversion is not necessary. The format conversion unit 1019 may also control the amount of data transmitted according to the specified transmission range.

[0551] The communication unit 1020 communicates with the server 901 and receives data transmission requests (sensor information transmission requests), etc., from the server 901.

[0552] The transmission control unit 1021 exchanges information such as the supported format with the communication destination via the communication unit 1020 and establishes communication.

[0553] The data transmission unit 1022 transmits sensor information 1037 to the server 901. The sensor information 1037 includes information acquired by multiple sensors 1015, such as information acquired by LiDAR, brightness images acquired by a visible light camera, infrared images acquired by an infrared camera, depth images acquired by a depth sensor, sensor position information, and velocity information.

[0554] Next, the configuration of server 901 will be described. Figure 47 is a block diagram showing an example configuration of server 901. Server 901 receives sensor information transmitted from client device 902 and creates three-dimensional data based on the received sensor information. Server 901 updates the three-dimensional map it manages using the created three-dimensional data. In addition, in response to a request from client device 902 to transmit the three-dimensional map, server 901 transmits the updated three-dimensional map to client device 902.

[0555] Server 901 comprises a data receiving unit 1111, a communication unit 1112, a reception control unit 1113, a format conversion unit 1114, a three-dimensional data creation unit 1116, a three-dimensional data synthesis unit 1117, a three-dimensional data storage unit 1118, a format conversion unit 1119, a communication unit 1120, a transmission control unit 1121, and a data transmission unit 1122.

[0556] The data receiving unit 1111 receives sensor information 1037 from the client device 902. The sensor information 1037 includes, for example, information acquired by LiDAR, brightness images acquired by a visible light camera, infrared images acquired by an infrared camera, depth images acquired by a depth sensor, sensor position information, and velocity information.

[0557] The communication unit 1112 communicates with the client device 902 and sends data transmission requests (for example, requests to transmit sensor information) to the client device 902.

[0558] The receiving control unit 1113 exchanges information such as the supported format with the communication destination via the communication unit 1112 and establishes communication.

[0559] The format conversion unit 1114 generates sensor information 1132 by decompressing or decoding the received sensor information 1037 if it is compressed or encoded. However, the format conversion unit 1114 does not perform decompression or decoding if the sensor information 1037 is uncompressed data.

[0560] The three-dimensional data creation unit 1116 creates three-dimensional data 1134 of the area around the client device 902 based on the sensor information 1132. For example, the three-dimensional data creation unit 1116 uses information acquired by LiDAR and visible light images obtained by a visible light camera to create point cloud data with color information of the area around the client device 902.

[0561] The three-dimensional data synthesis unit 1117 updates the three-dimensional map 1135 managed by the server 901 by synthesizing the three-dimensional data 1134, which was created based on the sensor information 1132, with the three-dimensional map 1135.

[0562] The three-dimensional data storage unit 1118 stores three-dimensional maps 1135, etc.

[0563] The format conversion unit 1119 generates a three-dimensional map 1031 by converting the three-dimensional map 1135 to a format supported by the receiving side. The format conversion unit 1119 may also reduce the amount of data by compressing or encoding the three-dimensional map 1135. Furthermore, the format conversion unit 1119 may omit processing if format conversion is not necessary. The format conversion unit 1119 may also control the amount of data transmitted according to the specified transmission range.

[0564] The communication unit 1120 communicates with the client device 902 and receives data transmission requests (such as requests to transmit a three-dimensional map) from the client device 902.

[0565] The transmission control unit 1121 exchanges information such as the supported format with the communication destination via the communication unit 1120 and establishes communication.

[0566] The data transmission unit 1122 transmits the three-dimensional map 1031 to the client device 902. The three-dimensional map 1031 is data that includes point clouds such as WLD or SWLD. The three-dimensional map 1031 may contain either compressed or uncompressed data.

[0567] Next, we will describe the operation flow of the client device 902. Figure 48 is a flowchart showing the operation of the client device 902 when acquiring a three-dimensional map.

[0568] First, the client device 902 requests the server 901 to transmit a three-dimensional map (such as a point cloud) (S1001). At this time, the client device 902 may also transmit its own location information obtained by GPS or the like, and request the server 901 to transmit a three-dimensional map related to that location information.

[0569] Next, the client device 902 receives a three-dimensional map from the server 901 (S1002). If the received three-dimensional map is compressed data, the client device 902 decodes the received three-dimensional map to generate an uncompressed three-dimensional map (S1003).

[0570] Next, the client device 902 creates three-dimensional data 1034 of the area around the client device 902 from sensor information 1033 obtained from multiple sensors 1015 (S1004). Then, the client device 902 estimates its own position using the three-dimensional map 1032 received from the server 901 and the three-dimensional data 1034 created from the sensor information 1033 (S1005).

[0571] Figure 49 is a flowchart showing the operation of the client device 902 when transmitting sensor information. First, the client device 902 receives a request to transmit sensor information from the server 901 (S1011). Upon receiving the transmission request, the client device 902 transmits the sensor information 1037 to the server 901 (S1012). If the sensor information 1033 includes multiple pieces of information obtained from multiple sensors 1015, the client device 902 may generate the sensor information 1037 by compressing each piece of information using a compression method suitable for each piece of information.

[0572] Next, the operation flow of server 901 will be described. Figure 50 is a flowchart showing the operation of server 901 when acquiring sensor information. First, server 901 requests client device 902 to send sensor information (S1021). Next, server 901 receives sensor information 1037 sent from client device 902 in response to the request (S1022). Next, server 901 creates three-dimensional data 1134 using the received sensor information 1037 (S1023). Next, server 901 reflects the created three-dimensional data 1134 in the three-dimensional map 1135 (S1024).

[0573] Figure 51 is a flowchart illustrating the operation of server 901 when transmitting a three-dimensional map. First, server 901 receives a request to transmit a three-dimensional map from client device 902 (S1031). Upon receiving the request to transmit a three-dimensional map, server 901 transmits the three-dimensional map 1031 to client device 902 (S1032). At this time, server 901 may extract a three-dimensional map of the vicinity of client device 902 according to its location information and transmit the extracted three-dimensional map. Alternatively, server 901 may compress the three-dimensional map composed of a point cloud using, for example, an octave tree compression method, and transmit the compressed three-dimensional map.

[0574] Modifications of this embodiment will be described below.

[0575] Server 901 uses sensor information 1037 received from client device 902 to create three-dimensional data 1134 of the area around client device 902. Next, server 901 calculates the difference between the created three-dimensional data 1134 and the three-dimensional map 1135 of the same area managed by server 901 by matching them. If the difference is greater than or equal to a predetermined threshold, server 901 determines that some kind of abnormality has occurred around client device 902. For example, when ground subsidence occurs due to a natural disaster such as an earthquake, a large difference may occur between the three-dimensional map 1135 managed by server 901 and the three-dimensional data 1134 created based on sensor information 1037.

[0576] The sensor information 1037 may include information indicating at least one of the following: the type of sensor, the performance of the sensor, and the model number of the sensor. Furthermore, a class ID corresponding to the sensor's performance may be added to the sensor information 1037. For example, if the sensor information 1037 is information acquired by a LiDAR, it is conceivable to assign identifiers to the sensor's performance, such as class 1 for sensors that can acquire information with accuracy in the millimeter range, class 2 for sensors that can acquire information with accuracy in the centimeter range, and class 3 for sensors that can acquire information with accuracy in the meter range. The server 901 may also estimate the sensor's performance information from the model number of the client device 902. For example, if the client device 902 is mounted in a vehicle, the server 901 may determine the sensor's specifications from the vehicle's make and model. In this case, the server 901 may have previously acquired information about the vehicle's make and model, or this information may be included in the sensor information. The server 901 may also use the acquired sensor information 1037 to switch the degree of correction applied to the three-dimensional data 1134 created using the sensor information 1037. For example, if the sensor performance is high precision (Class 1), the server 901 does not perform any correction on the three-dimensional data 1134. If the sensor performance is low precision (Class 3), the server 901 applies a correction to the three-dimensional data 1134 according to the accuracy of the sensor. For example, the lower the accuracy of the sensor, the stronger the degree (intensity) of the correction applied by the server 901.

[0577] Server 901 may simultaneously send requests for the transmission of sensor information to multiple client devices 902 located in a given space. When Server 901 receives multiple sensor information from multiple client devices 902, it is not necessary to use all of the sensor information to create the three-dimensional data 1134. For example, it may select which sensor information to use depending on the performance of the sensors. For example, when updating the three-dimensional map 1135, Server 901 may select high-precision sensor information (Class 1) from the multiple sensor information received and use the selected sensor information to create the three-dimensional data 1134.

[0578] Server 901 is not limited to servers such as traffic monitoring clouds, but may also be other client devices (in-vehicle). Figure 52 shows the system configuration in this case.

[0579] For example, client device 902C requests sensor information from a nearby client device 902A and obtains the sensor information from client device 902A. Then, client device 902C uses the obtained sensor information from client device 902A to create three-dimensional data and updates the three-dimensional map of client device 902C. In this way, client device 902C can generate a three-dimensional map of the space obtainable from client device 902A, taking advantage of the performance of client device 902C. For example, this case is likely to occur when client device 902C has high performance.

[0580] In this case, client device 902A, which provided the sensor information, is granted the right to acquire the high-precision three-dimensional map generated by client device 902C. Client device 902A receives the high-precision three-dimensional map from client device 902C in accordance with that right.

[0581] Furthermore, client device 902C may send requests for the transmission of sensor information to multiple nearby client devices 902 (client devices 902A and 902B). If the sensor of client device 902A or client device 902B is high-performance, client device 902C can create three-dimensional data using the sensor information obtained from this high-performance sensor.

[0582] Figure 53 is a block diagram showing the functional configuration of the server 901 and the client device 902. The server 901 includes, for example, a three-dimensional map compression / decoding processing unit 1201 that compresses and decodes three-dimensional maps, and a sensor information compression / decoding processing unit 1202 that compresses and decodes sensor information.

[0583] The client device 902 comprises a three-dimensional map decoding processing unit 1211 and a sensor information compression processing unit 1212. The three-dimensional map decoding processing unit 1211 receives encoded data of the compressed three-dimensional map, decodes the encoded data, and obtains the three-dimensional map. The sensor information compression processing unit 1212 compresses the sensor information itself instead of the three-dimensional data created from the acquired sensor information, and sends the encoded data of the compressed sensor information to the server 901. With this configuration, the client device 902 only needs to internally store a processing unit (device or LSI) that performs the processing of decoding the three-dimensional map (point cloud, etc.), and does not need to internally store a processing unit that performs the processing of compressing the three-dimensional data of the three-dimensional map (point cloud, etc.). This reduces the cost and power consumption of the client device 902.

[0584] As described above, the client device 902 according to this embodiment is mounted on a mobile body and creates three-dimensional data 1034 of the surrounding area of ​​the mobile body from sensor information 1033 indicating the surrounding conditions of the mobile body obtained by a sensor 1015 mounted on the mobile body. The client device 902 estimates the self-position of the mobile body using the created three-dimensional data 1034. The client device 902 transmits the acquired sensor information 1033 to the server 901 or another mobile body 902.

[0585] According to this, the client device 902 transmits sensor information 1033 to the server 901, etc. This may reduce the amount of data transmitted compared to transmitting three-dimensional data. In addition, since the client device 902 does not need to perform processing such as compression or encoding of three-dimensional data, the processing load on the client device 902 can be reduced. Therefore, the client device 902 can achieve a reduction in the amount of data transmitted or a simplification of the device configuration.

[0586] Furthermore, the client device 902 sends a request to the server 901 to send a three-dimensional map, and receives the three-dimensional map 1031 from the server 901. In estimating its own position, the client device 902 uses the three-dimensional data 1034 and the three-dimensional map 1032 to estimate its own position.

[0587] Furthermore, the sensor information 1033 includes at least one of the following: information obtained from the laser sensor, brightness image, infrared image, depth image, sensor position information, and sensor velocity information.

[0588] Furthermore, sensor information 1033 includes information indicating the performance of the sensor.

[0589] Furthermore, the client device 902 encodes or compresses the sensor information 1033, and when transmitting the sensor information, it transmits the encoded or compressed sensor information 1037 to the server 901 or another mobile device 902. This allows the client device 902 to reduce the amount of data transmitted.

[0590] For example, the client device 902 includes a processor and memory, and the processor uses the memory to perform the above processing.

[0591] Furthermore, the server 901 according to this embodiment is capable of communicating with a client device 902 mounted on the mobile body, and receives sensor information 1037 from the client device 902 that indicates the surrounding conditions of the mobile body, obtained by a sensor 1015 mounted on the mobile body. The server 901 creates three-dimensional data 1134 of the surroundings of the mobile body from the received sensor information 1037.

[0592] According to this, the server 901 creates three-dimensional data 1134 using sensor information 1037 transmitted from the client device 902. This may reduce the amount of data transmitted compared to when the client device 902 transmits the three-dimensional data. In addition, since the client device 902 does not need to perform processing such as compression or encoding of the three-dimensional data, the processing load on the client device 902 can be reduced. Therefore, the server 901 can reduce the amount of data transmitted or simplify the configuration of the device.

[0593] Furthermore, the server 901 also sends a request to the client device 902 to transmit sensor information.

[0594] Furthermore, the server 901 updates the three-dimensional map 1135 using the created three-dimensional data 1134 and sends the three-dimensional map 1135 to the client device 902 in response to a request from the client device 902 to send the three-dimensional map 1135.

[0595] Furthermore, the sensor information 1037 includes at least one of the following: information obtained from the laser sensor, brightness image, infrared image, depth image, sensor position information, and sensor velocity information.

[0596] Furthermore, sensor information 1037 includes information indicating the performance of the sensor.

[0597] Furthermore, the server 901 corrects the three-dimensional data according to the performance of the sensor. This allows the three-dimensional data creation method to improve the quality of the three-dimensional data.

[0598] Furthermore, when receiving sensor information, the server 901 receives multiple pieces of sensor information 1037 from multiple client devices 902, and selects the sensor information 1037 to be used to create the three-dimensional data 1134 based on the multiple pieces of information indicating the performance of the sensors contained in the multiple pieces of sensor information 1037. In this way, the server 901 can improve the quality of the three-dimensional data 1134.

[0599] Furthermore, the server 901 decodes or decodes the received sensor information 1037 and creates three-dimensional data 1134 from the decoded or decoded sensor information 1132. This allows the server 901 to reduce the amount of data transmitted.

[0600] For example, server 901 is equipped with a processor and memory, and the processor uses the memory to perform the above processing.

[0601] (Embodiment 8) This embodiment describes a modified version of Embodiment 7 described above. Figure 54 is a diagram showing the configuration of the system according to this embodiment. The system shown in Figure 54 includes a server 2001, a client device 2002A, and a client device 2002B.

[0602] Client devices 2002A and 2002B are mounted on a moving object such as a vehicle and transmit sensor information to server 2001. Server 2001 transmits a three-dimensional map (point cloud) to client devices 2002A and 2002B.

[0603] Client device 2002A comprises a sensor information acquisition unit 2011, a storage unit 2012, and a data transmission feasibility determination unit 2013. The configuration of client device 2002B is similar. Furthermore, in the following description, unless otherwise specified, client device 2002A and client device 2002B will also be referred to as client device 2002.

[0604] Figure 55 is a flowchart showing the operation of the client device 2002 according to this embodiment.

[0605] The sensor information acquisition unit 2011 acquires various sensor information using sensors (sensor group) mounted on the mobile body. In other words, the sensor information acquisition unit 2011 acquires sensor information indicating the surrounding conditions of the mobile body obtained by sensors (sensor group) mounted on the mobile body. The sensor information acquisition unit 2011 also stores the acquired sensor information in the storage unit 2012. This sensor information includes at least one of LiDAR acquisition information, visible light images, infrared images, and depth images. The sensor information may also include at least one of sensor position information, velocity information, acquisition time information, and acquisition location information. Sensor position information indicates the position of the sensor that acquired the sensor information. Velocity information indicates the velocity of the mobile body when the sensor acquired the sensor information. Acquisition time information indicates the time when the sensor information was acquired by the sensor. Acquisition location information indicates the position of the mobile body or sensor when the sensor information was acquired by the sensor.

[0606] Next, the data transmission feasibility determination unit 2013 determines whether the mobile device (client device 2002) is in an environment where it can transmit sensor information to the server 2001 (S2002). For example, the data transmission feasibility determination unit 2013 may use information such as GPS to identify the location and time of the client device 2002 and determine whether data can be transmitted. Alternatively, the data transmission feasibility determination unit 2013 may determine whether data can be transmitted based on whether it can connect to a specific access point.

[0607] If the client device 2002 determines that the moving object is in an environment where it can transmit sensor information to the server 2001 (Yes in S2002), it transmits the sensor information to the server 2001 (S2003). In other words, as soon as the client device 2002 is in a situation where it can transmit sensor information to the server 2001, it transmits the sensor information it holds to the server 2001. For example, suppose a millimeter-wave access point capable of high-speed communication is installed at an intersection. When the client device 2002 enters the intersection, it transmits the sensor information it holds to the server 2001 at high speed using millimeter-wave communication.

[0608] Next, the client device 2002 deletes the sensor information already transmitted to the server 2001 from the storage unit 2012 (S2004). The client device 2002 may also delete sensor information that has not yet been transmitted to the server 2001 if it meets certain conditions. For example, the client device 2002 may delete the sensor information from the storage unit 2012 when the acquisition time of the stored sensor information becomes older than a certain time from the current time. In other words, the client device 2002 may delete the sensor information from the storage unit 2012 if the difference between the time the sensor information was acquired by the sensor and the current time exceeds a predetermined time. Furthermore, the client device 2002 may delete the sensor information from the storage unit 2012 when the acquisition location of the stored sensor information is more than a certain distance from the current location. In other words, the client device 2002 may delete the sensor information from the storage unit 2012 if the difference between the position of the moving object or sensor when the sensor information was acquired by the sensor and the current position of the moving object or sensor exceeds a predetermined distance. This makes it possible to reduce the capacity of the storage unit 2012 of the client device 2002.

[0609] If client device 2002 has not finished acquiring sensor information (No in S2005), client device 2002 repeats the processing from step S2001 onwards. If client device 2002 has finished acquiring sensor information (Yes in S2005), client device 2002 terminates processing.

[0610] Furthermore, the client device 2002 may select the sensor information to send to the server 2001 according to the communication status. For example, if high-speed communication is possible, the client device 2002 will prioritize sending sensor information with a large size stored in the memory unit 2012 (e.g., LiDAR acquisition information). If high-speed communication is difficult, the client device 2002 will send sensor information with a small size stored in the memory unit 2012 and a high priority (e.g., visible light images). This allows the client device 2002 to efficiently send the sensor information stored in the memory unit 2012 to the server 2001 according to the network status.

[0611] Furthermore, the client device 2002 may obtain time information indicating the current time and location information indicating the current location from the server 2001. The client device 2002 may also determine the acquisition time and location of sensor information based on the acquired time and location information. In other words, the client device 2002 may obtain time information from the server 2001 and generate acquisition time information using the acquired time information. Furthermore, the client device 2002 may obtain location information from the server 2001 and generate acquisition location information using the acquired location information.

[0612] For example, regarding time information, the server 2001 and client device 2002 synchronize their time using a mechanism such as NTP (Network Time Protocol) or PTP (Precision Time Protocol). This allows client device 2002 to obtain accurate time information. Furthermore, since the server 2001 can synchronize time between multiple client devices, the time in sensor information acquired by different client devices 2002 can be synchronized. Therefore, the server 2001 can handle sensor information that indicates the synchronized time. Note that any method other than NTP or PTP can be used for time synchronization. Also, GPS information may be used as the above time information and location information.

[0613] Server 2001 may obtain sensor information from multiple client devices 2002 by specifying the time or location. For example, in the event of an accident, to find clients that were nearby, Server 2001 broadcasts a sensor information transmission request to multiple client devices 2002, specifying the time and location of the accident. Client devices 2002 that have sensor information for the corresponding time and location then transmit the sensor information to Server 2001. In other words, Client device 2002 receives a sensor information transmission request from Server 2001 that includes specification information specifying the location and time. If Client device 2002 determines that the storage unit 2012 has stored sensor information obtained at the location and time indicated by the specification information, and that the moving object is in an environment where it can transmit sensor information to Server 2001, it transmits the sensor information obtained at the location and time indicated by the specification information to Server 2001. As a result, Server 2001 can obtain sensor information related to the occurrence of an accident from multiple client devices 2002 and use it for accident analysis, etc.

[0614] Furthermore, the client device 2002 may refuse to transmit sensor information when it receives a request from the server 2001 to transmit sensor information. Alternatively, the client device 2002 may pre-configure which of the multiple sensor information requests it can transmit. Or, the server 2001 may query the client device 2002 each time to determine whether or not to transmit sensor information.

[0615] Furthermore, client devices 2002 that transmit sensor information to server 2001 may be awarded points. These points can be used to pay for things like gasoline, electric vehicle (EV) charging fees, highway tolls, or rental car fees. Also, after acquiring sensor information, server 2001 may delete information that identifies the client device 2002 that sent the sensor information. For example, this information could be the network address of client device 2002. This anonymizes the sensor information, allowing users of client device 2002 to confidently transmit sensor information from client device 2002 to server 2001. Server 2001 may also consist of multiple servers. For example, by sharing sensor information among multiple servers, even if one server fails, other servers can communicate with client device 2002. This prevents service interruptions due to server failures.

[0616] Furthermore, the specified location in the sensor information transmission request indicates the location where the accident occurred, and may differ from the location of the client device 2002 at the specified time specified in the sensor information transmission request. Therefore, the server 2001 can request information acquisition from client devices 2002 located within a specified range, such as within XXm of the location, by specifying a range such as within XXm of the location. Similarly, for the specified time, the server 2001 may specify a range such as within N seconds before or after a certain time. This allows the server 2001 to acquire sensor information from client devices 2002 that were located within XXm of absolute position S at "time: tN to t+N". When the client device 2002 transmits three-dimensional data such as LiDAR, it may transmit data generated immediately after time t.

[0617] Furthermore, the server 2001 may separately specify information indicating the location of the client device 2002 from which sensor information is to be acquired, and the location where the sensor information is desired. For example, the server 2001 specifies that sensor information including at least the range YYm from absolute position S should be acquired from a client device 2002 located within XXm of absolute position S. When the client device 2002 selects the three-dimensional data to transmit, it selects one or more randomly accessible units of three-dimensional data so as to include at least the sensor information within the specified range. Also, when the client device 2002 transmits a visible light image, it may transmit multiple temporally consecutive image data, including at least the frame immediately before or after time t.

[0618] If the client device 2002 can utilize multiple physical networks, such as 5G, WiFi, or multiple modes in 5G, for transmitting sensor information, the client device 2002 may select the network to use according to the priority notified by the server 2001. Alternatively, the client device 2002 may select a network that can secure appropriate bandwidth based on the size of the data to be transmitted. Alternatively, the client device 2002 may select a network to use based on the cost of data transmission, etc. Furthermore, the transmission request from the server 2001 may include information indicating a transmission deadline, such as transmitting if the client device 2002 can start transmitting by time T. If sufficient sensor information is not obtained within the deadline, the server 2001 may issue another transmission request.

[0619] The sensor information may include compressed or uncompressed sensor data, along with header information indicating the characteristics of the sensor data. The client device 2002 may transmit the header information to the server 2001 via a different physical network or communication protocol than the sensor data. For example, the client device 2002 transmits the header information to the server 2001 prior to transmitting the sensor data. The server 2001 determines whether to acquire the sensor data from the client device 2002 based on the analysis results of the header information. For example, the header information may include information indicating the LiDAR point cloud acquisition density, elevation angle, or frame rate, or the resolution, signal-to-noise ratio, or frame rate of the visible light image. This allows the server 2001 to acquire sensor information from the client device 2002 that has sensor data of the determined quality.

[0620] As described above, the client device 2002 is mounted on the mobile body and acquires sensor information indicating the surrounding conditions of the mobile body obtained by sensors mounted on the mobile body, and stores the sensor information in the storage unit 2012. The client device 2002 determines whether the mobile body is in an environment where it can transmit sensor information to the server 2001, and if it determines that the mobile body is in an environment where it can transmit sensor information to the server, it transmits the sensor information to the server 2001.

[0621] Furthermore, the client device 2002 creates three-dimensional data of the surroundings of the moving object from the sensor information and uses the created three-dimensional data to estimate the self-position of the moving object.

[0622] Furthermore, the client device 2002 sends a request to the server 2001 to send a 3D map, and receives the 3D map from the server 2001. In estimating its own position, the client device 2002 uses the 3D data and the 3D map to estimate its own position.

[0623] Furthermore, the processing performed by the client device 2002 may be implemented as an information transmission method in the client device 2002.

[0624] Furthermore, the client device 2002 includes a processor and memory, and the processor may use the memory to perform the above processing.

[0625] Next, the sensor information collection system according to this embodiment will be described. Figure 56 is a diagram showing the configuration of the sensor information collection system according to this embodiment. As shown in Figure 56, the sensor information collection system according to this embodiment includes terminal 2021A, terminal 2021B, communication device 2022A, communication device 2022B, network 2023, data collection server 2024, map server 2025, and client device 2026. Note that terminal 2021A and terminal 2021B will also be referred to as terminal 2021 unless specifically distinguished. Similarly, communication device 2022A and communication device 2022B will also be referred to as communication device 2022 unless specifically distinguished.

[0626] The data collection server 2024 collects data such as sensor data obtained from sensors on terminal 2021 as position-related data that is associated with its position in three-dimensional space.

[0627] Sensor data refers to data acquired using sensors installed in terminal 2021, such as the surrounding conditions of terminal 2021 or the internal conditions of terminal 2021. Terminal 2021 transmits sensor data collected from one or more sensor devices located in a position where it can communicate directly with terminal 2021, or where it can communicate via one or more relay devices using the same communication method, to the data acquisition server 2024.

[0628] The location-related data may include, for example, information indicating the operating status of the terminal itself or the equipment installed on the terminal, operation logs, and service usage status. Furthermore, the location-related data may include information linking the identifier of terminal 2021 to the location or travel path of terminal 2021.

[0629] The location-related data contains location information that corresponds to location information in three-dimensional data, such as three-dimensional map data. Details of the location information will be described later.

[0630] Location-related data may include, in addition to location information which indicates a location, at least one of the following: time information as described above, and information indicating the attributes of the data included in the location-related data, or the type of sensor that generated the data (e.g., model number). Location information and time information may be stored in the header area of ​​the location-related data or in the header area of ​​the frame that stores the location-related data. Alternatively, location information and time information may be transmitted and / or stored separately from the location-related data as metadata associated with the location-related data.

[0631] The map server 2025 is connected to, for example, the network 2023 and transmits three-dimensional data, such as three-dimensional map data, in response to requests from other devices, such as the terminal 2021. Furthermore, as described in each of the embodiments above, the map server 2025 may also have a function to update the three-dimensional data using sensor information transmitted from the terminal 2021.

[0632] The data collection server 2024 is connected to network 2023, for example, and collects location-related data from other devices such as terminal 2021. It stores the collected location-related data in a storage device, either internally or on another server. The data collection server 2024 also transmits the collected location-related data or metadata of 3D map data generated based on the location-related data to terminal 2021 upon request from terminal 2021.

[0633] Network 2023 is a communication network, such as the Internet. Terminal 2021 is connected to Network 2023 via communication device 2022. Communication device 2022 communicates with terminal 2021 by switching between one or more communication methods. Communication device 2022 is, for example, (1) a base station such as LTE (Long Term Evolution), (2) an access point (AP) such as WiFi or millimeter wave communication, (3) a gateway for an LPWA (Low Power Wide Area) Network such as SIGFOX, LoRaWAN or Wi-SUN, or (4) a communication satellite that communicates using a satellite communication method such as DVB-S2.

[0634] The base station may communicate with terminal 2021 using a method classified as NB-IoT (Narrow Band-IoT) or LPWA such as LTE-M, or it may communicate with terminal 2021 while switching between these methods.

[0635] Here, we take an example where terminal 2021 has the function to communicate with communication device 2022 that uses two types of communication methods, and communicates with map server 2025 or data collection server 2024 using one of these communication methods, or by switching between multiple communication methods and the communication device 2022 that is the direct communication partner. However, the configuration of the sensor information collection system and terminal 2021 is not limited to this. For example, terminal 2021 may not have the function to communicate using multiple communication methods, but may have the function to communicate using one of the communication methods. Also, terminal 2021 may support three or more communication methods. Furthermore, each terminal 2021 may support different communication methods.

[0636] Terminal 2021 has the configuration of, for example, the client device 902 shown in Figure 46. Terminal 2021 performs position estimation, such as its own position, using the received three-dimensional data. Terminal 2021 also generates position-related data by associating sensor data acquired from sensors with position information obtained through position estimation processing.

[0637] The location information added to location-related data indicates, for example, the position in the coordinate system used in the three-dimensional data. For example, the location information is a coordinate value expressed as latitude and longitude. In this case, terminal 2021 may include information indicating the coordinate system on which the coordinate value is based, and the three-dimensional data used for position estimation, along with the coordinate value itself. The coordinate value may also include altitude information.

[0638] Furthermore, location information may be associated with data units or spatial units that can be used to encode the three-dimensional data described above. These units include, for example, WLD, GOS, SPC, VLM, or VXL. In this case, location information is represented by an identifier that identifies a data unit such as an SPC that corresponds to location-related data. In addition to the identifier that identifies a data unit such as an SPC, location information may also include information indicating three-dimensional data encoded from the three-dimensional space containing the data unit such as the SPC, or information indicating a detailed location within the SPC. Information indicating three-dimensional data is, for example, the file name of the three-dimensional data.

[0639] Thus, by generating location-related data associated with location information based on position estimation using three-dimensional data, this system can add more accurate location information to sensor information than when adding location information based on the self-position of a client device (terminal 2021) acquired using GPS. As a result, even when other devices use the location-related data in other services, it may be possible to more accurately identify the location corresponding to the location-related data in real space by performing position estimation based on the same three-dimensional data.

[0640] In this embodiment, the example of data transmitted from terminal 2021 being location-related data was used for explanation. However, data transmitted from terminal 2021 may not be associated with location information. In other words, the transmission and reception of three-dimensional data or sensor data described in other embodiments may be performed via the network 2023 described in this embodiment.

[0641] Next, we will describe different examples of location information that indicates a position in three-dimensional or two-dimensional real space or map space. The location information attached to location-related data may also be information that indicates the relative position to a feature point in three-dimensional data. Here, the feature point that serves as the basis for the location information is, for example, a feature point that is encoded as SWLD and notified to terminal 2021 as three-dimensional data.

[0642] Information indicating the relative position to a feature point may be represented, for example, by a vector from the feature point to the point indicated by the position information, and may also be information indicating the direction and distance from the feature point to the point indicated by the position information. Alternatively, information indicating the relative position to a feature point may be information indicating the displacement amounts of the X, Y, and Z axes from the feature point to the point indicated by the position information. Furthermore, information indicating the relative position to a feature point may be information indicating the distance from each of three or more feature points to the point indicated by the position information. Note that the relative position may not be the relative position of the point indicated by the position information expressed with respect to each feature point, but rather the relative position of each feature point expressed with respect to the point indicated by the position information. An example of position information based on the relative position to a feature point includes information for identifying the reference feature point and information indicating the relative position of the point indicated by the position information with respect to that feature point. Furthermore, if information indicating the relative position to a feature point is provided separately from the three-dimensional data, the information indicating the relative position to a feature point may include the coordinate axes used to derive the relative position, information indicating the type of three-dimensional data, and / or information indicating the magnitude per unit quantity (scale, etc.) of the value of the information indicating the relative position.

[0643] Furthermore, the location information may include information indicating the relative position of multiple feature points to each feature point. When the location information is represented by the relative position to multiple feature points, terminal 2021, which is trying to identify the location indicated by the location information in real space, may calculate candidate points for the location indicated by the location information from the position of each feature point estimated from the sensor data, and determine that the point obtained by averaging the calculated candidate points is the point indicated by the location information. With this configuration, the influence of errors when estimating the position of feature points from sensor data can be reduced, and thus the accuracy of estimating the point indicated by the location information in real space can be improved. In addition, if the location information includes information indicating the relative position to multiple feature points, even if there are feature points that cannot be detected due to constraints such as the type or performance of the sensors that terminal 2021 has, it is possible to estimate the value of the point indicated by the location information if even one of the multiple feature points can be detected.

[0644] Points identifiable from sensor data can be used as feature points. Points identifiable from sensor data are, for example, points or points within a region that satisfy predetermined conditions for feature point detection, such as the aforementioned three-dimensional features or visible light data features being above a threshold.

[0645] Furthermore, markers placed in real space may be used as feature points. In this case, the markers only need to be detectable and their location determined from data acquired using sensors such as LiDAR or cameras. For example, markers can be represented by changes in color or brightness values ​​(reflectance), or by three-dimensional shapes (such as bumps and depressions). Alternatively, coordinate values ​​indicating the position of the marker, or a two-dimensional code or barcode generated from the identifier of the marker may be used.

[0646] Furthermore, a light source that transmits optical signals may be used as a marker. When an optical signal light source is used as a marker, not only information for obtaining location, such as coordinate values ​​or identifiers, but also other data may be transmitted by the optical signal. For example, the optical signal may include information such as the content of the service corresponding to the location of the marker, an address such as a URL for obtaining the content, or an identifier of a wireless communication device for receiving the service, and a wireless communication method for connecting to the wireless communication device. By using an optical communication device (light source) as a marker, it becomes easier to transmit data other than location information, and it becomes possible to dynamically switch such data.

[0647] Terminal 2021 grasps the correspondence between feature points in different data sets, for example, by using an identifier commonly used between the data sets, or by using information or a table that shows the correspondence between feature points in the data sets. If there is no information showing the correspondence between feature points, terminal 2021 may determine that the feature point that is closest when the coordinates of a feature point in one three-dimensional data set are converted to a position in the three-dimensional data space of the other is the corresponding feature point.

[0648] As described above, when using location information based on relative position, even between terminals 2021 or services using different three-dimensional data, the location indicated by the location information can be identified or estimated based on common feature points included in or associated with each three-dimensional data. As a result, it becomes possible to identify or estimate the same location with higher accuracy between terminals 2021 or services using different three-dimensional data.

[0649] Furthermore, even when using map data or 3D data represented using different coordinate systems, the impact of errors associated with coordinate system transformations can be reduced, enabling the integration of services based on more accurate location information.

[0650] The following describes examples of functions provided by data collection server 2024. Data collection server 2024 may transfer the received location-related data to other data servers. If there are multiple data servers, data collection server 2024 will determine which data server to transfer the received location-related data to and will transfer the location-related data to the data server determined to be the destination.

[0651] The data collection server 2024 determines the transfer destination based, for example, on the destination server determination rules pre-configured on the data collection server 2024. The destination server determination rules are configured, for example, in a destination table that associates identifiers associated with each terminal 2021 with the destination data servers.

[0652] Terminal 2021 adds an identifier associated with terminal 2021 to the location-related data it transmits and sends it to the data collection server 2024. The data collection server 2024 identifies the destination data server corresponding to the identifier added to the location-related data based on destination server determination rules using a destination table or the like, and sends the location-related data to the identified data server. The destination server determination rules may also be specified using determination conditions such as the time or place where the location-related data was acquired. Here, the identifier associated with the sending terminal 2021 mentioned above is, for example, an identifier unique to each terminal 2021, or an identifier indicating the group to which terminal 2021 belongs.

[0653] Furthermore, the destination table does not necessarily have to directly associate identifiers associated with the source terminal with destination data servers. For example, data collection server 2024 maintains a management table that stores tag information assigned to each identifier unique to terminal 2021, and a destination table that associates said tag information with destination data servers. Data collection server 2024 may use the management table and the destination table to determine the destination data server based on the tag information. Here, the tag information is, for example, management control information or service provision control information assigned to the type, model number, owner, group to which it belongs, or other identifiers corresponding to the identifier of terminal 2021. Also, instead of identifiers associated with the source terminal 2021, a unique identifier for each sensor may be used in the destination table. Furthermore, the rules for determining the destination server may be set from client device 2026.

[0654] The data collection server 2024 may determine multiple data servers as destinations and transfer the received location-related data to those multiple data servers. With this configuration, for example, when automatically backing up location-related data, or when it is necessary to send location-related data to data servers that provide each service in order to use the location-related data in common across different services, the intended data transfer can be achieved by changing the settings of the data collection server 2024. As a result, the man-hours required for system construction and modification can be reduced compared to setting destinations for location-related data on individual terminals 2021.

[0655] The data acquisition server 2024 may, in response to a transfer request signal received from a data server, register the data server specified in the transfer request signal as a new transfer destination and transfer any subsequent location-related data to that data server.

[0656] The data collection server 2024 may store the location-related data received from terminal 2021 in a recording device, and in response to a transmission request signal received from terminal 2021 or the data server, it may transmit the location-related data specified in the transmission request signal to the requesting terminal 2021 or data server.

[0657] The data collection server 2024 may determine whether it is possible to provide location-related data to the requesting data server or terminal 2021, and if it is determined that it is possible to provide the data, it may transfer or transmit the location-related data to the requesting data server or terminal 2021.

[0658] If the data collection server 2024 receives a request for current location-related data from client device 2026, it may request terminal 2021 to send location-related data even if it is not the timing for terminal 2021 to send location-related data, and terminal 2021 may send location-related data in response to that request.

[0659] In the above description, it was assumed that terminal 2021 transmits location information data to the data collection server 2024. However, the data collection server 2024 may also have functions necessary for collecting location-related data from terminal 2021, such as functions for managing terminal 2021, or functions used when collecting location-related data from terminal 2021.

[0660] The data collection server 2024 may also have the function of sending a data request signal to terminal 2021 requesting the transmission of location information data and collecting location-related data.

[0661] The data collection server 2024 has pre-registered management information, such as an address for communicating with the terminal 2021 that is the target of data collection, or a terminal 2021-specific identifier. Based on the registered management information, the data collection server 2024 collects location-related data from terminal 2021. The management information may also include information such as the type of sensor that terminal 2021 has, the number of sensors that terminal 2021 has, and the communication method that terminal 2021 supports.

[0662] The data collection server 2024 may collect information from terminal 2021, such as its operating status or current location.

[0663] The registration of management information may be performed by the client device 2026, or the registration process may be initiated when terminal 2021 sends a registration request to the data collection server 2024. The data collection server 2024 may have a function to control communication with terminal 2021.

[0664] The communication between the data collection server 2024 and the terminal 2021 may be a dedicated line provided by a service provider such as an MNO (Mobile Network Operator) or MVNO (Mobile Virtual Network Operator), or a virtual dedicated line configured with a VPN (Virtual Private Network). This configuration allows for secure communication between the terminal 2021 and the data collection server 2024.

[0665] The data collection server 2024 may have a function to authenticate terminal 2021 or a function to encrypt data transmitted to and from terminal 2021. Here, the authentication process of terminal 2021 or the data encryption process is performed using an identifier unique to terminal 2021 or an identifier unique to a group of terminals including multiple terminals 2021, which has been shared in advance between the data collection server 2024 and terminal 2021. This identifier is, for example, the IMSI (International Mobile Subscriber Identity), which is a unique number stored on a SIM (Subscriber Identity Module) card. The identifier used for authentication and the identifier used for data encryption may be the same or different.

[0666] Authentication or data encryption between the data collection server 2024 and terminal 2021 can be provided if both the data collection server 2024 and terminal 2021 have the functionality to perform such processing, and is independent of the communication method used by the relaying communication device 2022. Therefore, a common authentication or encryption process can be used without considering the communication method used by terminal 2021, improving the convenience of system construction for users. However, "independent of the communication method used by the relaying communication device 2022" means that it is not essential to change it according to the communication method. In other words, for the purpose of improving transmission efficiency or ensuring security, the authentication or data encryption process between the data collection server 2024 and terminal 2021 may be switched according to the communication method used by the relaying device.

[0667] The data collection server 2024 may provide the client device 2026 with a UI that manages data collection rules, such as the type of location-related data to be collected from terminal 2021 and the data collection schedule. This allows the user to specify the terminal 2021 from which to collect data, as well as the data collection time and frequency, using the client device 2026. The data collection server 2024 may also specify a map area from which to collect data and collect location-related data from terminal 2021 included in that area.

[0668] When data collection rules are managed on a per-terminal basis (terminal 2021), the client device 2026 displays a list of the terminals or sensors to be managed on its screen. The user then sets whether data collection is necessary or the collection schedule for each item in the list.

[0669] When specifying an area on a map from which data should be collected, the client device 2026 displays, for example, a two-dimensional or three-dimensional map of the area to be managed. The user selects the area from which data should be collected on the displayed map. The area selected on the map may be a circular or rectangular area centered on a point specified on the map, or a circular or rectangular area that can be identified by dragging. Alternatively, the client device 2026 may select an area using predefined units such as a city, an area within a city, a block, or a major road. Furthermore, instead of specifying an area using a map, an area may be set by entering latitude and longitude values, or an area may be selected from a list of candidate areas derived based on entered text information. The text information may include, for example, the names of regions, cities, or landmarks.

[0670] Furthermore, data collection may be performed while dynamically changing the specified area by having the user specify one or more terminals 2021 and set conditions such as within a 100-meter radius around those terminals 2021.

[0671] Furthermore, if the client device 2026 is equipped with sensors such as a camera, a map area may be specified based on the real-world position of the client device 2026 obtained from the sensor data. For example, the client device 2026 may estimate its own position using the sensor data and specify a data collection area within a predetermined distance from a point on the map corresponding to the estimated position, or within a distance specified by the user. Alternatively, the client device 2026 may specify the sensing area of ​​the sensor, i.e., the area corresponding to the acquired sensor data, as the data collection area. Or, the client device 2026 may specify a data collection area based on the position corresponding to the sensor data specified by the user. The estimation of the map area or position corresponding to the sensor data may be performed by the client device 2026 or by the data collection server 2024.

[0672] When specifying an area on a map, the data collection server 2024 may collect the current location information of each terminal 2021 to identify terminals 2021 within the specified area and request the identified terminals 2021 to transmit location-related data. Alternatively, instead of the data collection server 2024 identifying terminals 2021 within the area, the data collection server 2024 may send information indicating the specified area to terminals 2021, and terminals 2021 may determine whether they are within the specified area and transmit location-related data if they are determined to be within the specified area.

[0673] The data collection server 2024 transmits data such as lists or maps to the client device 2026 to provide the aforementioned UI (User Interface) in the application run by the client device 2026. The data collection server 2024 may transmit not only data such as lists or maps, but also the application program to the client device 2026. Furthermore, the aforementioned UI may be provided as content created in HTML or similar format that can be displayed in a browser. Note that some data, such as map data, may be provided by servers other than the data collection server 2024, such as the map server 2025.

[0674] When the client device 2026 receives input that indicates completion of input, such as when the user presses a settings button, it sends the entered information as configuration information to the data collection server 2024. Based on the configuration information received from the client device 2026, the data collection server 2024 sends a signal to each terminal 2021 requesting location-related data or notifying them of location-related data collection rules, and then collects location-related data.

[0675] Next, we will describe an example of controlling the operation of terminal 2021 based on additional information added to three-dimensional or two-dimensional map data.

[0676] In this configuration, object information indicating the location of a power supply unit, such as a wireless power supply antenna or power supply coil embedded in a road or parking lot, is included in or associated with three-dimensional data and provided to terminal 2021, such as a car or drone.

[0677] A vehicle or drone that has acquired object information in order to charge will automatically move its position so that the charging part, such as the charging antenna or charging coil, is positioned opposite the area indicated by the object information, and will begin charging. In the case of a vehicle or drone that does not have an autonomous driving function, the driver or operator will be presented with the direction to move or the operation to perform using images or sounds displayed on the screen. When it is determined that the position of the charging part, calculated based on the estimated self-position, has entered the area indicated by the object information or within a predetermined distance from that area, the images or sounds presented will switch to instructions to stop driving or operation, and charging will begin.

[0678] Furthermore, the object information may not be information indicating the location of the power supply unit, but rather information indicating a region where, if the charging unit is placed within that region, a charging efficiency of a predetermined threshold or higher can be obtained. The location of the object information may be represented by the center point of the region indicated by the object information, or by a region or line in a two-dimensional plane, or by a region, line or plane in three-dimensional space.

[0679] This configuration allows for the determination of the power supply antenna's location, which cannot be determined from LiDAR sensing data or camera footage. This enables more precise alignment between the wireless charging antenna on a vehicle or other device (2021) and the wireless power supply antenna embedded in the road or other infrastructure. As a result, the charging speed during wireless charging can be shortened, and charging efficiency can be improved.

[0680] The object information may include objects other than the power supply antenna. For example, the three-dimensional data may include the location of the AP for millimeter-wave wireless communication as object information. This allows terminal 2021 to know the AP's location in advance, and then direct the beam's direction towards the object information to initiate communication. As a result, improvements in communication quality can be achieved, such as increased transmission speed, reduced time to initiate communication, and extended communication period.

[0681] The object information may include information indicating the type of object corresponding to the object information. Furthermore, the object information may include information indicating the processing that terminal 2021 should perform if terminal 2021 is located within a real-space region corresponding to the three-dimensional data position of the object information, or within a predetermined distance from that region.

[0682] Object information may be provided from a server different from the server providing the 3D data. When object information is provided separately from the 3D data, object groups containing object information used in the same service may be provided as separate data depending on the type of service or equipment.

[0683] The three-dimensional data used in combination with object information may be either point cloud data from a WLD or feature point data from a SWLD.

[0684] The above describes the server and client devices, etc., according to the embodiments of this disclosure, but this disclosure is not limited to these embodiments.

[0685] Furthermore, each processing unit included in the server and client devices, etc., according to the above embodiment is typically implemented as an LSI (Large-Scale Integrated Circuit). These may be individually integrated into a single chip, or some or all of them may be integrated into a single chip.

[0686] Furthermore, integrated circuit implementation is not limited to LSIs; it may also be achieved using dedicated circuits or general-purpose processors. Field-Programmable Gate Arrays (FPGAs), which can be programmed after LSI manufacturing, or reconfigurable processors, which allow for the reconfiguration of the connections and settings of circuit cells within the LSI, may also be used.

[0687] Furthermore, in each of the above embodiments, each component may be implemented by being composed of dedicated hardware or by executing a software program suitable for each component. Each component may also be implemented by a program execution unit such as a CPU or processor reading and executing a software program recorded on a recording medium such as a hard disk or semiconductor memory.

[0688] Furthermore, this disclosure may be implemented as a method for creating three-dimensional data, etc., executed by a server and client device, etc.

[0689] Furthermore, the division of functional blocks in the block diagram is just one example; multiple functional blocks can be implemented as a single functional block, a single functional block can be divided into multiple parts, or some functions can be moved to other functional blocks. In addition, the functions of multiple functional blocks with similar functions can be processed in parallel or time-sharing by a single piece of hardware or software.

[0690] Furthermore, the order in which each step in the flowchart is performed is illustrative for the purpose of specifically illustrating this disclosure, and may be in a different order. Also, some of the above steps may be performed simultaneously (in parallel) with other steps.

[0691] Although one or more embodiments of server and client devices have been described above based on the embodiments, this disclosure is not limited to these embodiments. Without departing from the spirit of this disclosure, various modifications that a person skilled in the art could conceive of may be applied to these embodiments, and forms constructed by combining components from different embodiments may also be included within the scope of one or more embodiments. [Industrial applicability]

[0692] This disclosure is applicable to client devices, etc. [Explanation of Symbols]

[0693] 100, 400 3D data encoding device 101, 201, 401, 501 Acquisition Department 102, 402 Coding area determination unit 103 Split part 104, 644 Encoding section 111,607 3D data 112, 211, 413, 414, 511, 634 Encoded three-dimensional data 200, 500 Three-Dimensional Data Decoding Devices 202 Decryption Start GOS Determination Unit 203 Decryption SPC Determination Unit 204, 625 Decoding section 212, 512, 513 decoded 3D data 403 SWLD extraction part 404 WLD encoder 405 SWLD encoding section 411 Input 3D data 412 Extracted 3D data 502 Header Analysis Unit 503 WLD Decoding Unit 504 SWLD Decoding Unit 600 Own vehicle 601 Surrounding vehicles 602, 605 Sensor detection range 603, 606 area 604 Occlusion Region 620, 620A Three-Dimensional Data Creation Device 621, 641 Three-dimensional data creation department 622 Request Scope Determination Unit 623 Search Department 624, 642 Receiver 626 Synthesis section 627 Detection area determination unit 628 Surrounding Condition Detection Unit 629 Autonomous Operation Control Unit 631, 651 Sensor Information 632 First 3D Data 633 Request Scope Information 635 Second 3D Data 636 Third-Dimensional Data 637 Request signal 638 Transmitted data 639 Surrounding Conditions Detection Results 640, 640A Three-Dimensional Data Transmission Device 643 Extraction part 645 Transmitter 646 Transmission feasibility determination unit 652 Fifth 3D Data 654 Sixth 3D Data 700 Three-dimensional information processing device 701 Three-dimensional map acquisition unit 702 Vehicle detection data acquisition unit 703 Abnormal Case Determination Unit 704 Action Determination Unit 705 Operation Control Unit 711 Three-dimensional map 712 Vehicle detection 3D data 801 Vehicle 802 Space 810 Three-dimensional data creation device 811 Data receiving unit 812, 819 Communications Department 813 Receiving Control Unit 814, 821 Format conversion section 815 Sensor 816 Three-Dimensional Data Creation Department 817 Three-dimensional data synthesis unit 818 Three-dimensional data storage unit 820 Transmission Control Unit 822 Data Transmission Unit 831, 832, 834, 835, 836, 837 Three-dimensional data 833 Sensor Information 901 Server 902, 902A, 902B, 902C client devices 1011, 1111 Data receiving section 1012, 1020, 1112, 1120 Communications Department 1013, 1113 Receiving control unit 1014, 1019, 1114, 1119 Format conversion section 1015 Sensor 1016, 1116 Three-dimensional data creation department 1017 Three-dimensional image processing unit 1018, 1118 Three-dimensional data storage unit 1021, 1121 Transmitting control unit 1022, 1122 Data transmission section 1031, 1032, 1135 Three-dimensional maps Sensor information for 1033, 1037, and 1132. 1034, 1035, 1134 3D data 1117 Three-dimensional data synthesis unit 1201 Three-dimensional map compression / decryption processing unit 1202 Sensor Information Compression / Decoding Processing Unit 1211 Three-dimensional map decoding processing unit 1212 Sensor Information Compression Processing Unit 2001 Server 2002, 2002A, 2002B client devices 2011 Sensor Information Acquisition Unit 2012 Memory Department 2013 Data transmission eligibility determination unit 2021, 2021A, 2021B terminals 2022, 2022A, 2022B Communication equipment 2023 Network 2024 Data Collection Server 2025 Map Server 2026 Client Devices

Claims

1. An information provision system that provides a three-dimensional map showing the situation in three-dimensional space to a client device mounted on a mobile vehicle, A map database that holds a first three-dimensional map in which space is divided hierarchically by geographical region, A data generation unit generates a second three-dimensional map which has a smaller data size than the first three-dimensional map, and which is composed of spatial elements among the multiple spatial elements included in the first three-dimensional map that contain feature quantities above a predetermined threshold, A data transmission unit that receives a request from the client device, identifies either the first three-dimensional map or the second three-dimensional map based on the request information contained in the request, and transmits the identified first three-dimensional map or second three-dimensional map to the client device as the three-dimensional map, Equipped with, Information provision system.

2. The requested information indicates at least one of the following: the intended use of the three-dimensional map, the communication environment of the client device, and the speed of the moving object. The information provision system according to claim 1.

3. The client device transmits the request to the data transmission unit if (1) the data for the geographical area of ​​the three-dimensional map used by the client device for self-position estimation is outdated, (2) it is a certain time before the client device leaves the geographical route it is currently traveling, or (3) the error in alignment between the three-dimensional data generated by the client device and the received three-dimensional map is greater than a certain amount. The information provision system according to claim 1.

4. The first three-dimensional map held in the map database is created by the server using multiple sensor information transmitted to the server from multiple sensors mounted on multiple mobile bodies, and is automatically updated by the server in real time. The information provision system according to claim 1.

5. The client device comprises a processor and memory, The aforementioned processor, The first three-dimensional map or the second three-dimensional map transmitted from the data transmission unit is stored in the memory. Using the stored first three-dimensional map or the second three-dimensional map, at least one of the following is performed: self-position estimation, surrounding environment detection, and autonomous operation control. The information provision system according to claim 1.

6. A client device mounted on a mobile vehicle, A receiving unit that receives either a first three-dimensional map in which space is hierarchically divided into geographical regions, or a second three-dimensional map which is composed of spatial elements of the first three-dimensional map that contain feature quantities above a predetermined threshold, and has a smaller data size than the first three-dimensional map. A processing unit that uses the first three-dimensional map or the second three-dimensional map received by the receiving unit to perform at least one of the following: self-position estimation, surrounding environment detection, and autonomous operation control. A request unit requests either the first three-dimensional map or the second three-dimensional map based on at least one of the execution requirements among the self-position estimation, surrounding environment detection, and autonomous operation control. Equipped with, Client device.

7. A method for generating and updating a three-dimensional map that shows the situation in a three-dimensional space, The steps include: collecting multiple sensor information transmitted to the server from multiple sensors mounted on multiple mobile bodies by the server; The steps include: using the collected sensor information, the server creates a first three-dimensional map in which the space is hierarchically divided into geographical areas; The first three-dimensional map created above is automatically updated in real time by the server, The steps include: generating a second three-dimensional map using the server, which is composed of multiple spatial elements from among the multiple spatial elements included in the updated first three-dimensional map, each containing a feature quantity above a predetermined threshold, and having a smaller data size than the first three-dimensional map; The steps include transmitting the first three-dimensional map or the second three-dimensional map to a client device, including, Method for generating and updating three-dimensional maps.