Information processing device and method, and program

WO2026204442A1PCT designated stage Publication Date: 2026-10-01SONY SEMICON SOLUTIONS CORP
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
PCT/JP2026/009788
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2025-03-28
Filing Date
2026-03-12
Publication Date
2026-10-01

Smart Images

  • Figure JP2026009788_01102026_PF_FP_ABST
    Figure JP2026009788_01102026_PF_FP_ABST
Patent Text Reader

Abstract

The present disclosure relates to an information processing device and method, and a program that make it possible to more easily generate a training data set for generating a map generation AI model. According to the present disclosure: matching is performed between an existing point group representing a three-dimensional shape of an object of which the position is indicated by an existing map and a generation point group representing a three-dimensional shape of an object detected by a sensor; a generation map is generated on the basis of the matching result and the existing map; and a training data set including a generation image, the generation point group, and the generation map is generated. The present disclosure can be applied to, for example, an information processing device, an information processing method, a program, or the like.
Need to check novelty before this filing date? Find Prior Art

Description

Information processing apparatus and method, and program

[0001] The present disclosure relates to an information processing apparatus and method, and a program, and particularly relates to an information processing apparatus and method, and a program that enable easier generation of a learning dataset for generating a map generation AI model.

[0002] Conventionally, for example, there has been a method of controlling automatic driving of vehicles and the like using an HD (High Definition) map, which is map information including traffic information necessary for automatic driving (for example, paths indicating lanes, center lines, intersection connections, etc., traffic signs, connections between paths, associations between traffic signals and paths, etc.) (see, for example, Patent Document 1). However, the cost for generation and information management of HD maps is high, and it has been difficult to prepare HD maps for all roads.

[0003] Therefore, a method of generating a map near the vehicle using an in-vehicle sensor and controlling automatic driving using the map has been conceived. For example, a method of inferring a map (AI inference map) from the output (images, point clouds, etc.) of an in-vehicle sensor using an AI (Artificial Intelligence) model has been proposed.

[0004] In order to generate such an AI model, for example, it is necessary to perform learning such that the difference between the AI inference map (predicted map) obtained by the AI model and an actual map (for example, an existing HD map) is minimized, and to derive parameters of the AI model. In order to perform such learning, a learning dataset including images, point clouds, and a correct HD map (corresponding to the aforementioned "actual map") was required.

[0005] Japanese Unexamined Patent Publication No. 2020-34906

[0006] However, generating a correct HD map corresponding to images and point clouds requires complicated work, which may increase costs.

[0007] This disclosure is made in light of these circumstances and aims to make it easier to generate training datasets for creating map generation AI models.

[0008] One aspect of this technology is an information processing device comprising: a matching unit that matches an existing point cloud representing the three-dimensional shape of an object whose position is indicated by an existing map, which is pre-prepared map information, with a generated point cloud representing the three-dimensional shape of an object detected by a sensor; and a dataset generation unit that generates a generated map, which is map information corresponding to the generated point cloud, based on the matching result and the existing map, and generates a training dataset using the generated map.

[0009] One aspect of this technology is an information processing method that includes matching an existing point cloud representing the three-dimensional shape of an object whose location is indicated by an existing map, which is pre-prepared map information, with a generated point cloud representing the three-dimensional shape of an object detected by a sensor; generating a generated map, which is map information corresponding to the generated point cloud, based on the matching result and the existing map; and generating a training dataset using the generated map.

[0010] One aspect of this technology is a program that causes a computer to perform the following processes: matching an existing point cloud representing the three-dimensional shape of an object whose location is indicated by an existing map, which is pre-prepared map information, with a generated point cloud representing the three-dimensional shape of an object detected by a sensor; generating a generated map, which is map information corresponding to the generated point cloud, based on the matching result and the existing map; and generating a training dataset using the generated map.

[0011] In one aspect of this technology, the information processing device, method, and program, a matching is performed between an existing point cloud representing the three-dimensional shape of an object whose location is indicated by an existing map, which is pre-prepared map information, and a generated point cloud representing the three-dimensional shape of an object detected by a sensor. Based on the result of this matching and the existing map, a generated map, which is map information corresponding to the generated point cloud, is generated, and a training dataset is generated using this generated map.

[0012] This is a diagram illustrating the map generation AI model. This is a diagram illustrating the method for generating training datasets. This is a diagram showing the main configuration example of an autonomous driving system. This is a block diagram showing the main configuration example of a server. This is a block diagram showing the configuration example of a vehicle control system. This is a diagram showing an example of the sensing area of ​​the external recognition sensor of the vehicle control system in Figure 5. This is a functional block diagram showing the main configuration example of functions realized by the server. This is a flowchart showing an example of the detection process flow. This is a flowchart showing an example of the learning process flow. This is a flowchart showing an example of the inference process flow. This is a functional block diagram showing the main configuration example of functions realized by the self-localization unit. This is a flowchart showing an example of the detection learning process flow. This is a flowchart showing an example of the learning process flow. This is a flowchart showing an example of the detection learning process flow. This is a flowchart showing an example of the learning process flow. This is a flowchart showing an example of the detection learning process flow.

[0013] The following describes the embodiments for implementing this disclosure. The explanation will be given in the following order: 1. Technical content and supporting literature, etc. 2. Automated driving control using HD maps 3. Map prediction using point cloud matching 4. First embodiment (autonomous driving system) 5. Appendix

[0014] <1. Supporting Documents for Technical Content and Terminology> The scope disclosed in this technology includes not only the contents described in the embodiments, but also the contents described in the following patent documents, which were publicly known at the time of filing, and the contents of other documents referenced in the following patent documents.

[0015] Patent Document 1: (mentioned above)

[0016] In other words, the contents described in the aforementioned patent documents, as well as the contents of other documents referenced in those patent documents, will also serve as a basis for determining the support requirements.

[0017] <2. Automated Driving Control Using HD Maps> <HD Maps> Conventionally, there have been methods for controlling the automated driving of vehicles, etc., using a full-area HD (High Definition) map, such as the method described in Patent Document 1. An HD map is map information that includes traffic information necessary for automated driving, and is also called high-precision three-dimensional map data. An HD map may include information such as Paths that show lanes, center lines, intersection connections, traffic signs, connections between Paths, and the linking of traffic lights to Paths. In this specification, an HD map is defined as map information that includes a relatively wide variety of road information, including at least lanes, and is of relatively high precision. Furthermore, full-area map information is referred to as a global map, and a full-area HD map (HD map as a global map) is also referred to as a global HD map.

[0018] However, HD maps require a large amount of high-precision information, and their generation is costly. Furthermore, maintaining the information contained in the HD maps in an up-to-date state also incurs high costs. Therefore, even with the creation of a global HD map, it was practically impossible to cover information on all existing roads. Consequently, in the case of controlling autonomous driving based on the global HD map described above, control could not be achieved for roads for which an HD map was not available.

[0019] Therefore, a method was conceived to generate a map of the area around the vehicle using on-board sensors and to control autonomous driving using that map. For example, as shown in Figure 1A, a method was conceived to infer a map (AI inference map) from the output of on-board sensors (images, point clouds, etc.) using an AI (Artificial Intelligence) model. In the example shown in Figure 1A, on-board sensors such as image sensors mounted on the vehicle 11 generate generated images and generated point clouds. The generated images are captured images of the area around the vehicle 11, and the generated point clouds are point cloud data that represent the three-dimensional shapes of objects around the vehicle 11 with point-by-point position information (geometry) and attribute information (attributes). When these generated images and point clouds are input as input data 12 to the map generation AI model 13, an AI inference map 14 with information equivalent to an HD map is generated through inference. In other words, this AI inference map 14 is local map information of the area around the vehicle 11 that corresponds to the generated images and point clouds used as input data 12. In this specification, a local HD map (HD map as a local map) is also referred to as a local HD map. By using a local HD map based on generated images and point clouds generated in the vehicle, automated driving control becomes possible even on roads where a global HD map does not exist.

[0020] To generate such a map generation AI model 13, for example, as shown in Figure 1B, a training dataset 20 is input to the AI ​​training model 24 and trained. This training dataset 20 consists of input data 21 (i.e., generated images and generated point clouds) and a generated map 22 for the map generation AI model 13. This generated map 22 corresponds to the output data of the map generation AI model 13 when the input data 21 is input. This training is performed to minimize the difference between the AI ​​inference map (predicted map) obtained by the map generation AI model and the actual map (e.g., an existing HD map), so the generated map 22 is map information that corresponds to or approximates the "actual map". In other words, the generated map 22 is a correct local HD map of the area around the vehicle.

[0021] As shown in Figure 1C, a generated map 33 (local HD map) can be generated using the generated image 31 and the generated point cloud 32. However, conventional methods require complicated and time-consuming work, such as converting the generated point cloud 32 to a bird's-eye view (BEV), performing inverse perspective mapping (IPM) on the generated image 31 to convert it to a bird's-eye view (BEV), and then manually annotating these based on the results, which could increase the cost of generating the training dataset.

[0022] Furthermore, in order to obtain better learning results (to further reduce the difference between the AI ​​inference map obtained by the map generation AI model and the actual map), it is necessary to improve the accuracy of the training dataset (i.e., improve the accuracy of the generated map). In the conventional method described above, the accuracy of the generated map depends on the accuracy of the generated images and point clouds used, so higher accuracy generated images and point clouds were required to suppress the reduction in the accuracy of the generated map. In other words, it was necessary to generate generated images and point clouds using higher accuracy in-vehicle sensors. Also, since the accuracy of the generated map depends on the ability of the operator performing this work, higher level of know-how and skills were required of the operator to suppress the reduction in the accuracy of the generated map. In short, obtaining better learning results could potentially increase the cost of generating the training dataset even further.

[0023] Furthermore, if there were any changes to the specifications of the in-vehicle sensors used to generate the generated images and point clouds (for example, the sensor mounting position or sensor type), the generated images and point clouds would change, requiring the creation of the generation map again, which could further increase the workload. In other words, there was a risk that the cost of generating the training dataset would further increase.

[0024] Furthermore, even at the same location, generated images and point clouds can change due to changes in the domains involved in image and point cloud generation, such as weather (sunny, rainy, cloudy, etc.), time of day (morning, noon, night, etc.), season (spring, summer, autumn, winter, etc.), and imaging conditions (backlight, etc.). In order to make the map generation AI model adapt to these changes (domain shifts) and output a correct AI inference map regardless of which domain's generated images or point clouds are input, conventional methods required training using training datasets that corresponded to those domain shifts. In other words, it was necessary to prepare a training dataset for each domain, including generated images, generated point clouds, and generated maps, as described above, which could further increase the cost of generating training datasets.

[0025] <3. Map prediction using point cloud matching> <Utilization of existing maps based on matching results between generated point clouds and existing point clouds> Therefore, by matching existing point clouds with generated point clouds, a generated map is generated using the portion of the existing map corresponding to the existing point cloud that corresponds to the generated point cloud (local map), and a training dataset (generated images, generated point clouds, generated maps) is generated.

[0026] For example, as shown in Figure 2A, suppose a global HD map is prepared in advance as an existing map 44, and a point cloud corresponding to that global HD map is also prepared in advance as an existing point cloud 43. In this specification, the pre-prepared map information is also referred to as an existing map. Furthermore, the point cloud that corresponds to the pre-prepared existing map (that is, a point cloud representing the three-dimensional shape of an object whose position is indicated by the existing map) is also referred to as an existing point cloud.

[0027] Furthermore, it is assumed that the generated image 41 and the generated point cloud 42 are generated in the vehicle using on-board sensors, etc. In this specification, the image of the area around the vehicle captured using on-board sensors such as an image sensor is also referred to as the generated image. The point cloud representing the three-dimensional shape of objects around the vehicle detected by on-board sensors such as a Lidar (Light Detection And Ranging) sensor is also referred to as the generated point cloud. The generated image 41 and the generated point cloud 42 correspond to each other and contain information within the same range.

[0028] In such cases, matching is performed between the generated point cloud 42 and the existing point cloud 43 to identify the portion of the existing point cloud 43 that matches the generated point cloud 42. Then, the portion corresponding to the matched portion is extracted from the existing map 44 and becomes the generated map 45 (map information generated from the generated image 41 and the generated point cloud 42) corresponding to the generated image 41 and the generated point cloud 42. In this specification, the map information corresponding to the generated image and the generated map (i.e., generated from the generated image and the generated map) is also referred to as the generated map.

[0029] For example, the information processing device may include a matching unit that matches an existing point cloud representing the three-dimensional shape of an object whose location is indicated by an existing map, which is pre-prepared map information, with a generated point cloud representing the three-dimensional shape of an object detected by a sensor, and a dataset generation unit that generates a generated map, which is map information corresponding to the generated point cloud, based on the matching result and the existing map, and generates a training dataset using the generated map.

[0030] For example, the information processing method executed by the information processing device may include matching an existing point cloud, which represents the three-dimensional shape of an object whose location is indicated by an existing map (which is pre-prepared map information), with a generated point cloud, which represents the three-dimensional shape of an object detected by a sensor; generating a generated map, which is map information corresponding to the generated point cloud, based on the result of the matching and the existing map; and generating a training dataset using the generated map.

[0031] For example, the program may be designed to cause a computer to perform the following processes: matching an existing point cloud, which represents the 3D shape of an object whose location is indicated by an existing map (pre-prepared map information), with a generated point cloud, which represents the 3D shape of an object detected by a sensor; generating a generated map, which is map information corresponding to the generated point cloud, based on the matching results and the existing map; and generating a training dataset using that generated map.

[0032] By utilizing existing maps (global HD maps) in this way, generated maps can be created simply by matching point clouds and extracting HD maps corresponding to the matching results (i.e., without requiring complex tasks such as conversion to bird's-eye views or annotation, or the know-how and skills required for such tasks). In other words, it becomes easier to generate generated images and corresponding generated maps. Consequently, it becomes easier to generate training datasets for creating map generation AI models. To put it another way, the increase in the generation cost of those training datasets can be suppressed.

[0033] Furthermore, the accuracy of the generated map depends on the accuracy of the existing global HD map. Also, the accuracy of the generated map does not depend on the generated image or point cloud, as long as the result of point cloud matching remains unchanged (i.e., the regions extracted from the existing map as the generated map are the same). In other words, this method allows for the more stable generation of higher-accuracy generated maps. Therefore, training datasets for generating map generation AI models can be generated with higher accuracy and greater stability. Additionally, this method can suppress the increase in workload due to rework in generating maps caused by changes in vehicle sensor specifications, etc. Therefore, it can suppress the increase in the cost of generating training datasets.

[0034] The training dataset may include any information that is used for training. For example, as shown in Figure 1B, the training dataset 20 may include generated images generated using a sensor, a generated point cloud representing the three-dimensional shape of an object detected by the sensor, and a generated map generated based on the matching results between the existing point cloud and the generated point cloud and an existing map.

[0035] Furthermore, both the existing map and the generated map can be map information, and may contain any type of information. For example, the map area of ​​the existing map (the area where the map information is displayed) can be of any size, as long as it covers a wider area than the map area of ​​the generated map. The map area of ​​the generated map can be of any size, as long as it is included in the map area of ​​the existing map and includes the area corresponding to the generated image or generated point cloud. In other words, for example, the existing map may be a global map containing map information covering a relatively wide area, and the generated map may be a local map containing map information covering a narrower area than that global map. To put it another way, the map area shown by the global map can be of any size, as long as it covers a wider area than the local map.

[0036] Furthermore, the existing map and the generated map may each be an HD map or an SD (Standard Definition) map with less information than an HD map. In this specification, an SD map is defined as a map with relatively low accuracy, containing fewer types of road information than an HD map. For example, the generated map may have the same amount of information as the existing map, or it may have less or more information than the existing map. If the amount of information in the generated map is less than or equal to the amount of information in the existing map, it is easier to generate the generated map because no additional information is required. For example, the existing map and the generated map may both be HD maps, which contain a relatively wide variety of road information, including at least lanes, and are relatively high-accuracy map information. For example, the existing map may be a global HD map and the generated map may be a local HD map.

[0037] Furthermore, the generated image can be any 2D data generated using a sensor, such as a visible light image (RGB image), an infrared (IR) image, a radar image produced by radio wave irradiation, a thermographic image representing heat distribution, or any other 2D data. In other words, the sensor used to generate the generated image can be an image sensor that detects visible light (RGB image sensor), an infrared sensor that detects infrared rays, a radar sensor that detects reflected waves after irradiating radio waves, a temperature sensor that detects heat distribution, or any other sensor.

[0038] Furthermore, the generated point cloud can be any 3D data generated using a sensor, and may be a lidar point cloud generated based on lidar data, a radar point cloud generated based on radar data, or a point cloud generated based on other information. In other words, the sensor used to generate the generated point cloud may be a lidar sensor that emits laser light and detects the reflected light, a radar sensor, or any other sensor.

[0039] Note that existing point clouds and generated point clouds are not limited to point clouds; they can be any 3D data. For example, they may be 3D data other than point clouds, such as meshes. In this specification, point clouds will be used as an example of this 3D data.

[0040] <Point Cloud Matching> The method for matching the existing point cloud with the generated point cloud can be any method. For example, matching may be performed using a sub-region of the generated point cloud. For example, in an information processing device, the matching unit may perform matching using a sub-region of the generated point cloud.

[0041] For example, matching may be performed using the entire generated point cloud 51 shown in Figure 2B with an existing point cloud, or using the point cloud of a partial region 52 (within the dotted line frame), which is a part of the generated point cloud 51, or using the point cloud of the region excluding the partial region 53 (outside the dotted line frame), which is a part of the generated point cloud 51, with respect to the existing point cloud.

[0042] By performing matching using only a portion of the generated point cloud in this way, the amount of information used for matching can be reduced compared to using the entire generated point cloud. Therefore, the increase in processing load and processing time for matching can be suppressed. In other words, it becomes easier to generate training datasets for creating map generation AI models. To put it another way, the increase in the cost of generating those training datasets can be suppressed.

[0043] The region of the generated point cloud to be applied to this matching can be set by any method. For example, this region may be set based on the distribution pattern of points in the generated point cloud. For example, in an information processing device, the matching unit may set a region to be applied to matching with an existing point cloud based on the distribution pattern of points in the generated point cloud. Alternatively, image recognition may be performed on the generated image, and this region may be set based on the image recognition results. For example, in an information processing device, the matching unit may set a region to be applied to matching with an existing point cloud based on the image recognition results of the generated image generated by an image sensor.

[0044] For example, a region with a dense number of data may be applied to matching. In other words, a region with a sparse number of data may be excluded from matching targets (not applied to matching). For example, in an information processing apparatus, a matching unit may set a region in a generated point cloud where the point density is higher than a predetermined reference as a partial region to be applied to matching with an existing point cloud. For example, under conditions such as heavy rain, heavy snow, and fog, the detection accuracy of a sensor generally decreases, the obtained point cloud tends to decrease, and there is a risk that matching accuracy may decrease. However, even under such adverse conditions, the shape of nearby objects can be detected with relatively high accuracy, and a dense point cloud can be generated. Therefore, as described above, by applying a region with a dense number of data to matching (in other words, removing a region with a sparse number of data before applying to matching), a reduction in matching accuracy can be suppressed even under the aforementioned adverse conditions.

[0045] For example, in the example of FIG. 2B, a sufficiently high-density point cloud is obtained in the partial region 52 where the shape of a nearby object is detected. Therefore, by applying such a partial region 52 as a partial region to matching, a reduction in matching accuracy can be suppressed. Further, in the partial region 53, a point cloud with sufficient density is not obtained. Therefore, by removing such a partial region 53 and applying other regions to matching, a reduction in matching accuracy can be suppressed.

[0046] Furthermore, point clouds corresponding to objects whose shape or other characteristics may change over time may be excluded from the matching target. In other words, point clouds corresponding to objects whose shape or other characteristics do not change easily over time may be applied to the matching. For example, in an information processing device, the matching unit may set the region of the generated point cloud excluding the point cloud representing the three-dimensional shape of a deformable object as a part of the region to be applied to matching with the existing point cloud. In the real world, there are objects whose shape may change over time, such as plants, as they grow. The shape of such objects may reduce the accuracy of the matching due to differences in the generation timing between the existing point cloud and the generated point cloud. Therefore, as described above, by excluding point clouds corresponding to deformable objects and applying them to the matching (in other words, applying point clouds corresponding to objects that do not change easily to the matching), the reduction in matching accuracy can be suppressed.

[0047] Furthermore, point clouds corresponding to the road surface (point clouds with low height) may be excluded from the matching target. In other words, point clouds with sufficient height may be applied to the matching. For example, in an information processing device, the matching unit may set the area of ​​the generated point cloud excluding the point cloud representing the three-dimensional shape of the road surface as a part of the area to be applied to matching with the existing point cloud. For example, road surface conditions are easily affected by environmental factors such as heavy rain, heavy snow, and fog, making it relatively difficult to obtain a correct point cloud, which could reduce the accuracy of the matching. Therefore, as described above, by excluding point clouds corresponding to such road surfaces from the matching target and applying only point clouds with a certain height to the matching, the reduction in matching accuracy can be suppressed.

[0048] Furthermore, point clouds corresponding to moving objects (point clouds that move in the time direction) may be removed from matching targets. In other words, point clouds corresponding to static objects (point clouds that do not move in the time direction) may be applied to matching. For example, in an information processing apparatus, the matching unit may set a region excluding the point cloud representing the three-dimensional shape of a moving object in the generated point cloud as a partial region to be applied to matching with the existing point cloud. For example, when a generated point cloud includes a point cloud corresponding to a moving object such as a person or a vehicle, there is a risk that the accuracy of matching with an existing point cloud (which does not include such a moving object) may be reduced. Therefore, as described above, by removing the point cloud corresponding to the moving object from the matching targets and applying only the point cloud corresponding to the static object to matching, it is possible to suppress a decrease in matching accuracy.

[0049] <Correspondence to domain shift> The map generation AI model may be trained to adapt to domain shift. In that case, a training data set for each domain may be generated. For example, in an information processing apparatus, the data set generation unit may generate a training data set for each domain by associating a plurality of generated images and generated point clouds corresponding to different domains with a generated map.

[0050] For example, by performing matching between the generated point cloud and the existing point cloud for a daytime generated image and a nighttime generated image respectively, a training data set corresponding to the "daytime" time period and a training data set corresponding to the "night" time period can be generated. In addition, for generated images under various weather conditions such as sunny, cloudy, backlight, rain, and snow, by performing matching between the generated point cloud and the existing point cloud, training data sets corresponding to each weather condition can be generated. Furthermore, for generated images in each season such as spring, summer, autumn and winter, by performing matching between the generated point cloud and the existing point cloud, training data sets corresponding to each season can be generated.

[0051] By training using such domain-specific training datasets, it is possible to generate an AI model for map generation that can handle domain shifts. In this technology, as mentioned above, generated maps are created from existing maps using point cloud matching, so if the generated point cloud does not change for each domain, generated maps can be easily generated regardless of the domain of the generated image. Furthermore, even if the generated point cloud changes in response to a domain shift, if the matching can be performed in a way that suppresses the change in the generated point cloud due to that domain shift, for example by applying a portion of the generated point cloud to the matching process as described above, then generated maps can be easily generated regardless of the domain of the generated image. In other words, training datasets for each domain can be easily generated.

[0052] <Training Dataset> As mentioned above, the training dataset may include any information used for training, and is not limited to the example of combinations of generated images, generated point clouds, and generated maps of HD maps described above. For example, the training dataset may also include SD maps, which are map information with fewer types of road information than HD maps and are of relatively low accuracy. In other words, the training dataset may include generated images, generated point clouds, generated maps of HD maps, and SD maps corresponding to those HD maps (generated maps).

[0053] The method for generating this SD map can be anything. For example, the existing map corresponding to the existing point cloud may include the existing HD map of the HD map and the existing SD map of the SD map. Then, using the existing HD map and the existing SD map, a generated HD map of the HD map and a generated SD map of the SD map may be generated. For example, in an information processing device, a dataset generation unit may generate a generated HD map of the HD map and a generated SD map of the SD map as generated maps based on the matching results, the existing HD map and the existing SD map, and generate a training dataset including the generated image, the generated point cloud, the generated HD map and the generated SD map.

[0054] In this way, generated SD maps can be easily created. Typically, SD maps have more coordinate shift information than HD maps. Therefore, by intentionally introducing a positional shift into the generated SD map, it can be made suitable as an SD map for training. Note that existing SD maps may contain some of the information from existing HD maps. In that case, existing SD maps can contain any information from the information contained in existing HD maps. Also, generated SD maps may contain some of the information from generated HD maps. In that case, generated SD maps can contain any information from the information contained in generated HD maps.

[0055] Furthermore, the existing map may only include the existing HD map, and the generated SD map may be generated using the generated HD map generated by point cloud matching or the like. For example, in an information processing device, the dataset generation unit may generate a generated HD map based on the matching results and the existing HD map, extract some information from the generated HD map to generate a generated SD map, and generate a training dataset that includes the generated image, the generated point cloud, the generated HD map, and the generated SD map.

[0056] In this way, the generated SD map can be easily generated. Note that the existing SD map may contain some of the information from the existing HD map. In that case, the existing SD map can contain any information from the existing HD map. Similarly, the generated SD map may contain some of the information from the generated HD map. In that case, the generated SD map can contain any information from the generated HD map.

[0057] By training with a training dataset configured in this way, it is possible to generate a map generation AI model that takes generated images, generated point clouds, and SD maps as input and outputs an AI inference map.

[0058] Furthermore, the information processing device may also include a learning unit that uses the generated training dataset to train a map generation AI model, which is an AI model that infers map information from generated images and generated point clouds. In this way, a map generation AI model can be generated.

[0059] Furthermore, the information processing device may further include an acquisition unit that acquires generated images and generated point clouds generated by an image sensor. A matching unit may then perform matching between the acquired generated point clouds and existing point clouds, and a dataset generation unit may generate a training dataset including the acquired generated images and generated point clouds. In this way, a training dataset can be generated using generated images and generated point clouds generated by other devices.

[0060] Furthermore, the information processing device may further include an image generation unit that generates generated images and a point cloud generation unit that generates the generated point cloud. A matching unit may perform matching between the generated point cloud and an existing point cloud, and a dataset generation unit may generate a training dataset including the generated images and point cloud. In this way, a training dataset can be generated using the generated images and point cloud generated by the device itself.

[0061] <4. First Embodiment> <Automated Driving System> The technology described above can be applied to any configuration (device, system, unit, processing unit, etc.). Figure 3 is a diagram showing an example of the main configuration of an automated driving system, which is one embodiment of an information processing system to which this technology is applied.

[0062] The automated driving system 100 shown in Figure 3 comprises a server 101, vehicles 102-1, 102-2, and 102-3, which are connected to each other via a network 110 so as to be able to communicate with one another. When it is not necessary to distinguish between vehicles 102-1, 102-2, and 102-3, they will be referred to as vehicle 102.

[0063] The automated driving system 100 is a system that performs processing related to the control of the automated driving of the vehicle 102. In Figure 3, three vehicles 102 (vehicle 102-1, vehicle 102-2, and vehicle 102-3) are shown, but the number of vehicles 102 that the automated driving system 100 has may be any number, for example, two or fewer, or four or more. Also, in Figure 3, one server 101 is shown, but the number of servers 101 that the automated driving system 100 has may be any number, for example, two or more. Similarly, the automated driving system 100 may have multiple networks 110.

[0064] This network 110 is a communication network composed of any communication medium. Communication conducted through network 110 may be wired communication, wireless communication, or both. In other words, network 110 may be a communication network for wired communication, a communication network for wireless communication, or a communication network composed of both. Furthermore, network 110 may be composed of a single communication network or of multiple communication networks.

[0065] For example, the internet may be included in this network 110. Public telephone network may also be included in this network 110. Furthermore, wide-area communication networks for wireless mobile devices, such as so-called 3G and 4G networks, may also be included in this network 110. For example, LPWA (Low Power Wide Area) communication networks such as LTE-M, which enable long-distance data communication and have low power consumption, may also be included in this network 110. Furthermore, WAN (Wide Area Network) and LAN (Local Area Network) may also be included in this network 110. Furthermore, wireless communication networks that perform communication compliant with the Bluetooth® standard may also be included in this network 110. Near-field communication (NFC) communication channels may also be included in this network 110. Furthermore, infrared communication channels may also be included in this network 110. Finally, wired communication networks compliant with standards such as HDMI (High-Definition Multimedia Interface)® and USB (Universal Serial Bus)® may also be included in this network 110. Thus, the network 110 may include communication networks and communication channels of any communication standard. Furthermore, these communication networks and communication channels may include not only communication media such as cables, but also devices and circuits necessary for communication, such as communication equipment and relay equipment.

[0066] The server 101 and the vehicle 102 can communicate with other devices and exchange information via this network 110. In other words, the server 101 and the vehicle 102 may communicate in any manner conforming to any of the various communication standards described above. For example, the server 101 and the vehicle 102 may use wired communication, wireless communication, or both.

[0067] <Server> Server 101 performs processing related to the control of the autonomous driving of the vehicle 102. For example, Server 101 may perform processing related to learning to generate a map generation AI model used for autonomous driving in the vehicle 102. Server 101 may also supply the map generation AI model generated by the learning process to the vehicle 102.

[0068] Figure 4 is a block diagram showing an example of the hardware configuration of server 101. As shown in Figure 4, in server 101, the CPU (Central Processing Unit) 201, ROM (Read Only Memory) 202, and RAM (Random Access Memory) 203 are interconnected via bus 204.

[0069] An input / output interface 210 is also connected to the bus 204. An input / output interface 210 is connected to an input unit 211, an output unit 212, a storage unit 213, a communication unit 214, and a drive 215.

[0070] The input unit 211 consists of, for example, a keyboard, mouse, microphone, touch panel, and input terminals. The output unit 212 consists of, for example, a display, speaker, and output terminals. The storage unit 213 consists of, for example, a hard disk, RAM disk, and non-volatile memory. The communication unit 214 consists of, for example, a network interface. The drive 215 drives removable media 221 such as a magnetic disk, optical disk, magneto-optical disk, or semiconductor memory.

[0071] In the server 101 configured as described above, the CPU 201 loads, for example, a program stored in the memory unit 213 into the RAM 203 via the input / output interface 210 and the bus 204, and executes it, thereby performing a series of processes described later. The RAM 203 also appropriately stores data necessary for the CPU 201 to perform various processes.

[0072] The program executed by the server 101 can be recorded and applied on a removable media 221, such as a package media. In this case, the program can be installed on the storage unit 213 via the input / output interface 210 by inserting the removable media 221 into the drive 215.

[0073] Furthermore, this program can also be provided via wired or wireless transmission media such as local area networks, the internet, or digital satellite broadcasting. In that case, the program can be received by the communication unit 214 and installed in the storage unit 213.

[0074] In addition, this program can be pre-installed in ROM 202 or storage unit 213.

[0075] The hardware configuration of server 101 shown in Figure 3 is described assuming that server 101 is comprised of a single device. In reality, the functions of server 101 (i.e., the configuration shown in Figure 3) may be implemented by a single device or by multiple devices. Furthermore, server 101 may be configured as a so-called cloud server.

[0076] <Vehicle> Vehicle 102 performs processing related to autonomous driving. For example, vehicle 102 may generate generated images and generated point clouds using on-board sensors, etc. Vehicle 102 may also supply the generated images and generated point clouds to server 101. Vehicle 102 may also acquire a map generation AI model supplied from server 101. Vehicle 102 may also perform inference using the map generation AI model and generate a generated map from the generated images and generated point clouds. Vehicle 102 may also perform processing related to autonomous driving using the generated map.

[0077] The vehicle 102 may have a configuration as shown in Figure 5, for example. Figure 5 is a block diagram showing an example configuration of a vehicle control system 311, which is a non-limiting example of a mobile device control system to which the present technology is applied.

[0078] The vehicle control system 311 is installed in the vehicle 102 and performs processing related to the automation of the vehicle's operation. This automation includes driving automation from Level 1 to Level 5, as well as remote driving and / or remote assistance of the vehicle 102 by a remote driver. The levels of driving automation may refer to the Society of Automotive Engineers (SAE) J3016™ APL2021 Levels of Driving Automation, where SAE Level 0 represents the lowest level of driving automation and SAE Level 5 represents the highest level of driving automation. For example, SAE Level 1 driving automation may consist of driver assistance functions that provide the driver with steering or brake / acceleration support, and SAE Level 5 driving automation may consist of automated driving functions that enable the vehicle to be driven under all conditions.

[0079] The vehicle control system 311 includes a vehicle control ECU (Electronic Control Unit) 321, a communication unit 322, a map information storage unit 323, a location information acquisition unit 324, an external recognition sensor 325, an in-vehicle sensor 326, a vehicle sensor 327, a memory unit 328, an automated driving control unit 329, a DMS (Driver Monitoring System) 330, an HMI (Human Machine Interface) 331, and a vehicle control unit 332.

[0080] Two or more (or, in some cases, all) of the following are connected to communicate with each other via a communication network 341: the vehicle control ECU 321, the communication unit 322, the map information storage unit 323, the location information acquisition unit 324, the external recognition sensor 325, the in-vehicle sensor 326, the vehicle sensor 327, the memory unit 328, the driving automation control unit 329, the DMS 330, the HMI 331, and the vehicle control unit 332. The communication network 341 is composed of an in-vehicle communication network or bus that conforms to digital bidirectional communication standards such as CAN (Controller Area Network), LIN (Local Interconnect Network), LAN (Local Area Network), FlexRay®, and Ethernet®. In some embodiments, the communication network 341 may have two or more types of communication networks, and different types of communication networks may be used depending on the type of data being transmitted. For example, CAN may be applied to data related to vehicle control, and Ethernet may be applied to large-capacity data. In some embodiments, two or more (or possibly all) units of the vehicle control system 311 may be directly connected using wireless communication (e.g., relatively short-range communication) without going through the communication network 341. In some embodiments, the wireless communication may use near-field wireless communication technology. Non-limiting examples of near-field wireless communication technology include Near Field Communication (NFC) and Bluetooth®. In some embodiments, two or more (or possibly all) units of the vehicle control system 311 may be connected using the communication network 341 and wireless communication technology (e.g., near-field wireless communication technology).

[0081] In the following embodiment, where two or more units of the vehicle control system 311 communicate via the communication network 341, the description of the communication network 341 will be omitted. For example, in an embodiment where the vehicle control ECU 321 and the communication unit 322 communicate via the communication network 341, it will simply be described as the vehicle control ECU 321 and the communication unit 322 communicating.

[0082] The vehicle control ECU 321 is composed of various processors, such as a CPU (Central Processing Unit) and an MPU (Micro Processing Unit). The vehicle control ECU 321 controls the functions of the entire vehicle control system 311 or a part of it.

[0083] The communication unit 322 communicates with various devices inside the vehicle 102 (hereinafter referred to as in-vehicle devices), various devices outside the vehicle 102 (hereinafter referred to as external devices), other vehicles, base stations, etc., and transmits and receives various types of data. In some embodiments, the communication unit 322 may use multiple communication technologies to perform communication.

[0084] A non-limiting example of communication between the communication unit 322 and external equipment will be briefly described. In some embodiments, the communication unit 322 may communicate with servers (hereinafter referred to as "external servers") located on an external network via a base station or access point using wireless communication technology. Examples of non-limiting wireless communication technologies include 5G (fifth-generation mobile communication system), LTE (Long Term Evolution), and DSRC (Dedicated Short Range Communications). External networks that the communication unit 322 can communicate with include, for example, the internet, a cloud network, or a network specific to a carrier. The communication technology used by the communication unit 322 to communicate with an external network is not particularly limited, as long as it is a wireless communication technology that enables digital two-way communication at a predetermined communication speed and over a predetermined distance.

[0085] In some embodiments, the communication unit 322 may communicate with terminals located near the vehicle using P2P (Peer To Peer) technology. Terminals located near the vehicle include, for example, terminals worn by relatively slow-moving objects such as pedestrians and cyclists, terminals installed in fixed locations such as stores, and / or MTC (Machine Type Communication) terminals. In some embodiments, the communication unit 322 may perform V2X (Vehicle to Everything) communication. V2X communication generally refers to communication between the vehicle and other entities. Non-exclusive examples of V2X communication include vehicle-to-vehicle communication with other vehicles, vehicle-to-infrastructure communication with roadside devices, etc., vehicle-to-home communication with homes, and vehicle-to-pedestrian communication with terminals carried or worn by pedestrians.

[0086] In some embodiments, the communication unit 322 may receive a program from outside the vehicle 102 to update the software that controls the operation of the vehicle control system 311 (for example, over the air). In some embodiments, the communication unit 322 may receive map information, traffic information, information about the vehicle 102's surroundings, etc., from outside the vehicle 102. In some embodiments, the communication unit 322 may transmit information about the vehicle 102, information about the vehicle 102's surroundings, etc., to an external device or external network. Non-limiting examples of information about the vehicle 102 that the communication unit 322 transmits to an external device or external network include data indicating the status of the vehicle 102, recognition results from the recognition unit 373, etc. In some embodiments, the communication unit 322 may communicate with a vehicle emergency call system. Non-limiting examples of a vehicle emergency call system include e-Call, etc.

[0087] In some embodiments, the communication unit 322 may receive electromagnetic waves transmitted by a road traffic information communication system. In some embodiments, such electromagnetic waves may be transmitted using radio beacons, optical beacons, FM multiplex broadcasting, etc.

[0088] A non-limiting example of communication with in-vehicle equipment that the communication unit 322 can perform will be outlined below. In some embodiments, the communication unit 322 may communicate with in-vehicle equipment using wireless communication. For example, in some embodiments, the communication unit 322 may communicate with in-vehicle equipment wirelessly using wireless communication technology that enables digital bidirectional communication at a predetermined or higher communication speed. Non-limiting examples of wireless communication technologies include wireless LAN, Bluetooth®, NFC, and WUSB (Wireless USB). Not limited to these, the communication unit 322 may also communicate with in-vehicle equipment using wired communication (in addition to or as an alternative to wireless communication). For example, in some embodiments, the communication unit 322 may communicate with in-vehicle equipment via wired communication through a cable connected to a connection terminal (not shown). In some embodiments, the communication unit 322 may communicate with in-vehicle equipment using wired communication technology that enables digital bidirectional communication at a predetermined or higher communication speed. Non-exclusive examples of wired communication technologies include USB (Universal Serial Bus), HDMI (High-Definition Multimedia Interface) (registered trademark), and MHL (Mobile High-definition Link).

[0089] Here, in-vehicle equipment refers to, for example, equipment located inside the vehicle 102 that is not connected to the communication network 341. In-vehicle equipment is divided into equipment that constitutes the vehicle control system 311 and equipment that does not. Non-exclusive examples of in-vehicle equipment that does not constitute the vehicle control system 311 include mobile devices and wearable devices owned by users of the vehicle 102 (e.g., the driver, passengers), and information equipment temporarily installed inside the vehicle 102. These devices can, for example, be moved outside the vehicle 102 and become external equipment.

[0090] The map information storage unit 323 stores maps acquired from external devices or external networks and / or maps created by the vehicle 102. For example, the map information storage unit 323 may store three-dimensional high-precision maps, global maps with lower precision than high-precision maps but covering a wide area, etc.

[0091] High-precision maps include, for example, dynamic maps, point cloud maps, and vector maps. A dynamic map may be a map consisting of four layers: dynamic information, semi-dynamic information, semi-static information, and static information, and may be provided to the vehicle 102 from an external server or the like. A point cloud map may be a map composed of point clouds (point cloud data). A vector map may be a map adapted for automated driving by associating traffic information, such as the locations of lanes and traffic lights, with a point cloud map.

[0092] The point cloud map and vector map may be provided from, for example, an external server, or they may be created in the vehicle 102 as maps for matching with the local map described later, based on sensing results from the camera 351, radar 352, LiDAR 353, etc., and stored in the map information storage unit 323. In addition, if high-precision maps are provided from an external server, in order to reduce communication capacity, map data of, for example, several hundred square meters relating to the planned route that the vehicle 102 will travel may be obtained from the external server.

[0093] The location information acquisition unit 324 acquires location information of the vehicle 102. The acquired location information may be supplied to the driving automation control unit 329. In some embodiments, the location information acquisition unit 324 may receive GNSS (Global Navigation Satellite System) signals from GNSS satellites. In some embodiments, the location information acquisition unit 324 may receive signals from beacons or the like.

[0094] The external recognition sensor 325 is equipped with various sensors used to recognize the external conditions of the vehicle 102, and supplies sensor data from one or more (or, in some cases, all) sensors to one or more (or, in some cases, all) units of the vehicle control system 311. The types and number of sensors equipped in the external recognition sensor 325 are arbitrary.

[0095] In some embodiments, the external recognition sensor 325 may include a camera 351, a radar 352, a LiDAR (Light Detection and Ranging, Laser Imaging Detection and Ranging) 353, and an ultrasonic sensor 354. However, the external recognition sensor 325 may also be configured to include one or more of the cameras 351, radar 352, LiDAR 353, and ultrasonic sensor 354. The number of cameras 351, radar 352, LiDAR 353, and ultrasonic sensors 354 is not particularly limited as long as it is a number that can be realistically installed in the vehicle 102. Furthermore, the types of sensors included in the external recognition sensor 325 are not limited to this example, and the external recognition sensor 325 may include other types of sensors. Examples of the sensing areas of each sensor included in the external recognition sensor 325 will be described later.

[0096] Camera 351 can use any suitable shooting method. In some embodiments, camera 351 may use a shooting method capable of distance measurement. Non-limiting examples of cameras using a shooting method capable of distance measurement include ToF (Time of Flight) cameras, stereo cameras, monocular cameras, and infrared cameras. However, camera 351 may not be capable of distance measurement and may simply be for acquiring images.

[0097] In some embodiments, the external recognition sensor 325 may include environmental sensors for detecting characteristics of the environment around the vehicle 102. Non-limiting examples of detectable environmental characteristics include weather, climate, brightness, etc. In some embodiments, the environmental sensors may include various sensors such as raindrop sensors, fog sensors, sunshine sensors, snow sensors, and illuminance sensors.

[0098] In some embodiments, the external recognition sensor 325 may include a microphone used for detecting sounds around the vehicle 102 and the location of sound sources.

[0099] The in-vehicle sensor 326 is equipped with various sensors for detecting information inside the vehicle 102 and supplies sensor data from one or more (or, in some cases, all) sensors to one or more (or, in some cases, all) units of the vehicle control system 311. The types and number of sensors equipped with the in-vehicle sensor 326 are not particularly limited, as long as they are of a type and number that can be realistically installed in the vehicle 102.

[0100] In some embodiments, the in-vehicle sensor 326 may include one or more sensors from among a camera, radar, seat sensor, microphone, and biosensor. In some embodiments, the camera included in the in-vehicle sensor 326 may use a distance-measuring shooting method. Non-limiting examples of cameras using a distance-measuring shooting method include ToF cameras, stereo cameras, monocular cameras, and infrared cameras. However, the camera included in the in-vehicle sensor 326 may not be for distance measurement and may simply be for acquiring captured images. The biosensor included in the in-vehicle sensor 326 may be provided, for example, on the seat or steering wheel, and may detect various biometric information of the user.

[0101] The vehicle sensor 327 is equipped with various sensors for detecting the state of the vehicle 102 and supplies sensor data from one or more (or, in some cases, all) sensors to one or more (or, in some cases, all) units of the vehicle control system 311. The types and number of sensors equipped with the vehicle sensor 327 are not particularly limited, as long as they are of a type and number that can realistically be installed on the vehicle 102.

[0102] In some embodiments, the vehicle sensor 327 may include a speed sensor, an acceleration sensor, an angular velocity sensor (gyro sensor), and / or an inertial measurement unit (IMU) integrating them. In some embodiments, the vehicle sensor 327 may include a steering angle sensor for detecting the steering angle of the steering wheel, a yaw rate sensor, an accelerator sensor for detecting the amount of operation of the accelerator pedal (e.g., pedal force, pedal stroke), and / or a brake sensor for detecting the amount of operation of the brake pedal (e.g., pedal force, pedal stroke). In some embodiments, the vehicle sensor 327 may include a rotation sensor for detecting the rotational speed of the engine or motor, an air pressure sensor for detecting the air pressure of the tires, a slip ratio sensor for detecting the slip ratio of the tires, and / or a wheel speed sensor for detecting the rotational speed of the wheels. In some embodiments, the vehicle sensor 327 may include a battery sensor for detecting the remaining charge and temperature of the battery, and / or an impact sensor capable of detecting external impacts.

[0103] The storage unit 328 includes at least one of a non-volatile storage medium and a volatile storage medium, and stores data and programs. Non-limiting examples of storage media include magnetic storage devices such as EEPROM (Electrically Erasable Programmable Read Only Memory), RAM (Random Access Memory), and / or HDD (Hard Disc Drive), semiconductor storage devices, optical storage devices, and magneto-optical storage devices. The storage unit 328 stores various programs and data used by one or more (or, in some cases, all) units of the vehicle control system 311. In some embodiments, the storage unit 328 may include an EDR (Event Data Recorder) or a DSSAD (Data Storage System for Automated Driving) to store information about the vehicle 102 before and after an event such as an accident, and information acquired by the in-vehicle sensor 326.

[0104] The automated driving control unit 329 controls the automated driving functions of the vehicle 102. In some embodiments, the automated driving control unit 329 may also include an analysis unit 361, an action planning unit 362, and an operation control unit 363.

[0105] The analysis unit 361 performs analysis processing of the vehicle 102 and / or the surrounding conditions. The analysis unit 361 includes a self-position estimation unit 371, a sensor fusion unit 372, and a recognition unit 373.

[0106] In some embodiments, the self-position estimation unit 371 may estimate the vehicle 102's position based on sensor data from an external recognition sensor 325 and a high-precision map stored in a map information storage unit 323. For example, the self-position estimation unit 371 may generate a local map based on sensor data from an external recognition sensor 325 and estimate the vehicle 102's position by matching the local map with a high-precision map. The position of the vehicle 102 may be based on, for example, the center of the rear wheel relative to the axle.

[0107] In some embodiments, the local map may be a three-dimensional high-precision map, an occupancy grid map, or the like, created using technologies such as SLAM (Simultaneous Localization and Mapping). The three-dimensional high-precision map may be, for example, the point cloud map described above. The occupancy grid map may be a map that divides the three-dimensional or two-dimensional space around the vehicle 102 into grids of a predetermined size and shows the occupancy status of objects on a grid-by-grid basis. The occupancy status of objects may be indicated, for example, by the presence or absence or probability of existence of an object. In some embodiments, the local map may also be used, for example, for detection and / or recognition processing of the external conditions of the vehicle 102 by the recognition unit 373.

[0108] In some embodiments, the self-position estimation unit 371 may estimate the self-position of the vehicle 102 based on position information acquired by the position information acquisition unit 324 and / or sensor data from the vehicle sensor 327.

[0109] The sensor fusion unit 372 performs sensor fusion processing to obtain information by combining multiple different types of sensor data (for example, image data supplied from the camera 351 and sensor data supplied from the radar 352). Methods for combining different types of sensor data are not limited to these, but include composite, integrated, fused, and combined methods.

[0110] The recognition unit 373 performs a detection process to detect the external conditions of the vehicle 102, and / or a recognition process to recognize the external conditions of the vehicle 102.

[0111] For example, the recognition unit 373 may perform detection and / or recognition processing of the external conditions of the vehicle 102 based on information from the external recognition sensor 325, information from the self-position estimation unit 371, information from the sensor fusion unit 372, etc.

[0112] Specifically, for example, the recognition unit 373 may perform detection and / or recognition processing of objects around the vehicle 102. Object detection processing may include, for example, detecting the presence, size, shape, position, and movement of an object. Object recognition processing may include, for example, recognizing attributes such as the type of object or identifying a specific object. Detection processing and recognition processing are not necessarily clearly separated, and at least some overlap may occur.

[0113] In some embodiments, the recognition unit 373 may detect objects around the vehicle 102 by performing clustering, which classifies the point cloud based on sensor data from the radar 352 and / or LiDAR 353 into clusters of points. This allows for the detection of the presence, size, shape, and position of objects around the vehicle 102.

[0114] In some embodiments, the recognition unit 373 may detect the movement of objects around the vehicle 102 by tracking the movement of clusters of points classified by clustering. This allows for the detection of the velocity and / or direction of travel (movement vector) of objects around the vehicle 102.

[0115] In some embodiments, the recognition unit 373 may detect and / or recognize vehicles (including bicycles), people, obstacles, structures, roads, traffic lights, traffic signs, road markings, etc., based on image data supplied from the camera 351. In some embodiments, the recognition unit 373 may recognize the types of objects around the vehicle 102 by performing recognition processing such as semantic segmentation.

[0116] In some embodiments, the recognition unit 373 may perform a recognition process of traffic rules around the vehicle 102 based on the map stored in the map information storage unit 323, the self-position estimation result by the self-position estimation unit 371, and / or the recognition result of objects around the vehicle 102 by the recognition unit 373. Through this process, the recognition unit 373 may recognize the location and / or status of traffic signals, the content of traffic signs and / or road markings, the content of traffic regulations, and / or drivable lanes, etc.

[0117] In some embodiments, the recognition unit 373 may perform recognition processing of the environment surrounding the vehicle 102. In some embodiments, the recognition unit 373 may recognize weather characteristics (temperature, humidity, brightness), and / or road surface conditions, etc.

[0118] The action planning unit 362 creates an action plan for the vehicle 102. For example, the action planning unit 362 may create an action plan by performing route planning and route following.

[0119] In some embodiments, the path planning may include global path planning and local path planning. Global path planning may include the process of planning a rough route from the start to the goal. Local path planning, also called track planning, may include the generation of a track that allows the vehicle 102 to travel safely and smoothly along the planned route in the vicinity of the vehicle 102, taking into account the vehicle's motion characteristics and the presence of any obstacles.

[0120] In some embodiments, route following may involve planning actions to safely and accurately travel along the route planned by the route planner within a planned time. The action planning unit 362 may, for example, calculate the target speed and / or target angular velocity of the vehicle 102 based on the results of this route following process.

[0121] The motion control unit 363 controls the operation of the vehicle 102 in order to realize the action plan created by the action planning unit 362.

[0122] For example, in some embodiments, the motion control unit 363 may control the steering control unit 381, brake control unit 382, ​​and / or drive control unit 383, which are included in the vehicle control unit 332 described later, to perform lateral vehicle motion control and / or longitudinal vehicle motion control so that the vehicle 102 travels along the trajectory calculated by the trajectory plan. For example, the motion control unit 363 may perform one or more driver assistance functions and / or control for the purpose of driving automation (e.g., lateral vehicle motion control, longitudinal vehicle motion control). Non-limited examples of driver assistance functions include collision avoidance or impact mitigation, inter-vehicle distance control (e.g., control to maintain a specific distance from a vehicle traveling in front of the vehicle 102), vehicle speed control (e.g., control to maintain a specific speed), vehicle collision warning, and lane departure warning. Non-limited examples of driving automation include driving without operation by the driver or remote driver.

[0123] In some embodiments, the DMS 330 may perform driver authentication processing and / or driver status recognition processing based on sensor data from the in-vehicle sensor 326 and / or input data input to the HMI 331, which will be described later. Non-limited examples of driver status that may be recognized include physical condition, alertness level, concentration level, fatigue level, gaze direction, intoxication level, driving operation, posture, etc.

[0124] In some embodiments, the DMS 330 may perform authentication processing for users other than the driver (e.g., passengers) and / or recognition processing for the status of such users. In some embodiments, the DMS 330 may perform recognition processing for the internal conditions of the vehicle 102 based on sensor data from the in-vehicle sensor 326. Non-limiting examples of characteristics of the internal conditions of the vehicle 102 that may be recognized include temperature, humidity, brightness, odor, etc.

[0125] The HMI331 receives various data and instructions as input and presents various data to the user.

[0126] A general overview of data input to the HMI 331 is provided. The HMI 331 is equipped with an input device for a person to input data, instructions, etc. Based on the data, instructions, etc., input by the input device, the HMI 331 generates an input signal and supplies it to one or more (or, in some cases, all) units of the vehicle control system 311. In some embodiments, the HMI 331 may be equipped with a touch panel, buttons, switches, and / or levers as input devices. Not limited to these, the HMI 331 may be equipped with an input device that allows information to be input by methods other than manual operation, such as voice or gestures. In some embodiments, the HMI 331 may be equipped with a remote control device using infrared and / or radio waves, or an external connection device that corresponds to the operation of the vehicle control system 311, as an input device. Non-limited examples of external connection devices include mobile devices (e.g., smartphones) and wearable devices (e.g., smartwatches).

[0127] A brief explanation of data presentation by HMI331 is provided below. HMI331 generates visual, auditory, and / or tactile information for the user and / or people outside the vehicle 102. HMI331 may also perform output control to control the output, output content, output timing, and / or output method of each generated piece of information. Non-limited examples of visual information that can be generated and output by HMI331 include information shown by images and light, such as operation screens, vehicle status displays, warning displays, and monitor images showing the surroundings of the vehicle 102. Non-limited examples of auditory information that can be generated and output by HMI331 include voice guidance, warning sounds, and warning messages. Non-limited examples of tactile information that can be generated and output by HMI331 include information that is given to the user's sense of touch through force, vibration, movement, etc.

[0128] In some embodiments, the HMI 331 may include, as an output device capable of outputting visual information, a display device that presents visual information by displaying an image itself, or a projector device that presents visual information by projecting an image. In some embodiments, the display device may be a device that displays visual information within the user's field of view, such as a head-up display, a transparent display, or a wearable device with AR (Augmented Reality) functionality, in addition to or as an alternative to a normal display device. In some embodiments, the HMI 331 may include, as an output device capable of outputting visual information, a display device provided in the vehicle 102, such as a navigation device, instrument panel, CMS (Camera Monitoring System), electronic mirror, lamp, etc.

[0129] In some embodiments, the HMI331 may include an audio speaker, headphones, or earphones as an output device capable of outputting auditory information.

[0130] In some embodiments, the HMI331 may include a haptic element using haptic technology as an output device capable of outputting tactile information. The haptic element may be provided, for example, on parts of the vehicle 102 that the user comes into contact with, such as the steering wheel or the seat.

[0131] The vehicle control unit 332 controls one or more (or, in some cases, all) units of the vehicle 102. The vehicle control unit 332 includes a steering control unit 381, a brake control unit 382, ​​a drive control unit 383, a body system control unit 384, a light control unit 385, and a horn control unit 386.

[0132] The steering control unit 381 detects and / or controls the state of the steering system of the vehicle 102. The steering system includes, for example, a steering mechanism with a steering wheel, an electric power steering system, etc. The steering control unit 381 includes, for example, a steering ECU that controls the steering system, an actuator that drives the steering system, etc.

[0133] The brake control unit 382 detects and / or controls the state of the brake system of the vehicle 102. The brake system includes, for example, a brake mechanism including a brake pedal, an ABS (Antilock Brake System), a regenerative braking mechanism, etc. The brake control unit 382 includes, for example, a brake ECU that controls the brake system, an actuator that drives the brake system, etc.

[0134] The drive control unit 383 detects and / or controls the state of the vehicle 102's drive system. The drive system includes, for example, an accelerator pedal, a drive force generating device for generating driving force such as an internal combustion engine or drive motor, and a drive force transmission mechanism for transmitting driving force to the wheels. The drive control unit 383 also includes, for example, a drive ECU for controlling the drive system and an actuator for driving the drive system.

[0135] The body system control unit 384 detects and / or controls the state of the body system of the vehicle 102. The body system includes, for example, a keyless entry system, a smart key system, power window devices, power seats, an air conditioning system, airbags, seat belts, a shift lever, etc. The body system control unit 384 also includes, for example, a body system ECU that controls the body system, actuators that drive the body system, etc.

[0136] The light control unit 385 detects and / or controls the state of various lights on the vehicle 102. Non-exclusive examples of lights that can be controlled by the light control unit 385 include headlights, taillights, fog lights, turn signals, brake lights, projector lights, bumper indicators, etc. The light control unit 385 includes a light ECU for controlling the lights, actuators for driving the lights, etc.

[0137] The horn control unit 386 detects and / or controls the state of the vehicle's car horn. The horn control unit 386 includes, for example, a horn ECU for controlling the car horn, an actuator for driving the car horn, and the like.

[0138] Figure 6 shows an example of the sensing area of ​​the external recognition sensor 325 in Figure 5, including the camera 351, radar 352, LiDAR 353, and ultrasonic sensor 354. In Figure 6, a schematic view of the vehicle 102 from above is shown.

[0139] Sensing regions 401F and 401B show examples of sensing regions for ultrasonic sensors 354. Sensing region 401F (for example, the sensing region of multiple ultrasonic sensors 354) covers the area around the front end of the vehicle 102. Sensing region 401B (for example, the sensing region of multiple ultrasonic sensors 354) covers the area around the rear end of the vehicle 102.

[0140] The sensing results in sensing region 401F and / or sensing region 401B may be used, for example, to assist in parking the vehicle 102.

[0141] Sensing areas 402F, 402B, 402L, and 402R illustrate examples of sensing areas for short-range or medium-range radar 352. Sensing area 402F covers an area in front of vehicle 102 that is further away than sensing area 401F. Sensing area 402B covers an area behind vehicle 102 that is further away than sensing area 401B. Sensing area 402L covers the area around the left rear of vehicle 102. Sensing area 402R covers the area around the right rear of vehicle 102.

[0142] The sensing results in sensing region 402F may be used, for example, to detect vehicles or pedestrians in front of vehicle 102. The sensing results in sensing region 402B may be used, for example, to prevent collisions behind vehicle 102. The sensing results in sensing region 402L and / or sensing region 402R may be used, for example, to detect one or more objects in the blind spots on the left and / or right sides of vehicle 102.

[0143] Sensing areas 403F, 403B, 403L, and 403R show examples of sensing areas by camera 351. Sensing area 403F covers a position further in front of vehicle 102 than sensing area 402F. Sensing area 403B covers a position further behind vehicle 102 than sensing area 402B. Sensing area 403L covers the left side of vehicle 102. Sensing area 403R covers the right side of vehicle 102.

[0144] The sensing results in sensing region 403F may be used, for example, for recognition of traffic lights and traffic signs, lane departure prevention support systems, and automatic headlight control systems. The sensing results in sensing region 403B may be used, for example, for parking assistance and / or surround view systems. The sensing results in sensing region 403L and / or sensing region 403R may be used, for example, for surround view systems.

[0145] Sensing area 404 shows an example of the sensing area of ​​LiDAR 353. Sensing area 404 covers a position further in front of the vehicle 102 than sensing area 403F. On the other hand, sensing area 404 has a narrower range in the left-right direction of the vehicle 102 than sensing area 403F.

[0146] The sensing results in the sensing region 404 may be used, for example, to detect objects such as surrounding vehicles.

[0147] Sensing area 405 shows an example of the sensing area of ​​the long-range radar 352. Sensing area 405 covers a position further in front of the vehicle 102 than sensing area 404. On the other hand, sensing area 405 has a narrower range in the lateral direction of the vehicle 102 than sensing area 404.

[0148] The sensing results in the sensing region 405 may be used, for example, for ACC (Adaptive Cruise Control), emergency braking, collision avoidance, etc.

[0149] In some embodiments, the sensing areas of each sensor of the external recognition sensor 325 (e.g., camera 351, radar 352, LiDAR 353, ultrasonic sensor 354) may take various configurations other than those shown in Figure 6. Specifically, in some embodiments, the ultrasonic sensor 354 may also sense the sides of the vehicle 102, or the LiDAR 353 may be configured to sense the rear of the vehicle 102. Furthermore, the installation positions of each sensor are not limited to the examples described above. Also, the number of sensors may be one or multiple.

[0150] <Application of this technology> The above-described technology may be applied to the automated driving system 100 with the above configuration in <3. Map prediction using point cloud matching>. In that case, Figure 7 shows a functional block diagram that shows the functions realized by the server 101 (CPU 201) executing the processing as blocks.

[0151] As shown in Figure 7, the server 101 (CPU 201) in this case has a matching unit 511, a dataset generation unit 512, and a learning unit 513 as functional blocks.

[0152] The matching unit 511 performs matching between the generated point cloud and the existing point cloud, and supplies the matching result to the dataset generation unit 512.

[0153] The dataset generation unit 512 obtains the matching results supplied from the matching unit 511. Based on the matching results and the existing map, the dataset generation unit 512 generates a generated map and uses the generated map to generate a training dataset. The dataset generation unit 512 supplies the training dataset to the training unit 513.

[0154] The learning unit 513 acquires a training dataset supplied by the dataset generation unit 512. Using this training dataset, the learning unit 513 performs training to generate a map generation AI model that takes generated images and generated point clouds as input and outputs corresponding generated maps. In other words, the learning unit 513 uses this training dataset to train the AI ​​learning model and generate the map generation AI model.

[0155] For example, the matching unit 511 may perform a matching between an existing point cloud representing the three-dimensional shape of an object whose position is indicated by an existing map, which is pre-prepared map information, and a generated point cloud representing the three-dimensional shape of an object detected by a sensor. Then, the dataset generation unit 512 may generate a generated map, which is map information corresponding to the generated point cloud, based on the matching result and the existing map, and generate a training dataset using that generated map.

[0156] In this case, the matching unit 511 may use a portion of the generated point cloud to perform matching with an existing point cloud. For example, the matching unit 511 may set a portion of the generated point cloud based on the distribution pattern of points. Alternatively, the matching unit 511 may set a portion of the generated image based on the image recognition result of the generated image generated by the image sensor.

[0157] For example, the matching unit 511 may set a region in the generated point cloud where the density of points is higher than a predetermined standard as a part of its region and perform matching with an existing point cloud. Alternatively, the matching unit 511 may set a region in the generated point cloud that excludes the point cloud representing the three-dimensional shape of a deformable object as a part of its region and perform matching with an existing point cloud. Alternatively, the matching unit 511 may set a region in the generated point cloud that excludes the point cloud representing the three-dimensional shape of a road surface as a part of its region and perform matching with an existing point cloud. Alternatively, the matching unit 511 may set a region in the generated point cloud that excludes the point cloud representing the three-dimensional shape of an animal body as a part of its region and perform matching with an existing point cloud.

[0158] The dataset generation unit 512 may generate a training dataset that includes generated images and point clouds generated using sensors, and generated maps generated using point cloud matching. For example, the dataset generation unit 512 may generate a training dataset for each domain by associating multiple generated images and point clouds corresponding to different domains with a generated map. The existing map may be a global map, which is map information covering a relatively wide area, and the generated map may be a local map, which is map information covering a narrower area than the global map. The existing map and the generated map may also be HD maps, which are map information with relatively high accuracy and include a relatively wide variety of road information, including at least lanes.

[0159] Furthermore, the dataset generation unit 512 may generate a training dataset that includes, in addition to the generated images, generated point clouds, and generated maps described above, an SD map, which is map information with relatively low accuracy and contains fewer types of road information than the HD map. For example, if the existing map includes an existing HD map and an existing SD map, the dataset generation unit 512 may generate a generated HD map and a generated SD map as generated maps based on the matching results, the existing HD map, and the existing SD map, and generate a training dataset that includes generated images, generated point clouds, generated HD maps, and generated SD maps. Alternatively, if the existing map is an HD map, the dataset generation unit 512 may generate a generated HD map as a generated map based on the matching results and the existing map, extract some information from that generated HD map to generate a generated SD map, and generate a training dataset that includes generated images, generated point clouds, generated HD maps, and generated SD maps.

[0160] The generated images and point clouds supplied from the vehicle 102, etc., may be acquired by the communication unit 214. In that case, the matching unit 511 may perform matching between the generated point clouds acquired by the communication unit 214 and existing point clouds, and the dataset generation unit 512 may generate a training dataset including the generated images and point clouds acquired by the communication unit 214. In other words, the communication unit 214 can also be said to be an acquisition unit that acquires generated images and point clouds.

[0161] Furthermore, the communication unit 214 may supply the generated training dataset and map generation AI model to the vehicle 102, etc. In other words, the communication unit 214 can also be considered a supply unit that supplies the training dataset and map generation AI model.

[0162] The learning performed by the learning unit 513 may be carried out by another device, such as the vehicle 102. In that case, the learning dataset generated by the dataset generation unit 512 should be supplied to the other device (vehicle 102, etc.) that performs the learning. In other words, the learning unit 513 may be omitted in that case.

[0163] With this configuration, the server 101 can more easily generate training datasets for generating a map generation AI model, as described above in <3. Map prediction using point cloud matching>.

[0164] <Detection Process Flow> Next, we will explain the processes performed by the server 101 and the vehicle 102. Referring to the flowchart in Figure 8, we will explain an example of the detection process flow performed by the vehicle 102. This detection process involves the vehicle 102 using on-board sensors to detect objects around the vehicle 102 and generating generated images and point clouds.

[0165] When the detection process is started, the vehicle 102's self-position estimation unit 371 generates a generated image and a generated point cloud in step S101, using various sensors of the external recognition sensor 325 (from camera 351 to ultrasonic sensor 354) as appropriate.

[0166] In step S102, the self-position estimation unit 371 uploads the generated image and generated point cloud to the server 101 via the communication unit 322.

[0167] The detection process ends when the process in step S102 is completed.

[0168] <Learning Process Flow 1> Next, an example of the learning process flow executed by the server 101 will be explained with reference to the flowchart in Figure 9. When the learning process starts, in step S121, the communication unit 214 acquires the generated image and generated point cloud uploaded from the vehicle 102, etc.

[0169] In step S122, the matching unit 511 sets a portion of the generated point cloud as the target for matching. This step may be omitted if the entire generated point cloud is to be used for matching.

[0170] In step S123, the matching unit 511 matches the target of the generated point cloud with the existing point cloud. In other words, the matching unit 511 matches the existing point cloud, which represents the three-dimensional shape of an object whose position is indicated by the existing map, which is pre-prepared map information, with the generated point cloud, which represents the three-dimensional shape of an object detected by the sensor.

[0171] In step S124, the dataset generation unit 512 extracts the portion corresponding to the generated point cloud from the existing map based on the matching result and uses it as a generated map. In step S125, the dataset generation unit 512 generates a training dataset that includes the generated image, the generated point cloud, and the generated map it has generated. In other words, the dataset generation unit 512 generates a generated map, which is map information corresponding to the generated point cloud, based on the matching result and the existing map, and uses that generated map to generate a training dataset.

[0172] In step S126, the learning unit 513 performs training using the training dataset and generates (constructs) a map generation AI model.

[0173] In step S127, the communication unit 214 supplies its map generation AI model to the vehicle 102, etc.

[0174] The learning process ends when the processing in step S127 is completed.

[0175] <Inference Processing Flow> Next, an example of the inference processing flow performed by vehicle 102 will be explained with reference to the flowchart in Figure 10. This inference processing is the process by which vehicle 102 performs inference on a map generation AI model using generated images and generated point clouds.

[0176] When the inference process begins, in step S151, the communication unit 322 of the vehicle 102 acquires a map generation AI model supplied from the server 101 or the like. The self-position estimation unit 371 holds (stores) that map generation AI model.

[0177] In step S152, the self-position estimation unit 371 generates a generated image and a generated point cloud using various sensors of the external recognition sensor 325 (from the camera 351 to the ultrasonic sensor 354) as appropriate.

[0178] In step S153, the self-localization unit 371 inputs the generated image and generated point cloud into the map generation AI model and outputs an AI inference map corresponding to them.

[0179] In step S154, the automated driving control unit 329 (for example, the action planning unit 362 or the motion control unit 363) executes processing related to automated driving based on its AI inference map.

[0180] The inference process ends when the processing in step S154 is completed.

[0181] By executing each process as described above, the server 101 can more easily generate a training dataset for generating a map generation AI model, as described above in <3. Map prediction using point cloud matching>.

[0182] <Self-Position Estimation Unit> The generation of the training dataset described above may also be performed in the vehicle 102. In that case, the vehicle 102 may generate the training dataset using the generated images and generated point clouds it generates itself. In other words, the various sensors of the vehicle 102's external recognition sensor 325 (from the camera 351 to the ultrasonic sensor 354) are used as appropriate to generate the generated images and generated point clouds. Therefore, the external recognition sensor 325 can also be called an image generation unit that generates generated images, or a point cloud generation unit that generates generated point clouds.

[0183] Figure 11 shows a functional block diagram illustrating the functions that the vehicle 102 (self-position estimation unit 371) implements as blocks.

[0184] As shown in Figure 11, the vehicle 102 (self-position estimation unit 371) in this case has a matching unit 611, a dataset generation unit 612, and a learning unit 613 as functional blocks.

[0185] The matching unit 611 is a processing unit similar to the matching unit 511 and performs the same processing. However, the matching unit 611 performs matching between the generated point cloud generated by the external recognition sensor 325 and the existing point cloud. The matching unit 611 also supplies the obtained matching results to the data set generation unit 612.

[0186] The dataset generation unit 612 is a processing unit similar to the dataset generation unit 512 and performs the same processing. However, the dataset generation unit 612 generates a generated map using the matching results supplied from the matching unit 611 and the existing map, and generates a training dataset using the generated images generated by the external recognition sensor 325. The dataset generation unit 612 also supplies the generated training dataset to the learning unit 613.

[0187] The learning unit 613 is a processing unit similar to the learning unit 513 and performs the same processing. However, the learning unit 613 performs learning using a learning dataset supplied by the dataset generation unit 612.

[0188] With this configuration, the vehicle 102 can more easily generate a training dataset for generating a map generation AI model, as described above in <3. Map prediction using point cloud matching>.

[0189] <Detection Learning Process Flow 1> Next, an example of the detection learning process flow executed by the vehicle 102 will be explained with reference to the flowchart in Figure 12. When the detection learning process is started, in step S181, the self-position estimation unit 371 generates a generated image and a generated point cloud using various sensors of the external recognition sensor 325 (from the camera 351 to the ultrasonic sensor 354) as appropriate.

[0190] Each process from step S182 to step S186 is executed in the same manner as each process from step S122 to step S126 in Figure 9.

[0191] In step S187, the self-position estimation unit 371 stores (memorizes) its map generation AI model.

[0192] When the process in step S187 is completed, the detection learning process is terminated.

[0193] By performing each process as described above, the vehicle 102 can more easily generate a training dataset for generating a map generation AI model, as described above in <3. Map prediction using point cloud matching>.

[0194] <Learning Process Flow 2> As mentioned above in <3. Map Prediction Using Point Cloud Matching>, the training dataset may also include generated SD maps. In that case, the server 101 may generate generated HD maps and generated SD maps from existing HD maps and existing SD maps.

[0195] An example of the learning process flow in that case will be explained with reference to the flowchart in Figure 13. When the learning process starts, each process from step S301 to step S303 is executed in the same way as each process from step S121 to step S123 in Figure 9.

[0196] In step S304, the dataset generation unit 512 extracts portions corresponding to the generated point cloud from the existing HD map based on the matching results, and uses them as generated HD maps. Similarly, the dataset generation unit 512 extracts portions corresponding to the generated point cloud from the existing SD map based on the matching results, and uses them as generated SD maps. In step S305, the dataset generation unit 512 generates a training dataset containing the generated images and generated point clouds, as well as the generated HD maps and SD maps.

[0197] Steps S306 and S307 are performed in the same manner as steps S126 and S127 in Figure 9.

[0198] When the process in step S307 is completed, the learning process ends.

[0199] By performing each process as described above, the server 101 can more easily generate a training dataset including generated images, generated point clouds, generated HD maps, and generated SD maps, as described above in <3. Map prediction using point cloud matching>.

[0200] <Detection and Learning Process Flow 2> Even if such a generated SD map is included, the vehicle 102 may generate a training dataset. In that case, the vehicle 102 may generate a generated HD map and a generated SD map from the existing HD map and the existing SD map.

[0201] An example of the detection learning process flow in that case will be explained with reference to the flowchart in Figure 14. When the detection learning process starts, each process from step S321 to step S323 is executed in the same way as each process from step S181 to step S183 in Figure 12.

[0202] In step S324, the dataset generation unit 612 extracts the portion corresponding to the generated point cloud from the existing HD map based on the matching result and uses it as the generated HD map. Similarly, the dataset generation unit 612 extracts the portion corresponding to the generated point cloud from the existing SD map based on the matching result and uses it as the generated SD map. In step S325, the dataset generation unit 612 generates a training dataset including the generated images and generated point cloud, as well as the generated generated HD map and generated SD map.

[0203] Steps S326 and S327 are performed in the same manner as steps S186 and S187 in Figure 12.

[0204] When the process in step S327 is completed, the detection learning process ends.

[0205] By performing each process as described above, the vehicle 102 can more easily generate a training dataset including generated images, generated point clouds, generated HD maps, and generated SD maps, as described above in <3. Map prediction using point cloud matching>.

[0206] <Learning Process Flow 3> Furthermore, as described above in <3. Map Prediction Using Point Cloud Matching>, the training dataset may also include generated SD maps. In that case, the server 101 may generate the generated SD map from the generated HD map. Typically, SD maps have larger coordinate position shift information than HD maps. Therefore, by deliberately introducing a position shift into the generated SD map, it can be made suitable as a training SD map.

[0207] An example of the learning process flow in that case will be explained with reference to the flowchart in Figure 15. When the learning process starts, each of the processes from step S341 to step S343 is executed in the same way as each of the processes from step S121 to step S123 in Figure 9.

[0208] In step S344, the dataset generation unit 512 extracts a portion corresponding to the generated point cloud from the existing HD map based on the matching result, and uses it as the generated HD map. In step S345, the dataset generation unit 512 extracts some information from the generated HD map to generate the generated SD map. In step S346, the dataset generation unit 512 generates a training dataset containing the generated image and generated point cloud, as well as the generated generated HD map and generated SD map.

[0209] Steps S347 and S348 are performed in the same manner as steps S126 and S127 in Figure 9.

[0210] When the process in step S348 is completed, the learning process ends.

[0211] By performing each process as described above, the server 101 can more easily generate a training dataset including generated images, generated point clouds, generated HD maps, and generated SD maps, as described above in <3. Map prediction using point cloud matching>.

[0212] <Detection and Learning Process Flow 3> Even if such a generated SD map is included, the vehicle 102 may generate a training dataset. In that case, the vehicle 102 may generate a generated SD map from the generated HD map.

[0213] An example of the detection learning process flow in that case will be explained with reference to the flowchart in Figure 16. When the detection learning process starts, each process from step S361 to step S363 is executed in the same way as each process from step S181 to step S183 in Figure 12.

[0214] In step S364, the dataset generation unit 612 extracts a portion corresponding to the generated point cloud from the existing HD map based on the matching result, and uses it as the generated HD map. In step S365, the dataset generation unit 612 extracts some information from the generated HD map to generate the generated SD map. In step S366, the dataset generation unit 612 generates a training dataset containing the generated image and generated point cloud, as well as the generated generated HD map and generated SD map.

[0215] Steps S367 and S368 are performed in the same manner as steps S186 and S187 in Figure 12.

[0216] When the process in step S368 is completed, the detection learning process ends.

[0217] By performing each process as described above, the vehicle 102 can more easily generate a training dataset including generated images, generated point clouds, generated HD maps, and generated SD maps, as described above in <3. Map prediction using point cloud matching>.

[0218] <5. Addendum> <Software> The series of processes described above can be executed by hardware or by software. When the series of processes described above are executed by software, the programs that make up the software are installed from a network or storage medium.

[0219] This recording medium, as shown in Figure 4, for example, consists of a removable media 221 on which a program is recorded, which is distributed separately from the main device body to the user for the purpose of delivering the program. This removable media 221 includes magnetic disks (including flexible disks) and optical disks (including CD-ROMs and DVDs). Furthermore, it also includes magneto-optical disks (including MDs (Mini Discs)) and semiconductor memory. In this case, for example, by inserting the removable media 221 into the drive 215, the program stored on the removable media 221 can be read and installed into the storage unit 213.

[0220] Furthermore, this program can also be provided via wired or wireless transmission media such as local area networks, the internet, or digital satellite broadcasting. In that case, the program can be received by the communication unit 214 and installed in the storage unit 213.

[0221] In addition, this program can be pre-installed in memory units or ROMs. For example, the program can be pre-installed in memory unit 213 or ROM 202.

[0222] The same principle applies to vehicle 102. For example, in Figure 5, a program stored on a removable media (not shown) mounted on a drive (not shown) may be read and installed in the storage unit 328, etc. Alternatively, the program may be received by the communication unit 322, etc. and installed in the storage unit 328, etc. Furthermore, the program may be pre-installed in the storage unit 328, etc.

[0223] <Applicable Subjects of This Technology> Furthermore, this technology can be applied to any configuration. For example, this technology can be applied to various electronic devices. In addition, this technology can be applied to the automatic driving control technology of any mobile body, not limited to vehicles, but including, for example, ships and airplanes. This mobile body may be a manned aircraft with a person on board, or an unmanned aircraft without a person on board, such as a so-called drone. Furthermore, this technology can be applied to any driving assistance technology, not limited to automatic driving control, but including, for example, route guidance and driving operation assistance.

[0224] Furthermore, this technology can also be implemented as part of a device, such as a processor as a system LSI (Large Scale Integration) (e.g., a video processor), a module using multiple processors (e.g., a video module), a unit using multiple modules (e.g., a video unit), or a set with additional functions added to a unit (e.g., a video set).

[0225] Furthermore, this technology can also be applied to network systems composed of multiple devices. For example, this technology may be implemented as cloud computing, where multiple devices share and collaborate on processing via a network. For example, this technology may be implemented in a cloud service that provides image (video) related services to any terminal such as computers, AV (Audio Visual) equipment, portable information processing terminals, and IoT (Internet of Things) devices.

[0226] In this specification, a system refers to a collection of multiple components (devices, modules (parts), etc.), regardless of whether all components are located in the same enclosure. Therefore, multiple devices housed in separate enclosures and connected via a network, and a single device containing multiple modules within a single enclosure, are both considered systems.

[0227] <Applicable Fields and Applications of This Technology> Systems, devices, and processing units incorporating this technology can be used in any field, such as transportation, medical care, security, agriculture, livestock farming, mining, beauty, factories, home appliances, weather, and nature monitoring. Furthermore, the applications are entirely arbitrary.

[0228] For example, this technology can be applied to systems and devices used to provide entertainment content. Furthermore, for example, this technology can be applied to systems and devices used for traffic management, such as traffic condition monitoring and automated driving control. In addition, for example, this technology can be applied to systems and devices used for security. Furthermore, for example, this technology can be applied to systems and devices used for automatic control of machinery, etc. Furthermore, for example, this technology can be applied to systems and devices used for agriculture and livestock farming. Furthermore, for example, this technology can be applied to systems and devices that monitor natural conditions such as volcanoes, forests, and oceans, as well as wildlife. Furthermore, for example, this technology can be applied to systems and devices used for sports.

[0229] <Other> In this specification, terms such as "combine," "multiplex," "add," "integrate," "include," "store," "insert," "insert," and "place" mean combining multiple things into one, such as combining encoded data and metadata into a single data, and represent one method of "associating" as described above.

[0230] Furthermore, the embodiments of this technology are not limited to those described above, and various modifications are possible without departing from the gist of this technology.

[0231] For example, the configuration described as a single device (or processing unit) may be divided and configured as multiple devices (or processing units). Conversely, the configurations described above as multiple devices (or processing units) may be combined and configured as a single device (or processing unit). Furthermore, it is also possible to add configurations other than those described above to the configuration of each device (or each processing unit). In addition, if the overall system configuration and operation are substantially the same, a part of the configuration of one device (or processing unit) may be included in the configuration of another device (or other processing unit).

[0232] Furthermore, for example, the program described above may be executed on any device. In that case, the device should have the necessary functions (such as functional blocks) and be able to obtain the necessary information.

[0233] Furthermore, for example, each step of a flowchart may be executed by one device, or it may be divided among multiple devices. Additionally, if a single step includes multiple processes, these processes may be executed by one device, or they may be divided among multiple devices. In other words, multiple processes included in a single step can be executed as multiple steps. Conversely, processes described as multiple steps can be combined and executed as a single step.

[0234] Furthermore, for example, a program executed by a computer may be structured so that the steps of the program are executed chronologically in the order described herein, or they may be executed in parallel or individually at necessary times, such as when a call is made. In other words, the steps may be executed in an order different from the order described above, as long as no inconsistencies arise. Moreover, the steps of this program may be executed in parallel with the processing of other programs, or in combination with the processing of other programs.

[0235] Furthermore, for example, multiple technologies relating to this technology can be implemented independently, as long as they do not create a contradiction. Of course, any multiple technologies can also be implemented in combination. For example, some or all of the technologies described in one embodiment can be implemented in combination with some or all of the technologies described in another embodiment. Also, some or all of the above-mentioned technologies can be implemented in combination with other technologies not mentioned above.

[0236] Furthermore, this technology can also take the following configurations: (1) An information processing device comprising: a matching unit that matches an existing point cloud representing the three-dimensional shape of an object whose position is indicated by an existing map, which is pre-prepared map information, with a generated point cloud representing the three-dimensional shape of an object detected by a sensor; and a dataset generation unit that generates a generated map, which is map information corresponding to the generated point cloud, based on the result of the matching and the existing map, and generates a training dataset using the generated map. (2) The information processing device according to (1), wherein the matching unit performs the matching using a part of the generated point cloud. (3) The information processing device according to (2), wherein the matching unit sets the part of the cloud based on the distribution pattern of points in the generated point cloud. (4) The information processing device according to (2) or (3), wherein the matching unit sets the part of the cloud based on the image recognition result of a generated image generated by an image sensor. (5) The information processing device according to any one of (2) to (4), wherein the matching unit sets a part of the cloud where the density of points in the generated point cloud is higher than a predetermined standard, and performs the matching. (6) The matching unit sets the region of the generated point cloud excluding the point cloud representing the three-dimensional shape of a deformable object as the partial region and performs the matching, as described in any of (2) to (5). (7) The matching unit sets the region of the generated point cloud excluding the point cloud representing the three-dimensional shape of a road surface as the partial region and performs the matching, as described in any of (2) to (6). (8) The matching unit sets the region of the generated point cloud excluding the point cloud representing the three-dimensional shape of an animal body as the partial region and performs the matching, as described in any of (2) to (7). (9) The learning dataset includes generated images generated using a sensor, the generated point cloud, and the generated map, as described in any of (1) to (8).(10) The data set generation unit generates the training data set for each domain by associating a plurality of generated images and generated point clouds corresponding to different domains with the generated map, as described in (9). (11) The existing map is a global map which is map information covering a relatively wide area, and the generated map is a local map which is map information covering a narrower area than the global map, as described in (9) or (10). (12) The existing map and the generated map are HD maps, and the HD map includes a relatively wide variety of road information, including at least lanes, and is relatively high-precision map information, as described in (11). (13) The training data set further includes an SD map, and the SD map includes fewer types of road information than the HD map and is relatively low-precision map information, as described in (12). (14) The existing map includes the existing HD map of the HD map and the existing SD map of the SD map, and the dataset generation unit generates the generated HD map of the HD map and the generated SD map of the SD map as generated maps based on the matching result, the existing HD map and the existing SD map, and generates the training dataset including the generated image, the generated point cloud and the generated HD map and the generated SD map, as the generated map. The information processing device according to (13). (15) The existing map is the HD map, and the dataset generation unit generates the generated HD map of the HD map as the generated map based on the matching result and the existing map, extracts some information from the generated HD map and generates the generated SD map which is the generated map of the SD map, and generates the training dataset including the generated image and the generated point cloud and the generated HD map and the generated SD map, as the information processing device according to (13).(16) The information processing device according to any one of (1) to (15), further comprising a learning unit that performs training to generate a map generation AI model using the generated training dataset, wherein the map generation AI model is an AI model that infers map information from generated images generated by an image sensor and the generated point cloud. (17) The information processing device according to any one of (1) to (16), further comprising an acquisition unit that acquires generated images generated by an image sensor and the generated point cloud, wherein the matching unit performs matching between the acquired generated point cloud and the existing point cloud, and the dataset generation unit generates the training dataset including the acquired generated images and the generated point cloud. (18) The information processing device according to any one of (1) to (16), further comprising an image generation unit that generates generated images and a point cloud generation unit that generates the generated point cloud, wherein the matching unit performs matching between the generated generated point cloud and the existing point cloud, and the dataset generation unit generates the training dataset including the generated generated images and the generated point cloud. (19) An information processing method comprising: matching an existing point cloud representing the three-dimensional shape of an object whose position is indicated by an existing map, which is pre-prepared map information, with a generated point cloud representing the three-dimensional shape of an object detected by a sensor; generating a generated map, which is map information corresponding to the generated point cloud, based on the result of the matching and the existing map; and generating a training dataset using the generated map. (20) A program for causing a computer to perform a process comprising: matching an existing point cloud representing the three-dimensional shape of an object whose position is indicated by an existing map, which is pre-prepared map information, with a generated point cloud representing the three-dimensional shape of an object detected by a sensor; generating a generated map, which is map information corresponding to the generated point cloud, based on the result of the matching and the existing map; and generating a training dataset using the generated map.

[0237] 100 Automated driving system, 101 Server, 102 Vehicle, 110 Network, 201 CPU, 202 RAM, 203 ROM, 204 Bus, 210 Input / Output interface, 211 Input unit, 212 Output unit, 213 Storage unit, 214 Communication unit, 215 Drive, 221 Removable media, 311 Vehicle control system, 321 Vehicle control ECU, 322 Communication unit, 323 Map information storage unit, 324 Location information acquisition unit, 325 External recognition sensor, 326 In-vehicle sensor, 327 Vehicle sensor, 328 Storage unit, 329 Driving automation control unit, 330 DMS, 331 HMI, 332 Vehicle control unit, 341 Communication network, 351 Camera, 352 Radar, 353 LiDAR, 354 Ultrasonic sensor, 361 Analysis unit, 362 Action planning unit, 363 Motion control unit, 371 Self-position estimation unit, 372 Sensor fusion unit, 373 Recognition unit, 381 Steering control unit, 382 Brake control unit, 383 Drive control unit, 384 Body system control unit, 385 Light control unit, 386 Horn control unit, 511 Matching unit, 512 Data set generation unit, 513 Learning unit, 611 Matching unit, 612 Data set generation unit, 613 Learning unit

Claims

1. An information processing device comprising: a matching unit that matches an existing point cloud representing the three-dimensional shape of an object whose position is indicated by an existing map, which is pre-prepared map information, with a generated point cloud representing the three-dimensional shape of an object detected by a sensor; and a dataset generation unit that generates a generated map, which is map information corresponding to the generated point cloud, based on the matching result and the existing map, and generates a training dataset using the generated map.

2. The information processing apparatus according to claim 1, wherein the matching unit performs the matching using a portion of the generated point cloud.

3. The information processing apparatus according to claim 2, wherein the matching unit sets a partial region based on the distribution pattern of points in the generated point cloud.

4. The information processing apparatus according to claim 2, wherein the matching unit sets a portion of the region based on the image recognition result of the generated image generated by the image sensor.

5. The information processing apparatus according to claim 2, wherein the matching unit sets a region in the generated point cloud where the density of points is higher than a predetermined standard as a partial region, and performs the matching.

6. The matching unit sets a region in the generated point cloud excluding the point cloud representing the three-dimensional shape of a deformable object as the partial region, and performs the matching, as described in claim 2.

7. The matching unit sets the region of the generated point cloud excluding the point cloud representing the three-dimensional shape of the road surface as the partial region and performs the matching, as described in claim 2.

8. The matching unit sets the region of the generated point cloud excluding the point cloud representing the three-dimensional shape of the animal body as the partial region and performs the matching, as described in claim 2.

9. The information processing apparatus according to claim 1, wherein the training dataset includes generated images generated using a sensor, the generated point cloud, and the generated map.

10. The information processing apparatus according to claim 9, wherein the dataset generation unit generates a training dataset for each domain by associating a plurality of generated images and generated point clouds corresponding to different domains with the generated map.

11. The information processing apparatus according to claim 9, wherein the existing map is a global map which is map information covering a relatively wide area, and the generated map is a local map which is map information covering a narrower area than the global map.

12. The information processing apparatus according to claim 11, wherein the existing map and the generated map are HD maps, and the HD map includes a relatively wide variety of road information, including at least lanes, and is relatively high-precision map information.

13. The information processing apparatus according to claim 12, wherein the training dataset further includes an SD map, the SD map includes fewer types of road information than the HD map, and is relatively low-precision map information.

14. The information processing apparatus according to claim 13, wherein the existing map includes the existing HD map of the HD map and the existing SD map of the SD map, the dataset generation unit generates the generated HD map of the HD map and the generated SD map of the SD map as the generated map, based on the matching result, the existing HD map and the existing SD map, and generates the training dataset including the generated image, the generated point cloud, the generated HD map and the generated SD map.

15. The information processing apparatus according to claim 13, wherein the existing map is the HD map, the dataset generation unit generates a generated HD map of the HD map as the generated map based on the matching result and the existing map, extracts some information from the generated HD map to generate a generated SD map which is the generated map of the SD map, and generates the training dataset which includes the generated image, the generated point cloud, the generated HD map and the generated SD map.

16. The information processing device according to claim 1, further comprising a learning unit that performs training to generate a map generation AI model using the generated training dataset, wherein the map generation AI model is an AI model that infers map information from generated images generated by an image sensor and the generated point cloud.

17. The information processing apparatus according to claim 1, further comprising an acquisition unit that acquires a generated image generated by an image sensor and the generated point cloud, wherein the matching unit performs matching between the acquired generated point cloud and the existing point cloud, and the data set generation unit generates the learning data set including the acquired generated image and the generated point cloud.

18. The information processing apparatus according to claim 1, further comprising: an image generation unit for generating generated images; and a point cloud generation unit for generating the generated point cloud, wherein the matching unit performs matching between the generated point cloud and the existing point cloud; and the dataset generation unit generates the training dataset including the generated generated images and the generated point cloud.

19. An information processing method comprising: matching an existing point cloud representing the three-dimensional shape of an object whose location is indicated by an existing map, which is pre-prepared map information, with a generated point cloud representing the three-dimensional shape of an object detected by a sensor; generating a generated map, which is map information corresponding to the generated point cloud, based on the result of the matching and the existing map; and generating a training dataset using the generated map.

20. A program for causing a computer to perform a process that includes matching an existing point cloud representing the three-dimensional shape of an object whose location is indicated by an existing map, which is pre-prepared map information, with a generated point cloud representing the three-dimensional shape of an object detected by a sensor; generating a generated map, which is map information corresponding to the generated point cloud, based on the result of the matching and the existing map; and generating a training dataset using the generated map.