Information processing device and method, and program

WO2026204251A1PCT designated stage Publication Date: 2026-10-01SONY SEMICON SOLUTIONS CORP
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
PCT/JP2026/008634
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2025-03-28
Filing Date
2026-03-06
Publication Date
2026-10-01

Smart Images

  • Figure JP2026008634_01102026_PF_FP_ABST
    Figure JP2026008634_01102026_PF_FP_ABST
Patent Text Reader

Abstract

The present disclosure relates to an information processing device, a method, and a program that make it possible to more accurately determine the state of a desired traffic light while suppressing cost increases. As regards an AI inference map which includes information equivalent to an HD map and is generated by inference by a map generation AI model using a generation image and a generation point group as input data, the position of the AI inference map is aligned with the coordinate system of an SD map to which traffic light position information indicating the position of a traffic light is linked. A target traffic light is designated on the basis of the traffic light position information. The state of the target traffic light is determined by image analysis of the generation image. The operation of a vehicle is controlled on the basis of the state determination result and the aligned AI inference map. The present disclosure can be applied to, for example, an information processing device, an information processing method, or a program.
Need to check novelty before this filing date? Find Prior Art

Description

Information processing apparatus, method, and program

[0001] The present disclosure relates to an information processing apparatus, method, and program, and particularly relates to an information processing apparatus, method, and program that can determine the state of a desired traffic light with higher accuracy while suppressing an increase in cost.

[0002] Conventionally, for example, there has been a method of controlling automatic driving of vehicles and the like using an HD (High Definition) map, which is map information including traffic information necessary for automatic driving (for example, paths indicating lanes, center lines, intersection connections, etc., traffic signs, connections between paths, associations between traffic lights and paths, etc.) (see, for example, Patent Document 1). However, the generation and information management of HD maps involve high costs, and it has been difficult to prepare HD maps for all roads.

[0003] Accordingly, a method of generating a map of the vicinity of a vehicle using an in-vehicle sensor and controlling automatic driving using the map has been conceived. For example, a method of inferring a map (AI inference map) from outputs (images, point clouds, etc.) of in-vehicle sensors using an AI (Artificial Intelligence) model has been considered.

[0004] However, it has been difficult to grasp the current state of a traffic light from the information of the AI inference map generated in this way. In response to this, for example, there has been a method of determining the state of a traffic light by analyzing a captured image (see, for example, Non-Patent Document 1).

[0005] Japanese Unexamined Patent Application Publication No. 2020-34906

[0006] Tianyu Li, Li Chen, Huijie Wang, Yang Li, Jiazhi Yang, Xiangwei Geng, Shengyin Jiang, Yuting Wang, Hang Xu, Chunjing Xu, Junchi Yan, Ping Luo, Hongyang Li, "Graph-based Topology Reasoning for Driving Scenes", arXiv:2304.05277v2 [cs.CV] 28 Aug 2023

[0007] However, such image analysis tends to be computationally intensive and time-consuming. Performing this type of image analysis without knowing the location of the traffic signals that the vehicle should follow could not only unnecessarily increase the processing load and time, but also reduce the accuracy of determining the desired traffic signal status.

[0008] This disclosure is made in view of these circumstances and aims to enable the determination of the desired traffic signal state with higher accuracy while suppressing cost increases.

[0009] One aspect of this technology is an information processing device comprising: an alignment unit that aligns the position of an AI inference map having information equivalent to an HD map, which is generated by inference of a map generation AI (Artificial Intelligence) model that takes generated images and generated point clouds as input data, with the coordinate system of an SD map associated with traffic light position information indicating the position of traffic lights; a traffic light designation unit that designates a target traffic light whose state is to be determined based on the traffic light position information; a state determination unit that determines the state of the designated target traffic light by performing image analysis on the generated image; and a control unit that controls the operation of a vehicle based on the state determination result of the target traffic light and the aligned AI inference map. The HD map is relatively high-precision map information that includes a relatively wide variety of road information, including at least lanes; the SD map is relatively low-precision map information that includes fewer types of road information than the HD map; the generated image is an image detected by a sensor; and the generated point cloud is a point cloud representing the three-dimensional shape of an object detected by a sensor.

[0010] One aspect of this technology is an information processing method which includes: aligning the position of an AI inference map having information equivalent to an HD map, generated by inference of a map generation AI (Artificial Intelligence) model that takes generated images and generated point clouds as input data, with the coordinate system of an SD map associated with traffic light position information indicating the location of traffic lights; specifying a target traffic light whose state to be determined based on the traffic light position information; determining the state of the specified target traffic light by performing image analysis on the generated image; and controlling the operation of a vehicle based on the state determination result of the target traffic light and the aligned AI inference map. The HD map is relatively high-precision map information that includes a relatively wide variety of road information, including at least lanes; the SD map is relatively low-precision map information that includes fewer types of road information than the HD map; the generated image is an image detected by a sensor; and the generated point cloud is a point cloud representing the three-dimensional shape of an object detected by a sensor.

[0011] One aspect of this technology is a program that causes a computer to perform the following processes: aligning the position of an AI inference map, which has information equivalent to an HD map and is generated by inference of a map generation AI (Artificial Intelligence) model using generated images and generated point clouds as input data, with the coordinate system of an SD map associated with traffic light position information indicating the location of traffic lights; specifying a target traffic light whose state to be determined based on the traffic light position information; determining the state of the specified target traffic light by performing image analysis on the generated images; and controlling the operation of a vehicle based on the state determination result of the target traffic light and the aligned AI inference map. The HD map is relatively high-precision map information that includes a relatively wide variety of road information, including at least lanes; the SD map is relatively low-precision map information that includes fewer types of road information than the HD map; the generated images are images detected by sensors; and the generated point clouds are point clouds representing the three-dimensional shapes of objects detected by sensors.

[0012] In one aspect of this technology, the information processing device and method, as well as the program, the following are performed: the position of an AI inference map, which has information equivalent to an HD map and is generated by inference of a map generation AI (Artificial Intelligence) model that takes generated images and generated point clouds as input data, is aligned with the coordinate system of an SD map to which traffic light position information indicating the location of traffic lights is linked; a target traffic light whose state is to be determined is specified based on the traffic light position information; the state of the specified target traffic light is determined by image analysis of the generated images; and the operation of the vehicle is controlled based on the state determination result of the target traffic light and the aligned AI inference map.

[0013] This is a diagram illustrating the determination of the state of a traffic light. This is a diagram illustrating the alignment of a map. This is a diagram illustrating the setting of a bounding box. This is a diagram illustrating the identification of a target traffic light. This is a diagram showing an example of the main configuration of an autonomous driving system. This is a block diagram showing an example of the main configuration of a server. This is a block diagram showing an example of the configuration of a vehicle control system. This is a diagram showing an example of the sensing area of ​​an external recognition sensor in the vehicle control system of Figure 7. This is a functional block diagram showing an example of the main configuration of functions related to this technology. This is a flowchart showing an example of the traffic light position information processing flow. This is a flowchart showing an example of the acquisition processing flow. This is a flowchart showing an example of the driving control processing flow. This is a flowchart showing an example of the traffic light state determination processing flow.

[0014] The following describes the embodiments for implementing this disclosure. The explanation will be given in the following order: 1. Technical content and supporting literature, etc. 2. Automated driving control 3. Traffic signal status determination using SD map 4. First embodiment (automated driving system) 5. Appendix

[0015] <1. Supporting Documents for Technical Content and Terminology> The scope disclosed in this technology includes not only the contents described in the embodiments, but also the contents described in the following patent and non-patent documents that were publicly known at the time of filing, as well as the contents of other documents referenced in the following patent and non-patent documents.

[0016] Patent document 1: (described above) Non-patent document 1: (described above)

[0017] In other words, the contents described in the aforementioned patent and non-patent documents, as well as the contents of other documents referenced in those patent and non-patent documents, will also serve as a basis for determining the support requirements.

[0018] <2. Automated Driving Control> <Use of HD Maps> Conventionally, there have been methods for controlling the automated driving of vehicles, etc., using a full-area HD (High Definition) map, such as the method described in Patent Document 1. An HD map is map information that includes traffic information necessary for automated driving, and is also called high-precision three-dimensional map data. An HD map may include information such as Paths that show lanes, center lines, intersection connections, traffic signs, connections between Paths, and the linking of traffic lights to Paths. In this specification, an HD map is assumed to be map information that includes a relatively wide variety of road information, including at least lanes, and is of relatively high precision. Furthermore, full-area map information is referred to as a global map, and a full-area HD map (HD map as a global map) is also referred to as a global HD map.

[0019] However, HD maps require a large amount of high-precision information, and their generation is costly. Furthermore, maintaining the information contained in the HD maps in an up-to-date state also incurs high costs. Therefore, even with the creation of a global HD map, it was practically impossible to cover information on all existing roads. Consequently, in the case of controlling autonomous driving based on the global HD map described above, control could not be achieved for roads for which an HD map was not available.

[0020] Therefore, a method was conceived to generate a map of the area around the vehicle using on-board sensors and to control autonomous driving using that map. For example, a method was considered in which an AI (Artificial Intelligence) model is used to infer a map (AI inference map) from the output of the on-board sensors (generated images, generated point clouds, etc.). The generated images are images taken of the area around the vehicle, and the generated point clouds are point cloud data that represent the three-dimensional shape of objects around the vehicle with point-by-point position information (geometry) and attribute information (attributes). These are generated using on-board sensors, etc. When such generated images and generated point clouds are input into a map generation AI model, an AI inference map with information equivalent to an HD map is generated through inference. In other words, this AI inference map is local map information of the area around the vehicle that corresponds to the generated images and generated point clouds. In this specification, a local HD map (HD map as a local map) is also referred to as a local HD map.

[0021] The generated image can be any 2D data generated using a sensor, such as a visible light image (RGB image), an infrared (IR) image, a radar image produced by radio wave irradiation, a thermographic image representing heat distribution, or any other 2D data. In other words, the sensor used to generate the generated image can be an image sensor that detects visible light (RGB image sensor), an infrared sensor that detects infrared rays, a radar sensor that detects reflected waves after irradiating with radio waves, a temperature sensor that detects heat distribution, or any other sensor.

[0022] Furthermore, the generated point cloud can be any 3D data generated using a sensor, and may be a lidar point cloud generated based on lidar data, a radar point cloud generated based on radar data, or a point cloud generated based on other information. In other words, the sensor used to generate the generated point cloud may be a lidar sensor that emits laser light and detects the reflected light, a radar sensor, or any other sensor. Note that the generated point cloud can be any 3D data and is not limited to point clouds. For example, it may be 3D data other than a point cloud, such as a mesh. In this specification, a point cloud will be used as an example of this 3D data.

[0023] <Understanding the Status of Traffic Lights> For example, as shown in Figure 1A, generated images and point clouds can be generated in a vehicle 11 traveling on an actual road, and these can be used to generate an AI inference map 12 (local HD map) of the area around the vehicle 11. By using such an AI inference map 12 that has information equivalent to an HD map, automated driving control becomes possible even on roads where a global HD map does not exist.

[0024] However, conventional methods typically used separate AI models for map inference (lanes, drivable areas, pedestrian crossings, driving paths, etc.) and for recognizing the location and state of traffic lights. As a result, the information in this AI inference map 12 cannot capture the state of traffic lights 13, which changes over time (for example, whether they are blue, yellow, or red). Therefore, it was difficult to correctly control autonomous driving using only the information in this AI inference map. In contrast, there is a method, such as the method described in Non-Patent Document 1, that determines the state of traffic lights by analyzing captured images. By detecting traffic lights and determining their state using such a method, it is possible to add information indicating the state of traffic lights to the information in the AI ​​inference map, thereby enabling more accurate control of autonomous driving.

[0025] However, such image analysis tends to be computationally intensive and time-consuming. If this type of image analysis is performed without knowing the location of the traffic lights that the vehicle should obey, it is necessary to analyze the generated image for every frame, which could lead to an unrealistic increase in processing load and time. Furthermore, even if the traffic lights that the vehicle should obey are not present in the generated image, image analysis is still performed, which could unnecessarily increase processing load and time. Also, even if traffic lights are present in the generated image, they are not necessarily the desired traffic lights (the traffic lights that the vehicle should obey). With this method, it is difficult to correctly identify which signals the vehicle should obey, and there is a risk of misjudging the state of traffic lights that the vehicle should not obey. In other words, there is a risk of reduced accuracy in determining the state of the desired traffic lights.

[0026] For example, it is conceivable to use information indicating the location of traffic lights in map data to identify the traffic light that a vehicle should obey. However, as in example A in Figure 1, the AI ​​inference map 12 generated using the generated image and generated point cloud does not include information indicating the location of the traffic light 13. Therefore, even if the information from this AI inference map is combined with a method such as the one described in Non-Patent Document 1, it is difficult to suppress the increase in processing load and processing time for determining the state of the desired traffic light, or the decrease in the accuracy of that state determination.

[0027] For example, in services that provide route guidance, such as so-called car navigation systems, there was a method of using SD (Standard Definition) map information. SD maps are map information with less information than HD maps. In this specification, SD maps are assumed to contain fewer types of road information than HD maps and to be relatively low-precision map information. For example, such an SD map may include information indicating the location of traffic lights. Or, information indicating the location of traffic lights may be linked to an SD map. In this specification, these cases are collectively described as "information indicating the location of traffic lights is linked to an SD map." In other words, "information indicating the location of traffic lights is linked to an SD map" also includes "it is included in the SD map." To put it another way, in this case, "information indicating the location of traffic lights" may be included in the SD map or may be managed as separate data from the SD map. It is conceivable to use such "information indicating the location of traffic lights" linked to an SD map to identify the traffic lights that a vehicle should obey.

[0028] However, the accuracy of location information in SD maps is generally low. For example, as shown in Figure 1B, the location of road 20 on the SD map may differ from the actual location of road 10. Therefore, the location of traffic lights 23 included in or linked to the SD map may also differ from the actual location of the traffic lights. Consequently, even when this traffic light location information is combined with an AI inference map, it is difficult to correctly identify the traffic light that a vehicle should obey. In other words, it is difficult to suppress the increased processing load and processing time for determining the state of the desired traffic light, as well as the decrease in the accuracy of that state determination.

[0029] <3. Determining the state of traffic lights using SD maps> <Shifting the position of the AI ​​inference map> Therefore, the AI ​​inference map is aligned with the coordinate system of the SD map.

[0030] For example, the information processing device may include: an alignment unit that aligns the position of an AI inference map, which has information equivalent to an HD map and is generated by inference of a map generation AI (Artificial Intelligence) model that takes generated images and generated point clouds as input data, with the coordinate system of an SD map to which traffic light position information indicating the location of traffic lights is associated; a traffic light designation unit that designates a target traffic light whose state is to be determined based on the traffic light position information; a state determination unit that determines the state of the designated target traffic light by performing image analysis on the generated image; and a control unit that controls the operation of the vehicle based on the state determination result of the target traffic light and the aligned AI inference map.

[0031] Furthermore, the information processing method executed by the information processing device includes aligning the position of an AI inference map, which has information equivalent to an HD map and is generated by inference of a map generation AI (Artificial Intelligence) model using generated images and generated point clouds as input data, with the coordinate system of an SD map to which traffic light position information indicating the location of traffic lights is linked; specifying a target traffic light whose state to be determined based on that traffic light position information; determining the state of the specified target traffic light by performing image analysis on the generated image; and controlling the operation of the vehicle based on the state determination result of the target traffic light and the aligned AI inference map.

[0032] Furthermore, the program is designed to cause a computer to perform the following processes: aligning the position of an AI inference map, which has information equivalent to an HD map and is generated by inference of a map generation AI (Artificial Intelligence) model using generated images and generated point clouds as input data, with the coordinate system of an SD map associated with traffic light position information indicating the location of traffic lights; specifying a target traffic light whose state to be determined based on that traffic light position information; determining the state of the specified target traffic light by performing image analysis on the generated image; and controlling the operation of the vehicle based on the state determination result of the target traffic light and the aligned AI inference map.

[0033] HD maps are assumed to be relatively high-precision map information that includes a relatively wide variety of road information, including at least lanes. SD maps are assumed to be relatively low-precision map information that includes fewer types of road information than HD maps. Generated images are assumed to be images detected by sensors. The number of generated images input to the map generation AI model may be any number. For example, multiple generated images with different directions, fields of view, resolutions, or data types may be included in the input data to the map generation AI model. Generated point clouds are assumed to be point clouds that represent the three-dimensional shapes of objects detected by sensors. Traffic light position information is two-dimensional position information that indicates the position of traffic lights in the map information, that is, the position of traffic lights on a plane parallel to the ground (i.e., two-dimensional position).

[0034] For example, the GPS signal received by vehicle 11 indicates the vehicle's actual position, but it is misaligned with the position of road 20 on the SD map. Therefore, as shown in Figure 2, the GPS signal is shifted onto the SD map coordinates, taking into account the distance between the GPS signal and the position on the SD map (the distance between the position of vehicle 11 and the position of vehicle 11'), the vehicle's orientation, the direction of travel on the road, etc., so that vehicle 11 is positioned on road 20 on the SD map. In other words, the position of vehicle 11 indicated by the GPS signal is shifted to the position of vehicle 11'. This shift causes the position of the AI ​​inference map 12 to shift onto the SD map coordinates (the position of AI inference map 12'). In other words, the AI ​​inference map 12 (AI inference map 12') represents map information around vehicle 11'.

[0035] By doing so, the position of traffic light 23 in the SD map can be correctly mapped to the AI ​​inference map 12 (AI inference map 12'). In other words, the information from the AI ​​inference map 12 and the traffic light position information linked to the SD map can be used in combination, making it possible to correctly identify the traffic light (position) that vehicle 11 should obey. Since the position of the traffic light that vehicle 11 should obey is identified, that traffic light can be correctly identified in the generated image. Therefore, by analyzing the generated image, the state of the traffic light that vehicle 11 should obey can be grasped more easily and accurately. In other words, the state of the desired traffic light can be determined with higher accuracy while suppressing an increase in cost.

[0036] <Bounding Box> A bounding box may be set within the generated image to indicate a portion of that image, and the image within that bounding box may be analyzed. For example, the information processing device may further include a bounding box setting unit that sets a bounding box for the generated image. The state determination unit may then determine the state of the target traffic signal by analyzing the area within the bounding box of the generated image.

[0037] For example, suppose a traffic light 42 exists in the generated image 41, as shown in Figure 3. A bounding box 43, indicated by a dotted line frame, is set for such a generated image 41, and the traffic light 42 is detected and its state is determined by image analysis only within the bounding box 43. By doing this, the range of image analysis is narrowed compared to when the entire generated image 41 is image analyzed to detect the traffic light 42 and determine its state, thus suppressing the increase in processing load and processing time for the image analysis.

[0038] In other words, it is desirable that the bounding box encloses the area in the generated image where the desired traffic light is more likely to exist, and excludes the area where it is less likely to exist. The more the bounding box is set to enclose the desired traffic light and narrow its range, the more the increase in processing load and processing time for image analysis can be suppressed.

[0039] Therefore, the three-dimensional position of the desired traffic light may be estimated, and a bounding box may be set in the generated image to include the estimated three-dimensional position. For example, in an information processing device, a bounding box setting unit may estimate three-dimensional (including height) traffic light position information based on traffic light position information linked to an SD map, and set a bounding box in the generated image to include the estimated three-dimensional position.

[0040] For example, the two-dimensional position of a traffic light (its position in map information) can be determined based on traffic light position information linked to an SD map. Therefore, by estimating the height of the traffic light by some method, the three-dimensional position of the desired traffic light can be estimated. Note that any method can be used to estimate the height of the traffic light. For example, it may suffice to estimate only the range in the height direction where there is a sufficiently high probability that a traffic light exists.

[0041] Then, based on the vehicle's position and imaging direction aligned to the SD map coordinate system, the estimated 3D position of the desired traffic light (or a predetermined range including the 3D position of that traffic light) is identified within the generated image, and a bounding box is set within the generated image to include the identified position (or range).

[0042] For example, traffic lights are subject to height restrictions due to legal regulations. These regulations can be used to limit the position and range of the bounding box in the height direction. For instance, in an information processing device, the bounding box setting unit may set the bounding box in the generated image to a height that complies with the legal regulations regarding traffic lights. By doing so, the bounding box can be set in the part of the generated image where the presence of a traffic light is more likely. Therefore, the processing load and processing time for image analysis to determine the state of traffic lights can be suppressed.

[0043] Furthermore, a traffic signal portion may be estimated based on the generated point cloud, and a bounding box may be set to include a portion corresponding to the estimated portion in the generated image. For example, in the generated point cloud, a portion having the shape of a traffic signal or a shape similar to a traffic signal is highly likely to be a traffic signal. Therefore, in the generated image, a bounding box may be set to include a region corresponding to such a portion. For example, in an information processing apparatus, a bounding box setting unit may set a bounding box based on a distribution pattern of points in the generated point cloud. By this configuration, a bounding box can be set in a portion of the generated image where a traffic signal is more likely to be present. Accordingly, an increase in the processing load and processing time of image analysis for determining the state of a traffic signal can be suppressed.

[0044] Conversely, a bounding box may be set by excluding regions where a traffic signal is unlikely to be present based on the generated point cloud. For example, a bounding box may be set by excluding portions where points are sparse or whose distribution pattern is clearly different from that of a traffic signal. For example, in an information processing apparatus, a bounding box setting unit may set a bounding box by excluding regions where the target traffic signal is unlikely to be present based on the distribution pattern. By this configuration, a bounding box can be set by excluding portions of the generated image where a traffic signal is less likely to be present. Accordingly, an increase in the processing load and processing time of image analysis for determining the state of a traffic signal can be suppressed.

[0045] Furthermore, a confidence level (likelihood of being a traffic light) may be set based on the distribution pattern of points in the generated point cloud, and this confidence level may be used to identify traffic lights. For example, parts with a higher confidence level (parts that are more likely to be traffic lights) may be given higher priority in being identified as traffic lights. For example, in an information processing device, the bounding box setting unit may further set a confidence level based on the distribution pattern. Then, the state determination unit may identify the target traffic light within the bounding box based on the image analysis results and their confidence levels, and determine the state of the identified target traffic light. By doing so, traffic lights can be identified more accurately.

[0046] <Identifying a Desired Traffic Signal within a Bounding Box> For example, there may be multiple traffic signals within a bounding box. In such cases, it may be possible to identify a desired traffic signal (the traffic signal that the vehicle should obey). For example, in an information processing device, a state determination unit may identify a target traffic signal from among the multiple traffic signals present within the bounding box and determine the state of the identified target traffic signal. By doing so, the state of the desired traffic signal can be determined more accurately.

[0047] The identification of the desired traffic signal may be performed by any method. For example, the desired traffic signal may be identified based on whether the light-emitting unit is visible in the generated image. For example, in the information processing apparatus, the state determination unit may identify the target traffic signal based on whether the light-emitting unit is visible in the generated image. For example, when the light-emitting unit is visible in the generated image, the traffic signal having that light-emitting unit may be identified as the target traffic signal. For example, as shown in FIG. 4A, traffic signals 51 and 52 exist within the bounding box 43. The surface of traffic signal 51 including the light-emitting unit faces toward the near side, so the light-emitting unit can be visually recognized. In contrast, the surface of traffic signal 52 including the light-emitting unit faces toward the far side, so the light-emitting unit cannot be visually recognized. In such a case, traffic signal 52 is clearly a traffic signal for the opposite lane, and not a traffic signal that the vehicle should comply with. Therefore, a traffic signal whose light-emitting unit is visible in the generated image (in the example of FIG. 4A, traffic signal 51) may be identified as the target traffic signal (the desired traffic signal). By doing this, the desired traffic signal can be identified more accurately.

[0048] In addition, the desired traffic signal may be identified according to the size of the light-emitting unit in the generated image. For example, in the information processing apparatus, the state determination unit may identify the target traffic signal based on the size of the light-emitting unit in the generated image. For example, the traffic signal with the largest light-emitting unit in the generated image may be identified as the target traffic signal. For example, as shown in FIG. 4B, traffic signals 61, 62 and 63 exist within the bounding box 43. Among these, the light-emitting unit of traffic signal 61 is the largest. In such a case, it is estimated that among traffic signals 61 to 63, traffic signal 61 is located closest to the near side. Therefore, traffic signal 61 may be identified as the traffic signal that the vehicle should comply with (the desired traffic signal). By doing this, the desired traffic signal can be identified more accurately.

[0049] Furthermore, the desired traffic light may be identified by the orientation of the light-emitting part in the generated image. For example, in an information processing device, a state determination unit may identify the target traffic light based on the orientation of the light-emitting part in the generated image. For example, a traffic light whose light-emitting part faces forward (towards the user) in the generated image may be identified as the target traffic light. For example, as shown in Figure 4C, traffic lights 71 and 72 exist within the bounding box 43, but the light-emitting part of traffic light 71 faces forward (towards the user). In contrast, the light-emitting part of traffic light 72 faces sideways (to the left). In such a case, it is clear that the traffic light 71 facing forward is the traffic light that the vehicle should obey (the desired traffic light). Therefore, that traffic light 71 may be identified as the traffic light that the vehicle should obey (the desired traffic light). By doing so, the desired traffic light can be identified more accurately.

[0050] Furthermore, a desired traffic light may be identified based on the state changes of the traffic lights. For example, in an information processing device, a state determination unit may identify a target traffic light based on the state changes of the traffic lights. For example, even if a traffic light is identified as a "traffic light that the vehicle should obey" based on its position, the vehicle control may not be able to react in time if the state of that traffic light changes to "red" just before the vehicle passes that "traffic light that the vehicle should obey". In such cases, that traffic light may be excluded from the target traffic lights, and the vehicle may not obey that traffic light. For example, if the state of the traffic light changes to "red" after the vehicle has entered an intersection corresponding to a "traffic light that the vehicle should obey", the automated driving control unit may control the vehicle to pass through the intersection without stopping it. By doing so, the automated driving of the vehicle can be controlled in a practical way.

[0051] <Designation of Target Traffic Lights> For example, even if a traffic light that a vehicle should follow next exists in the generated image, if that traffic light is located far from the vehicle, the determination of the traffic light's state cannot be used for controlling the automatic driving system. In other words, unnecessary processing may increase the processing load. Therefore, the target traffic light may be designated at a more appropriate time so that the state determination is performed at a more appropriate time. For example, the traffic light may be designated as a target traffic light when the vehicle approaches within a predetermined range from the traffic light (i.e., when the distance between the vehicle and the traffic light becomes shorter than a predetermined standard). For example, in an information processing device, the traffic light designation unit may designate a traffic light located within a predetermined range from the vehicle as a target traffic light based on the traffic light position information. By doing so, the increase in processing load and processing time can be suppressed.

[0052] The information processing device may further include an image generation unit that generates generated images. This image generation unit may generate any number of generated images. For example, this image generation unit may generate multiple generated images with different orientations, angles of view, resolutions, or data types as input data to be input to a map generation AI model. The information processing device may further include a point cloud generation unit that generates generated point clouds. Furthermore, the information processing device may further include an AI inference map generation unit that generates an AI inference map using the generated images and generated point clouds with the map generation AI model.

[0053] Furthermore, the information processing device may further include a storage unit that stores linked SD maps and traffic signal location information. The alignment unit may then align the position of the AI ​​inference map to the coordinate system of the SD map. Additionally, the traffic signal designation unit may determine the state based on the traffic signal location information.

[0054] Furthermore, guidance may be provided based on the status determination results of traffic signals, for example, as in a navigation system. For example, the information processing device may further include a guidance unit that provides guidance regarding vehicle operation based on the status determination results of the target traffic signals and a aligned AI inference map. In this way, the status determination results of traffic signals may be applied not only to the control of autonomous driving but also to any other processing.

[0055] <4. First Embodiment> <Automated Driving System> The technology described above can be applied to any configuration (device, system, unit, processing unit, etc.). Figure 5 shows an example of the main configuration of an automated driving system, which is one form of an information processing system to which this technology is applied.

[0056] The automated driving system 100 shown in Figure 5 comprises a server 101, vehicles 102-1, 102-2, and 102-3, which are connected to each other via a network 110 so as to be able to communicate with one another. When it is not necessary to distinguish between vehicles 102-1, 102-2, and 102-3, they will be referred to as vehicle 102.

[0057] The automated driving system 100 is a system that performs processing related to the control of the automated driving of the vehicle 102. In Figure 5, three vehicles 102 (vehicle 102-1, vehicle 102-2, and vehicle 102-3) are shown, but the number of vehicles 102 that the automated driving system 100 has may be any number, for example, two or fewer, or four or more. Also, in Figure 5, one server 101 is shown, but the number of servers 101 that the automated driving system 100 has may be any number, for example, two or more. Similarly, the automated driving system 100 may have multiple networks 110.

[0058] This network 110 is a communication network composed of any communication medium. Communication conducted through network 110 may be wired communication, wireless communication, or both. In other words, network 110 may be a communication network for wired communication, a communication network for wireless communication, or a communication network composed of both. Furthermore, network 110 may be composed of a single communication network or of multiple communication networks.

[0059] For example, the internet may be included in this network 110. Public telephone network may also be included in this network 110. Furthermore, wide-area communication networks for wireless mobile devices, such as so-called 3G and 4G networks, may also be included in this network 110. For example, LPWA (Low Power Wide Area) communication networks such as LTE-M, which enable long-distance data communication and have low power consumption, may also be included in this network 110. Furthermore, WAN (Wide Area Network) and LAN (Local Area Network) may also be included in this network 110. Furthermore, wireless communication networks that perform communication compliant with the Bluetooth® standard may also be included in this network 110. Near-field communication (NFC) communication channels may also be included in this network 110. Furthermore, infrared communication channels may also be included in this network 110. Finally, wired communication networks compliant with standards such as HDMI (High-Definition Multimedia Interface)® and USB (Universal Serial Bus)® may also be included in this network 110. Thus, the network 110 may include communication networks and communication channels of any communication standard. Furthermore, these communication networks and communication channels may include not only communication media such as cables, but also devices and circuits necessary for communication, such as communication equipment and relay equipment.

[0060] The server 101 and the vehicle 102 can communicate with other devices and exchange information via this network 110. In other words, the server 101 and the vehicle 102 may communicate in any manner conforming to any of the various communication standards described above. For example, the server 101 and the vehicle 102 may use wired communication, wireless communication, or both.

[0061] <Server> Server 101 performs processing related to the control of the vehicle 102's automatic driving. For example, Server 101 may link SD maps with traffic signal location information. Server 101 may also supply the linked SD maps and traffic signal location information to the vehicle 102 as a result of this processing.

[0062] Figure 6 is a block diagram showing an example of the hardware configuration of server 101. As shown in Figure 6, in server 101, the CPU (Central Processing Unit) 201, ROM (Read Only Memory) 202, and RAM (Random Access Memory) 203 are interconnected via a bus 204.

[0063] An input / output interface 210 is also connected to the bus 204. An input / output interface 210 is connected to an input unit 211, an output unit 212, a storage unit 213, a communication unit 214, and a drive 215.

[0064] The input unit 211 consists of, for example, a keyboard, mouse, microphone, touch panel, and input terminals. The output unit 212 consists of, for example, a display, speaker, and output terminals. The storage unit 213 consists of, for example, a hard disk, RAM disk, and non-volatile memory. The communication unit 214 consists of, for example, a network interface. The drive 215 drives removable media 221 such as a magnetic disk, optical disk, magneto-optical disk, or semiconductor memory.

[0065] In the server 101 configured as described above, the CPU 201 loads, for example, a program stored in the memory unit 213 into the RAM 203 via the input / output interface 210 and the bus 204, and executes it, thereby performing a series of processes described later. The RAM 203 also appropriately stores data necessary for the CPU 201 to perform various processes.

[0066] The program executed by the server 101 can be recorded and applied on a removable media 221, such as a package media. In this case, the program can be installed on the storage unit 213 via the input / output interface 210 by inserting the removable media 221 into the drive 215.

[0067] Furthermore, this program can also be provided via wired or wireless transmission media such as local area networks, the internet, or digital satellite broadcasting. In that case, the program can be received by the communication unit 214 and installed in the storage unit 213.

[0068] In addition, this program can be pre-installed in ROM 202 or memory unit 213.

[0069] The hardware configuration of server 101 shown in Figure 6 is described assuming that server 101 is comprised of a single device. In reality, the functions of server 101 (i.e., the configuration shown in Figure 6) may be implemented by a single device or by multiple devices. Furthermore, server 101 may be configured as a so-called cloud server.

[0070] <Vehicle> Vehicle 102 performs processes related to autonomous driving. For example, vehicle 102 may generate generated images or generated point clouds using on-board sensors, etc. Vehicle 102 may also input the generated images or generated point clouds into a map generation AI model for inference and generate a generated map. Vehicle 102 may also perform processes related to autonomous driving using the generated map.

[0071] The vehicle 102 may have a configuration as shown in Figure 7, for example. Figure 7 is a block diagram showing an example configuration of a vehicle control system 311, which is a non-limiting example of a mobile device control system to which this technology is applied.

[0072] The vehicle control system 311 is installed in the vehicle 102 and performs processing related to the automation of the vehicle's operation. This automation includes driving automation from Level 1 to Level 5, as well as remote driving and / or remote assistance of the vehicle 102 by a remote driver. The levels of driving automation may refer to the Society of Automotive Engineers (SAE) J3016™ APL2021 Levels of Driving Automation, where SAE Level 0 represents the lowest level of driving automation and SAE Level 5 represents the highest level of driving automation. For example, SAE Level 1 driving automation may consist of driver assistance functions that provide the driver with steering or brake / acceleration support, and SAE Level 5 driving automation may consist of automated driving functions that enable the vehicle to be driven under all conditions.

[0073] The vehicle control system 311 includes a vehicle control ECU (Electronic Control Unit) 321, a communication unit 322, a map information storage unit 323, a location information acquisition unit 324, an external recognition sensor 325, an in-vehicle sensor 326, a vehicle sensor 327, a memory unit 328, an automated driving control unit 329, a DMS (Driver Monitoring System) 330, an HMI (Human Machine Interface) 331, and a vehicle control unit 332.

[0074] Two or more (or, in some cases, all) of the following are connected to communicate with each other via a communication network 341: the vehicle control ECU 321, the communication unit 322, the map information storage unit 323, the location information acquisition unit 324, the external recognition sensor 325, the in-vehicle sensor 326, the vehicle sensor 327, the memory unit 328, the driving automation control unit 329, the DMS 330, the HMI 331, and the vehicle control unit 332. The communication network 341 is composed of an in-vehicle communication network or bus that conforms to digital bidirectional communication standards such as CAN (Controller Area Network), LIN (Local Interconnect Network), LAN (Local Area Network), FlexRay®, and Ethernet®. In some embodiments, the communication network 341 may have two or more types of communication networks, and different types of communication networks may be used depending on the type of data being transmitted. For example, CAN may be applied to data related to vehicle control, and Ethernet may be applied to large-capacity data. In some embodiments, two or more (or possibly all) units of the vehicle control system 311 may be directly connected using wireless communication (e.g., relatively short-range communication) without going through the communication network 341. In some embodiments, the wireless communication may use near-field wireless communication technology. Non-limiting examples of near-field wireless communication technology include Near Field Communication (NFC) and Bluetooth®. In some embodiments, two or more (or possibly all) units of the vehicle control system 311 may be connected using the communication network 341 and wireless communication technology (e.g., near-field wireless communication technology).

[0075] In the following embodiment, where two or more units of the vehicle control system 311 communicate via the communication network 341, the description of the communication network 341 will be omitted. For example, in an embodiment where the vehicle control ECU 321 and the communication unit 322 communicate via the communication network 341, it will simply be described as the vehicle control ECU 321 and the communication unit 322 communicating.

[0076] The vehicle control ECU 321 is composed of various processors, such as a CPU (Central Processing Unit) and an MPU (Micro Processing Unit). The vehicle control ECU 321 controls the functions of the entire vehicle control system 311 or a part of it.

[0077] The communication unit 322 communicates with various devices inside the vehicle 102 (hereinafter referred to as in-vehicle devices), various devices outside the vehicle 102 (hereinafter referred to as external devices), other vehicles, base stations, etc., and transmits and receives various types of data. In some embodiments, the communication unit 322 may use multiple communication technologies to perform communication.

[0078] A non-limiting example of communication between the communication unit 322 and external equipment will be briefly described. In some embodiments, the communication unit 322 may communicate with servers (hereinafter referred to as "external servers") located on an external network via a base station or access point using wireless communication technology. Examples of non-limiting wireless communication technologies include 5G (fifth-generation mobile communication system), LTE (Long Term Evolution), and DSRC (Dedicated Short Range Communications). External networks that the communication unit 322 can communicate with include, for example, the internet, a cloud network, or a network specific to a carrier. The communication technology used by the communication unit 322 to communicate with an external network is not particularly limited, as long as it is a wireless communication technology that enables digital two-way communication at a predetermined communication speed and over a predetermined distance.

[0079] In some embodiments, the communication unit 322 may communicate with terminals located near the vehicle using P2P (Peer To Peer) technology. Terminals located near the vehicle include, for example, terminals worn by relatively slow-moving objects such as pedestrians and cyclists, terminals installed in fixed locations such as stores, and / or MTC (Machine Type Communication) terminals. In some embodiments, the communication unit 322 may perform V2X (Vehicle to Everything) communication. V2X communication generally refers to communication between the vehicle and other entities. Non-exclusive examples of V2X communication include vehicle-to-vehicle communication with other vehicles, vehicle-to-infrastructure communication with roadside devices, etc., vehicle-to-home communication with homes, and vehicle-to-pedestrian communication with terminals carried or worn by pedestrians.

[0080] In some embodiments, the communication unit 322 may receive a program from outside the vehicle 102 to update the software that controls the operation of the vehicle control system 311 (for example, over the air). In some embodiments, the communication unit 322 may receive map information, traffic information, information about the vehicle 102's surroundings, etc., from outside the vehicle 102. In some embodiments, the communication unit 322 may transmit information about the vehicle 102, information about the vehicle 102's surroundings, etc., to an external device or external network. Non-limiting examples of information about the vehicle 102 that the communication unit 322 transmits to an external device or external network include data indicating the status of the vehicle 102, recognition results from the recognition unit 373, etc. In some embodiments, the communication unit 322 may communicate with a vehicle emergency call system. Non-limiting examples of a vehicle emergency call system include e-Call, etc.

[0081] In some embodiments, the communication unit 322 may receive electromagnetic waves transmitted by a road traffic information communication system. In some embodiments, such electromagnetic waves may be transmitted using radio beacons, optical beacons, FM multiplex broadcasting, etc.

[0082] A non-limiting example of communication with in-vehicle equipment that the communication unit 322 can perform will be outlined below. In some embodiments, the communication unit 322 may communicate with in-vehicle equipment using wireless communication. For example, in some embodiments, the communication unit 322 may communicate with in-vehicle equipment wirelessly using wireless communication technology that enables digital bidirectional communication at a predetermined or higher communication speed. Non-limiting examples of wireless communication technologies include wireless LAN, Bluetooth®, NFC, and WUSB (Wireless USB). Not limited to these, the communication unit 322 may also communicate with in-vehicle equipment using wired communication (in addition to or as an alternative to wireless communication). For example, in some embodiments, the communication unit 322 may communicate with in-vehicle equipment via wired communication through a cable connected to a connection terminal (not shown). In some embodiments, the communication unit 322 may communicate with in-vehicle equipment using wired communication technology that enables digital bidirectional communication at a predetermined or higher communication speed. Non-exclusive examples of wired communication technologies include USB (Universal Serial Bus), HDMI (High-Definition Multimedia Interface) (registered trademark), and MHL (Mobile High-definition Link).

[0083] Here, in-vehicle equipment refers to, for example, equipment located inside the vehicle 102 that is not connected to the communication network 341. In-vehicle equipment is divided into equipment that constitutes the vehicle control system 311 and equipment that does not. Non-exclusive examples of in-vehicle equipment that does not constitute the vehicle control system 311 include mobile devices and wearable devices owned by users of the vehicle 102 (e.g., the driver, passengers), and information equipment temporarily installed inside the vehicle 102. These devices can, for example, be moved outside the vehicle 102 and become external equipment.

[0084] The map information storage unit 323 stores maps acquired from external devices or external networks and / or maps created by the vehicle 102. For example, the map information storage unit 323 may store three-dimensional high-precision maps, global maps with lower precision than high-precision maps but covering a wide area, etc.

[0085] High-precision maps include, for example, dynamic maps, point cloud maps, and vector maps. A dynamic map may be a map consisting of four layers: dynamic information, semi-dynamic information, semi-static information, and static information, and may be provided to the vehicle 102 from an external server or the like. A point cloud map may be a map composed of point clouds (point cloud data). A vector map may be a map adapted for automated driving by associating traffic information, such as the locations of lanes and traffic lights, with a point cloud map.

[0086] The point cloud map and vector map may be provided from, for example, an external server, or they may be created in the vehicle 102 as maps for matching with the local map described later, based on sensing results from the camera 351, radar 352, LiDAR 353, etc., and stored in the map information storage unit 323. In addition, if high-precision maps are provided from an external server, in order to reduce communication capacity, map data of, for example, several hundred square meters relating to the planned route that the vehicle 102 will travel may be obtained from the external server.

[0087] The location information acquisition unit 324 acquires location information of the vehicle 102. The acquired location information may be supplied to the driving automation control unit 329. In some embodiments, the location information acquisition unit 324 may receive GNSS (Global Navigation Satellite System) signals from GNSS satellites. In some embodiments, the location information acquisition unit 324 may receive signals from beacons or the like.

[0088] The external recognition sensor 325 is equipped with various sensors used to recognize the external conditions of the vehicle 102, and supplies sensor data from one or more (or, in some cases, all) sensors to one or more (or, in some cases, all) units of the vehicle control system 311. The types and number of sensors equipped in the external recognition sensor 325 are arbitrary.

[0089] In some embodiments, the external recognition sensor 325 may include a camera 351, a radar 352, a LiDAR (Light Detection and Ranging, Laser Imaging Detection and Ranging) 353, and an ultrasonic sensor 354. However, the external recognition sensor 325 may also be configured to include one or more of the cameras 351, radar 352, LiDAR 353, and ultrasonic sensor 354. The number of cameras 351, radar 352, LiDAR 353, and ultrasonic sensors 354 is not particularly limited as long as it is a number that can be realistically installed in the vehicle 102. Furthermore, the types of sensors included in the external recognition sensor 325 are not limited to this example, and the external recognition sensor 325 may include other types of sensors. Examples of the sensing areas of each sensor included in the external recognition sensor 325 will be described later.

[0090] Camera 351 can use any suitable shooting method. In some embodiments, camera 351 may use a shooting method capable of distance measurement. Non-limiting examples of cameras using a shooting method capable of distance measurement include ToF (Time of Flight) cameras, stereo cameras, monocular cameras, and infrared cameras. However, camera 351 may not be capable of distance measurement and may simply be for acquiring images.

[0091] In some embodiments, the external recognition sensor 325 may include environmental sensors for detecting characteristics of the environment around the vehicle 102. Non-limiting examples of detectable environmental characteristics include weather, climate, brightness, etc. In some embodiments, the environmental sensors may include various sensors such as raindrop sensors, fog sensors, sunshine sensors, snow sensors, and illuminance sensors.

[0092] In some embodiments, the external recognition sensor 325 may include a microphone used for detecting sounds around the vehicle 102 and the location of sound sources.

[0093] The in-vehicle sensor 326 is equipped with various sensors for detecting information inside the vehicle 102 and supplies sensor data from one or more (or, in some cases, all) sensors to one or more (or, in some cases, all) units of the vehicle control system 311. The types and number of sensors equipped with the in-vehicle sensor 326 are not particularly limited, as long as they are of a type and number that can be realistically installed in the vehicle 102.

[0094] In some embodiments, the in-vehicle sensor 326 may include one or more sensors from among a camera, radar, seat sensor, microphone, and biosensor. In some embodiments, the camera included in the in-vehicle sensor 326 may use a distance-measuring shooting method. Non-limiting examples of cameras using a distance-measuring shooting method include ToF cameras, stereo cameras, monocular cameras, and infrared cameras. However, the camera included in the in-vehicle sensor 326 may not be for distance measurement and may simply be for acquiring captured images. The biosensor included in the in-vehicle sensor 326 may be provided, for example, on the seat or steering wheel, and may detect various biometric information of the user.

[0095] The vehicle sensor 327 is equipped with various sensors for detecting the state of the vehicle 102 and supplies sensor data from one or more (or, in some cases, all) sensors to one or more (or, in some cases, all) units of the vehicle control system 311. The types and number of sensors equipped with the vehicle sensor 327 are not particularly limited, as long as they are of a type and number that can realistically be installed on the vehicle 102.

[0096] In some embodiments, the vehicle sensor 327 may include a speed sensor, an acceleration sensor, an angular velocity sensor (gyro sensor), and / or an inertial measurement unit (IMU) integrating them. In some embodiments, the vehicle sensor 327 may include a steering angle sensor for detecting the steering angle of the steering wheel, a yaw rate sensor, an accelerator sensor for detecting the amount of operation of the accelerator pedal (e.g., pedal force, pedal stroke), and / or a brake sensor for detecting the amount of operation of the brake pedal (e.g., pedal force, pedal stroke). In some embodiments, the vehicle sensor 327 may include a rotation sensor for detecting the rotational speed of the engine or motor, an air pressure sensor for detecting the air pressure of the tires, a slip ratio sensor for detecting the slip ratio of the tires, and / or a wheel speed sensor for detecting the rotational speed of the wheels. In some embodiments, the vehicle sensor 327 may include a battery sensor for detecting the remaining charge and temperature of the battery, and / or an impact sensor capable of detecting external impacts.

[0097] The storage unit 328 includes at least one of a non-volatile storage medium and a volatile storage medium, and stores data and programs. Non-limiting examples of storage media include magnetic storage devices such as EEPROM (Electrically Erasable Programmable Read Only Memory), RAM (Random Access Memory), and / or HDD (Hard Disc Drive), semiconductor storage devices, optical storage devices, and magneto-optical storage devices. The storage unit 328 stores various programs and data used by one or more (or in some cases all) units of the vehicle control system 311. In some embodiments, the storage unit 328 may include an EDR (Event Data Recorder) or a DSSAD (Data Storage System for Automated Driving) to store information about the vehicle 102 before and after an event such as an accident, and information acquired by the in-vehicle sensor 326.

[0098] The automated driving control unit 329 controls the automated driving functions of the vehicle 102. In some embodiments, the automated driving control unit 329 may also include an analysis unit 361, an action planning unit 362, and an operation control unit 363.

[0099] The analysis unit 361 performs analysis processing of the vehicle 102 and / or the surrounding conditions. The analysis unit 361 includes a self-position estimation unit 371, a sensor fusion unit 372, and a recognition unit 373.

[0100] In some embodiments, the self-position estimation unit 371 may estimate the vehicle 102's position based on sensor data from an external recognition sensor 325 and a high-precision map stored in a map information storage unit 323. For example, the self-position estimation unit 371 may generate a local map based on sensor data from an external recognition sensor 325 and estimate the vehicle 102's position by matching the local map with a high-precision map. The position of the vehicle 102 may be based on, for example, the center of the rear wheel relative to the axle.

[0101] In some embodiments, the local map may be a three-dimensional high-precision map, an occupancy grid map, or the like, created using technologies such as SLAM (Simultaneous Localization and Mapping). The three-dimensional high-precision map may be, for example, the point cloud map described above. The occupancy grid map may be a map that divides the three-dimensional or two-dimensional space around the vehicle 102 into grids of a predetermined size and shows the occupancy status of objects on a grid-by-grid basis. The occupancy status of objects may be indicated, for example, by the presence or absence or probability of existence of an object. In some embodiments, the local map may also be used, for example, for detection and / or recognition processing of the external conditions of the vehicle 102 by the recognition unit 373.

[0102] In some embodiments, the self-position estimation unit 371 may estimate the self-position of the vehicle 102 based on position information acquired by the position information acquisition unit 324 and / or sensor data from the vehicle sensor 327.

[0103] The sensor fusion unit 372 performs sensor fusion processing to obtain information by combining multiple different types of sensor data (for example, image data supplied from the camera 351 and sensor data supplied from the radar 352). Methods for combining different types of sensor data are not limited to these, but include composite, integrated, fused, and combined methods.

[0104] The recognition unit 373 performs a detection process to detect the external conditions of the vehicle 102, and / or a recognition process to recognize the external conditions of the vehicle 102.

[0105] For example, the recognition unit 373 may perform detection and / or recognition processing of the external conditions of the vehicle 102 based on information from the external recognition sensor 325, information from the self-position estimation unit 371, information from the sensor fusion unit 372, etc.

[0106] Specifically, for example, the recognition unit 373 may perform detection and / or recognition processing of objects around the vehicle 102. Object detection processing may include, for example, detecting the presence, size, shape, position, and movement of an object. Object recognition processing may include, for example, recognizing attributes such as the type of object or identifying a specific object. Detection processing and recognition processing are not necessarily clearly separated, and at least some overlap may occur.

[0107] In some embodiments, the recognition unit 373 may detect objects around the vehicle 102 by performing clustering, which classifies the point cloud based on sensor data from the radar 352 and / or LiDAR 353 into clusters of points. This allows for the detection of the presence, size, shape, and position of objects around the vehicle 102.

[0108] In some embodiments, the recognition unit 373 may detect the movement of objects around the vehicle 102 by tracking the movement of clusters of points classified by clustering. This allows for the detection of the velocity and / or direction of travel (movement vector) of objects around the vehicle 102.

[0109] In some embodiments, the recognition unit 373 may detect and / or recognize vehicles (including bicycles), people, obstacles, structures, roads, traffic lights, traffic signs, road markings, etc., based on image data supplied from the camera 351. In some embodiments, the recognition unit 373 may recognize the types of objects around the vehicle 102 by performing recognition processing such as semantic segmentation.

[0110] In some embodiments, the recognition unit 373 may perform a recognition process of traffic rules around the vehicle 102 based on the map stored in the map information storage unit 323, the self-position estimation result by the self-position estimation unit 371, and / or the recognition result of objects around the vehicle 102 by the recognition unit 373. Through this process, the recognition unit 373 may recognize the location and / or status of traffic signals, the content of traffic signs and / or road markings, the content of traffic regulations, and / or drivable lanes, etc.

[0111] In some embodiments, the recognition unit 373 may perform recognition processing of the environment surrounding the vehicle 102. In some embodiments, the recognition unit 373 may recognize weather characteristics (temperature, humidity, brightness), and / or road surface conditions, etc.

[0112] The action planning unit 362 creates an action plan for the vehicle 102. For example, the action planning unit 362 may create an action plan by performing route planning and route following.

[0113] In some embodiments, the path planning may include global path planning and local path planning. Global path planning may include the process of planning a rough route from the start to the goal. Local path planning, also called track planning, may include the generation of a track that allows the vehicle 102 to travel safely and smoothly along the planned route in the vicinity of the vehicle 102, taking into account the vehicle's motion characteristics and the presence of any obstacles.

[0114] In some embodiments, route following may involve planning actions to safely and accurately travel along the route planned by the route planner within a planned time. The action planning unit 362 may, for example, calculate the target speed and / or target angular velocity of the vehicle 102 based on the results of this route following process.

[0115] The motion control unit 363 controls the operation of the vehicle 102 in order to realize the action plan created by the action planning unit 362.

[0116] For example, in some embodiments, the motion control unit 363 may control the steering control unit 381, brake control unit 382, ​​and / or drive control unit 383, which are included in the vehicle control unit 332 described later, to perform lateral vehicle motion control and / or longitudinal vehicle motion control so that the vehicle 102 travels along the trajectory calculated by the trajectory plan. For example, the motion control unit 363 may perform one or more driver assistance functions and / or control for the purpose of driving automation (e.g., lateral vehicle motion control, longitudinal vehicle motion control). Non-limited examples of driver assistance functions include collision avoidance or impact mitigation, inter-vehicle distance control (e.g., control to maintain a specific distance from a vehicle traveling in front of the vehicle 102), vehicle speed control (e.g., control to maintain a specific speed), vehicle collision warning, and lane departure warning. Non-limited examples of driving automation include driving without operation by the driver or remote driver.

[0117] In some embodiments, the DMS 330 may perform driver authentication processing and / or driver status recognition processing based on sensor data from the in-vehicle sensor 326 and / or input data input to the HMI 331, which will be described later. Non-limited examples of driver status that may be recognized include physical condition, alertness level, concentration level, fatigue level, gaze direction, intoxication level, driving operation, posture, etc.

[0118] In some embodiments, the DMS 330 may perform authentication processing for users other than the driver (e.g., passengers) and / or recognition processing for the status of such users. In some embodiments, the DMS 330 may perform recognition processing for the internal conditions of the vehicle 102 based on sensor data from the in-vehicle sensor 326. Non-limiting examples of characteristics of the internal conditions of the vehicle 102 that may be recognized include temperature, humidity, brightness, odor, etc.

[0119] The HMI331 receives various data and instructions as input and presents various data to the user.

[0120] A general overview of data input to the HMI 331 is provided. The HMI 331 is equipped with an input device for a person to input data, instructions, etc. Based on the data, instructions, etc., input by the input device, the HMI 331 generates an input signal and supplies it to one or more (or, in some cases, all) units of the vehicle control system 311. In some embodiments, the HMI 331 may be equipped with a touch panel, buttons, switches, and / or levers as input devices. Not limited to these, the HMI 331 may be equipped with an input device that allows information to be input by methods other than manual operation, such as voice or gestures. In some embodiments, the HMI 331 may be equipped with a remote control device using infrared and / or radio waves, or an external connection device that corresponds to the operation of the vehicle control system 311, as an input device. Non-limited examples of external connection devices include mobile devices (e.g., smartphones) and wearable devices (e.g., smartwatches).

[0121] A brief explanation of data presentation by HMI331 is provided below. HMI331 generates visual, auditory, and / or tactile information for the user and / or people outside the vehicle 102. HMI331 may also perform output control to control the output, output content, output timing, and / or output method of each generated piece of information. Non-limited examples of visual information that can be generated and output by HMI331 include information shown by images and light, such as operation screens, vehicle status displays, warning displays, and monitor images showing the surroundings of the vehicle 102. Non-limited examples of auditory information that can be generated and output by HMI331 include voice guidance, warning sounds, and warning messages. Non-limited examples of tactile information that can be generated and output by HMI331 include information that is given to the user's sense of touch through force, vibration, movement, etc.

[0122] In some embodiments, the HMI 331 may include, as an output device capable of outputting visual information, a display device that presents visual information by displaying an image itself, or a projector device that presents visual information by projecting an image. In some embodiments, the display device may be a device that displays visual information within the user's field of view, such as a head-up display, a transparent display, or a wearable device with AR (Augmented Reality) functionality, in addition to or as an alternative to a normal display device. In some embodiments, the HMI 331 may include, as an output device capable of outputting visual information, a display device provided in the vehicle 102, such as a navigation device, instrument panel, CMS (Camera Monitoring System), electronic mirror, lamp, etc.

[0123] In some embodiments, the HMI331 may include an audio speaker, headphones, or earphones as an output device capable of outputting auditory information.

[0124] In some embodiments, the HMI331 may include a haptic element using haptic technology as an output device capable of outputting tactile information. The haptic element may be provided, for example, on parts of the vehicle 102 that the user comes into contact with, such as the steering wheel or the seat.

[0125] The vehicle control unit 332 controls one or more (or, in some cases, all) units of the vehicle 102. The vehicle control unit 332 includes a steering control unit 381, a brake control unit 382, ​​a drive control unit 383, a body system control unit 384, a light control unit 385, and a horn control unit 386.

[0126] The steering control unit 381 detects and / or controls the state of the steering system of the vehicle 102. The steering system includes, for example, a steering mechanism with a steering wheel, an electric power steering system, etc. The steering control unit 381 includes, for example, a steering ECU that controls the steering system, an actuator that drives the steering system, etc.

[0127] The brake control unit 382 detects and / or controls the state of the brake system of the vehicle 102. The brake system includes, for example, a brake mechanism including a brake pedal, an ABS (Antilock Brake System), a regenerative braking mechanism, etc. The brake control unit 382 includes, for example, a brake ECU that controls the brake system, an actuator that drives the brake system, etc.

[0128] The drive control unit 383 detects and / or controls the state of the vehicle 102's drive system. The drive system includes, for example, an accelerator pedal, a drive force generating device for generating driving force such as an internal combustion engine or drive motor, and a drive force transmission mechanism for transmitting driving force to the wheels. The drive control unit 383 also includes, for example, a drive ECU for controlling the drive system and an actuator for driving the drive system.

[0129] The body system control unit 384 detects and / or controls the state of the body system of the vehicle 102. The body system includes, for example, a keyless entry system, a smart key system, power window devices, power seats, an air conditioning system, airbags, seat belts, a shift lever, etc. The body system control unit 384 also includes, for example, a body system ECU that controls the body system, actuators that drive the body system, etc.

[0130] The light control unit 385 detects and / or controls the state of various lights on the vehicle 102. Non-exclusive examples of lights that can be controlled by the light control unit 385 include headlights, taillights, fog lights, turn signals, brake lights, projector lights, bumper indicators, etc. The light control unit 385 includes a light ECU for controlling the lights, actuators for driving the lights, etc.

[0131] The horn control unit 386 detects and / or controls the state of the vehicle's car horn. The horn control unit 386 includes, for example, a horn ECU for controlling the car horn, an actuator for driving the car horn, and the like.

[0132] Figure 8 shows an example of the sensing area of ​​the external recognition sensor 325 in Figure 7, including the camera 351, radar 352, LiDAR 353, and ultrasonic sensor 354. In Figure 8, a schematic view of the vehicle 102 from above is shown.

[0133] Sensing regions 401F and 401B show examples of sensing regions for ultrasonic sensors 354. Sensing region 401F (for example, the sensing region of multiple ultrasonic sensors 354) covers the area around the front end of the vehicle 102. Sensing region 401B (for example, the sensing region of multiple ultrasonic sensors 354) covers the area around the rear end of the vehicle 102.

[0134] The sensing results in sensing region 401F and / or sensing region 401B may be used, for example, to assist in parking the vehicle 102.

[0135] Sensing areas 402F, 402B, 402L, and 402R illustrate examples of sensing areas for short-range or medium-range radar 352. Sensing area 402F covers an area in front of vehicle 102 that is further away than sensing area 401F. Sensing area 402B covers an area behind vehicle 102 that is further away than sensing area 401B. Sensing area 402L covers the area around the left rear of vehicle 102. Sensing area 402R covers the area around the right rear of vehicle 102.

[0136] The sensing results in sensing region 402F may be used, for example, to detect vehicles or pedestrians in front of vehicle 102. The sensing results in sensing region 402B may be used, for example, to prevent collisions behind vehicle 102. The sensing results in sensing region 402L and / or sensing region 402R may be used, for example, to detect one or more objects in the blind spots on the left and / or right sides of vehicle 102.

[0137] Sensing areas 403F, 403B, 403L, and 403R show examples of sensing areas by camera 351. Sensing area 403F covers a position further in front of vehicle 102 than sensing area 402F. Sensing area 403B covers a position further behind vehicle 102 than sensing area 402B. Sensing area 403L covers the left side of vehicle 102. Sensing area 403R covers the right side of vehicle 102.

[0138] The sensing results in sensing region 403F may be used, for example, for recognition of traffic lights and traffic signs, lane departure prevention support systems, and automatic headlight control systems. The sensing results in sensing region 403B may be used, for example, for parking assistance and / or surround view systems. The sensing results in sensing region 403L and / or sensing region 403R may be used, for example, for surround view systems.

[0139] Sensing area 404 shows an example of the sensing area of ​​LiDAR 353. Sensing area 404 covers a position further in front of the vehicle 102 than sensing area 403F. On the other hand, sensing area 404 has a narrower range in the left-right direction of the vehicle 102 than sensing area 403F.

[0140] The sensing results in the sensing region 404 may be used, for example, to detect objects such as surrounding vehicles.

[0141] Sensing area 405 shows an example of the sensing area of ​​the long-range radar 352. Sensing area 405 covers a position further in front of the vehicle 102 than sensing area 404. On the other hand, sensing area 405 has a narrower range in the lateral direction of the vehicle 102 than sensing area 404.

[0142] The sensing results in the sensing region 405 may be used, for example, for ACC (Adaptive Cruise Control), emergency braking, collision avoidance, etc.

[0143] In some embodiments, the sensing areas of each sensor of the external recognition sensor 325 (e.g., camera 351, radar 352, LiDAR 353, ultrasonic sensor 354) may take various configurations other than those shown in Figure 8. Specifically, in some embodiments, the ultrasonic sensor 354 may also sense the sides of the vehicle 102, or the LiDAR 353 may be configured to sense the rear of the vehicle 102. Furthermore, the installation positions of each sensor are not limited to the examples described above. Also, there may be one or more sensors.

[0144] <Application of this technology> The above-described technology may be applied to the automated driving system 100 with the above configuration in <3. Determination of the state of traffic signals using SD maps>. In that case, a functional block diagram showing the functions realized by the server 101 (CPU 201) executing processing is shown in Figure 9A. As shown in Figure 9A, the server 101 (CPU 201) has a traffic signal position information generation unit 511, an association unit 512, and a provision unit 513 as functional blocks.

[0145] The traffic light position information generation unit 511 performs processing related to the generation of traffic light position information indicating the location of traffic lights. For example, the traffic light position information generation unit 511 may estimate the location of traffic lights by analyzing an aerial photograph and generate traffic light position information indicating that location. Alternatively, the traffic light position information generation unit 511 may supply the generated traffic light position information to the association unit 512.

[0146] The association unit 512 performs processing related to linking (associating) the SD map and the traffic signal location information. For example, the association unit 512 may acquire traffic signal location information supplied from the traffic signal location information generation unit 511. The association unit 512 may also associate the acquired traffic signal location information with the SD map. In other words, it may be possible to indicate the location of the traffic signals in the coordinate system of the SD map. The association unit 512 may also supply the mutually associated SD map and traffic signal location information to the provision unit 513.

[0147] The provisioning unit 513 performs processing related to the provision of linked SD maps and signal location information. For example, the provisioning unit 513 may acquire SD maps and signal location information supplied from the association unit 512. The provisioning unit 513 may also supply the acquired SD maps and signal location information to, for example, a vehicle 102, etc., via the communication unit 214.

[0148] Furthermore, Figure 9B shows a functional block diagram illustrating the functions realized by the vehicle 102 (recognition unit 373) executing processes as blocks. As shown in Figure 9B, the vehicle 102 (recognition unit 373) has, as functional blocks, an alignment unit 531, a signal light designation unit 532, a bounding box setting unit 533, and a state determination unit 534.

[0149] The alignment unit 531 performs processing related to the alignment of the AI ​​inference map. For example, the alignment unit 531 may align the AI ​​inference map to the coordinate system of the SD map.

[0150] The signal designation unit 532 performs processing related to the designation of the target signal. For example, the signal designation unit 532 may designate the target signal based on signal location information associated with the SD map.

[0151] The bounding box setting unit 533 performs processing related to the setting of bounding boxes. For example, the bounding box setting unit 533 may set bounding boxes in the generated image based on traffic light position information associated with the SD map.

[0152] The state determination unit 534 performs processing related to determining the state of the traffic signal. For example, the state determination unit 534 may determine the state of the target traffic signal within the bounding box.

[0153] For example, the alignment unit 531 may align the position of an AI inference map, which has information equivalent to an HD map and is generated by inference of a map generation AI model using the generated image and generated point cloud as input data, to the coordinate system of an SD map associated with traffic light position information indicating the location of traffic lights. The traffic light designation unit 532 may then designate a target traffic light whose status to be determined based on that traffic light position information. The status determination unit 534 may then determine the status of the designated target traffic light by performing image analysis on the generated image.

[0154] Furthermore, the bounding box setting unit 533 may set a bounding box for the generated image. The state determination unit 534 may then determine the state of the target traffic light by performing image analysis within the bounding box of the generated image. For example, the bounding box setting unit 533 may set the bounding box of the generated image at a height corresponding to legal regulations regarding traffic lights. Alternatively, the bounding box setting unit 533 may set the bounding box based on the distribution pattern of points in the generated point cloud. Alternatively, the bounding box setting unit 533 may set the bounding box by excluding areas where the target traffic light is unlikely to exist, based on the distribution pattern. Furthermore, the bounding box setting unit 533 may set a confidence level based on the distribution pattern. The state determination unit 534 may then identify the target traffic light within the bounding box and determine the state of the identified target traffic light based on the image analysis results and the confidence level.

[0155] The state determination unit 534 may identify a target traffic light from among multiple traffic lights within the bounding box and determine the state of the identified target traffic light. For example, the state determination unit 534 may identify a target traffic light based on whether the light-emitting part is visible in the generated image. Alternatively, the state determination unit 534 may identify a target traffic light based on the size of the light-emitting part in the generated image. Furthermore, the state determination unit 534 may identify a target traffic light based on the orientation of the light-emitting part in the generated image. Finally, the state determination unit 534 may identify a target traffic light based on the state changes of the traffic light.

[0156] For example, the signal designation unit 532 may designate a signal located within a predetermined range from the vehicle as the target signal based on the signal location information.

[0157] Furthermore, the camera 351 of the vehicle 102 (information processing device) may generate generated images. The state determination unit 534 may then determine the state of the designated target traffic light by performing image analysis on the generated images generated by the camera 351. In other words, the camera 351 can also be described as an image generation unit that generates generated images. The camera 351 may generate any number of generated images. For example, the camera 351 may generate multiple generated images with different directions, angles of view, resolutions, or data types as input data to be input to the map generation AI model. Also, the LiDAR 353 of the vehicle 102 (information processing device) may generate a point cloud. In other words, the LiDAR 353 can also be described as a point cloud generation unit that generates a point cloud. Furthermore, the self-position estimation unit 371 of the vehicle 102 (information processing device) may generate an AI inference map using the generated images and the generated point cloud with the map generation AI model. Furthermore, the alignment unit 531 may align the position of the AI ​​inference map with the coordinate system of the SD map to which traffic light position information indicating the position of the traffic lights is linked. In other words, the self-position estimation unit 371 can also be called an AI inference map generation unit that generates the AI ​​inference map.

[0158] Furthermore, the storage unit 328 of the vehicle 102 (information processing device) may store the SD map and the traffic signal location information. Also, the HMI 331 of the vehicle 102 (information processing device) may provide guidance regarding the operation of the vehicle based on the status determination result of the target traffic signal and the aligned AI inference map. In other words, the HMI 331 can also be said to be a guidance unit that provides guidance regarding the operation of the vehicle.

[0159] With this configuration, the vehicle 102 can determine the desired traffic signal state with higher accuracy while suppressing cost increases, as described above in <3. Traffic Signal State Determination Using SD Map>.

[0160] <Traffic Signal Location Information Processing Flow> Next, we will explain the processes performed by the server 101 and the vehicle 102. Referring to the flowchart in Figure 10, we will explain an example of the traffic signal location information processing flow performed by the server 101. This traffic signal location information processing is the process of generating traffic signal location information and linking it to the SD map.

[0161] When traffic signal location information processing is started, the traffic signal location information generation unit 511 of the server 101 estimates the traffic signal location based on aerial photographs and generates traffic signal location information in step S101.

[0162] In step S102, the association unit 512 associates the SD map with the traffic signal location information.

[0163] In step S103, the providing unit 513 stores and manages the SD map and signal location information. In step S104, the providing unit 513 provides the SD map and signal location information to, for example, a vehicle 102.

[0164] Once the process in step S104 is completed, the traffic signal position information processing is finished.

[0165] By executing each process in this manner, the server 101 can generate traffic signal location information from aerial photographs and link it to the SD map.

[0166] <Acquisition Process Flow> Next, an example of the acquisition process flow executed by vehicle 102 will be explained with reference to the flowchart in Figure 11. This acquisition process is the process of acquiring linked SD maps and traffic signal location information provided by server 101.

[0167] When the acquisition process is started, in step S201, the communication unit 322 acquires the linked SD map and traffic signal location information provided by the server 101.

[0168] In step S202, the storage unit 328 stores the acquired SD map and traffic signal location information.

[0169] The acquisition process ends when the processing in step S202 is completed.

[0170] <Flow of Driving Control Processing> Next, an example of the flow of driving control processing performed in vehicle 102 will be explained with reference to the flowchart in Figure 12.

[0171] When the driving control process is started, in step S301, the camera 351 of the vehicle 102 generates a generated image. Also, the LiDAR 352 generates a generated point cloud.

[0172] In step S302, the self-localization unit 371 generates a local HD map (AI inference map) based on the generated image and the generated point cloud.

[0173] In step S303, the recognition unit 373 executes a traffic light state determination process and determines the state of the target traffic light using the traffic light location information associated with the SD map. In areas where an HD map exists, the recognition unit 373 may recognize the location of the traffic light according to the information in the HD map and perform traffic light state determination. In other words, if the vehicle 102 is located in an area where no HD map exists, the recognition unit 373 may execute a traffic light state determination process and determine the state of the target traffic light using the traffic light location information associated with the SD map.

[0174] In step S304, the action planning unit 362 plans an action based on the local HD map and the traffic light status.

[0175] In step S305, the operation control unit 363 controls the operation based on the local HD map, the traffic signal status, and the action plan. For example, the operation control unit 363 controls the vehicle's operation based on the status determination result of the target traffic signal and the aligned AI inference map.

[0176] When the process in step S305 is completed, the operation control process ends.

[0177] <Traffic Light Status Determination Process Flow> Next, with reference to the flowchart in Figure 13, an example of the traffic light status determination process flow executed in step S303 of Figure 12 will be explained.

[0178] When the traffic light status determination process is started, the alignment unit 531 aligns the local HD map with the SD map in step S321. For example, the GPS signal received by vehicle 102 indicates the actual position of vehicle 102, but it is misaligned with the road position on the SD map. Therefore, the alignment unit 531 shifts the GPS signal onto the coordinates of the SD map, taking into account the distance between the GPS signal and the position on the SD map, the orientation of vehicle 102, the direction of travel on the road, etc., so that vehicle 102 is positioned on the road on the SD map. In other words, the alignment unit 531 shifts the position of vehicle 102 indicated by the GPS signal onto the road on the SD map. This shift causes the position of the AI ​​inference map, which has information equivalent to an HD map and is generated by the inference of a map generation AI model that uses generated images and generated point clouds as input data, to shift onto the coordinate system of the SD map to which traffic light position information indicating the position of traffic lights is linked.

[0179] In step S322, the signal designation unit 532 designates a target signal based on the signal location information associated with the SD map. For example, the signal designation unit 532 designates a target signal whose status is to be determined based on the signal location information.

[0180] In step S323, the bounding box setting unit 533 sets bounding boxes within the generated image based on the traffic signal position information associated with the SD map.

[0181] In step S324, the state determination unit 534 determines the state of the target traffic signal within the set bounding box. For example, the state determination unit 534 identifies the target traffic signal to be followed by performing image analysis on the region within the bounding box of the generated image, and determines its state.

[0182] Once the process in step S324 is completed, the traffic light status determination process ends, and the process returns to Figure 12.

[0183] By performing each process as described above, the vehicle 102 can determine the desired traffic signal state with higher accuracy while suppressing cost increases, as described above in <3. Traffic Signal State Determination Using SD Map>.

[0184] <5. Addendum> <Software> The series of processes described above can be executed by hardware or by software. When the series of processes described above are executed by software, the programs that make up the software are installed from a network or storage medium.

[0185] This recording medium, as shown in Figure 6, for example, consists of a removable media 221 on which a program is recorded, which is distributed separately from the main unit of the device to the user for the purpose of delivering the program. This removable media 221 includes magnetic disks (including flexible disks) and optical disks (including CD-ROMs and DVDs). Furthermore, it also includes magneto-optical disks (including MDs (Mini Discs)) and semiconductor memory. In this case, for example, by inserting the removable media 221 into the drive 215, the program stored on the removable media 221 can be read and installed into the storage unit 213.

[0186] Furthermore, this program can also be provided via wired or wireless transmission media such as local area networks, the internet, or digital satellite broadcasting. In that case, the program can be received by the communication unit 214 and installed in the storage unit 213.

[0187] In addition, this program can be pre-installed in memory units or ROMs. For example, the program can be pre-installed in memory unit 213 or ROM 202.

[0188] The same principle applies to vehicle 102. For example, in Figure 7, a program stored on a removable media (not shown) mounted on a drive (not shown) may be read and installed in the storage unit 328, etc. Alternatively, the program may be received by the communication unit 322, etc. and installed in the storage unit 328, etc. Furthermore, the program may be pre-installed in the storage unit 328, etc.

[0189] <Applicable Subjects of This Technology> Furthermore, this technology can be applied to any configuration. For example, this technology can be applied to various electronic devices. In addition, this technology can be applied to the automatic driving control technology of any mobile body, not limited to vehicles, but including, for example, ships and airplanes. This mobile body may be a manned aircraft with a person on board, or an unmanned aircraft without a person on board, such as a so-called drone. Furthermore, this technology can be applied to any driving assistance technology, not limited to automatic driving control, but including, for example, route guidance and driving operation assistance.

[0190] Furthermore, this technology can also be implemented as part of a device, such as a processor as a system LSI (Large Scale Integration) (e.g., a video processor), a module using multiple processors (e.g., a video module), a unit using multiple modules (e.g., a video unit), or a set with additional functions added to a unit (e.g., a video set).

[0191] Furthermore, this technology can also be applied to network systems composed of multiple devices. For example, this technology may be implemented as cloud computing, where multiple devices share and collaborate on processing via a network. For example, this technology may be implemented in a cloud service that provides image (video) related services to any terminal such as computers, AV (Audio Visual) equipment, portable information processing terminals, and IoT (Internet of Things) devices.

[0192] In this specification, a system refers to a collection of multiple components (devices, modules (parts), etc.), regardless of whether all components are located in the same enclosure. Therefore, multiple devices housed in separate enclosures and connected via a network, and a single device containing multiple modules within a single enclosure, are both considered systems.

[0193] <Applicable Fields and Applications of This Technology> Systems, devices, and processing units incorporating this technology can be used in any field, such as transportation, medical care, security, agriculture, livestock farming, mining, beauty, factories, home appliances, weather, and nature monitoring. Furthermore, the applications are entirely arbitrary.

[0194] For example, this technology can be applied to systems and devices used to provide entertainment content. Furthermore, for example, this technology can be applied to systems and devices used for traffic management, such as traffic condition monitoring and automated driving control. In addition, for example, this technology can be applied to systems and devices used for security. Furthermore, for example, this technology can be applied to systems and devices used for automatic control of machinery, etc. Furthermore, for example, this technology can be applied to systems and devices used for agriculture and livestock farming. Furthermore, for example, this technology can be applied to systems and devices that monitor natural conditions such as volcanoes, forests, and oceans, as well as wildlife. Furthermore, for example, this technology can be applied to systems and devices used for sports.

[0195] <Other> In this specification, terms such as "combine," "multiplex," "add," "integrate," "include," "store," "insert," "insert," and "place" mean combining multiple things into one, such as combining encoded data and metadata into a single data, and represent one method of "associating" as described above.

[0196] Furthermore, the embodiments of this technology are not limited to those described above, and various modifications are possible without departing from the gist of this technology.

[0197] For example, the configuration described as a single device (or processing unit) may be divided and configured as multiple devices (or processing units). Conversely, the configurations described above as multiple devices (or processing units) may be combined and configured as a single device (or processing unit). Furthermore, it is also possible to add configurations other than those described above to the configuration of each device (or each processing unit). In addition, if the overall system configuration and operation are substantially the same, a part of the configuration of one device (or processing unit) may be included in the configuration of another device (or other processing unit).

[0198] Furthermore, for example, the program described above may be executed on any device. In that case, the device should have the necessary functions (such as functional blocks) and be able to obtain the necessary information.

[0199] Furthermore, for example, each step of a flowchart may be executed by one device, or it may be divided among multiple devices. Additionally, if a single step includes multiple processes, these processes may be executed by one device, or they may be divided among multiple devices. In other words, multiple processes included in a single step can be executed as multiple steps. Conversely, processes described as multiple steps can be combined and executed as a single step.

[0200] Furthermore, for example, a program executed by a computer may be structured so that the steps of the program are executed chronologically in the order described herein, or they may be executed in parallel or individually at necessary times, such as when a call is made. In other words, the steps may be executed in an order different from the order described above, as long as no inconsistencies arise. Moreover, the steps of this program may be executed in parallel with the processing of other programs, or in combination with the processing of other programs.

[0201] Furthermore, for example, multiple technologies relating to this technology can be implemented independently, as long as they do not create a contradiction. Of course, any multiple technologies can also be implemented in combination. For example, some or all of the technologies described in one embodiment can be implemented in combination with some or all of the technologies described in another embodiment. Also, some or all of the above-mentioned technologies can be implemented in combination with other technologies not mentioned above.

[0202] Furthermore, this technology can also take the following configuration: (1) An information processing device comprising: an alignment unit that aligns the position of an AI inference map having information equivalent to an HD map, which is generated by inference of a map generation AI (Artificial Intelligence) model that takes generated images and generated point clouds as input data, with the coordinate system of an SD map to which traffic light position information indicating the position of traffic lights is associated; a traffic light designation unit that designates a target traffic light whose state is to be determined based on the traffic light position information; a state determination unit that determines the state of the designated target traffic light by performing image analysis on the generated image; and a control unit that controls the operation of a vehicle based on the state determination result of the target traffic light and the aligned AI inference map, wherein the HD map is relatively high-precision map information that includes a relatively wide variety of road information, including at least lanes; the SD map is relatively low-precision map information that includes fewer types of road information than the HD map; the generated image is an image detected by a sensor; and the generated point cloud is a point cloud representing the three-dimensional shape of an object detected by a sensor. (2) The information processing device according to (1), further comprising a bounding box setting unit for setting a bounding box for the generated image, wherein the state determination unit determines the state of the target traffic signal by performing image analysis of the bounding box within the generated image. (3) The information processing device according to (2), wherein the bounding box setting unit sets the bounding box in the generated image at a height corresponding to legal regulations concerning the traffic signal. (4) The information processing device according to (2) or (3), wherein the bounding box setting unit sets the bounding box based on the distribution pattern of points in the generated point cloud. (5) The information processing device according to (4), wherein the bounding box setting unit sets the bounding box by excluding areas where the target traffic signal is unlikely to exist based on the distribution pattern.(6) The bounding box setting unit further sets a confidence level based on the distribution pattern, and the state determination unit identifies the target signal within the bounding box and determines the state of the identified target signal, as described in (4) or (5). (7) The state determination unit identifies the target signal from among a plurality of signal lights within the bounding box and determines the state of the identified target signal, as described in any of (4) to (6). (8) The state determination unit identifies the target signal based on whether the light-emitting part is visible in the generated image, as described in (7). (9) The state determination unit identifies the target signal based on the size of the light-emitting part in the generated image, as described in (7) or (8). (10) The state determination unit identifies the target signal based on the orientation of the light-emitting part in the generated image, as described in any of (7) to (9). (11) The information processing device according to any one of (7) to (10), wherein the state determination unit identifies the target signal based on the state change of the signal. (12) The information processing device according to any one of (1) to (11), wherein the signal designation unit designates a signal located within a predetermined range from a vehicle as the target signal based on the signal location information. (13) The information processing device according to any one of (1) to (12), further comprising an image generation unit for generating the generated image. (14) The information processing device according to (13), further comprising a point cloud generation unit for generating the generated point cloud. (15) The information processing device according to (14), further comprising an AI inference map generation unit for generating the AI ​​inference map using the map generation AI model with the generated image and the generated point cloud. (16) The information processing device according to any one of (1) to (15), further comprising a storage unit for storing the SD map and the signal location information. (17) The information processing device according to any one of (1) to (16), further comprising a guidance unit that provides guidance regarding vehicle operation based on the status determination result of the target traffic signal and the aligned AI inference map.(18) An information processing method comprising: aligning the position of an AI inference map having information equivalent to an HD map, which is generated by inference of a map generation AI (Artificial Intelligence) model that takes generated images and generated point clouds as input data, with the coordinate system of an SD map to which traffic light position information indicating the position of traffic lights is associated; specifying a target traffic light whose state is to be determined based on the traffic light position information; determining the state of the specified target traffic light by performing image analysis on the generated images; and controlling the operation of a vehicle based on the state determination result of the target traffic light and the aligned AI inference map, wherein the HD map is relatively high-precision map information that includes a relatively wide variety of road information, including at least lanes; the SD map is relatively low-precision map information that includes fewer types of road information than the HD map; the generated images are images detected by a sensor; and the generated point clouds are point clouds representing the three-dimensional shape of an object detected by a sensor. (19) A program for causing a computer to perform the following processes: (19) Aligning the position of an AI inference map having information equivalent to an HD map, which is generated by inference of a map generation AI (Artificial Intelligence) model that takes generated images and generated point clouds as input data, with the coordinate system of an SD map to which traffic light position information indicating the position of traffic lights is associated; specifying a target traffic light whose state to be determined based on the traffic light position information; determining the state of the specified target traffic light by performing image analysis on the generated images; and controlling the operation of a vehicle based on the state determination result of the target traffic light and the aligned AI inference map, wherein the HD map is relatively high-precision map information that includes a relatively wide variety of road information, including at least lanes; the SD map is relatively low-precision map information that includes fewer types of road information than the HD map; the generated images are images detected by a sensor; and the generated point clouds are point clouds representing the three-dimensional shape of an object detected by a sensor.

[0203] 100 Automated driving system, 101 Server, 102 Vehicle, 110 Network, 201 CPU, 202 RAM, 203 ROM, 204 Bus, 210 Input / Output interface, 211 Input unit, 212 Output unit, 213 Storage unit, 214 Communication unit, 215 Drive, 221 Removable media, 311 Vehicle control system, 321 Vehicle control ECU, 322 Communication unit, 323 Map information storage unit, 324 Location information acquisition unit, 325 External recognition sensor, 326 In-vehicle sensor, 327 Vehicle sensor, 328 Storage unit, 329 Driving automation control unit, 330 DMS, 331 HMI, 332 Vehicle control unit, 341 Communication network, 351 Camera, 352 Radar, 353 LiDAR, 354 Ultrasonic sensor, 361 Analysis unit, 362 Action planning unit, 363 Motion control unit, 371 Self-position estimation unit, 372 Sensor fusion unit, 373 Recognition unit, 381 Steering control unit, 382 Brake control unit, 383 Drive control unit, 384 Body system control unit, 385 Light control unit, 386 Horn control unit, 511 Traffic light position information generation unit, 512 Association unit, 513 Provision unit, 531 Alignment unit, 532 Traffic light designation unit, 533 Bounding box setting unit, 534 State determination unit

Claims

1. An information processing device comprising: an alignment unit that aligns the position of an AI inference map having information equivalent to an HD map, generated by inference of a map generation AI (Artificial Intelligence) model that takes generated images and generated point clouds as input data, to the coordinate system of an SD map associated with traffic light position information indicating the location of traffic lights; a traffic light designation unit that designates a target traffic light whose state is to be determined based on the traffic light position information; a state determination unit that determines the state of the designated target traffic light by performing image analysis on the generated image; and a control unit that controls the operation of a vehicle based on the state determination result of the target traffic light and the aligned AI inference map, wherein the HD map is relatively high-precision map information that includes a relatively wide variety of road information, including at least lanes; the SD map is relatively low-precision map information that includes fewer types of road information than the HD map; the generated image is an image detected by a sensor; and the generated point cloud is a point cloud representing the three-dimensional shape of an object detected by a sensor.

2. The information processing apparatus according to claim 1, further comprising a bounding box setting unit for setting a bounding box for the generated image, wherein the state determination unit determines the state of the target signal by performing image analysis of the bounding box within the generated image.

3. The information processing apparatus according to claim 2, wherein the bounding box setting unit sets the bounding box of the generated image to a height in accordance with the legal regulations relating to the traffic signal.

4. The information processing apparatus according to claim 2, wherein the bounding box setting unit sets the bounding box based on the distribution pattern of points in the generated point cloud.

5. The information processing apparatus according to claim 4, wherein the bounding box setting unit sets the bounding box based on the distribution pattern, excluding areas where the target signal is unlikely to exist.

6. The information processing apparatus according to claim 4, wherein the bounding box setting unit further sets a confidence level based on the distribution pattern, and the state determination unit identifies the target signal within the bounding box and determines the state of the identified target signal based on the results of the image analysis and the confidence level.

7. The information processing apparatus according to claim 4, wherein the state determination unit identifies a target signal from among a plurality of signal lights present in the bounding box and determines the state of the identified target signal.

8. The information processing apparatus according to claim 7, wherein the state determination unit identifies the target signal based on whether the light-emitting unit is visible in the generated image.

9. The information processing apparatus according to claim 7, wherein the state determination unit identifies the target signal based on the size of the light-emitting part in the generated image.

10. The information processing apparatus according to claim 7, wherein the state determination unit identifies the target signal based on the orientation of the light-emitting part in the generated image.

11. The information processing device according to claim 7, wherein the state determination unit identifies the target signal based on the state change of the signal.

12. The information processing device according to claim 1, wherein the signal designation unit designates a signal located within a predetermined range from the vehicle as the target signal based on the signal location information.

13. The information processing apparatus according to claim 1, further comprising an image generation unit for generating the generated image.

14. The information processing apparatus according to claim 13, further comprising a point cloud generation unit for generating the generated point cloud.

15. The information processing apparatus according to claim 14, further comprising an AI inference map generation unit that generates the AI ​​inference map using the map generation AI model with respect to the generated image and the generated point cloud.

16. The information processing apparatus according to claim 1, further comprising a storage unit for storing the SD map and the signal location information.

17. The information processing device according to claim 1, further comprising a guidance unit that provides guidance regarding vehicle operation based on the status determination result of the target traffic signal and the aligned AI inference map.

18. An information processing method comprising: aligning the position of an AI inference map having information equivalent to an HD map, which is generated by inference of a map generation AI (Artificial Intelligence) model that takes generated images and generated point clouds as input data, with the coordinate system of an SD map associated with traffic light position information indicating the position of traffic lights; specifying a target traffic light whose state to be determined based on the traffic light position information; determining the state of the specified target traffic light by performing image analysis on the generated image; and controlling the operation of a vehicle based on the state determination result of the target traffic light and the aligned AI inference map, wherein the HD map is relatively high-precision map information that includes a relatively wide variety of road information, including at least lanes; the SD map is relatively low-precision map information that includes fewer types of road information than the HD map; the generated image is an image detected by a sensor; and the generated point cloud is a point cloud representing the three-dimensional shape of an object detected by a sensor.

19. A program for causing a computer to perform the following processes: aligning the position of an AI inference map, which has information equivalent to an HD map and is generated by inference of a map generation AI (Artificial Intelligence) model using generated images and generated point clouds as input data, with the coordinate system of an SD map associated with traffic light position information indicating the location of traffic lights; specifying a target traffic light whose state to be determined based on the traffic light position information; determining the state of the specified target traffic light by performing image analysis on the generated images; and controlling the operation of a vehicle based on the state determination result of the target traffic light and the aligned AI inference map, wherein the HD map is relatively high-precision map information that includes a relatively wide variety of road information, including at least lanes; the SD map is relatively low-precision map information that includes fewer types of road information than the HD map; the generated images are images detected by sensors; and the generated point clouds are point clouds representing the three-dimensional shapes of objects detected by sensors.