Methods for training an AI model for online generation of HD maps, methods for operating a vehicle that is at least partially automated, and vehicles

By training an AI model with artificially error-enriched SD card data, the method addresses the issue of discrepancies between SD card and sensor data, improving the robustness and safety of HD map generation and vehicle control.

DE102024004074B3Active Publication Date: 2026-04-23MERCEDES BENZ GROUP AG
View PDF 4 Cites 0 Cited by

Patent Information

Authority / Receiving Office
DE · DE
Patent Type
Patents
Current Assignee / Owner
MERCEDES BENZ GROUP AG
Filing Date
2024-12-05
Publication Date
2026-04-23

AI Technical Summary

Technical Problem

Existing methods for generating high-definition (HD) maps online in vehicles are prone to errors due to discrepancies between SD card data and real-time sensor data, leading to potential unsafe vehicle operation.

Method used

A method for training an AI model using SD card data enriched with artificial errors, processed by an encoder-decoder model and discriminator, to adapt to deviations between SD card and sensor data representations.

Benefits of technology

Enhances the robustness of the AI model to handle environmental discrepancies, ensuring safer and more reliable HD map generation and control command derivation for automated vehicles.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

The invention relates to a method for training an AI model (1) for the online generation of HD maps (HD-Map), wherein the AI ​​model (1) is supplied with SD map data (SD-Map) and sensor data (2) from environmental sensors of a vehicle as input data during training (101), the sensor data (2) being correlated with sections in the SD map data (SD-Map). The invention is characterized in that the SD map data (SD-Map) is enriched with artificial errors, at least partially, in the sections correlated with the sensor data (2) before being supplied to the AI ​​model (1), such that the infrastructure topology in the initial SD map data (SD-Map-ini) differs at least partially from the infrastructure topology in the enriched SD map data (SD-Map-aug).
Need to check novelty before this filing date? Find Prior Art

Description

[0001] The invention relates to a method for training an AI model for the online generation of HD maps according to the preamble of claim 1, a method for operating a vehicle that is at least partially automated and controllable according to the preamble of claim 7, and a correspondingly controllable vehicle.

[0002] The level of automation in vehicles is constantly increasing. Today's production vehicles are already capable of automatically maintaining a safe distance from the vehicle in front, changing lanes, initiating emergency braking maneuvers, and so on. In the future, it is expected that fully autonomous vehicles will be approved for road use.

[0003] For the safe operation of automated and autonomously controlled vehicles, it is essential that they can perceive their surroundings in order to detect static and dynamic objects. This allows for the generation of appropriate control commands for safe vehicle operation. Dynamic objects include other road users such as pedestrians, cyclists, cars, trucks, and the like. Static objects include infrastructure such as lanes, lane markings, traffic lights, traffic signs, and the like. These infrastructure objects are interrelated. For example, certain traffic lights are assigned to specific lanes, or certain traffic signs apply only to specific lanes.

[0004] Autonomous vehicles, in particular, use so-called HD maps, also known as high-definition maps, to derive control commands. These HD maps depict the environment with high resolution, allowing the course of individual lanes to be described with centimeter-level accuracy. Furthermore, HD maps contain the relationships between infrastructure objects, also known as the topology. By processing HD maps, typically using artificial intelligence, reliable control commands for vehicles can be generated, ensuring safe vehicle operation. The information depth of low-resolution maps, also known as SD maps (standard-definition maps), is generally insufficient for this purpose.

[0005] Due to the high information content of HD maps, their creation is very complex and costly. HD maps are generally processed manually and are therefore only available for a few select regions, mostly cities. This jeopardizes the widespread use of autonomous vehicles.

[0006] There are known approaches to generating high-definition (HD) maps live while a vehicle is in motion. In this process, the vehicle's surroundings are mapped onto the HD map. This approach is called online HD map generation. Using online HD maps, automated or autonomous driving modes can be made available more frequently and in a wider geographical area.

[0007] The online generation of HD maps by processing SD card material and sensor data from the environmental sensors of a respective vehicle is known, for example, from: Augmenting Lane Perception and Topology Understanding with Standard Definition Navigation Maps, Katie Z. Luo et al, Nvidia, arXiv:2311.04079v1 [cs.CV] 7 Nov 2023, https: / / arxiv.org / abs / 2311.04079.

[0008] A similar approach for generating online HD maps, where the prediction result of a corresponding AI module is refined by a masked auto-encoder module located at the end of the processing pipeline and pre-trained on a masked HD map, is known from: P-MapNet: Far-seeing Map Generator Enhanced by both SDMap and HDMap Priors, Zhou Jiang et al, arXiv:2403.10521v3 [cs.CV] 29 Mar 2024, https: / / arxiv.org / abs / 2403.10521.

[0009] Furthermore, it is known to adapt existing HD maps online by processing HD map and sensor data so that the HD maps better represent the real environment. This involved manually correcting errors in existing HD maps and removing content. See: Mind the map! Accounting for existing maps when estimating online HDMaps from Sensors., Rémy Sun, et al., Université Côte d'Azur, Inria, CNRS, I3S, Maasai, Nice, France, arXiv:2311.10517vs [cs.LG] 14 Mar 2024, https: / / arxiv.org / abs / 2311.10517.

[0010] One problem that arises is that SD cards sometimes represent reality incorrectly or inaccurately. For example, the environment may have been inaccurately mapped during the SD card creation process, or the road layout may have changed due to construction work. The semantic content of the SD card and the sensor data therefore differ. This can lead to errors in the online generation of HD maps.

[0011] Furthermore, DE 10 2019 216 836 A1 discloses a method for training at least one algorithm for a motor vehicle control unit, a computer program product, and a motor vehicle. In this process, a neural network for an autonomous driving function of the motor vehicle is further trained to detect malfunctions or abnormal behavior of other road users.

[0012] Furthermore, DE 10 2021 100 791 A1 discloses a method for determining training data for model improvement and a data processing device. Training data for a machine learning model are examined for errors that could lead to the machine learning model failing to process the underlying raw data. If the error is large enough, the underlying raw data are used for more in-depth training of the machine learning model.

[0013] Furthermore, DE 10 2021 208 187 A1 discloses a method for providing training data for training an artificial neural network, a method for training an artificial neural network, a computer program product, and a data structure. In this process, training data for the artificial neural network is artificially manipulated to better represent extreme values.

[0014] Furthermore, US patent 12 073 329 B2 discloses a method for detecting hostile interference with input data for an artificial neural network. By training the artificial neural network on manipulated input data, attacks on the neural network aimed at achieving a desired output result by an attacker can be made more difficult.

[0015] The present invention is based on the objective of providing means by which at least partially automated vehicles can be operated more safely by processing online generated HD maps.

[0016] According to the invention, this problem is solved by a method for training an AI model for the online generation of HD maps with the features of claim 1, and by a method for operating a vehicle that is at least partially automated and controlled with the features of claim 7. Advantageous embodiments and further developments, as well as a correspondingly operable vehicle, are described in the dependent claims.

[0017] A generic method for training an AI model for the online generation of HD maps, wherein the AI ​​model is supplied with SD card material and sensor data from environmental sensors of a vehicle as input data during training, wherein the sensor data are correlated to sections in the SD card material, is further developed according to the invention in that the SD card material is enriched with artificial errors at least partially in the sections correlated with the sensor data before being supplied to the AI ​​model, so that the infrastructure topology in the initial SD card material differs at least partially from the infrastructure topology in the enriched SD card material.

[0018] The applicant has recognized that by deliberately introducing differences between the SD card data used to train the AI ​​model and the sensor data, the robustness of the AI ​​model in the later operational phase is increased with respect to deviations encountered in reality between the stored SD card data and the environment sampled by the sensors. As mentioned at the outset, the actual environment encountered by the vehicle will typically not perfectly match the infrastructure topology stored on the SD card.If the AI ​​model for online generation of HD maps based on SD card data and sensor data supplied by the environmental sensors is trained exclusively on training data where the environment described by the SD card data and the environment depicted by the sensor data are nearly identical, the risk increases that, if corresponding differences actually occur between the SD card data and the sampled environment, the HD map generated by the AI ​​model will not be suitable for deriving appropriate control commands for the vehicle. The HD map could deviate too significantly from the real environment, potentially generating control commands that could lead to unsafe vehicle operation. For example, the vehicle might assume a curve radius that is too large or too small, or assume that the stop line of a traffic light is 1-5 meters further back than it actually is, and so on.Accordingly, the vehicle might over- or under-steer in the curve, over-exert the stop line, and so on. However, by applying the method according to the invention, such erratic behavior can be avoided, or at least the risk of such behavior occurring can be reduced, thus increasing road safety.

[0019] According to the invention, the AI ​​model for online HD map generation learns to handle corresponding deviations between the environment depicted by the SD card material and the environment depicted by the sensor data, so that a realistic HD map can be generated despite any existing deviations. The training of the AI ​​model can otherwise proceed as in the prior art cited above. In particular, camera images from one or more environmental cameras and / or point clouds generated using one or more LiDARe sensors are used as sensor data. Other sensor modalities such as ultrasonic sensor systems or radar sensors are also suitable. As known from the prior art, the image of the environment is transformed into a bird's-eye view, also known as a bird's-eye view (BEV). This can be done as described in: Lift, Splat, Shoot: Encoding Images From Arbitrary Camera Rigs by Implicitly Unprojecting to 3D, Jonah Philion, et. al, https: / / arxiv.org / abs / 2008.05711; oder: BEVFormer: Learning Bird's-Eye-View Representation from Multi-Camera Images via Spatiotemporal Transformers, Zhiqi Li, et. al, https: / / arxiv.org / abs / 2203.17270.

[0020] All conceivable SD cards can be used, such as OpenStreetMap (OSM). The relevant map section can be determined by comparing its position with the currently determined vehicle position. For this purpose, the vehicle can determine its geolocation using a navigation system. GNSS satellite signals from various systems can be evaluated, such as GPS, Galileo, GLONASS, Beidou, and the like. The SD map data used in training can be enhanced with various artificial errors, such as different radii of curvature and / or dimensions for the lane layout, a different or missing number of lanes, missing infrastructure elements, an incorrect or missing direction of travel, and the like.

[0021] An advantageous further development of the method according to the invention provides that the SD card data is enriched with artificial errors by processing the initial SD card data and / or at least a subset of the sensor data using artificial intelligence. This allows the generation of the enriched SD card data to be automated in a simple and reliable manner. Thus, even faster and simpler generation of large data sets is possible, which in turn allows for faster and more comprehensive training of the AI ​​model.

[0022] According to a further advantageous embodiment of the method according to the invention, it is further provided that the initial SD card data and / or at least the subset of the sensor data are processed by an encoder-decoder model, wherein the encoder maps input data onto a latent space. The encoder-decoder model is a deep neural network or a cascade of several neural networks, which consists of two interconnected parts, namely the encoder and the decoder. With the aid of the encoder, the read input data can be mapped onto a higher-dimensional space, which is read and processed by the decoder to generate the desired data. This higher-dimensional space is referred to here as the "latent space".The generation of the artificially error-enriched SD card material can be done solely using the original SD card material, solely using a subset or the complete sensor data, or a combination of original SD card material and sensor data.

[0023] The SD card data is ideally segmented into a grid so that a convolutional neural network (CNN) can process it. Recognizing relevant environmental features in the sensor data can also be done using a CNN. The environmental features depicted in the BEV perspective, based on the SD card data and the sensor data, are concatenated and fed into a further cascade of CNNs and a feedforward neural network (FNN). This cascade of neural networks ultimately yields the desired variable in latent space, here denoted as "Z".

[0024] The variable Z is then fed to the decoder, which is also a cascade of neural networks, specifically an FNN and a deconvolutional network. The decoder outputs the SD card data artificially enriched with errors. This SD card data can be in the form of a single card or multiple cards. Thus, several cards can be generated based on the same input data, or "embeddings." These cards contain randomly generated, different errors. This allows for even larger and faster data generation for training.

[0025] Preferably, the inventive method provides that the size of the latent space is artificially limited to a defined value. The latent space is thus preferably kept intentionally small to prevent the encoder-decoder model from perfectly replicating the environment described by the input data. This ensures that the generated SD card data differs from the read SD card data. This is a particularly efficient and therefore computationally efficient solution. Furthermore, this is important so that the AI ​​model for online HD card generation can learn to adapt to imperfect SD card data, which is highly likely to be processed later in the operational phase.

[0026] According to a further advantageous embodiment of the method according to the invention, the size of the latent space is initially chosen to be an order of magnitude smaller than the average size of the SD card encodings. If the latent space is too large, the risk increases that all relevant information from the read SD card data will be captured, so that no errors occur in the enriched SD card data. If, on the other hand, the latent space is too small, the risk increases that only random errors will occur, which are independent of the infrastructure topology of the read SD card data or the sensor data. If the initial size of the latent space is chosen to be an order of magnitude smaller than the average size of the SD card encodings, adequate enrichment of the SD card data with errors is usually possible.Based on this, adjustments can be made for further iterations to increase or decrease the size of the latent space as needed.

[0027] A further advantageous embodiment of the method according to the invention provides that the enriched SD card material is processed by a discriminator, wherein the discriminator classifies the enriched SD card material as real or artificial, wherein SD card material classified as real is released for training the AI ​​model and SD card material classified as artificial is either re-enriched with errors in a further iteration loop or discarded. This further improves the robustness of the AI ​​model for online generation of HD cards. The AI ​​model cannot adapt its behavior to the information of whether an artificially generated or a real SD card is being processed, since all SD cards appear real to the AI ​​model.

[0028] The discriminator also represents a cascade of artificial neural networks, such as a CNN and an FNN. The discriminator reads the SD card data fed to and generated by the encoder-decoder model and determines a difference between the two. This difference is referred to as the loss. The loss is compared to a defined threshold. Depending on the comparison, the discriminator classifies the enriched SD card data as "real" or "artificial." The discriminator is capable of evaluation through appropriate training. Thus, the discriminator ensures that the enriched SD card data corresponds to plausible variations of real SD card data.

[0029] A generic method for operating a vehicle that is at least partially automated, wherein control commands for the vehicle are generated by processing an HD card, wherein the HD card is generated online in the vehicle by processing sensor data generated by means of environmental sensors of the vehicle and SD card material by a computer model, is further developed according to the invention in that the computer model is trained with a method described above.

[0030] A vehicle of this type, comprising an environmental sensor system for generating sensor data and a computing unit for generating control commands for the vehicle based on an HD map, provides according to the invention that a computer model trained with a method described above for the online generation of HD maps is implemented in the computing unit.

[0031] As explained in the training procedure, the robustness of the AI ​​model to deviations between the real environment and its representation in SD card data is improved. This enables the AI ​​model to generate more suitable HD maps online, which in turn has a positive effect on deriving control commands for the vehicle. Consequently, the vehicle according to the invention is safer to operate, which also improves road safety. Thus, the invention provides a perception system for automated or autonomously controlled vehicles that is better able to detect and handle errors in SD card data, leading to greater reliability and safety in autonomous navigation.

[0032] Further advantageous embodiments of the inventive method for training a computer model for online generation of HD maps and of the inventive method for operating a vehicle that is at least partially automated are also evident from the exemplary embodiments which are described in more detail below with reference to the figures.

[0033] This shows: Fig. 1 A schematic representation of the processing pipeline in a vehicle for generating automated or autonomous control commands for the vehicle based on HD maps generated online in the vehicle; Fig. 2 A schematic representation of a first process step of a method according to the invention for training a computer model for online generation of HD maps; and Fig. 3 A schematic representation of a second process step of the inventive method for training the Kl model.

[0034] Fig. Figure 1 shows a known processing pipeline in a vehicle capable of at least partial or even autonomous control for generating control commands 5. The dashed area describes the training 101 of an AI model 1 designed to process sensor data 2 and SD card data SD-Map. The trained AI model 1 delivers an online generated HD map HD-Map, which is read and processed by a processing module 6. The processing module 6 then provides control commands 5 for the vehicle, represented by an accelerator pedal, a brake pedal, and a steering wheel to symbolize acceleration and steering commands. The processing module 6 can employ a wide variety of algorithms, particularly those based on artificial intelligence, preferably using artificial neural networks.

[0035] Currently, the training of AI model 1 (101) is performed using sensor data 2 and SD card data (SD-Map), which contains or describes an identical representation of the vehicle's surroundings. This results in AI model 1 being relatively insensitive to discrepancies between the content of sensor data 2 and the SD card data (SD-Map). In practice, i.e., online, differences will occur between the representation of the environment described by sensor data 2 and the representation described by the SD card data (SD-Map). The SD card data (SD-Map) typically remains unchanged for extended periods, while the real environment can change, for example, due to construction work. Furthermore, the SD card data (SD-Map) may be based on estimates or a rough survey of the environment. In contrast, sensor data 2 is acquired in real time, increasingly using state-of-the-art and therefore more precise measurement technology.This results, for example, in a different lane alignment, curve radius, the position of stop lines, and the like. Since the AI ​​model 1 has so far been trained under the premise that the content between the SD card data (SD-Map) and the sensor data 2 is identical, the inference quality of AI model 1 is adversely affected if these deviations occur. In other words, the online-generated HD map is not suitable for deriving appropriate control commands 5 for the vehicle.

[0036] To prevent this problem, a method according to the invention for training the Kl model 1 is provided. The core idea is in Fig. Figure 2 illustrates this. In training 101, the SD card data (SD-Map) is artificially enriched with errors so that the AI ​​model 1 learns to handle them appropriately, thus increasing the quality of the HD map output. This, in turn, allows the processing module 6 to generate suitable control commands 5. Fig. Figure 2 shows an advantageous embodiment in which an encoder-decoder model 3 is used to generate error-enriched SD map material SD-Map-aug from original SD map material SD-Map-ini.

[0037] The encoder-decoder model 3 is referred to here as the "generator". The generator reads the original SD card data SD-Map-ini and corresponding sensor data 2, such as camera images, LiDAR point clouds, and the like. The encoder-decoder model 3 comprises an encoder ENC, which maps the input data into a latent space 4, here in the form of the variable "Z". The decoder DEC, in turn, reads the variable Z and generates the error-enriched SD card data SD-Map-aug from it.

[0038] Advantageously, the enriched SD map data SD-Map-aug is fed to a discriminator DISC, which serves to classify the enriched SD map data SD-Map-aug as real or artificial. This is done in step 201. The corresponding result can be fed back to update the discriminator DISC in step 202. The discriminator DISC can access a map data database 7, which stores various SD map data SD-Map. The SD map data SD-Map-ini, originally fed to the encoder-devoder model 3, is preferentially stored in the map data database 7. Through appropriate training, the discriminator DISC is able to classify the enriched SD map data SD-Map-aug accordingly.Enriched SD map data classified as "real" (SD-Map-aug) is released for training 101 of the AI ​​model 1, and enriched SD map data classified as "artificial" (SD-Map-aug) is newly generated. This further increases the robustness of the trained AI model 1.

[0039] Fig. Figure 3 shows the use of the error-laden SD map material SD-Map-aug in training 101 for the Kl model 1.

[0040] Accordingly, the AI ​​model 1 trained using the inventive method is used in a vehicle for the online generation of HD maps (HD-Map). The vehicle's control behavior is then less susceptible to differences between the representation of the environment described by sensor data 2 and the representation of the environment described by SD card data (SD-Map). The vehicle is thus more reliable and safer to operate.

Claims

[1] Method for training an AI model (1) for online generation of HD maps (HD-Map), wherein the AI ​​model (1) in training (101) is supplied with SD map material (SD-Map) and sensor data (2) from environmental sensors of a vehicle as input data, wherein the sensor data (2) are correlated to sections in the SD map material (SD-Map), characterized by , that the SD map material (SD-Map) is enriched with artificial errors at least partially in the sections correlating with the sensor data (2) before being fed to the AI ​​model (1), so that the infrastructure topology in the initial SD map material (SD-Map-ini) differs at least partially from the infrastructure topology in the enriched SD map material (SD-Map-aug). [2] Method according to claim 1, characterized by, that the SD card material (SD-Map) is enriched with the artificial errors by processing the initial SD card material (SD-Map-ini) and / or at least a subset of the sensor data (2) using artificial intelligence. [3] Method according to claim 2, characterized by , that the initial SD card material (SD-Map-ini) and / or at least the subset of sensor data (2) are processed by an encoder-decoder model (3), wherein the encoder (ENC) maps input data to a latent space (4). [4] Method according to claim 3, characterized by , that the size of the latent space (4) is artificially limited to a fixed size value. [5] Method according to claim 4, characterized by , that the size of the latent space (4) is initially chosen to be an order of magnitude smaller than the average size of the SD card encodings. [6] Method according to any one of claims 2 to 5, characterized by, that the enriched SD map material (SD-Map-aug) is processed by a discriminator (DISC), wherein the discriminator (DISC) classifies the enriched SD map material (SD-Map-aug) as real or artificial, wherein SD map material (SD-Map-aug) classified as real is released for training the AI ​​model (1) and SD map material (SD-Map-aug) classified as artificial is re-enriched with errors in a further iteration loop or discarded. [7] Method for operating a vehicle that is at least partially automated, wherein control commands (5) for the vehicle are generated by processing an HD map, wherein the HD map is generated online in the vehicle by processing sensor data (2) generated by means of an environmental sensor system of the vehicle and SD map material by a computer model (1), characterized by, that the AI ​​model (1) is trained using a method according to any one of claims 1 to 6. [8] Vehicle comprising an environmental sensor system for generating sensor data (2) and a computing unit for generating control commands (5) for the vehicle based on an HD map (HD-Map), characterized by , that an AI model (1) trained using a method according to one of claims 1 to 6 for online generation of HD maps (HD-Map) is implemented in the computing unit.

Citation Information

Patent Citations

  • Method for training at least one algorithm for a motor vehicle control unit, computer program product and motor vehicle

    DE102019216836A1

  • Method for determining training data for a model improvement and data processing device

    DE102021100791A1

  • Method for providing training data for teaching an artificial neural network, method for teaching an artificial neural network, computer program product and data structure

    DE102021208187A1

  • US000012073329B2