Techniques for generating road map for automated driving
By using a visual base model to generate road maps, the problem of lack of contextual information in road maps in autonomous driving is solved, accurate road layout and traffic regulations information is provided, and safe driving and vehicle-to-vehicle communication of autonomous vehicles are supported, thus realizing the generation and applicability of high-definition maps.
Patent Information
- Application Number
- CN202510424004.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2024-04-05
- Filing Date
- 2025-04-07
- Publication Date
- 2025-10-14
AI Technical Summary
Existing basic models have not yet been able to effectively utilize rich contextual information in road map generation for autonomous driving, especially the accurate representation of applicable traffic regulations and road layout, resulting in the generated road maps being inaccurate and insufficiently applicable.
A visual base model is used to generate road maps by receiving image input data, including road layout and locally assigned context information. Deep neural networks and graph generation models such as autoregressive models, variational autoencoders, and generative adversarial networks are used. The training dataset contains annotated images of road layout and context information to generate high-definition maps to meet the needs of autonomous vehicles.
The generated road map can provide accurate road layout and applicable traffic regulations information, support the safe driving of autonomous vehicles, improve the accuracy and applicability of the road map, adapt to different vehicle types and environments, and support vehicle-to-vehicle communication and fine-grained adjustment.
Smart Images

Figure CN120778089A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to a technology for generating a road map particularly suitable for use in automated driving (AD) of vehicles. A method and a computing device for generating a road map, a method and a computing device for training a visual base model for generating a road map, a system for generating a road map, a controller for an AD vehicle, a computer program product and a computer-readable storage medium are provided. BACKGROUND
[0002] In one aspect, conventional base models have been used for many different modalities, such as images, text or audio, as proposed for example by R. Girdhar et al. in “ImageBind: One Embedding Space To Bind Them all” arXiv.org:2305.05665v2 [cs.CV], which is incorporated herein by reference. In another aspect, a map contains rich information, such as road layout, but also contextual information, such as speed limits. Vector graphics is one possible way to represent road layout. A. Jain et al. in “VectorFusion: Text-to-SVG by Abstracting Pixel-Based Diffusion Models” arxiv.org:2211.11319v1 [cs.CV] for example propose to fine-tune a text-to-image synthesizer to generate vector graphics for buildings, objects, animals and landscapes, which is incorporated herein by reference.
[0003] As shown by recent applications (LLAMA, ChatGPT, Stable Diffusion), base models (trained with seemingly unlimited amounts of data) can reach human-like capabilities in synthesizing and / or generating realistic data. However, currently, how to utilize base models for use cases in automated, in particular autonomous, driving is still an open question. SUMMARY
[0004] In the following, the solution presented in this text is described with respect to the claimed computer-implemented method, in particular for generating a road map, and with respect to the claimed computing device. Features, advantages or alternative embodiments explained with respect to the method can be assigned to the other claimed objects (e.g. device, system, computer program or computer program product) and vice versa. In other words, claims for a computing device and / or a system comprising a computing device can be improved with features described or claimed in the context of the respective method. In this case, the functional features of the method are embodied by structural units of the computing device (and / or system), respectively, and vice versa.
[0005] With respect to the first method aspect, a computer-implemented method for generating a road map particularly suitable for autonomous driving (AD) of a vehicle or robot is provided. The method comprises a step of receiving image input data. The image input data comprises acquired image data representing at least one region drivable by the vehicle. The method further comprises a step of generating a road map based on the received image input data. The generation of the road map is performed by a trained visual base model for road map generation, particularly suitable for AD. The generated road map comprises a road layout within the at least one region drivable by the vehicle and contextual information assigned locally taking into account applicable traffic regulations.
[0006] The road map (short: map) with locally assigned contextual information generated with the visual base model (also referred to as: visual base model) can provide an improved and highly accurate road layout (e.g. with resolution and / or landmark (such as road side edge) positioning accurate to a few centimeters, e.g. 3 centimeters) and applicable traffic regulations (e.g. speed limits, right of way and / or permissibility of lane changes) suitable for fully or partially autonomous driving (AD) vehicles. Alternatively or additionally, the generated road map can be used for knowledge transfer to downstream or subsequent control systems. In particular, the performance of the AD vehicle itself can be improved by training and testing the AD system of the AD vehicle using the generated road map. Further, alternatively or additionally, the (e.g. further) visual base model can be improved in its counting capabilities, e.g. when using several lanes within the road map as ground truth results for counting lanes within the image input data.
[0007] The road map can particularly comprise a high definition (HD) map. The HD map can be particularly suitable for autonomous driving.
[0008] An AD vehicle can be a car, bus or truck equipped with an autonomous driving (AD) functionality. The AD driving functionality can comprise full automation (also denoted as autonomous driving, in particular L5), at least partial automation (e.g. high automation L4, conditional automation L3 or partial automation L2) and / or assisted driving automation (e.g. LI, e.g. including adaptive cruise control ACC). A fully automated vehicle can also be denoted as an autonomous vehicle and / or an autonomous driving vehicle.
[0009] The vehicle can be (or can comprise) in particular a motorized passenger vehicle (e.g. requiring registration), a car and / or a utility vehicle. The vehicle can be a street- driving vehicle. Alternatively or additionally, the vehicle can be comprised in a robotic system (e.g. a robot in a manufacturing environment).
[0010] The applicable traffic regulations can comprise different regulations for different types of vehicles. Alternatively or additionally, the applicable traffic regulations can define (and / or limit) a manner in which a vehicle can drive or use an area (or path) (e.g. speed limit, weight limit, width limit and / or height limit defining parking requirements of a parking lot). For example, a road or area can be restricted for use by vehicles below a predetermined weight threshold (e.g. for driving on a bridge) and / or below a predetermined height threshold (e.g. for driving under a bridge). Alternatively or additionally, the applicable traffic regulations can comprise a predetermined width threshold, beyond which a vehicle can not use the road (e.g. due to passage width, lane width and / or a barrier outside the lane and / or road).
[0011] The applicable traffic regulations can be denoted as a permissible path trajectory (also denoted as path and / or trajectory) and / or a non-permissible path trajectory (e.g. in a pedestrian zone) for the vehicle. Further, alternatively or additionally, the applicable traffic regulations can be denoted (and / or by) road markings (e.g. comprising solid lines and / or dashed lines), traffic signs and / or traffic lights.
[0012] The acquired image data can be acquired by a sensor system, in particular an image capturing device. The sensor system and / or the image capturing device can be based on at least one of the following technologies: video, radar, LiDAR, ultrasound, motion, thermal imaging. Alternatively or additionally, the acquired image data can be acquired by means of (e.g. ground) satellites and / or by means of drones. For example, satellite data can advantageously provide an overview of an extended area, in particular comprising at least one area drivable by the vehicle.
[0013] The image input data can comprise at least the acquired image data and / or can particularly comprise digital image data. In particular, the image input data can comprise an aerial view of at least one region. Alternatively or additionally, the image input data can comprise a map, e.g. an (in particular open) street map and / or an (in particular open) navigation map, e.g. in a graphical format.
[0014] A region drivable by a vehicle (also simply referred to as: drivable region) can also be denoted as trafficable, passable, trafficable, traversable and / or reachable by a vehicle. A drivable region (e.g. comprising a road) can particularly comprise a region (e.g. a road) that is open to traffic, but is not limited to a region (e.g. a road) that is legally open to the public and / or specifically designed for driving a vehicle. For example, a private path can be drivable (e.g. in a technical sense), but not designated as “open to traffic”.
[0015] A region can be defined by positional and / or geographical data and / or can for example comprise a road (also denoted as street) and / or a lane of a street.
[0016] The generated road map can comprise a (e.g. simple) geometric representation of the road layout comprising for example lines, curves and / or intersections. Alternatively or additionally, the generated road map can be represented in Scalable Vector Graphics (SVG) or can comprise SVG. Any representation of the road layout and / or any graph can comprise nodes and edges between nodes (e.g. a subset thereof). Nodes can represent (e.g. equidistant) points of a structure (e.g. visible) and edges can represent a relationship between nodes. For example, adjacent nodes of a line between lanes are connected by an edge.
[0017] The generated road map can comprise and / or indicate a road layout. The generated road map comprises contextual information. The contextual information can be used for automatically analyzing applicable traffic regulations. The contextual information can be assigned locally to a specific position in the road map. For example, a speed limit or a one-way traffic regulation for a specific section of a street can be represented with a respective contextual information that is assigned to the part of the road map in which the street section is represented. The contextual information can be provided in the generated road map as an annotation and / or as an overlay graph and / or as an expandable box and / or as a thumbnail image and / or as text embedding.
[0018] In general, the road layout can comprise any (in particular road-related) structure (and / or representation, and / or information) of at least one area drivable by the vehicle, in particular visible in an (e.g. realistic) overhead view. For example, the road layout can indicate roads and / or streets drivable by the vehicle, and / or passages or intersections for pedestrian or bicycle lines. The road layout can comprise information necessary for analyzing a traffic scenario. It can refer to static data (in particular without changes over time). Alternatively or additionally, the road layout can be represented by an SVG or any other graphical representation (and / or graphics). For example, a lane can be attributed to allow a left turn. Alternatively or additionally, an arrow indicating the left turn allowance can be represented graphically.
[0019] For example, the road layout can comprise a shape and / or number of lanes, a type of lane (e.g. for use by motorized vehicles such as cars and / or for use by bicycles), a width of a lane, a type of separation of adjacent lanes (e.g. with solid and / or dashed lines), arrows indicating a driving direction of a lane, stop lines, zebra crossings, parking lots (e.g. boundaries thereof), and / or any (e.g. drawn) features of a road visible in an overhead view. For example, a number of lanes, a width of each lane, and / or a type of lines (in particular solid or dashed lines) between adjacent lanes can be visible from above. Alternatively or additionally, a type of lane for motorized vehicles (e.g. cars) can be distinguished from a bicycle lane and / or a pedestrian road by its width and / or color (e.g. width and / or color of asphalt and / or hard pavement).
[0020] According to the specification of the road layout, which comprises in particular any type of information visible from above (and / or in an overhead view), the generated road map comprises visible cues necessary for (e.g. path) planning of safe driving (in particular AD). Thus, the generated road map can be applied for automatic planning and subsequent control of the vehicle.
[0021] Alternatively or additionally, the road layout can comprise a set of predetermined primitives, e.g. comprising simple geometric forms, in particular line elements. The line elements can be solid or dashed lines, respectively, for separating adjacent lanes between which a lane change of a vehicle is not allowed, or a lane change of a vehicle is allowed. Alternatively or additionally, the line elements can comprise straight lines (e.g. of a predetermined length) and / or line elements defined by simple functions (e.g. circular segments comprising a circle of a predetermined radius).
[0022] The context information comprised in the generated road map can comprise applicable traffic laws (also referred to as: traffic laws and / or traffic rules in short), which can depend on a country and / or a region. Alternatively or additionally, the context information can comprise one or more road conditions.
[0023] The locally assigned context information can relate to one or more nodes and / or edges of the graph (e.g., SVG), e.g., a simple graph representation, and / or one or more graph elements (in particular their (multiple) interrelations with each other), the graph elements representing a road layout (e.g., a section and / or a portion thereof) of at least one area drivable by the vehicle, e.g., within city limits, outside city limits, on a highway, at a crossroad, at a pedestrian crossing, at a bridge, at a road with a significant inclination (e.g., 4% or above).
[0024] For example, the locally assigned context information taking into account applicable traffic regulations can include weight constraints of a lane, height constraints of a lane, road structure (in particular including road surface and / or road conditions, e.g., paved, unpaved, bumpy, and / or potholed), speed limits, right of way, no-go, traffic signs, traffic lights, and / or accuracy indicators.
[0025] The locally assigned context information can be specific to a vehicle type. For example, no-go and / or (e.g., low) speed limits can only apply to vehicles above a predetermined weight and / or length (e.g., buses and / or trucks). Alternatively or additionally, the locally assigned context information can be time-specific. The time-specificity can relate to a certain time of day and / or a certain day. For example, lower speed limits can apply to reduce noise at night, and / or access to a lane can be reserved for a group of vehicles (e.g., buses and / or taxis) during working days, in particular during rush hours.
[0026] For example, the accuracy indicators can include expected accuracy of locations of lines forming a boundary of a lane and / or other landmarks in the road layout. Alternatively or additionally, the accuracy indicators can include a confidence level of a correct assignment of applicable traffic regulations, e.g., a confidence level of a correct assignment of a speed limit (in particular obtained from image input data, e.g., after semantic segmentation).
[0027] The specification of the locally assigned context information in combination with the road layout (and / or visible cues) further facilitates safe driving, in particular (e.g., path) planning of the AD.
[0028] The visual base model can be or comprise a generative artificial intelligence system (AI), i.e. a system that generates or produces content, like images. The visual base model can be based on a deep neural network. The visual base model can comprise: autoregressive base models, which generate an input block by block; and denoising base models, which corrupt, then restore, an input. Alternatively or additionally, the visual base model can comprise graph generation models (e.g. autoregressive models, variational autoencoders, regularizing flows, generative adversarial networks, GANs, and / or diffusion models) for (deep and / or nested) graph generation, as described in Y. Zhu, “A Survey of Deep Graph Generation: Methods and Applications”, arXiv:2203.06714v3, which is incorporated herein by reference.
[0029] The visual base model can be (pre-)trained (training phase) from a broad range of data, e.g. a huge dataset, and can be adapted or fine-tuned (adaptation phase) for various different specific downstream tasks. The visual base model can be trained by means of self-supervision, transfer, and / or active learning.
[0030] The training dataset can comprise images (e.g. acquired by an optical sensor system, e.g. like a satellite image / imaging system) of regions as image input data and annotations (in particular ground truth) of road maps (e.g. locally associated) comprising road layouts and locally assigned context information. The annotations can be embedded in the images or can be locally associated with the images. Alternatively or additionally, the training dataset can comprise (e.g. artificially generated) vector graphics (e.g. scalable vector graphics, SVG) representing road layouts (and / or resulting images of road layouts) with attributes representing locally assigned context information (in particular ground truth). The image input data associated with the vector graphics can be generated in a synthetic manner (e.g. by means of a GAN). Further, alternatively or additionally, the training dataset can comprise images (e.g. acquired by an optical sensor system, e.g. like a satellite image / imaging system) as image input data and (e.g. artificially generated) vector graphics (e.g. SVG) representing road layouts (and / or resulting images of road layouts) of the images with attributes representing locally assigned context information (in particular ground truth).
[0031] The method according to the first aspect can further comprise the step of providing the generated road map to a controller of a vehicle configured for automated (e.g. autonomous) driving (AD), or to a robot.
[0032] The provision of the generated road map in combination with sensors for detecting other current road users and / or (particularly moving) obstacles (e.g. parked vehicles such as cars and / or bicycles) may enable AD of the vehicle.
[0033] The generated road map or parts thereof can be used for vehicle-to-vehicle communication in order to provide information to other participants in the traffic scene.
[0034] The road layout of the generated road map may be adjustable in terms of granularity, in particular for presenting locally assigned context information and / or depending on locally assigned context information.
[0035] The (e.g. presented) road layout may be adjustable in granularity by means of user input (in particular by a user and / or driver of the vehicle for AD). The user input may be received via a user interface (UI), which for example comprises a rotatable button and / or a graphical user interface (GUI). The UI may be deployed in the vehicle for AD, for example at a center console, and / or in particular within reach of the user and / or driver. Alternatively or additionally, the (e.g. presented) road layout may be adjustable (in particular without user input) depending on the (e.g. planned) speed of the vehicle for AD.
[0036] The scalability of the road layout allows for detailed adaptation as needed. For example, a coarse granularity can be suitable for high-speed driving on a straight highway. A fine granularity can be suitable for navigating challenging traffic situations, such as within city limits, multiple lanes for different types of vehicles, and / or intersections.
[0037] The image input data may include overhead image data of an area, the area including at least one area in which the vehicle is drivable. Alternatively or additionally, the image input data may be acquired by an optical sensor system (including a satellite imaging system, an (particularly airborne) camera system, and / or an aerial photography system), for example via one or more drones.
[0038] Bird's-eye view image data, satellite images and / or aerial images can advantageously provide reliable data about lane structures, intersections and / or zebra crossings, for example, regarding width and / or length.
[0039] The generation of a road map may include generating a (e.g., deep and / or nested) graph comprising nodes and edges. The edges may connect, link, and / or relate nodes (particularly subsets thereof) representing the road layout. The nodes and edges may be supplemented with attributes representing locally assigned contextual information. Optionally, the (e.g., deep) graph comprises scalable vector graphics (SVG).
[0040] For example, a node can represent a point on a line. An edge can relate nodes comprised in a line. For example, an edge can connect a point in the middle of a line to its nearest two neighbors, each neighbor being in a different direction along the line (e.g., one node to the left and one node to the right of the middle node). Alternatively or additionally, edges between nodes associated with two lines forming a boundary of a lane can represent a stop line. Alternatively or additionally, on a higher level in the (e.g., depth and / or nesting) graph, nodes can for example represent lines, and edges can for example represent their role when describing a piece of a lane (e.g., a left boundary, a right boundary, or a center line). Further alternatively or additionally, on a still higher level in the (e.g., depth and / or nesting) graph, nodes can for example represent pieces of a lane, and edges can for example connect pieces of a lane and indicate whether it is possible to cross and whether it is in compliance with traffic rules. Still further alternatively or additionally, lines can form a lane, and / or lanes can form a road network (e.g., depending on the traffic agent).
[0041] In particular, the SVG can be generated by a (e.g., general-purpose) depth (and / or nesting) graph generator. The (e.g., general-purpose) depth (and / or nesting) graph generator can in particular generate attributes of nodes and / or edges.
[0042] The (e.g., generated) graph can comprise a nested structure. Alternatively or additionally, the graph can be encoded as or converted into textual data for generating a road map.
[0043] With the graph representation with attributes, a particularly simple, memory-efficient, and / or fast rendering of a road map can be provided.
[0044] The generation of a road map can comprise generating a representation of a road layout using graph primitives. The graph primitives can comprise simple geometric forms, in particular lines (e.g., solid lines and / or dashed lines) and / or line elements parameterized by simple functions, in particular linear functions and / or functions representing circular segments.
[0045] The graph primitives can be combined with the representation in terms of nodes and edges. For example, a line can be represented by nodes related to edges (e.g., according to a simple function).
[0046] The representation in terms of graph primitives can provide a road map in a simple, memory-efficient, and / or fast-to-load format to a renderer.
[0047] The trained visual basis model can comprise a graph generation model. Optionally, the graph generation model can comprise an autoregressive model, a variational autoencoder, a regularizing flow, a generative adversarial network, and / or a diffusion model.
[0048] The graph generation model can be configured for (e.g., deep and / or nested) graph generation with nodes, edge-related nodes (e.g., included in the SVG), and supplemental attributes of at least a subset of the nodes and edges.
[0049] By means of the trained visual base model, in particular the graph generation model, attributes of nodes and / or edges can be generated based on the image input data. For example, speed limits can be derived from environmental information (e.g., depending on the country, depending on the curvature and / or radius of the lane, depending on the setting within or outside of a city, and / or depending on a highway bridge exposed on a valley). Alternatively or additionally, based on the image input data, the trained visual base model, in particular the graph generation model, can classify a road as a traffic-busy road or a traffic-unbusy road (e.g., based on a number of vehicles in the image input data and / or based on its environment including indicators such as a densely populated area or a sparsely populated area).
[0050] The trained visual base model can be configured for performing a classification and / or semantic segmentation of the image input data. Alternatively, a pre-processing can be performed on the image input data, e.g., by means of a segmentation algorithm for segmenting relevant structures in the image input data.
[0051] The classification and / or semantic segmentation can in particular include detecting objects and / or structures, in particular traffic lights, traffic signs, and / or road structures (e.g., including road surfaces).
[0052] With respect to the second method aspect, a method of training a visual base model for generating a road map based on received image input data is provided. The method comprises a step of receiving a training data set. The training data set can comprise annotated aerial view images, in particular acquired by means of an optical sensor system, in particular a satellite imaging system. The annotations can comprise a road layout with locally assigned context information. Alternatively or additionally, the training data set can comprise rendered images, which are based on (in particular artificially generated) graphs (e.g., SVGs) or are in the format of graphs, which represent a road map or a road layout (and / or a resulting image of a road layout). The training data set can in particular comprise artificially generated graphs (and / or resulting road layout images) as ground truth (at least in part), and artificially generated (in particular based on artificially generated graphs) aerial view images. Further, alternatively or additionally, the training data set can comprise images (in particular aerial view images of traffic scenes, e.g., acquired by means of an optical sensor system) and rendered images, which are based on (in particular artificially generated) graphs (e.g., SVGs) or are in the format of graphs, which represent a road map or a road layout, in particular for (aerial view) images.
[0053] The method further comprises a step of training a visual base model for generating a road map based on received image input data. The generated road map comprises a road layout as well as contextual information taking into account local assignments of applicable traffic regulations. Optionally, the training is self-supervised and / or based on a reconstruction loss between ground truth comprised in a received training data set (e.g., its annotations or presented images) and a road map generated by the visual base model.
[0054] The training data set can be represented as where x n is an (e.g., satellite) image, and y n is its graph (in particular SVG) representation, which can be obtained, for example, from OpenStreetMap.
[0055] For example, the computing device and / or model architecture can comprise a conventional encoder-decoder architecture, wherein the encoder learns a mapping from images to a latent space, and the decoder learns a mapping from the latent space to a graph (e.g., SVG) representation.
[0056] For the encoder, for example, a pre-trained image transformer can be used.
[0057] The graph representation (e.g., SVG) can correspond to (and / or can comprise) a language for describing two-dimensional graphs, in particular in XML. For example, a deep (and / or nested) graph generator as described by Y. Zhu in “A Survey of Deep Graph Generation: Methods and Applications” arXiv:2203.06714v3 can be used. Alternatively or additionally, since the graph representation (e.g., SVG) is text-based, the decoder can be implemented as a language decoder, e.g., a pre-trained text transformer.
[0058] For example, cross-entropy can be used as training loss, in particular since cross-entropy is commonly used for language modeling tasks.
[0059] The weights of the visual base model can be initialized from a pre-trained default image-to-text model (e.g., TrOCR, https: / / arxiv.org / pdf / 2109.1022.pdf , the document of which is incorporated herein).
[0060] In a first extension, once the (e.g., visual base) model has been trained, it can also be used as a basis for a specialist model that is only available with a small training data set D’. For example, a standard image-to-text model can first be used, then fine-tuned on image-to-OpenStreetMap data, and then fine-tuned on HD maps.
[0061] In a second extension, if also text information outside the image is given as input, a second encoder, in particular a text encoder, can be added to the (e.g., visual base) model.
[0062] The training phase can comprise testing, validating and / or verifying (e.g., initially) the trained visual base model (and / or can precede testing, validating and / or verifying (e.g., initially) the trained visual base model).
[0063] The training of the visual base model can comprise utilizing a generated understanding of traffic rules and / or public road layout.
[0064] By the training of the visual base model, a powerful tool for generating road maps based on a small amount of image input data (e.g., a single satellite image) and / or without the need for human interaction is provided.
[0065] With respect to a further aspect, a use of the road map is provided. The road map is generated by the method according to the first method aspect, which method is used for AD vehicles or for trajectory prediction, path planning, collision avoidance and / or behavior prediction of traffic, in particular for AD vehicles.
[0066] The generated road map can be stored and / or used locally on a processing unit of the AD vehicle. Alternatively or additionally, the generated road map can be accessed via cloud technology of the AD vehicle (e.g., using wireless communication technology).
[0067] The generated road map can be communicated to units outside the vehicle, like other vehicles, servers and / or mobile devices, by using network technology.
[0068] The behavior prediction can relate to any traffic outside the AD vehicle (e.g., including further vehicles and / or pedestrians). The behavior prediction can be used for collision avoidance, in particular. For example, by predicting the behavior of traffic participants, the AD vehicle can be configured for anticipatory driving. For example, to avoid a collision with other traffic participants, the speed can be reduced and / or the lane can be changed.
[0069] Using the generated road map, the safety of the AD can be improved. Alternatively or additionally, the path to be taken by the AD vehicle can be optimized.
[0070] With respect to the first device aspect, a computing device for generating a road map, particularly suitable for use in AD for a vehicle, is provided. The computing device comprises an input image data reception interface configured for receiving image input data. The image input data comprises acquired image data representing at least one region drivable by the vehicle. The computing device further comprises a road map generation module for generating a road map based on the received image input data. The generation of the road map is performed by a trained visual base model for road map generation, particularly suitable for use in AD. The generated road map comprises a road layout within the at least one region drivable by the vehicle and contextual information assigned locally taking into account applicable traffic regulations.
[0071] The computing device can further be configured to perform any one of the steps and / or comprise any one of the features disclosed within the context of the first method aspect.
[0072] With respect to the second device aspect, a computing device for training a visual base model for generating a road map based on received image input data is provided. The computing device comprises a training data reception interface configured for receiving a training data set. The training data set can comprise annotated aerial view images, particularly acquired by means of an optical sensor system, particularly a satellite imaging system. The annotations can comprise a road layout with locally assigned contextual information. Alternatively or additionally, the training data set can comprise rendered images, the rendered images being based on (particularly artificially generated) graphics or being in the format of the graphics, the graphics representing a road layout (and / or a resulting image of a road layout). The training data set can particularly comprise artificially generated graphics (and / or resulting road layout images) as ground truth (at least in part), and artificially generated (particularly based on artificially generated graphics) aerial view images. Further, alternatively or additionally, the training data set can comprise (particularly aerial view) images and rendered images in the format of (particularly artificially generated) graphics, the graphics representing a road layout of the (particularly aerial view) images.
[0073] The computing device further comprises a visual base model training module configured for training a visual base model for generating a road map based on received image input data. The generated road map comprises a road layout and contextual information assigned locally taking into account applicable traffic regulations. Optionally, the training is self-supervised and / or based on a reconstruction loss between ground truth comprised in the received training data set (particularly in the annotations or the rendered images) and a road map generated by the visual base model.
[0074] The computing devices of the first device aspect and the second device aspect can be identical.
[0075] With respect to the system aspect, a system for generating a road map for AD, particularly for a vehicle, is provided. The system comprises a computing device according to the first method aspect and at least one sensor and / or image capturing device configured for providing image input data to an input image data receiving interface of the computing device.
[0076] With respect to the third device aspect, a controller for an AD vehicle is provided. The controller comprises a receiving interface for receiving a road map generated based on received image input data. The road map comprises a road layout within at least one region drivable by the vehicle and contextual information assigned considering applicable traffic regulations. The controller further comprises a processing unit configured for trajectory planning of the AD vehicle.
[0077] With respect to a further aspect, a computer program product is provided, the computer program product comprising program elements which, when loaded into the memory of a computing device, cause the computing device to perform the steps of the first method aspect and / or the second method aspect.
[0078] With respect to a still further aspect, a computer readable medium is provided, having stored thereon program elements which can be read and executed by a computing device to perform the method according to the first method aspect and / or according to the second method aspect when the program elements are executed by the computing device. BRIEF DESCRIPTION OF DRAWINGS
[0079] Exemplary embodiments of the method, device, system and controller according to the present disclosure are illustrated in the accompanying drawings. In particular: Figure 1 An exemplary flow chart of a computer implemented method for generating a road map according to the present disclosure is shown, Figure 2 An exemplary flow chart of a computer implemented method for training a visual base model for generating a road map according to the present disclosure is shown, Figure 3 An architecture of a computer device for generating a road map is shown in a schematic representation, Figure 4 An architecture of a computer device for training a visual base model is schematically illustrated, Figure 5 An exemplary graphical representation of an urban area comprising a region drivable by a vehicle is shown, Figure 6 An urban area according to Figure 5 is exemplarily shown as a vector graphic, Figure 7 An exemplary architecture of a computing device, in particular of a road map generation module 306, is schematically depicted, and An exemplary architecture of a computing device, in particular of a road map generation module 306, is schematically depicted, andFigure 8 The use of the method for automated driving is illustrated by way of example. DETAILED DESCRIPTION
[0080] Figure 1 An exemplary flow chart of a computer-implemented method 100 for generating a road map particularly suitable for automated driving (AD) of a vehicle is shown.
[0081] The method 100 includes a step S104 of receiving image input data. The image input data includes acquired image data representing at least one area where the vehicle can travel.
[0082] Method 100 further includes a step S106 of generating a road map based on the received S104 image input data. The road map generation S106 is performed by a trained visual base model for road map generation, particularly suitable for AD. The generated S106 road map includes a road layout within at least one area where the vehicle is drivable, as well as locally assigned contextual information that takes into account applicable traffic regulations.
[0083] Optionally, the method 100 comprises a step S108 of providing S108 the generated S106 road map to a controller of a vehicle configured for AD.
[0084] Figure 2 An exemplary flow chart of a computer-implemented method 200 for training a visual base model for generating a road map based on received image input data is shown.
[0085] Method 200 comprises a step S202 of receiving a training dataset. The training dataset may comprise an annotated top-view image, in particular acquired by means of satellite imaging and / or one or more drones. The annotations may comprise a road layout with locally assigned contextual information. Alternatively or additionally, the training dataset may comprise a rendered image, the rendered image being based on or in the format of a (in particular artificially generated) graphic (e.g. SVG) representing a road layout or a road map (and / or a resulting image of a road layout). The training dataset may in particular comprise the artificially generated graphic (and / or the resulting road layout image) as ground truth (at least part thereof), and an artificially generated (in particular based on the artificially generated graphic) top-view image. Further, alternatively or additionally, the training dataset may comprise a top-view image and a rendered image, in particular in the format of a (in particular artificially generated) graphic (e.g. SVG), representing an image of a road layout or a road map.
[0086] The method 200 further comprises a step S204 of training a visual base model for generating a road map based on received image input data. The generated road map comprises a road layout as well as contextual information taking into account local assignments of applicable traffic regulations. Optionally, the training S204 is self-supervised and / or based on a reconstruction loss between ground truth comprised in the received S202 training data set, in particular in elements of the training data set representing a road layout or a road map, and a road map generated by the visual base model, wherein the generation of the road map is based on images comprised in the received S202 training data set.
[0087] Figure 3 An architecture of a computing device 300 for generating a road map particularly suitable for AD of a vehicle is schematically illustrated.
[0088] The computing device 300 comprises an input image data reception interface 304 configured for receiving image input data. The image input data comprises acquired image data representing at least one region drivable by a vehicle.
[0089] The computing device 300 further comprises a road map generation module 306 for generating a road map based on received image input data. The generation of the road map is performed by a trained visual base model for road map generation, particularly suitable for AD. The generated road map comprises a road layout within the at least one region drivable by a vehicle as well as contextual information taking into account local assignments of applicable traffic regulations.
[0090] Optionally, the computing device 300 comprises an output interface 308 configured for providing the generated road map to a controller of a vehicle configured for AD.
[0091] The input image data reception interface 304, and optionally the output interface 308, can be embodied by an input-output interface 310. The road map generation module 306 can be embodied by a processing unit 312. The computing device 300 can further comprise at least one memory 314.
[0092] Figure 4 An architecture of a computing device 400 for training a visual base model for generating a road map based on received image input data is schematically illustrated.
[0093] The computing device 400 comprises a training data reception interface 402 configured for receiving a training data set. The training data set can comprise annotated aerial view images, in particular acquired by means of an optical sensor system, in particular a satellite imaging system and / or one or more drones. The annotations can comprise a road layout with locally assigned context information. Alternatively or additionally, the training data set can comprise rendered images based on (in particular artificially generated) graphics (e.g. SVG) comprising a road layout (and / or a resulting image of a road layout). The training data set can in particular comprise artificially generated graphics (and / or resulting road layout images) as ground truth (at least in part) and artificially generated (in particular based on artificially generated graphics) aerial view images. Further, alternatively or additionally, the training data set can comprise (in particular aerial view) images and rendered images based on (in particular artificially generated) graphics (e.g. SVG) comprising a road layout for the (in particular aerial view) images.
[0094] The computing device 400 further comprises a visual base model training module 404 configured for training a visual base model for generating a road map based on received image input data. The generated road map comprises a road layout and locally assigned context information taking into account applicable traffic regulations.
[0095] Optionally, the training is self-supervised and / or based on a reconstruction loss between ground truth comprised in the received training data set and a road map generated by the visual base model, wherein the generation of the road map is based on images comprised in the received training data set.
[0096] The training data reception interface 402 can be embodied by the input- output interface 406. The visual base model training module 404 can be embodied by the processing unit 408. The computing device 400 can further comprise at least one memory 410.
[0097] Any of the processing units 312; 408 can be embodied by a central processing unit (CPU) and / or a graphics processing unit (GPU).
[0098] The technology facilitates efficient and highly accurate generation of road maps in extensions of F. Poggenhans et al., “Lanelet2: A high-definition map framework for the future of automated driving” (2018), which is incorporated herein by reference. While the (e.g., nested) graph structure generated by the method 100 can (in particular potentially and / or among other options) resemble (and / or be identical to) the (e.g., nested) graph structure of Lanelet2, the graph is generated (in particular automatically) by a (e.g., learned and / or trained) visual basis model in accordance with the computer-implemented method 100. In contrast, Lanelet2 uses manual and / or semi-automatic “clicking” of the map and / or automatically generates the map (e.g., partially) from in-car measurements. For example, lane detection and / or (e.g., GPS) location can be used to generate the map, however, this is very costly, in particular in terms of time, effort, computational resources, and memory, due to the need to drive to all possible places. On the other hand, the method 100 can utilize satellite images and / or navigation maps (and / or Secure Digital, SD maps) (e.g., as image input data), e.g., which can be available as OpenStreetMap.
[0099] The technology relates to providing a generative model (in particular a visual basis model) of a road map, which can in turn be used to train and test various algorithms for AD, such as behavior prediction and / or planning.
[0100] By the technology, a visual basis model can be trained to generate a road map from image data (also referred to simply as: images). The road map comprises a road layout (also referred to as: road structure) represented as a graph and (in particular locally assigned) contextual information, such as speed limits, zebra crossings, and / or right-of-way.
[0101] Figure 5 An exemplary graph representation of an urban area is shown, which comprises areas drivable by (e.g., AD) vehicles. The exemplary graph representation comprises a road layout and locally assigned contextual information. For example, (e.g., car) lanes 502 of the two-way Karlstrasse and the one-way Karlstrasse are provided with arrows representing the driving direction of each lane. At the intersection of the two streets (Karlstrasse and Kaiserstrasse), traffic lights are installed, as shown at reference sign 504. Several types of parking areas are shown, i.e., a parking area for bicycles 506-1, a parking area for motorcycles 506, and a parking lot 506-3. Figure 5 An exemplary graph representation of an urban area is shown, which comprises areas drivable by (e.g., AD) vehicles. The exemplary graph representation comprises a road layout and locally assigned contextual information. For example, (e.g., car) lanes 502 of the two-way Karlstrasse and the one-way Karlstrasse are provided with arrows representing the driving direction of each lane. At the intersection of the two streets (Karlstrasse and Kaiserstrasse), traffic lights are installed, as shown at reference sign 504. Several types of parking areas are shown, i.e., a parking area for bicycles 506-1, a parking area for motorcycles 506, and a parking lot 506-3.Figure 5 The drivable area in FIG further schematically illustrates a bus lane 508 (eg, at a bus stop).
[0102] Figure 6 It is shown as an example Figure 5 The urban area is represented as a vector graphic, in particular an SVG. The (e.g. car) lane 502 and the bus lane 508 are represented by nodes and edges linking these nodes, the nodes being depicted by crosses. Figure 6 In the example, edges are represented by straight lines.
[0103] Positions where it is necessary to stop (e.g., at a stop line) or give way (e.g., corresponding to additional nodes other than the node indicated by the cross) are schematically indicated by squares. Alternatively or additionally, additional nodes (e.g., depicted by squares), such as a curved (and / or bent) road leading to parking lot 506-3, may divide (and / or split) the corresponding lane into straight segments.
[0104] Figure 7 An exemplary architecture of computing device 300, and in particular road map generation module 306 (and / or visual base model), is schematically depicted. The exemplary architecture includes an encoder 704 (e.g., a pre-trained image transformer) and a decoder 708. Encoder 704 is configured to receive image input data, as illustrated by a top-down image 702. Encoder 704 is further configured to output a latent representation at reference symbol 706 (e.g., represented by Z). Latent representation 706 predicts the correct sequence of shifted tokens and / or words of a graphic (and / or XML) based on the input.
[0105] For example, one can use cross-entropy as the training loss, especially since cross-entropy is commonly used in language modeling tasks.
[0106] Common training techniques may include predicting the next word-gram conditioned on the input (e.g., at reference symbol 702) and the past word-grams. An efficient implementation may include shifting the word-grams (e.g., at Figure 7 The ground truth sequence (denoted as "right shifted" in [ ]) is input to the decoder 708 and an initial sequence is predicted. The sequence may correspond to, for example, vector graphics (particularly SVG) and / or any other map format (e.g., which is intended for prediction and / or generation).
[0107] At reference symbol 710 , a rendered image is exemplarily shown, which may be used only for visualization (eg, as a display when the generated road map is used for AD and / or for manual supervision during a training phase). The rendered image 710 need not be used for training.
[0108] Figure 8 The use of the method for AD is exemplarily illustrated. At reference sign 802, an AD vehicle is schematically depicted. The AD vehicle 802 comprises a controller (and / or planning unit) 804 and, optionally, a display 806. The controller (and / or planning unit) 804 receives one or more generated road maps from the computing device 300 configured for generating road maps. On the optional display 806, a graphical representation of the road map (e.g., as depicted at reference sign 710 in FIG. 7) can be provided to the driver of the AD vehicle 802 for monitoring (e.g., for autonomous vehicles) and / or guidance (e.g., for vehicles with assisted driving functionality). Figure 7
[0109] Conventional maps, satellite images, and OpenStreetMap and other navigation maps (e.g., in various combinations) can provide a large (and / or seemingly infinite) amount of data for training the visual basis model for generating road maps based on received image input data and for generating road maps specifically for AD. Alternatively or additionally, for pre-training, artificially generated graphics (e.g., SVGs, in particular as at least part of the ground truth) and their respective renderings (in particular as image input data) can act as (e.g., truly infinite) training data pairs.
[0110] The computing device 300 (and / or a network, e.g., comprising an encoder and a decoder) trained to generate graphics from images (e.g., road graphics from satellite images, which are available on a large scale from all over the world; and / or SVGs from their renderings, which can be self-supervised for any number of generated SVGs) can improve the performance of various downstream tasks. For example, map data can be added to aerial view image input data (in short: images) for any data set for which it is available. Adding map data to aerial view images increases the amount of training and test data for standard routines in AD scenarios (e.g., trajectory prediction and / or planning) that require map information.
[0111] This technique significantly extends the approach used in VectorFusion, e.g., by training the computing device 300 (and / or network) with a reconstruction loss between real images and rendered images (e.g., generated by the visual base model). In contrast to VectorFusion, the graphs used for training and / or road map generation are generated using the general deep (and / or nested) graph generator described in Y. Zhu et al., arxiv.org:2203.06714v3. An advantage of using the deep graph generator is that further relationships can be added to the graph, such as node attributes (e.g., road conditions, such as paved and / or bumpy) and / or edges between nodes (e.g., stop signs related to specific stop lines).
[0112] Alternatively or additionally, the trained (in particular visual base) model can also be used in other contexts. According to a first embodiment, a model capable of generating a road map can require a general understanding of traffic rules and common road layouts. The knowledge of traffic rules and common road layouts can be transferred and help to further improve the performance of AD (in particular autonomous driving) vehicles. According to a second embodiment, which can be combined with the first embodiment, generating a road map can require (in particular “good”) counting capabilities (e.g., in terms of several parallel lanes). Conventional visual base models lack these counting capabilities. Training a visual base model on a road map generation process can help to improve the counting capabilities.
[0113] The techniques for training a visual base model for generating a road map based on received image input data and for generating a road map, in particular for AD, can be used for analyzing data obtained from sensors. The sensors can determine measurements of the environment in the form of sensor signals, which can be given by, in particular, digital images, e.g., including videos, radar, LiDAR, ultrasound, motion, and / or thermal images.
[0114] Downstream uses of the techniques for training a visual base model for generating a road map based on received image input data and for generating a road map can include virtual sensors, video, and / or audio analysis and / or classification. Visual sensors can include, based on the sensor signals, information can be obtained about elements encoded by the sensor signals (e.g., indirect measurements can be performed based on sensor signals used as direct measurements), e.g., taking into account locally assigned context information (in particular traffic rules) associated with the imaged road layout.
[0115] Video and / or audio analysis and / or classification may be used to classify sensor data, detect the presence of objects in sensor data, and / or perform semantic segmentation on sensor data, e.g., with respect to traffic signs and / or road surfaces.
[0116] The techniques for training a visual base model for generating a road map based on received image input data and for generating a road map are particularly suitable for any AD application, particularly autonomous vehicles. Alternatively or additionally, a robot at a manufacturing site (for which a road map is generated according to the techniques) can automatically and / or autonomously move within the manufacturing site (e.g., as a drivable area).
[0117] Upstream use of the technology for training a visual base model for generating road maps based on received image input data and for generating road maps particularly for AD may include active learning, testing and / or data curation in actively selecting data (particularly thereby reducing data traffic) that a technical system (e.g., comprising at least one sensor and / or image capture device) transmits to a back-end computer, which in turn may use such information to train a machine learning system (e.g., a visual base model) for testing, verifying and / or validating the machine learning system and / or for generating road layouts for training and testing AD algorithms.
[0118] Further upstream uses of the technology for training a visual base model for generating a road map based on received image input data and for generating road maps in particular for AD may include methods and / or data for training, for example by generating training data for training and / or by generating test, verification and / or validation data to check whether the trained ML system (e.g., including the visual base model and / or AD algorithm) can then be operated safely.
[0119] Further upstream uses of the technology for training a visual base model for generating road maps based on received image input data and for generating road maps specifically for AD can include generating a model to generate training or test data (e.g., for an AD algorithm). After being trained in this manner, the ML system (e.g., an AD algorithm) can then be put into downstream use.
[0120] Prior art cited: [1]Girdhar et al in "ImageBind: One Embedding Space To Bind Them all",arxiv.org:2305.05665v2[cs.CV] [2]“VectorFusion:Text-to-SVG by Abstracting Pixel-Based Diffusion Models” by A. Jain et al., arxiv.org:2211.11319v1 [cs.CV] [3] Y. Zhu in “A Survey of Deep Graph Generation: Methods and Applications”, arXiv:2203.06714v3 [4] M. Li et a.: “TrOCR: Transformer-based Optical Character Recognition with Pre-Trained Models”, https: / / arxiv.org / pdf / 2109.1022.pdf). [5] Poggenhans, Fabian & Pauls, Jan-Hendrik & Janosovits, Johannes & Orf, Stefan & Naumann, Maximilian & Kuhnt, Florian & Mayr, Matthias. (2018). Lanelet2: A high-definition map framework for the future of automated driving. 1672-1679. 10.1109 / ITSC.2018.8569929.
Claims
1. A computer-implemented method (100) for generating a road map particularly suitable for autonomous driving (AD) of a vehicle, the method comprising the following method steps: - receiving (S104) image input data, wherein the image input data includes acquired image data representing at least one area in which the vehicle is drivable; - Generating (S106) a road map based on the received (S104) image input data with the aid of a trained visual basis model for road map generation, in particular for AD, and wherein the generated (S106) road map comprises a road layout within said at least one area in which the vehicle is drivable and locally assigned contextual information taking into account applicable traffic regulations.
2. The method (100) according to claim 1, further comprising the steps of: - providing ( S108 ) the generated ( S106 ) road map to a controller of a vehicle configured for autonomous driving AD.
3. Method (100) according to any of the preceding claims, wherein the road layout of the generated (S106) road map is adjustable in terms of granularity, in particular for presentation and / or depending on locally assigned context information.
4. A method (100) according to any of the preceding claims, wherein the image input data comprises overhead view image data of an area, the area including the at least one area in which the vehicle is drivable, and / or wherein the image input data is acquired by means of an optical sensor system, in particular a satellite imaging system and / or preferably an airborne camera system, in particular via one or more drones.
5. The method (100) according to any of the preceding claims, wherein the generation (S106) of the road map, in particular the road layout, comprises generating a graph, in particular a deep and / or nested graph, comprising nodes and edges, wherein the edges connect the nodes, in particular subsets of the nodes, and wherein the nodes and edges can be supplemented by attributes representing locally assigned context information; optionally The graphics, in particular deep and / or nested graphics, comprise in particular scalable vector graphics SVG.
6. The method (100) according to any one of the preceding claims, wherein the generating (S106) of the road map comprises generating a representation of the road layout using primitives.
7. A method (100) according to any of the preceding claims, wherein the trained visual basis model comprises a graphics generation model, optionally wherein the graphics generation model comprises at least one of an autoregressive model, a variational autoencoder, a regularized flow, a generative adversarial network and / or a diffusion model.
8. The method (100) according to any of the preceding claims, wherein the trained visual basis model is configured to perform classification and / or semantic segmentation of image input data.
9. A computer-implemented method (200) for training a visual base model for generating a road map based on received image input data, the method comprising the steps of: - receiving (S202) a training data set, wherein the training data set includes at least one of the following: o annotated top-view images, in particular acquired with the aid of an optical sensor system, in particular by a satellite imaging system and / or by one or more drones, wherein the annotations include a road layout with locally assigned contextual information; o based on a rendered image, in particular an artificially generated graphic, representing a road layout; and / or o a top-view image and a rendered image based on a, in particular artificially generated, graphic of a road map representing said top-view image; as well as - training a visual base model for generating a road map based on the received image input data (S204), wherein the generated road map includes road layout and locally assigned context information taking into account applicable traffic regulations; Optionally, the training ( S204 ) is self-supervised and / or based on a reconstruction loss between ground truth values included in the received ( S202 ) training dataset and a road map generated by a visual base model.
10. Use of the road map generated by the method (100) according to any one of claims 1 to 8 for trajectory prediction, path planning, collision avoidance and / or behavior prediction, in particular trajectory prediction, path planning, collision avoidance and / or behavior prediction of traffic, AD vehicles and / or robotic systems.
11. A computing device (300) for generating a road map particularly suitable for autonomous driving (AD) of a vehicle, the computing device (300) comprising: - an input image data receiving interface (304), configured to receive image input data, wherein the image input data includes acquired image data representing at least one area in which the vehicle is drivable; - A road map generation module (306) configured to generate a road map based on the received image input data, wherein the generation of the road map is performed by a trained visual basis model for road map generation, and wherein the generated road map includes a road layout within the at least one area in which the vehicle is drivable and locally assigned contextual information taking into account applicable traffic regulations.
12. The computing device (300) of the immediately preceding claim, further configured to perform any of the steps of any of the methods (100) of claims 2 to 8 and / or include any of the features of any of the methods (100) of claims 2 to 8.
13. A computing device (400) for training a visual base model for generating a road map based on received image input data, the computing device (400) comprising: A training data receiving interface (402) configured to receive a training data set, wherein the training data set includes at least one of the following: o annotated top-view images, in particular acquired with the aid of an optical sensor system, in particular a satellite imaging system, and / or by one or more drones, wherein the annotations include a road layout with locally assigned contextual information; o based on a rendered image, in particular an artificially generated graphic, representing a road layout; and / or o a top-view image and a rendered image based on a, in particular artificially generated, graphic of a road map representing said top-view image; as well as - a visual base model training module (404) configured to train a visual base model for generating a road map based on received image input data, wherein the generated road map includes a road layout and locally assigned context information taking into account applicable traffic regulations; Optionally, the training is self-supervised and / or based on a reconstruction loss between ground truth values included in a received training dataset and a road map generated by a visual base model.
14. A system for generating a road map particularly suitable for autonomous driving (AD) of a vehicle, the system comprising: - A computing device (300) according to claim 11 or 12; as well as - at least one sensor system and / or image capture device configured to provide image input data to an input image data receiving interface (304) of said computing device (300).
15. A controller (804) for an autonomous driving (AD) vehicle, the controller comprising: - a receiving interface for receiving a generated road map, wherein the road map is generated by the method according to any one of the preceding claims 1 to 8 based on the received image input data, wherein the road map includes a road layout in at least one area in which the vehicle is drivable and locally assigned context information taking into account applicable traffic regulations; as well as - A processing unit configured for trajectory planning of the AD vehicle.