Technology for generating road maps for automated driving
A visual foundation model generates detailed road maps with traffic regulations for automated driving, addressing the challenge of integrating contextual information into road maps for enhanced vehicle navigation and safety.
Patent Information
- Application Number
- JP2025062403
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-04-05
- Filing Date
- 2025-04-04
- Publication Date
- 2025-10-17
AI Technical Summary
Existing models struggle to leverage underlying data for automated driving applications, particularly in generating accurate road maps with contextual information such as traffic regulations and road layouts.
A visual foundation model trained on image data generates road maps with high-definition road layouts and locally assigned contextual information, including traffic regulations, using a deep neural network and self-supervised learning, enabling accurate and efficient map generation.
The solution provides highly accurate road maps suitable for automated driving, enhancing vehicle performance and safety by integrating detailed road layouts and traffic regulations, facilitating safe navigation and route planning.
Smart Images

Figure 2025158963000001_ABST
Abstract
Description
[Technical Field]
[0001] The present application relates to techniques for generating road maps, particularly suitable for use in automated driving (AD) of vehicles. Provided are a method and computing device for generating road maps, a method and computing device for training a visual foundation model for generating road maps, a system for generating road maps, a controller for an AD vehicle, a computer program product, and a computer-readable storage medium. [Background technology]
[0002] Prior art On the one hand, traditional underlying models are used for many different modalities, such as images, text, or audio, as presented, for example, in R. Girdhar et al. in “Image Bind: One Embedding Space To Bind Them all,” arxiv.org:2305.05665v2 [cs.CV], which is incorporated herein by reference. On the other hand, maps contain not only rich information, such as road layouts, but also contextual information, such as speed limits. Vector graphics are one possible approach to represent road layouts. For example, A. Jain et al., “Vector Fusion: Text-to-SVG by Abstracting Pixel-Based Diffusion Models,” arxiv.org:2211.11319v1 [cs.CV], which is incorporated herein by reference, proposes fine-tuning a text-to-image synthesis circuit to generate vector graphics of buildings, objects, animals, and scenarios. [Prior art documents] [Non-patent literature]
[0003] [Non-Patent Document 1] R. Girdhar et al in “Image Bind: One Embedding Space To Bind Them all”, arxiv.org:2305.05665v2 [cs.CV] [Non-patent document 2] A. Jain et al., “Vector Fusion: Text-to-SVG by Abstracting Pixel-Based Diffusion Models” arxiv.org:2211.11319v1[cs.CV] Summary of the Invention [Problem to be solved by the invention]
[0004] As recent applications (LLAMA, ChatGPT, stable diffusion) have shown, the underlying model (trained on a seemingly infinite amount of data) can reach human-like capabilities in synthesizing and / or generating real-world data. However, how to leverage the underlying model for automated driving (especially autonomous driving) application cases remains an open question. [Means for solving the problem]
[0005] Disclosure of the Invention In the following, the solutions presented herein are described in terms of claimed computer-implemented methods, in particular methods for generating road maps, and claimed computing devices. Features, advantages, or alternative embodiments described in terms of methods may correspond to objects (e.g., devices, systems, computer programs, or computer program products) described in other claims, and vice versa. In other words, claims of computing devices and / or systems including computing devices may be improved by features described or claimed in the context of the respective corresponding methods. In this case, functional features of the methods may be embodied by structural units of the computing device (and / or system), and vice versa.
[0006] Regarding a first method aspect, a computer-implemented method for generating a road map, particularly suitable for use in automated driving (AD) of a vehicle or robot, is provided. The method includes receiving image input data. The image input data includes acquired image data representing at least one area navigable by the vehicle. The method further includes generating a road map based on the received image input data. The generation of the road map is performed using a visual basis model trained for generating road maps particularly suitable for use in AD. The generated road map includes a road layout within the at least one area navigable by the vehicle, together with locally assigned contextual information taking into account any applicable traffic regulations.
[0007] A road map (also referred to as a map for short) generated from a visual basis model (also referred to as a visual basis model) with locally assigned context information can provide an improved, highly accurate road layout (e.g., with a resolution of a few centimeters, e.g., down to 3 cm, and / or positioning of landmarks such as roadside strips) along with applicable traffic regulations (e.g., speed limits, right-of-way, and / or lane change allowability) suitable for use in fully or partially automated driving (AD) vehicles. Alternatively or additionally, the generated road map can be used for knowledge transfer to downstream or subsequent control systems. In particular, the performance of an AD vehicle itself can be improved by using the generated road map for training and testing the AD system of the AD vehicle. Further alternatively or additionally, for example, if multiple lanes in the road map are used as ground truth results for counting lanes in image input data, the (e.g., separate) visual basis model can be improved in its counting capabilities.
[0008] The road map may particularly include a high definition (HD) map, which may be particularly suitable for automated driving.
[0009] An AD vehicle may be a car, bus, or truck with automated driving (AD) capabilities. The AD driving capabilities may have full automation (specifically designated as L5), at least partial automation (e.g., high L4, conditional L3, or partial L2), and / or supplemental driving automation (e.g., L1, including, for example, adaptive cruise control (ACC)). Fully automated vehicles are also sometimes referred to as autonomous and / or automated vehicles.
[0010] The vehicle may be (or may include) a motor vehicle (e.g., subject to registration), an automobile, and / or a utility vehicle, among others. The vehicle may be a road vehicle. Alternatively or additionally, the vehicle may be included in a robotic system (e.g., a robot in a manufacturing environment).
[0011] Applicable traffic regulations may include various regulations for different types of vehicles. Alternatively or additionally, applicable traffic regulations may stipulate (and / or restrict) the manner in which vehicles may travel or use an area (or route) (e.g., speed limits, weight limits, width and / or height limits, defining parking requirements for parking lots). For example, a road or area may be restricted so that only vehicles below a certain weight threshold (e.g., for crossing bridges) and / or only vehicles below a certain height threshold (e.g., for underbridge travel). Alternatively or additionally, applicable traffic regulations may include a certain width threshold above which vehicles cannot use the road (e.g., due to aisle width, lane width, and / or off-lane and / or off-road guardrails).
[0012] Applicable traffic regulations may also be expressed as permissible route trajectories (also described as routes and / or trajectories) and / or non-permissible route trajectories (e.g., pedestrian zones) for vehicles. Further alternatively or additionally, applicable traffic regulations may be expressed in (and / or by) road markings (e.g., including solid and / or dashed lines), traffic signs, and / or traffic signals.
[0013] The acquired image data can be acquired by a sensor system, in particular an image acquisition device. The sensor system and / or the image acquisition device can be based on at least one of the following technologies: video, radar, LiDAR, ultrasound, motion, thermal imaging. Alternatively or additionally, the acquired image data can be acquired by means of a satellite (e.g., in Earth orbit) and / or by means of a drone. For example, satellite data can advantageously provide an aerial view of an extended area (in particular including at least one area drivable by a vehicle).
[0014] The image input data may at least comprise captured image data and / or may in particular comprise digital image data. In particular, the image input data may comprise an aerial image of at least one area. Alternatively or additionally, the image input data may for example comprise a map in graphical form, such as a (in particular open) street map and / or a (in particular open) navigation map.
[0015] A vehicular drivable area (also abbreviated as drivable area) may also describe traversable, passable, multidirectionally drivable, traversable, and / or accessible by vehicles. Drivable areas (e.g., including roads) may specifically include areas open to traffic (e.g., streets), but are not limited to areas legally open to the public (e.g., streets) and / or may include areas specifically designed for the operation of vehicles. For example, a private road may be drivable (e.g., in a technical sense) but is not designated as "open to traffic."
[0016] Areas may be defined by location data and / or geographic data and / or may include, for example, roads (also called streets) and / or lanes of streets.
[0017] The generated road map may include (e.g., simple) geometric representations of the road layout, such as lines, curves, and / or intersections. Alternatively or additionally, the generated road map may be represented in or may include Scalable Vector Graphics (SVG). SVG, any graph, and / or any representation of a road layout may include nodes and edges between nodes (e.g., between subsets of nodes). Nodes may represent (e.g., equidistant) points of (particularly visible) structures, and edges may represent relationships between nodes. For example, adjacent nodes of a line between lanes are connected by edges.
[0018] The generated road map may include and / or indicate a road layout. The generated road map may include context information. The context information may be used to automatically analyze traffic restrictions that may apply. The context information may be assigned locally to a specific location within the road map. For example, a speed limit or one-way restriction for a particular section of a street may be represented in corresponding context information assigned to the portion of the road map where that street section is represented. The context information may be provided in the generated road map as annotations and / or as overlay graphics and / or as expandable boxes and / or as thumbnail images and / or as text embeddings.
[0019] In general, a road layout may include any (especially road-related) structures (and / or representations and / or information thereof) of at least one area navigable by vehicles, especially as visible in (e.g., photorealistic) overhead images. A road layout may, for example, indicate roads and / or streets navigable by vehicles, and / or pedestrian or motorcycle paths or crosswalks. A road layout may include information necessary for analyzing a traffic scene. This may refer to static data (especially data that does not change over time). Alternatively or additionally, a road layout may be represented by SVG or any other graphical representation (and / or graph). For example, a lane may have an attribute that allows left turns. Alternatively or additionally, an arrow indicating the permissibility of a left turn may be displayed graphically.
[0020] For example, the road layout may include lane shapes and / or number of lanes, lane types (e.g., lanes used by motorized vehicles such as cars and / or lanes used by bicycles), lane widths, adjacent lane separation types (e.g., solid and / or dashed lines), arrows indicating the direction of travel for lanes, stop lines, zebra crossing markings, parking spots (e.g., their boundaries), and / or any (e.g., painted) features visible in the overhead image. For example, multiple lanes, the width of each lane, and / or the line type between adjacent lanes (particularly solid or dashed lines) may be visible from above. Alternatively or additionally, lane types for motorized vehicles (e.g., cars) may be distinguishable from lane types for bicycles and / or sidewalks by lane width and / or lane color (e.g., tarmac and / or pavement condition).
[0021] By specifying a road layout that includes any type of information that is visible, especially from above (and / or in an overhead image), the generated road map includes visual cues that are necessary for safe driving (e.g., route) planning, especially for ADs, and is therefore applicable to automatic planning and subsequent control of vehicles.
[0022] Alternatively or additionally, the road layout may include a set of predefined primitives, including, for example, simple geometric shapes, in particular line elements. The line elements may be solid or dashed lines separating adjacent lanes, including those that do not allow vehicles to change lanes or those that allow vehicles to change lanes, respectively. Alternatively or additionally, the line elements may include straight lines (e.g., of a predetermined length) and / or line elements defined by simple functions (e.g., including circular arcs of a predetermined radius).
[0023] The context information included in the generated road map may include applicable traffic regulations (also referred to as traffic regulations and / or traffic rules), which may be country-dependent and / or region-dependent. Alternatively or additionally, the context information may include one or more road conditions.
[0024] The local allocation of context information may relate to one or more nodes and / or edges of a graph (e.g., SVG), e.g., a simple geometric representation, and / or one or more primitives (especially their interrelationships) representing the road layout of at least one area drivable by a vehicle, e.g., within city boundaries, outside city boundaries, motorways, junctions, crosswalks, bridges, roads (especially segments and / or parts thereof) with significant slopes (e.g., 4% or more).
[0025] For example, locally assigned context information taking into account applicable traffic regulations may include lane weight constraints, lane height constraints, road structure (including, inter alia, road surface and / or road conditions, e.g., paved, unpaved, with bumps and / or depressions), speed limits, right-of-way, no-entry restrictions, traffic signs, traffic lights and / or accuracy indicators.
[0026] Locally assigned context information may be specific to a vehicle type. For example, road bans and / or (e.g., low) speed limits may only apply to vehicles (e.g., buses and / or trucks) above a certain weight and / or length. Alternatively or additionally, locally assigned context information may also be time-specific. Time-specificity may be related to the time of day and / or day. For example, lower speed limits may be applied at night to reduce noise, and / or access to certain lanes may be reserved for certain groups of vehicles (e.g., buses and / or taxis) on weekdays, especially during rush hour.
[0027] The accuracy indicator may include, for example, the predicted accuracy of the location of boundary lane lines and / or other landmarks in the road layout. Alternatively or additionally, the accuracy indicator may also include a confidence level of the correct assignment of traffic regulations that may be applied, such as speed limits (e.g., derived from the image input data after semantic segmentation).
[0028] The specification of locally assigned context information combined with road layout (and / or visual cues) further facilitates safe driving, especially AD (e.g., route) planning.
[0029] The visual foundation model may be or include a generative artificial intelligence system (AI), i.e., one that generates or shapes content such as images. The visual foundation model may be based on a deep neural network. The visual foundation model may include an autoregressive foundation model that generates inputs piecewise and a denoising foundation model that corrupts and then restores the input. Alternatively or additionally, the visual foundation model may include a graph generation model for (deep and / or nested) graph generation (e.g., an autoregressive model, a variational autoencoder, a regularized flow, a generative adversarial network, a GAN, and / or a diffusion model), such as those described by Y. Zhu in “A Survey of Deep Graph Generation: Methods and Applications,” arXiv:2203.06714v3, incorporated herein by reference.
[0030] The visual foundation model can be (pre-)trained based on a wide range of data (huge datasets) (training phase) and can be adapted or fine-tuned for a variety of specific downstream tasks (adaptation phase). The visual foundation model can be trained by self-supervised learning, transfer and / or active learning.
[0031] The training dataset may include, for example, images of an area as image input data acquired by an optical sensor system, e.g., a satellite imagery / imagery system, and annotations (particularly as ground truth) of a road map including (e.g., locally associated) road layout and locally assigned context information. The annotations may be embedded in the image or locally associated with the image. Alternatively or additionally, the training dataset may include (e.g., artificially generated) vector graphics (e.g., scalable vector graphics (SVG)) representing a road layout (and / or resulting images of the road layout) with attributes representing the locally assigned context information (particularly as ground truth). The image input data associated with the vector graphics may be artificially generated (e.g., by a GAN). Further alternatively or additionally, the training dataset may include images (e.g. acquired by an optical sensor system, e.g. satellite imagery / imagery system) as image input data and (e.g. artificially generated) vector graphics (e.g. SVG) representing the road layout of the images (and / or the resulting images of the road layout) with attributes representing locally assigned contextual information (in particular as ground truth).
[0032] The method according to the first aspect may further comprise providing the generated road map to a controller or robot of a vehicle configured for automated driving (e.g. autonomous driving) AD.
[0033] Providing the generated road map can enable AD of the vehicle in combination with sensors that detect other road users and / or (especially moving) obstacles at the time (e.g. parked vehicles, e.g. cars and / or bicycles).
[0034] The generated road map or parts thereof can also be used for vehicle-to-vehicle communication to provide information of other traffic participants in the traffic scene.
[0035] The road layout of the generated road map may be granularly adjustable, particularly for rendering purposes and / or depending on locally assigned context information.
[0036] The (e.g., rendered) road layout may be adjustable in granularity by user input (particularly by the user and / or driver of the AD vehicle). The user input may be received, for example, via a user interface (UI) including a rotary button and / or a graphic user interface (GUI). The UI may be deployed in the AD vehicle, for example, in a central console and / or particularly within reach of the user and / or driver. Alternatively or additionally, the (e.g., rendered) road layout may be adjustable (particularly without user input) depending on the (e.g., planned) speed of the AD vehicle.
[0037] The granularity of the road layout can be adjusted to achieve the required level of detail. For example, a coarse granularity may be appropriate for high-speed driving on straight motorway tracks. A fine granularity may be suitable for navigating difficult traffic situations with multiple lanes and / or intersections for various types of vehicles, for example within city boundaries.
[0038] The image input data may include aerial image data of an area including at least one area navigable by the vehicle. Alternatively or additionally, the image input data may be acquired by an optical sensor system including, for example, a satellite imaging system, a (particularly aerial) camera system, and / or an aerial photography system, for example via one or more drones.
[0039] Advantageously, overhead image data, satellite imagery and / or aerial imagery can provide reliable data on lane structures, zebra markings of cross streets and / or pedestrian crossings, for example in terms of width and / or length.
[0040] Generating a road map may include generating a (e.g., deep and / or nested) graph comprising nodes and edges. The edges may connect, link, and / or associate nodes (particularly a subset thereof) representing the road layout. The nodes and edges may be complemented by attributes representing locally assigned contextual information. Optionally, the (e.g., depth) graph comprises vector graphics, particularly Scalable Vector Graphics (SVG).
[0041] A node may represent, for example, a point on a line. An edge may be associated with a node included in a line. For example, an edge may connect a point in the center of a line to its two closest neighbors, each along a different direction of the line (e.g., one node to the left and one node to the right of the center node). Alternatively or additionally, an edge between nodes associated with two lines that are lane boundaries may represent a stop line. Alternatively or additionally, for higher levels in a (e.g., deep and / or nested) graph, a node may represent, for example, a line, and an edge may represent, for example, the role of the edge in describing a portion of a lane (e.g., a left boundary, a right boundary, or a center line). Further alternatively or additionally, for even higher levels in a (e.g., deep and / or nested) graph, a node may represent, for example, a lane portion, and an edge may connect, for example, the lane portions to indicate whether crossing is permitted and traffic rules are observed. Further alternatively or additionally, the lines may form lanes and / or the lanes may form a road network (eg, depending on the transportation operator).
[0042] In particular, SVG can be generated by a (e.g., general) deep (and / or nested) graph generator, which can, in particular, generate attributes of nodes and / or edges.
[0043] The (e.g., generated) graph may include nested structures. Alternatively or additionally, the graph can be encoded as or transferred to text data to generate a road map.
[0044] A graph representation with attributes can provide a particularly simple and efficient and / or fast memory for rendering a representation of a road map.
[0045] Generating a road map may include generating a representation of a road layout using primitives, which may include lines and / or line elements (e.g., solid and / or dashed) parameterized by functions representing simple geometric shapes, in particular simple functions, in particular linear functions and / or circular arcs.
[0046] Primitives can be combined with representations in terms of nodes and edges, for example a line can be represented by multiple nodes related by edges (which for example follow a simple function).
[0047] The representation in terms of primitives can provide a simple, memory-efficient and / or fast road map for loading into a rendering circuit's format.
[0048] The trained visual foundation model may include a graph generative model. Optionally, the graph generative model may include an autoregressive model, a variational autoencoder, a regularized flow, a generative adversarial network, and / or a diffusion model.
[0049] The graph generation model can be configured to generate a (e.g., deep and / or nested) graph having nodes, edges associated with the nodes (e.g., contained in an SVG), and complementary attributes for at least a subset of the nodes and edges.
[0050] The trained visual foundation model (particularly a graph generation model) can be used to generate attributes for nodes and / or edges based on the image input data. For example, speed limits can be derived from environmental information (e.g., country-dependent, lane curvature and / or radius-dependent, location within or outside city boundaries, and / or open-air roadway bridges spanning valleys). Alternatively or additionally, based on the image input data, the trained visual foundation model (particularly a graph generation model) can classify roads as congested or sparsely trafficked (e.g., based on a plurality of vehicles in the image input data and / or based on the vehicular environment, including indicators of densely or dispersedly formed areas).
[0051] The trained visual basis model can be configured to perform classification and / or semantic segmentation of the image input data. Alternatively, the image input data can be pre-processed, for example, using a segmentation algorithm to segment relevant structures within the image input data.
[0052] Classification and / or semantic segmentation includes in particular the detection of objects and / or structures, in particular traffic lights, traffic signs and / or road structures (including for example road surface conditions).
[0053] According to a second method aspect, there is provided a method for training a visual basis model for generating a road map based on received image input data. The method comprises receiving a training dataset. The training dataset may include annotated overhead images, particularly acquired by an optical sensor system, in particular a satellite imaging system. The annotations may include a road layout together with locally assigned context information. Alternatively or additionally, the training dataset may include, in particular, artificially generated graphics representing a road map or road layout, such as rendered images based on SVG (and / or resulting images of the road layout) or in a format of such graphics. The training dataset may in particular include, as (at least partial) ground truth, the artificially generated graphics (and / or resulting road layout images) together with the artificially generated overhead images (in particular based on the artificially generated graphics). Further alternatively or additionally, the training dataset may comprise images (in particular aerial images of traffic scenes acquired, for example, by an optical sensor system) and in particular artificially generated graphics representing road maps or road layouts, in particular in (aerial) images, for example rendered images based on SVG or in a format of such graphics.
[0054] The method further comprises training a visual basis model for generating a road map based on the received image input data, the generated road map including a road layout together with locally assigned context information taking into account any traffic restrictions that may apply. Optionally, the training is self-supervised and / or based on a reconstruction loss between ground truth contained in the received training dataset (e.g., its annotations or its rendered images) and the road map generated by the visual basis model.
[0055] The training dataset is
number
[0056] The computing device and / or model architecture may include, for example, a conventional encoder-decoder architecture, where the encoder learns a mapping from an image to a latent space and the decoder learns a mapping from the latent space to a graphical (e.g., SVG) representation.
[0057] For the encoder, for example, a pre-trained image transformer can be used.
[0058] The graphical representation, e.g., SVG, can correspond to (and / or include) a language for describing two-dimensional graphics, in particular XML. For example, a deep (and / or nested) graph generator, as described by Y. Zhu in "A Survey of Deep Graph Generation: Methods and Applications", arXiv:2203.06714v3, can be used. Alternatively or additionally, since the graphical representation (e.g., SVG) is text-based, the decoder can be implemented as a language decoder, e.g., a pre-trained text transformer.
[0059] As a training loss, one can use, for example, cross-entropy, which is routinely used for language modeling tasks.
[0060] The weights of the visual basis model can be initialized from a pre-trained default image-to-text model (e.g., TrOCR, described in https: / / arxiv.org / pdf / 2109.1022.pdf, incorporated herein by reference).
[0061] In a first extension, once a (e.g., visually based) model is trained, it can be used as the basis for an expert model with only a small training dataset D' available. For example, a standard image-to-text model is first used, then fine-tuned on image-to-open street map data, and then fine-tuned on HD maps.
[0062] In a second extension, if text information is also given as input in addition to images, a second encoder, specifically a text encoder, can be added to the (eg visual basis) model.
[0063] The training phase may include (and / or may be followed by) testing, validation and / or validation of the (e.g., initially) trained visual basis model.
[0064] Training the visual foundation model may involve utilizing a generative understanding of traffic regulations and / or common road layouts.
[0065] Training a visual basis model provides a powerful tool for generating road maps based on a small number of image inputs (e.g., a single satellite image) and / or without human interaction.
[0066] In another aspect, there is provided a use of a road map generated by a method according to the first method aspect for trajectory prediction, route planning, collision avoidance and / or behavior prediction of or for AD vehicles, in particular traffic.
[0067] The generated road map can be stored and / or used locally on a processing unit of the AD vehicle. Alternatively or additionally, the generated road map can be accessible via cloud technology (e.g., using wireless communication technology) of the AD vehicle. The generated road map can be transferred to other vehicles, units external to the vehicle, such as other vehicles, servers, and / or mobile devices, using network technology.
[0068] Behavior prediction can relate to any traffic outside the AD vehicle (e.g., including other vehicles and / or pedestrians). Behavior prediction can be used, among other things, for collision avoidance. For example, by predicting the behavior of other traffic participants, the AD vehicle can be configured to drive predictively. For example, it can reduce its speed and / or change lanes to avoid a collision with another traffic participant.
[0069] The generated road map can be used to improve AD safety, or alternatively or additionally, to optimize the routes taken by AD vehicles.
[0070] Regarding a first device aspect, a computing device for generating a road map particularly suitable for use in AD for a vehicle is provided. The computing device comprises an image input data receiving interface configured to receive image input data. The image input data includes captured image data representing at least one area navigable by the vehicle. The computing device further comprises a road map generation module configured to generate a road map based on the received image input data. The generation of the road map is performed by a visual foundation model trained specifically for generating road maps suitable for use in AD. The generated road map includes a road layout within the at least one area navigable by the vehicle, together with locally assigned context information taking into account any applicable traffic regulations.
[0071] The computing device may further be configured to perform any one of the steps of the first method aspect and / or include any one of the features disclosed within the context of the first method aspect.
[0072] Regarding a second device aspect, a computing device for training a visual basis model for generating a road map based on received image input data is provided. The computing device includes a training data receiving interface configured to receive a training dataset. The training dataset may include an annotated aerial imagery, particularly acquired using an optical sensor system, particularly a satellite imaging system. The annotations may include a road layout together with locally assigned context information. Alternatively or additionally, the training dataset may include rendered images, particularly based on artificially generated graphics representing a road layout (and / or a resulting image of the road layout) or in a format of such graphics. The training dataset may particularly include artificially generated graphics (and / or resulting road layout images) as ground truth, together with artificially generated aerial images (particularly based on artificially generated graphics). Further alternatively or additionally, the training dataset may include (particularly aerial) images and rendered images, particularly in a format of aerial images, representing the road layout of the (particularly artificially generated) images.
[0073] The computing device further includes a visual basis model training module configured to train a visual basis model for generating a road map based on the received image input data, the generated road map including a road layout together with locally assigned context information taking into account traffic regulations that may apply. Optionally, the training is self-supervised and / or based on a reconstruction loss between a ground truth included in a received training dataset (particularly in annotations or rendered images) and the road map generated by the visual basis model.
[0074] The computing device of the first device embodiment and the computing device of the second device embodiment may be the same.
[0075] With respect to a system aspect, there is provided a system for generating road maps particularly suitable for use in vehicle AD, the system comprising a computing device according to the first method aspect and at least one sensor and / or image capture device configured to provide image input data to an image input data receiving interface of the computing device.
[0076] According to a third apparatus aspect, a controller for an AD vehicle is provided. The controller comprises a receiving interface configured to receive a generated road map based on received image input data. The road map includes a road layout in at least one area navigable by the vehicle along with locally assigned context information taking into account any applicable traffic regulations. The controller further comprises a processing unit configured for trajectory planning of the AD vehicle.
[0077] In another aspect, there is provided a computer program product including program elements that, when loaded into a memory of a computing device, cause the computing device to perform the steps of the first method aspect and / or the steps of the second method aspect.
[0078] In yet another aspect, there is provided a computer-readable medium having stored thereon program elements that, when executed by a computing device, are readable and executable by the computing device to perform the method steps according to the first method aspect and / or to perform the method steps according to the second method aspect.
[0079] Exemplary embodiments of methods, apparatus, systems and controllers according to the present disclosure are illustrated in the drawings. In particular, the drawings show: [Brief explanation of the drawings]
[0080] [Figure 1] 1 is an exemplary flowchart illustrating a computer-implemented method for generating a road map according to the present disclosure. [Figure 2] 1 is an exemplary flowchart illustrating a computer-implemented method for training a visual foundation model for generating a road map according to the present disclosure. [Figure 3] FIG. 1 shows a schematic representation of the architecture of a computing device for generating road maps. [Figure 4] FIG. 1 is a diagram illustrating the architecture of a computing device for training a visual foundation model. [Figure 5] FIG. 1 illustrates an exemplary graphical representation of a city area including an area drivable by a vehicle. [Figure 6] FIG. 6 is an exemplary illustration of the city area of FIG. 5 as a vector graphic. [Figure 7] 3 is a diagram illustrating an exemplary architecture of a computing device, in particular a road map generation module 306. FIG. [Figure 8] FIG. 10 is an exemplary diagram illustrating the use of the method for autonomous driving. DETAILED DESCRIPTION OF THE INVENTION
[0081] Detailed Description FIG. 1 shows an exemplary flowchart of a computer-implemented method 100 for generating road maps particularly suitable for use in automated driving (AD) of vehicles.
[0082] The method 100 includes a step S104 of receiving image input data, which includes captured image data representative of at least one area navigable by a vehicle.
[0083] Method 100 further includes a step S106 of generating a road map based on the image input data received in S104. The step S106 of generating the road map is performed by a visual basis model trained specifically for generating road maps suitable for use in AD. The road map generated in S106 includes a road layout in at least one area navigable by the vehicle, along with locally assigned context information taking into account any applicable traffic regulations.
[0084] Optionally, the method 100 includes a step S108 of providing the road map generated in S106 to a controller of a vehicle configured for AD.
[0085] FIG. 2 illustrates an exemplary flowchart of a computer-implemented method 200 for training a visual basis model for generating a road map based on received image input data.
[0086] The method 200 includes a step S202 of receiving a training dataset. The training dataset may include, in particular, annotated aerial images acquired using a satellite imaging system and / or one or more drones. The annotations may include a road layout together with locally assigned context information. Alternatively or additionally, the training dataset may include, in particular, rendered images based on graphics (e.g., SVG) or in a format representing a road layout or road map (and / or an image resulting from the road layout), in particular artificially generated. The training dataset may include, in particular, artificially generated graphics (and / or resulting road layout images) as (at least partial) ground truth together with artificially generated aerial images (in particular based on artificially generated graphics). Further alternatively or additionally, the training dataset may include aerial images and rendered images, in particular artificially generated graphics (e.g., SVG) format representing a road layout or road map of the image.
[0087] Method 200 further includes a step S204 of training a visual basis model for generating a road map based on the received image input data. The generated road map includes a road layout together with locally assigned context information taking into account possible traffic regulations. Optionally, training step S204 is self-supervised and / or based on a reconstruction loss between ground truth included in the training dataset received in S202 (in particular elements of the training dataset representing a road layout or road map) and the road map generated by the visual basis model, where generation of the road map is based on images included in the training dataset received in S202.
[0088] FIG. 3 shows a schematic architecture of a computing device 300 for generating road maps particularly suitable for use in vehicular AD.
[0089] The computing device 300 includes an image input data receiving interface 304 configured to receive image input data. The image input data includes captured image data representative of at least one area navigable by a vehicle.
[0090] The computing device 300 further includes a road map generation module 306 configured to generate a road map based on the received image input data. The generation of the road map is performed by a visual foundation model trained for road map generation particularly suitable for use in AD. The generated road map includes a road layout in at least one area navigable by the vehicle, together with locally assigned context information taking into account any applicable traffic regulations.
[0091] Optionally, the computing device 300 includes an output interface 308 configured to provide the generated road map to a controller of a vehicle configured for AD.
[0092] The image input data receiving interface 304 and optionally the output interface 308 may be embodied by an input / output interface 310. The road map generating module 306 may be embodied by a processing unit 312. The computing device 300 may further comprise at least one memory 314.
[0093] FIG. 4 shows a schematic architecture of a computing device 400 for training a visual foundation model for generating a road map based on received image input data.
[0094] The computing device 400 includes a training data receiving interface 402 configured to receive a training dataset. The training dataset may include annotated aerial images acquired using an optical sensor system, in particular a satellite imaging system and / or one or more drones. The annotations may include road layouts along with locally assigned context information. Alternatively or additionally, the training dataset may include rendered images, in particular artificially generated graphics (e.g., SVG) including road layouts (and / or resulting images of the road layouts). The training dataset may include artificially generated graphics (and / or resulting road layout images) as (at least partially) ground truth, along with artificially generated aerial images (in particular based on the artificially generated graphics). Further alternatively or additionally, the training dataset may include (in particular aerial) images and rendered images, in particular artificially generated graphics (e.g., SVG) including road maps for the (in particular aerial) images.
[0095] The computing device 400 further includes a visual basis model training module 404 configured to train a visual basis model for generating a road map based on the received image input data. The generated road map includes a road layout along with locally assigned context information taking into account any applicable traffic regulations.
[0096] Optionally, the training is self-supervised and / or based on a reconstruction loss between ground truth included in the received training dataset and a road map generated by a visual basis model, where the generation of the road map is based on images included in the received training dataset.
[0097] The training data receiving interface 402 may be embodied by an input / output interface 406. The visual basis model training module 404 may be embodied by a processing unit 408. The computing device 400 may further comprise at least one memory 410.
[0098] Either one of the processing units 312; 408 may be implemented by a central processing unit (CPU) and / or a graphics processing unit (GPU).
[0099] The techniques herein extend "Lanelet2: A high-definition map framework for the future of automated driving" by F. Poggenhans et al. (2018), incorporated herein by reference, to facilitate efficient and accurate generation of road maps. The (e.g., nested) graph structure generated by method 100 can be similar (and / or identical) to the (e.g., nested) graph structure of Lanelet2 by computer-implemented method 100 (particularly, potentially and / or among other options), but the graph is generated (particularly automatically) by a visual foundation model (e.g., learned and / or trained). In contrast, Lanelet2 uses manual and / or semi-automated map "clicking" and / or (e.g., partial) automatic generation of maps from in-vehicle measurements. For example, maps can be generated using lane detection and / or (e.g., GPS) location, but maps are extremely costly in terms of time, effort, computational resources, and memory, especially due to the need to drive to every possible location. On the other hand, the method 100 may use (eg, as image input data) satellite images and / or navigation maps (and / or Secure Digital (SD) maps), which may be used, for example, as OpenStreetMaps.
[0100] The technology herein relates to generative models (particularly visual basis models) that provide road maps, which themselves can be used for training and testing various algorithms for AD, such as behavior prediction and / or planning.
[0101] With the techniques herein, a visual basis model can be trained to generate a road map from image data (also abbreviated as image) that includes a road layout (road structure) represented as a graph and (particularly locally assigned) contextual information such as speed limits, crosswalks, and / or rights-of-way.
[0102] An exemplary graphical representation of a city area with an area (e.g., drivable by AD vehicles) is shown in Figure 5. The exemplary graphical representation includes a road layout together with locally assigned context information. For example, the (e.g., car) lanes 502 of the two-way Karlstrasse and one-way Kaiserstrasse are provided with arrows indicating the direction of travel for each lane. At the intersection of the two streets (Karlstrasse and Kaiserstrasse), a traffic light is installed, indicated by reference numeral 504. Several types of parking areas are shown in the exemplary graphical representation of Figure 5: bicycle parking area 506-1, motorcycle parking area 506-2, and car parking area 506-3. The drivable area in Figure 5 further schematically shows bus lanes 508 (e.g., at a bus terminal).
[0103] Figure 6 shows an exemplary illustration of the cityscape of Figure 5 as a vector graphic, specifically an SVG. Traffic lanes 502 (e.g., for passenger cars) and bus lanes 508 are represented by nodes, represented by crosses, and edges connecting the nodes. Edges are represented by straight lines in the example of Figure 6.
[0104] Locations where stopping (e.g., at stop lines) or yielding (e.g., corresponding to additional nodes relative to the nodes indicated by crosses) are indicated generally by rectangles. Additionally or alternatively, additional nodes (e.g., indicated by rectangles), such as a curved (and / or tortuous) path to car park 506-3, may divide (and / or segment) the corresponding lane into multiple straight line segments.
[0105] 7 illustrates a schematic diagram of an exemplary architecture of the computing device 300, and in particular the road map generation module 306 (and / or the visual foundation model). The exemplary architecture includes an encoder 704 (e.g., a pre-trained image transformer) and a decoder 708. The encoder 704 is configured to receive image input data, as illustrated by the overhead image 702. The encoder 704 is further configured to output a latent representation at 706 (e.g., denoted by Z). The latent representation 706 predicts the correct sequence of shifted tokens and / or graph words (and / or XML) as a function of the input.
[0106] One can use, for example, cross-entropy as a training loss, especially since cross-entropy is routinely used for language modeling tasks.
[0107] A common training technique may include predicting the next token based on the input (e.g., at reference numeral 702) and past tokens. An efficient implementation includes inputting a ground truth sequence shifted by a token (e.g., indicated as "right shift" in FIG. 7) into the decoder 708 and predicting an initial sequence. The sequence may correspond, for example, to a vector graphic (e.g., SVG) and / or any other map format (e.g., intended for prediction and / or generation). At reference numeral 710, an exemplary rendered image is shown, which may be intended for visualization purposes only (e.g., for display when the generated road map is used for AD and / or for manual monitoring during the training phase). The rendered image 710 does not necessarily need to be used for training.
[0108] 8 exemplarily illustrates the use of the method for AD. An AD vehicle is shown schematically at 802. The AD vehicle 802 includes a controller (and / or planning unit) 804 and, optionally, a display 806. The controller (and / or planning unit) 804 receives one or more generated road maps from a computing device 300 configured to generate the road maps. On the optional display 806, a graphical representation of the road map (e.g., shown by reference numeral 710 in FIG. 7) can be provided to the driver of the AD vehicle 802 for monitoring (e.g., for an autonomous vehicle) and / or guidance (e.g., for a vehicle with assisted driving functions).
[0109] Conventional maps, satellite imagery, and OpenStreetMap and other navigation maps (e.g., in various combinations) can provide a large amount (and / or seemingly infinite) of data for training a visual basis model for generating road maps based on received image input data, and more particularly for generating road maps used in AD. Alternatively or additionally, for pre-training, artificially generated graphics (e.g., SVG, particularly as at least part of ground truth) and their corresponding renderings (particularly as image input data) can be utilized as (e.g., truly infinite) training data pairs.
[0110] A computing device 300 (and / or a network including, for example, an encoder and decoder) trained to generate graphs from images (e.g., road graphs from satellite imagery available on a large scale from around the world, and / or equivalently, rendering SVGs capable of self-supervised learning for any number of generated SVGs) can improve performance for a variety of downstream tasks. For example, map data can be added to any dataset where overhead imagery input data (abbreviated as images) is available. Adding map data to overhead imagery expands the amount of training and test data for standard routines in AD scenarios (e.g., trajectory prediction and / or trajectory planning) that require map information.
[0111] This technique significantly extends the methods used in VectorFusion, for example, by training the computing device 300 (and / or network) with a reconstruction loss between a true image and a rendered graph image (e.g., generated by a visual foundation model). In contrast to VectorFusion, the graphs used for training and / or generating the road map are generated using a general deep (and / or nested) graph generator, as described in Y. Zhu et al., in “A Survey of Deep Graph Generation: Methods and Applications,” arXiv:2203.06714v3. An advantage of using a deep graph generator is that additional relationships, such as node attributes (e.g., road conditions, e.g., paved and / or rough, etc.) and / or edges between nodes (e.g., stop signs associated with particular stop lines) can be added to the graph.
[0112] Alternatively or additionally, the trained (especially visual basis) model can also be used in other contexts. According to a first embodiment, a model capable of generating a road map may require a general understanding of traffic regulations and common road layouts. The knowledge of traffic regulations and common road layouts may be transferred and help to further improve the performance of AD (especially autonomous) vehicles. According to a second embodiment, which can be combined with the first embodiment, the generation of a road map may require (especially "good") counting capabilities (e.g., regarding multiple parallel lanes). Conventionally used visual basis models lack these counting capabilities. Training a visual basis model in the road map generation process may help to improve the counting capabilities.
[0113] Techniques for training a visual basis model for generating a road map based on received image input data, and in particular for generating a road map for use in AD, can be used to analyze data obtained from sensors that can determine measurements of the environment in the form of sensor signals, which can be provided by digital images, including, for example, video, radar, LiDAR, ultrasonic, motion, and / or thermal images.
[0114] The technique for training a visual basis model for generating a road map based on received image input data, and downstream uses of the technique for generating a road map, may include virtual sensors, video analysis and / or audio analysis and / or classification, which may include, for example, taking into account locally assigned context information (in particular traffic regulations) associated with the imaged road layout, such that information about elements encoded by the sensor signals can be obtained based on the sensor signals (e.g., indirect measurements can be performed based on sensor signals used as direct measurements).
[0115] Video analysis and / or audio analysis and / or classification can be used to classify the sensor data, to detect the presence of objects in the sensor data, and / or to perform semantic segmentation on the sensor data, e.g., with respect to traffic signs and / or road surfaces.
[0116] The techniques for training a visual basis model for generating a road map based on received image input data and the techniques for generating the road map are particularly suitable for any AD application, in particular for autonomous vehicles. Alternatively or additionally, a manufacturing robot for which a road map is generated according to the techniques can move automatically and / or autonomously within the manufacturing site (e.g., as a drivable area).
[0117] Upstream uses of the technology for training a visual basis model for generating a road map based on received image input data, and in particular for generating road maps used in AD, may include active learning, testing and / or data curation in which a technical system (e.g., including at least one sensor and / or image capture device) actively selects data to send to a backend computer (e.g., to reduce data traffic), which backend computer can further use the information therein for training a machine learning system (e.g., a visual basis model), for testing, validation and / or validation of the machine learning system, and / or for generating road layouts for training and testing AD algorithms.
[0118] Another upstream use of the techniques for training a visual basis model for generating a road map based on received image input data, and in particular for generating road maps used in AD, may include methods and / or training data, for example by generating training data for training and / or by generating test data, validation data and / or validation data for verifying whether a trained ML system (e.g., including the visual basis model and / or AD algorithm) is capable of operating safely.
[0119] Yet another upstream use of the technology for training a visual basis model for generating a road map based on received image input data, and particularly for generating road maps for use in AD, may include a generative model for generating training or test data (e.g., for an AD algorithm). After being trained in this manner, the ML system (e.g., an AD algorithm) can then be transitioned for downstream use.
[0120] Cited prior art documents [1]R. Girdhar et al in “Image Bind: One Embedding Space To Bind Themall”, arxiv.org:2305.05665v2 [cs.CV] [2]A. Jain et al., “Vector Fusion: Text-to-SVG by Abstracting Pixel-Based Diffusion Models” arxiv.org:2211.11319v1[cs.CV] [3]Y. Zhu in “A Survey of Deep Graph Generation: Methods and Applications”, arXiv:2203.06714v3 [4]M. Li et al.: “TrOCR: Transformer-based Optical Character Recognition with Pre-Trained Models”, https: / / arxiv.org / pdf / 2109.1022.pdf). [5]Poggenhans, Fabian & Pauls, Jan-Hendrik & Janosovits, Johannes & Orf,Stefan & Naumann, Maximilian & Kuhnt, Florian & Mayr, Matthias.(2018). Lanelet2: A high-definition map framework for the future of automated driving. 1672-1679.10.1109 / ITSC.2018.8569929.
Claims
1. A computer-implemented method (100) for generating a road map suitable for use in a vehicle, particularly in automated driving (AD), comprising: The method (100) comprises: a method step (S104) of receiving image input data comprising acquired image data representative of at least one area navigable by the vehicle; - a method step (S106) of generating a road map based on the received (S104) image input data using a visual basis model trained for the generation of road maps particularly suitable for use in AD, the road map generated (S106) comprising a road layout in at least one area navigable by the vehicle together with locally assigned context information taking into account any applicable traffic regulations; A method (100) comprising:
2. The method (100) further comprises: - Providing the road map generated (S106) to a controller of a vehicle configured for autonomous driving (AD) (S108). The method (100) of claim 1, comprising:
3. 3. The method (100) according to claim 1 or 2, wherein the road layout of the generated (S106) road map is granular and adjustable, in particular for rendering purposes and / or depending on locally assigned context information.
4. The image input data includes overhead image data of an area including at least one area navigable by a vehicle, and / or the image input data are acquired by an optical sensor system, in particular a satellite imaging system, and / or a camera system, preferably an aerial camera system, in particular an aerial camera system via one or more drones, The method (100) of any one of claims 1 to 3.
5. generating (S106) the road map, in particular the road layout, comprises generating a graph, in particular a deep graph and / or a nested graph, comprising nodes and edges, the edges connecting the nodes, in particular a subset of the nodes, the nodes and the edges being possibly complemented by attributes expressing locally assigned context information; Optionally, the graph, in particular the deep graph and / or the nested graph, comprises a vector graphic, in particular a scalable vector graphic SVG; The method (100) of any one of claims 1 to 4.
6. The method (100) of any one of claims 1 to 5, wherein generating (S106) the road map comprises generating a representation of the road layout using primitives.
7. 7. The method (100) of any one of claims 1 to 6, wherein the trained visual foundation model comprises a graph generative model, optionally comprising at least one of an autoregressive model, a variational autoencoder, a normalized flow, a generative adversarial network, and a diffusion model.
8. The method (100) of any one of claims 1 to 7, wherein the trained visual basis model is configured to perform classification and / or semantic segmentation of the image input data.
9. 1. A computer-implemented method (200) for training a visual basis model for generating a road map based on received image input data, comprising: The method (200) comprises: annotated aerial images, in particular acquired using optical sensor systems, in particular satellite imaging systems and / or one or more drones, the annotations comprising road layouts together with locally assigned context information, - graphics representing road layouts, in particular renderings based on artificially generated graphics, and / or Graphics representing road maps for overhead images, especially overhead images and rendering images based on artificially generated graphics receiving a training data set (S202) including at least one of: - training a visual basis model (S204) for generating a road map based on the received image input data, said generated road map including a road layout together with locally assigned context information taking into account any applicable traffic regulations; Including, Optionally, the training step (S204) is self-supervised and / or based on a reconstruction loss between a ground truth included in the received (S202) training dataset and a road map generated by the visual basis model. Computer-implemented method.
10. Use of a road map generated by the method (100) of any one of claims 1 to 8 for trajectory prediction, route planning, collision avoidance and / or behavior prediction of AD vehicles and / or robotic systems, in particular traffic.
11. A computing device (300) for generating a road map suitable for use in a vehicle, particularly in automated driving (AD), comprising: The computing device (300) an input image data receiving interface (304) adapted to receive image input data comprising acquired image data representative of at least one area navigable by the vehicle; a road map generation module (306) configured to generate a road map based on the received image input data, the generation of the road map being performed by a visual basis model trained for the generation of road maps, the generated road map comprising a road layout in at least one area navigable by the vehicle together with locally assigned context information taking into account traffic regulations that may apply; A computing device (300) comprising:
12. 12. The computing device (300) of claim 11, further configured to perform any one of the steps of the method (100) of any one of claims 2 to 8 and / or to include any one of the features of the method (100) of any one of claims 2 to 8.
13. 1. A computing device (400) for training a visual basis model for generating a road map based on received image input data, comprising: annotated aerial images, in particular acquired using optical sensor systems, in particular satellite imaging systems and / or one or more drones, the annotations comprising a road layout with locally assigned context information, - graphics representing road layouts, in particular renderings based on artificially generated graphics, and / or Graphics representing road maps for overhead images, especially overhead images and rendering images based on artificially generated graphics a training data receiving interface (402) configured to receive a training data set including at least one of: a visual basis model training module (404) adapted to train a visual basis model for generating a road map based on received image input data, said generated road map including a road layout together with locally assigned context information taking into account traffic regulations that may apply; Equipped with Optionally, the training is self-supervised and / or based on a reconstruction loss between ground truth included in the received training dataset and a road map generated by the visual basis model. A computing device (400).
14. 1. A system for generating road maps particularly suitable for use in automated driving (AD) of vehicles, comprising: - a computing device (300) according to claim 11 or 12, - at least one sensor system and / or image acquisition device configured to provide image input data to an input image data receiving interface (304) of said computing device (300); A system comprising:
15. A controller (804) for a vehicle for autonomous driving (AD), a receiving interface for receiving a generated road map, said road map being generated by a method according to any one of claims 1 to 8 on the basis of received image input data, said road map comprising a road layout in at least one area navigable by the vehicle together with locally assigned context information taking into account possible traffic regulations; a processing unit configured for trajectory planning of an AD vehicle; A controller (804) comprising: