Method for creating a bird's eye view representation of a vehicle and vehicle
The method synchronizes and fuses sensor data from multiple vehicles to create a unified bird's eye view, addressing inefficiencies in existing methods, enabling accurate and efficient perception and planning for automated driving.
Patent Information
- Application Number
- DE102023005165
- Authority / Receiving Office
- DE · DE
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2023-12-14
- Publication Date
- 2025-06-18
AI Technical Summary
Existing methods for creating a bird's eye view representation from vehicle surroundings using sensor data lack efficient end-to-end learning and synchronization, leading to inaccuracies and inefficiencies in perception and planning for automated driving.
A method involving time-synchronized image data acquisition and fusion of sensor data from multiple vehicles using a latent representation, combined in a unified coordinate system, enables end-to-end training and error propagation for improved perception and planning.
Enhances the accuracy and efficiency of automated driving by providing a synchronized and aligned bird's eye view representation, allowing for effective training and integration of learning-based components for perception and planning.
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
The present invention relates to a method for creating a latent representation of a vehicle and vehicle designed as a bird's eye view.The publication "Lift, Spat Shoot: Encoding Images from Arbitrary Camera Rigs by Explicitly Non-projecting to 3D; Jonah Philion, Sanja Fidler; NVIDIA University of Toronto Vector Institute; https: / / arxiv.org / pdf / 2008.05711.pdf" discloses a method in which, for the movement planning, data from an arbitrary number of cameras of an automatically driving vehicle are extracted with an end-to-end architecture of a representation of a scene from a bird's eye view.From US 2023 / 0053785 A1 a method is known in which a machine learning model from images of sensors arranged on a vehicle represents objects in the environment of the vehicle in a bird's eye view.It is an object of the present invention to provide a method and a device which uses images of sensor systems of further vehicles to generate a bird's eye view of a vehicle and its surroundings using an end-to-end architecture.The object is achieved by a method having the features of claim 1 and a device according to claim 9. The dependent claims define preferred and advantageous embodiments of the present invention.In the method according to the invention, a temporally synchronized generation of image data generated by the sensors thereof is requested by the ego vehicle from further vehicles located in the environment, the synchronized generated image data are received by the ego vehicle and extracted with the model to form a latent representation of a bird's eye view of the ego vehicle.A latent representation is understood in the present case as an abstract multidimensional space which represents a bird's eye view generated from image data. Advantageously, the generation of a latent representation of the bird's eye view allows learning-based components (e.g. prediction and planning) to be used subsequently, so that the perception of the environment can also be trained during the training. In other words, this means that the architectures can be trained end-to-end because the formation and fusion of specific object hypotheses can be differentiated and thus error propagation is made possible during training. The end-to-end trainable latent representation presented here, which is ideally divided by a plurality of vehicles, has the advantage that planning or prediction algorithms, for example, can also be used in a building manner, which planning or prediction algorithms are additionally learning-based and during the training of which all the components located upstream can be trained together. In this case, it is described how a latent representation can be obtained from a plurality of camera images of an automated vehicle, on which latent representation learning-based planning models, for example, can be trained end-to-end following. The latent representation is a representation that can be interpreted as a grid map from bird's eye view around the automated vehicle. It should be noted that the latent representation is not interpretable for the humanSince the data received from image data of the further and preferably that of the ego vehicle are unstructured, it is essential that the data received for generating the latent representation of all vehicles, i.e. when using image data of the ego vehicle, also the data thereof, are synchronized. In an advantageous manner, the ego vehicle requests, before receiving the data, generation of the image data, which is preferably synchronized in time with its own data acquisition.According to a further development of the invention, capture times of image data captured by the sensor system of the ego vehicle are used as a specification for a time-synchronized capture. The ego vehicle specifies a sampling rate and / or a start signal for synchronization to the other vehicles as master. Advantageously, the control of the ego vehicle can use its image data with the image data of further vehicles to generate the latent representation of the bird's eye view.In a further preferred embodiment, the time-synchronized detection takes place on the Global Positioning (GPS) time system. The ego vehicle specifies a timing based on the time measurement of the GPS system and a time start signal. All vehicles can advantageously access GPS signals at the same time, as a result of which it is possible to realize latency-free synchronization of the image data generation.In a further specific embodiment, only image data of vehicles situated up to a predefined distance in the surroundings of the ego vehicle are used as a basis for generating the latent representation of the bird's eye view. For example, only vehicles are considered that have a predefined Euclidean distance from the ego vehicle. The distance is preferably dynamically adapted as a function of parameters such as the route traveled, speed or a driving situation. At high speeds, the distance is increased in order to obtain a sufficient preview for driving maneuvers carried out by the automated ego vehicle; at low speeds, the distance may be reduced in order to reduce the data load. Advantageously, the information required for the automated driving state can be generated via the distance while optimizing the data rate.In a development of the invention, in order to generate the bird's eye view, the image data of each of the vehicles are mapped in a uniform coordinate system as a point cloud and are each merged to form a total point cloud, wherein localization errors of the total point clouds with one another are reduced by an optimization algorithm. To generate the point cloud in a uniform coordinate system, the camera calibration data (intrinsic and extrinsic) are also transmitted to the ego vehicle and taken into account. The position of a vehicle in a global georeferenced coordinate system is generally very error-prone. Even most modern geolocation methods (e.g. fusion from GPS, odometry, guidance post detection) do not achieve centimeter accurate localization. The ego vehicle is able to generate image data transformed into a global reference coordinate system from faulty vehicle positions and image calibration data. This is identical for all further vehicles and for the ego vehicle itself. For example, a center of the ego vehicle may be selected as the origin. Furthermore, in a similar manner to the publication "Lift, Plate Shot: Encoding Images from Arbitrary Camera Rigs by Incompletely Projecting to 3D", a so-called "lift" is carried out which converts image data of a scene into a 3D layout from a bird's eye view.The following procedure is adopted for each of the images defined by the image data: Extrinsic and intrinsic matrices (camera calibration data) are used, which describe how each image sensor views the scene. The target is to convert the 2D images into a 3D point cloud in the uniform coordinate system (by this is meant a georeferenced coordinate system). For this purpose, no depth sensors are used, but a latent feature and simultaneously a distribution over the depth are estimated for each pixel in the sensor image by means of a convolutional neural network. The distribution over depth enables the model to cope with situations where the depth is uncertain or ambiguous. The result is a point cloud with latent features for each point in the uniform coordinate system. For each sensor, the point cloud has the shape of a frustum. The training of the CNN is only performed with the training of the target task (because the overall architecture is trainable end-to-end). Nevertheless, pre-trained (e.g., on ImageNet) CNNs may also be used. The point clouds produced from the images of the ego vehicle are arranged slightly incorrectly (e.g. rotated or translated) in comparison to the point clouds produced from the images of the further vehicles on account of the inaccurate geolocation. Accordingly, the point clouds of the further vehicles behave among one another.To remedy this, the following procedure is adopted:The point clouds (frustums) of all images of the ego vehicle and those of the further vehicles are each merged into a total point cloud for each vehicle. This is possible in principle since all images of different sensors of the respective vehicle are subject to the same geolocation error.The plurality of point clouds of the respective vehicles that are incorrectly aligned with one another are correctly aligned with one another in the following step. This is an optimization problem which is already known, in particular, from SLAM technology. One possible implementation is the ICP algorithm (Iterative Closest Point Algorithm).After the optimization, the point clouds of the respective vehicles are located in a uniform and optimized coordinate system. A total point cloud can be created by merging all point clouds. Due to the time synchronization performed and the optimization performed (e.g. by means of ICP), both temporal errors and local geolocation errors are excluded. Subsequently, the latent grid map of the ego vehicle can be created from the bird's eye view (BEV). For this purpose, a CxHW tensor is applied. It is initialized with zeroes. C is identical to the latent feature size of each point of the point clouds. The features from each point in the overall point cloud are added to the next geometric cell in the tensor. This corresponds to sum pooling. This is followed by a tensor which, in the form of the latent representation of the bird's eye view, summarizes the entire overall point cloud (latent bird's eye view grid map).In one embodiment, based on the latent representation, the bird's-eye view pulldown decoders operate on object detection, prediction, trajectory planning, and / or lane extraction. In training these downstream decoders, an error measure is calculated (e.g., when predicting the errors between prediction and the motion of an agent being driven in truth). This error measure is propagated through the entire model and leads to not only the downstream decoder being trained but also to all upstream components, in particular the neural network, being trained.According to a further embodiment of the present invention, the image data comprise data from cameras, lidar sensors or a radar sensor. As a rule, these are first converted into a 3D grid by an encoder. Upsetting this 3D grid in the z-dimension also produces a bird's perspective grid map. This can be fused, for example, directly before the downstream decoders with the latent representation of the bird's eye view resulting from the cameras. A simple way of fusing is the concatenatenation of both latent bird's-eye view raster maps.The vehicle according to the invention comprises a model implemented in a computing unit for generating a representation of a bird's eye view from data of various sensors for detecting the environment, wherein the vehicle requests from other vehicles located in the environment a temporally synchronized generation by the ego vehicle of image data generated with the sensors thereof, receives the synchronously generated image data and is extracted with the model to form a representation of a bird's eye view of the ego vehicle.The present invention will be explained in more detail below with reference to the accompanying drawing in which the same or similar parts are denoted by the same reference numerals.The single FIGURE shows a traffic scenario with an ego vehicle and further vehicles.The automated driving ego vehicle 1 travels along a road. The ego vehicle 1 has sensors 3 which are directed into the environment and are designed as a camera. The ego vehicle 1 is connected via a wireless network, not shown, to further vehicles 5 via a data link. The further vehicles 5 likewise have the sensors 3 designed as a camera. Further vehicles 7 not bound in via the network travel in the environment of the ego vehicle 1.Environment data are required for the automated driving operation of the ego vehicle 1, wherein the ego vehicle 1 receives image data from the further vehicles 5 via the network in addition to image data generated by the own cameras 3. In addition to the image data, the camera calibration data (i.e. intrinsic and extrinsic) are also transmitted from each image to the ego vehicle 1. The camera calibration data relates to the respective vehicle coordinate system, the relationship of which is known accurately due to attachment to the rigid vehicle body. Therefore, the positions of the vehicles 5 in a global georeferenced coordinate system are also sent to the ego vehicle 1. The ego vehicle 1 is configured to use the image data to implement it using a model known from the prior art and preferably embodied as a neural network to ascertain a latent representation of the bird's eye view in the environment. The image data received by the further vehicles 5 can result in better detection of the environment around the ego vehicle 1.Based on the latent representation of the bird's eye view, planning or prediction algorithms required for the automated driving operation are used, the algorithms are learning-based, during whose training all the components arranged upstream can be trained.To create the latent representation of the bird's eye view, the ego vehicle 1 sends a signal to the surrounding vehicles 5 connected by data technology.The signal includes a request to transmit image data captured by the cameras of the surrounding vehicles. The signal further comprises a specification for synchronizing the recording time of the image data by the further vehicles. If the ego vehicle uses images captured using its own cameras and sensors to create the representation, the synchronization takes place with their recording times. In addition to recording times, which are preferably based on a GPS time signal, a repetition frequency is predefined with the signal for synchronization. With the transmission of the image data of the environment, the networked further vehicles 5 transmit their own positions and geometry data, so that these themselves can be included in the creation of the latent representation by the ego vehicle 1.In the present case, the signal is directed only to vehicles 5 which are located in a predefined radius 11 around the ego vehicle 1 in order to restrict the data traffic.To generate the bird's eye view comprising the vehicles 3, 5, 7, the image data of each of the vehicles 3, 5 are mapped in a uniform coordinate system as a point cloud and are each merged into a total point cloud, wherein localization errors of the total point clouds among one another are reduced by an optimization algorithm.References included in the specificationThis list of documents cited by the applicant has been produced in an automated manner and is only included for the better information of the reader. The list is not part of the German patent application or utility model application. The DPMA does not take any adhesion for any faults or omissions.Patent Literature citedUS 2023 / 0053785 A1
[0003] Cited Non-Patent LiteratureLift, Plate Shot: Encoding Images from Arbitrary Camera Rigs by Incompletely Projecting to 3D; Jonah Philion, Sanja Fidler; NVIDIA University of Toronto Vector Institute; https: / / arxiv.org / pdf / 2008.05711.pdf
[0002]
Claims
Method for creating a latent representation, designed as a bird's eye view, of an automatedly operated ego vehicle (1) from data of various sensors (3) with a model implemented in a computing unit (9) of the ego vehicle, characterized in that a temporally synchronized generation of image data generated with the sensors (3) thereof is requested by the ego vehicle (1) from further vehicles (5) located in the environment, and the synchronously generated image data is received and extracted with the model to form a latent representation of a bird's eye view of the ego vehicle (1).Method according to Claim 1, characterized in that capture times for image data captured by the sensor system (3) of the ego vehicle (1) are used as a specification for time-synchronized capture.Method according to Claim 1 or 2, characterized in that the time-synchronized detection takes place on the time system of the Global Positioning System (GPS).Method according to one of Claims 1 to 3, characterized in that only image data of vehicles (5) situated up to a predefined distance in the environment of the ego vehicle (1) are used as a basis for generating the bird's eye view.Method according to Claim 4, characterized in that the distance is dependent on the travel speed of the ego vehicle (1).Method according to one of Claims 1 to 5, characterized in that, in order to generate the bird's eye view, the image data of each of the vehicles (1, 5) are mapped in a uniform coordinate system as a point cloud and are in each case combined to form a total point cloud, wherein localization errors of the total point clouds with respect to one another are reduced by an optimization algorithm.Method according to one of Claims 1 to 6, characterized in that, on the basis of the representation, the bird's-eye view pulldown decoders are trained at least on one of the following functions: object detection prediction trajectory planning trace extractionMethod according to one of claims 1 to 7, characterised in that the image data comprise data from at least one of the following devices: - cameras - lidar - radarEgo vehicle for carrying out a method according to one of the preceding claims, having a model implemented in a computing unit (9) for generating a latent representation of a bird's eye view from data of various sensors (3) for detecting the environmental data, characterized in that the ego vehicle (1) requests, from further vehicles (5) located in the environment, a temporally synchronized generation of image data generated by the sensors thereof, receives the image data generated in a synchronized manner and extracts it with the model to form a representation of a bird's eye view of the ego vehicle (1).
Citation Information
Patent Citations
Display method and display system for a vehicle
DE102011084084A1
Vision-based machine learning model for aggregation of static objects and systems for autonomous driving
US20230053785A1