Sensor Virtualization
The method addresses the scalability issue in training neural networks for vehicle object prediction by mapping physical sensor data to virtual sensors and then to a 3D model space, allowing for unified training and improved efficiency and reliability.
Patent Information
- Application Number
- JP2024570804
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2022-05-30
- Filing Date
- 2023-05-30
- Publication Date
- 2025-06-05
AI Technical Summary
Current methods for training neural networks to predict objects in the vicinity of a vehicle are not scalable due to the need for separate models for each sensor configuration, leading to high time and cost expenses.
A method involving the acquisition of physical sensor data, performing a first mapping to virtual sensors, and a second mapping to 3D model space, followed by training a neural network using virtual sensor data and annotations in the 3D model space, allowing for sensor-independent perception and unified training.
This approach enables the combination of training data from different physical sensors, compensates for sensor failures, and improves the reliability and efficiency of training neural networks for 360-degree environment modeling.
Smart Images

Figure 2025517563000001_ABST
Abstract
Description
[Technical field]
[0001] The present invention relates to a method for training a neural network to predict objects in the vicinity of a vehicle and to a system for training a neural network to predict objects in the vicinity of a vehicle, the system being arranged to perform the method according to any one of claims 1 to 15.
[0002] The invention also relates to a computer readable storage medium storing program code including instructions for carrying out such a method. [Background technology]
[0003] A key element of autonomous driving is the ability to build a 360-degree environment model. The environment model can be obtained by utilizing different sensor modalities. However, this involves training a different model for each sensor configuration. Unfortunately, this solution is not scalable in terms of time and cost (e.g., training data collection, annotation, neural network modeling, and training). Therefore, a unified training method is needed to overcome the limitations mentioned above. Summary of the Invention
[0004] It is an object of the present invention to provide a method for training a neural network to predict objects in the vicinity of a vehicle and a system for training a neural network to predict objects in the vicinity of a vehicle, the system being configured to perform the method according to any one of claims 1 to 15, which overcomes one or more of the above-mentioned problems of the prior art.
[0005] A first aspect of the invention is a method for training a neural network to predict objects in the vicinity of a vehicle, comprising: acquiring physical sensor data from a physical sensor having one or more modalities; performing a first mapping from the physical sensor data to one or more virtual sensors to obtain virtual sensor data; performing a second mapping from the virtual sensor data to the 3D model space; and training a neural network based on the virtual sensor data and one or more annotations in the 3D model space.
[0006] The method of the first aspect has the advantage that a sensor-independent perception setup is used, so training data acquired using different physical sensors can be combined.
[0007] Preferably, the physical sensor has at least two modalities. Also, the physical sensor may comprise at least one sensor whose sensor data is not acquired directly in the 3D model space.
[0008] Training of the neural network can be performed, for example, based on ground truth training data that includes annotations, which are not limited herein, i.e., annotations may be object characteristics, objectness scores, and / or other predictions that are not directly related to one or more objects.
[0009] In a first implementation of the method according to the first aspect, the method further includes acquiring training data using a fleet of vehicles, the fleet using different physical sensors.
[0010] Since all physical sensor data from a fleet of vehicles is mapped to virtual sensor data (preferably reflecting the same virtual sensor), uniform training can be performed even if the initially acquired sensor data is heterogeneous. For example, the physical sensors may differ in their localization on the vehicle, field of view, range, and / or many other acquisition parameters. These differences can be compensated for by the first mapping, which can in fact correct the physical sensor data of multiple different physical sensors such that the corresponding virtual sensor data has properties as if it was acquired with one uniform virtual sensor (or a group of virtual sensors corresponding to different modalities, for example).
[0011] It has proven particularly useful in experiments to perform a virtualization of several modalities, for example having at least two virtual sensors corresponding to at least two virtual modalities.
[0012] The vehicle fleet preferably comprises at least 10 vehicles, particularly at least 100 or 1,000 vehicles, or more preferably at least 10,000 vehicles.
[0013] In a further implementation of the method according to the first aspect, performing the first mapping includes applying a transformation from the physical sensor data to obtain virtual sensor data, the transformation being based on a difference between actual physical properties of the physical sensor and virtual physical properties of the virtual sensor.
[0014] For example, a physical sensor may comprise one or more cameras with a given first focal length, and a virtual sensor may comprise one or more cameras with a second focal length different from the first focal length. Thus, a mathematical transformation may be used to convert the physical sensor data into virtual sensor data. In other words, the virtual sensor data appears as if it was acquired with a camera having the second focal length. Thus, a fleet of vehicles may use cameras with different focal lengths, but all acquired data will be available as virtual sensor data with a common second focal length.
[0015] In a further implementation of the method according to the first aspect, the 3D model space comprises an overhead raster. This has the advantage that many annotations can be better estimated based on a 3D model space that is an overhead raster.
[0016] In a further implementation of the method according to the first aspect, performing the second mapping includes, if a failure of a first one of the physical sensors is detected, the method embedding using virtual sensor data obtained from a second one of the physical sensors, the first and second sensors using different modalities.
[0017] This has the advantage that the method can compensate for failures of one or more physical sensors, thus significantly improving the reliability of the method in practical use. For example, in autonomous driving, it must be ensured that the vehicle can navigate safely even if one of the cameras is obstructed, for example by dirt on the camera lens.
[0018] Filling of missing data can be performed as follows: Since the different modalities are mapped onto the model space, the overlapping observations (feature maps) from the different sensor modalities can be stored in the neural network. This provides sufficient redundancy in case of a sensor drop (i.e., one of the sensors fails to provide a signal for a particular timestamp). For example, at time t, the front camera is unable to provide a camera observation, but the long-range lidar is working properly. In this case, the feature map corresponding to the camera in front of the ego vehicle in the model space is empty and would not be detected without additional sensor equipment. However, the lidar-provided observations (feature maps) are stored in the neural network, allowing to detect objects even if the camera observation is impaired.
[0019] In a further implementation of the method according to the first aspect, the virtual sensors consist of one virtual sensor for each virtual modality.
[0020] In other words, in this implementation, there is only one virtual sensor for each virtual modality. Experiments show that this mapping to one virtual sensor for each (virtual) modality simplifies the training data and allows for better training results.
[0021] In a further implementation of the method according to the first aspect, the method further comprises training encoder and decoder parameters of the transducer model, the encoder mapping from the virtual sensor data to a latent space and the decoder mapping from the latent space to the 3D model space.
[0022] This has the advantage that the mapping can be efficiently learned from training data, even for modalities where a well-defined mathematical transformation from virtual sensor data to 3D model space is not available for this modality. Experiments show that the transducer model can provide good results for the mapping from virtual sensor data to 3D model space.
[0023] In a further implementation of the method according to the first aspect, the neural network includes a feature mapping sub-network that maps from the virtual sensor data to a feature map in the 3D model space and a processing head that maps from the feature map to the annotation space.
[0024] In this implementation, the feature map in the 3D model space is an intermediate result that is used by the processing head to determine the annotation in the annotation space. Experiments show that this setup produces excellent results.
[0025] In this implementation, the annotation space may include annotations such as object labels and / or the (non-)existence of objects, although it is understood that there is no limitation on the types of annotations used. In particular, annotations may also include scalar and / or vector values.
[0026] In a further implementation of the method according to the first aspect, the feature map comprises feature sub-maps in the 3D model space for each of the at least two modalities, preferably each feature sub-map being supplied to the processing head.
[0027] For example, there may be a first feature submap for camera data, a second feature submap for lidar data, and a third feature submap for radar data. Each of the feature submaps may contain multiple values for each location in the 3D model space. The entire feature map can be thought of as being stored in one large tensor.
[0028] In a further implementation of the method according to the first aspect, the one or more annotations include the presence of an object and / or a label of the object. The prediction may also include an objectness score, such as, for example, a probability that the object is present at a given position in the 3D model space. The label of the object may include, for example, a pedestrian label, a road label, a vehicle label, etc.
[0029] In a further implementation of the method according to the first aspect, the method further comprises performing a fusion between at least two modalities in the 3D model space. A lightweight convolutional neural network can be used to perform this mapping.
[0030] In a further implementation of the method according to the first aspect, training the neural network includes training the neural network multiple times, with at least one intervening training data from one or more of the virtual sensors and / or physical sensors being omitted.
[0031] This implementation has the advantage that the robustness of the trained neural network is improved.
[0032] In a further implementation of the method of the first aspect, the first mapping includes one or more first parameters, the second mapping includes one or more second parameters, the processing head for obtaining the one or more annotations includes one or more third parameters, and training the neural network includes end-to-end training to obtain the first, second, and third parameters.
[0033] End-to-end training of all parameters has the advantage that it is simply necessary to obtain sufficient training data and all relevant parameters of the entire system can be determined automatically.
[0034] A further aspect of the present invention refers to a system for training a neural network to predict objects around a vehicle, the system being configured to execute the method of the first aspect or one of the implementation forms of the first aspect. The system can be implemented, for example, on a server, in particular on multiple servers or in the cloud. The system can be configured to continuously receive new training data, for example from a fleet of vehicles that can be connected to the system.
[0035] A further aspect of the present invention refers to a computer-readable storage medium storing program code comprising instructions which, when executed by a processor, perform the method of the second aspect or one of the implementations of the second aspect.
[0036] In order to more clearly describe the technical features of the embodiments of the present invention, the accompanying drawings provided for describing the embodiments are briefly introduced below. The accompanying drawings in the following description are only some embodiments of the present invention, and modifications to these embodiments are possible without departing from the scope of the present invention defined in the claims. [Brief description of the drawings]
[0037] [Figure 1] 1 is a flowchart of a method according to one embodiment of the present disclosure. [Diagram 2] FIG. 2 is a block diagram illustrating a first mapping from physical sensors to virtual sensor data according to one embodiment of the present disclosure. [Diagram 3] FIG. 13 is a block diagram illustrating a second mapping from virtual sensor data to a feature map in a 3D model space according to one embodiment of the present disclosure. [Figure 4] 4 is a flowchart for processing a feature map, for example a feature map determined by the system of FIG. 3, according to one embodiment of the disclosure. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS
[0038] The above description is merely an implementation method of the present invention, and the scope of the present invention is not limited thereto. Any modifications or replacements can be easily made by those skilled in the art. Therefore, the protection scope of the present invention should be subject to the protection scope of the attached claims.
[0039] Neural networks trained for 3D perception tasks are sensitive to the input sensor because they can implicitly learn the unique properties of a particular sensor. For example, when a neural network is employed to detect an object on a camera stream recorded at a different focal length than the one the network was trained on, the depth estimates will typically have noticeable errors due to the inherent misalignment of the cameras.
[0040] An embodiment of the present invention solves this problem by defining reference virtual sensors and mapping the real sensors onto these virtual sensors, preferably the virtual sensors reflecting different physical properties than at least one of the physical sensors.
[0041] 1 is a flow chart of a method 100 according to the invention. The method is used to train a neural network to predict objects in the surroundings of a vehicle. For example, the vehicle can be a (semi- or fully) autonomous vehicle.
[0042] The method includes a first step 110 of acquiring physical sensor data from a physical sensor having at least two modalities. For example, the physical sensor may comprise a camera and a lidar or radar sensor. In other embodiments, there may be only one modality.
[0043] The method includes a second step 120 of performing a first mapping of physical sensor data to virtual sensors to obtain virtual sensor data. For example, images (video streams) from a camera can be mapped to obtain virtual images (virtual video streams), and lidar and / or radar data can be mapped to obtain virtual lidar and / or virtual radar data. The mapping of camera data can be based on the difference between the real physical properties of the camera and the virtual physical properties of the virtual camera. Similarly, the mapping of data acquired by a lidar or radar can be mapped to a corresponding virtual lidar or radar. It is understood that further such mappings can be performed, for example, from multiple sensors of one modality to one or more virtual sensors corresponding to that modality.
[0044] For example, the virtual camera can have predefined characteristics (e.g., focal length, principal point, position relative to the car's coordinate system). Mapping from the physical camera to this virtual camera can be performed by obtaining a transformation from the physical camera to the virtual camera using extrinsic and intrinsic matrices. In the case of radar, the virtual sensor can have a predefined field of view and perception range. The physical radar signal is then preprocessed and projected onto the 3D model space raster grid, and this preprocessed signal is inserted into the virtual radar's overhead grid (which may have a longer perception range and a different angular resolution than the physical sensor). Physical-virtual lidar mapping works similarly.
[0045] In certain embodiments, for example when using several different cameras and one type of lidar, there may be only one virtual sensor: the different cameras are mapped to a virtual camera sensor, but the same type of physical lidar is installed in all recording vehicles, so the lidar can be used without virtualization.
[0046] The method includes a further step 130 of performing a second mapping from the virtual sensor data to the 3D model space.
[0047] The method includes a final step 140 of training a neural network based on the virtual sensor data and one or more annotations in the 3D model space.
[0048] In the case of cameras, this (first) mapping can be created using the projection geometry and camera matrices that guarantee invariance to translation and rotation. 3D sensing devices, on the other hand, typically emit their signals in model space by design. Due to this property, sensor virtualization can be performed using extrinsic matrices to provide translation and rotation invariance. A proper resolution match between the original 3D sensor and the virtual sensor can be guaranteed by scaling the input signal by a factor that can be derived from the sensor configuration. In this way, a sensor-independent perception system can be developed, i.e., avoiding the need to train a separate model when recordings from a new sensor configuration become accessible.
[0049] Sensor virtualization allows for the utilization of heterogeneous sensor configurations to train 360-degree perception models. However, it does not provide an efficient end-to-end solution for training said models. The most frequently used method for implementing 360-degree perception systems is to train separate models for each sensor modality and then combine them to form an environment model. This solution is time-consuming and expensive, as training convolutional neural networks requires a significant amount of power consumption and computational power.
[0050] Since the ground truth of the ego-vehicle's surroundings is explicitly defined in model (world) space and the input sensors are mapped onto a reference virtual sensor with known orientation and field of view, it is beneficial to output predictions in model space as an overhead raster. During training, the neural network can be fed by signals from all sensors, and predictions are defined in model space, which is also a representation of the ground truth. In this way, all inputs can be processed together in one step.
[0051] The presented solution to improve efficiency is partly inspired by a Siamese network that uses shared weights for the two images. In the presented embodiment, a dedicated sub-network is used for each sensor modality. For example, if a recording system has four pinhole cameras (front, back, left, right), only one backbone network is needed for feature extraction. If fisheye cameras are part of the sensor setup, their inputs can be used by the pinhole camera feature extractor network or a new fisheye encoder network can be defined. The number of trainable network parameters can be linearly reduced by the number of sensors corresponding to the same sensor group.
[0052] Since the ground truth is preferably defined in a 3D model space and we want to predict in 3D as well, the camera features are transformed from the 2D image space to the 3D world space. The dimensionality increase mapping can be performed in several ways. One option is to utilize a pseudo-lidar solution. The preferred embodiment uses an encoder-decoder architecture. The encoder-decoder network includes two sub-networks. The encoder is responsible for embedding the input into a latent space, which is typically a smaller dimensional space than the original input space. The decoder network is then fed by the encoded input and transforms it into the desired representation.
[0053] In the preferred embodiment, the encoder-decoder network is fed by the camera images. The encoder is a convolutional neural network with asymmetric strides, and the decoder uses transposed convolutions with asymmetric strides. As a first step, the encoder extracts relevant features from the image in a trainable way. The transition from image space to model space can be further enhanced by a vision transformer that uses an attention mechanism where the transformer module learns which image features are attended to by which model space features. The output of the decoder is a final raster representation of the model space around the ego-vehicle, ensuring a bijective mapping between the ground truth and the predictions.
[0054] Different sensors have inherently different signal representations. For example, a camera image is represented as a dense grid where every pixel has a color value. The input of a lidar is completely different from a camera image and can be represented as a point cloud consisting of the X, Y, Z coordinates of the reflectance and the intensity of the reflectance. A radar can have a similar representation with an additional velocity component. Both lidar and radar have a sparse representation compared to an image. Furthermore, 3D sensing devices operate in model space by design while camera images exist in 2D space. To avoid mixing substantially different input representations, in the preferred embodiment, one sub-network for each sensor modality is used. The sub-networks are populated by inputs from different sensors without introducing uncertainties due to heterogeneous sensor inputs. Uncertainties are handled by sensor virtualization, i.e., mapping to virtual sensors.
[0055] With the sensor virtualization and the known position and orientation of the reference virtual sensor, the environment model can be trained end-to-end. The training step operates on a set of sensor inputs, and the network outputs predictions in the model space. Every sensor-specific sub-network is fed by the corresponding sensor input and outputs its feature map in the model space relative to the reference sensor. A 360-degree environment model can be obtained by appropriately positioning and concatenating the feature maps of the sub-networks to form an overhead raster. For example, the feature map of the rear camera is placed on the raster in front of the ego vehicle. Similarly, the lidar feature map corresponding to the area behind the vehicle can be concatenated to the feature map of the rear camera.
[0056] FIG. 2 is a block diagram illustrating a first mapping from physical sensors to virtual sensor data.
[0057] The system 200 includes physical sensors, such as multiple cameras 210a, 210b, multiple lidar sensors 220a, 220b, and multiple radar sensors 230a, 230b, each of which is processed in a similar manner to arrive at virtual sensor data, in the illustrated embodiment.
[0058] The images 212a, 212b acquired from the multiple cameras 210a, 210b are processed by the virtual camera unit 214 to obtain virtual images 216a, 216b. The number of acquired virtual images 216a, 216b may correspond to the number of physical images 210a, 210b, but in other embodiments may be completely different. The mapping performed in the virtual camera unit 214 may include applying a geometric transformation to the physical images 210a, 210b, such as adjusting the field of view or focal length and / or adjusting the geometric distortion. Also, the physical images 212a, 212b may have a different resolution and / or size compared to the virtual images 216a, 216b. The virtual images 216a, 216b can be viewed as standardized and certified images.
[0059] A number of lidar sensors 220a, 220b generate physical lidar data 222a, 222b, which are processed by a virtual lidar module 224 to obtain virtual lidar data 226.
[0060] The plurality of radar sensors 230a, 230b generate physical radar data 232a, 232b. Preferably, a first lidar sensor 220a of the plurality of lidar sensors generates a first point cloud 222a and a second lidar sensor of the plurality of lidar sensors generates a second point cloud 222b, which are processed by a virtual radar module 234 to obtain virtual radar data 236.
[0061] The use cases for lidar and radar are similar to the camera example. Physical sensors with different parameters (e.g., perception range, field of view) are mapped onto a virtual sensor. Since lidar and radar operate in model space, it is preferable that there is only one virtual data combined. This can be seen in the case of radar, where a first radar sensor 230a generates four sensor signals in its first physical radar data 232a, a second physical radar sensor 230b generates three sensor signals in its second physical radar data 232b, and the virtual radar unit combines both physical signals into one virtual data with four signals.
[0062] Preferably, no manually operated blending is required since the mapping has already been performed to the 3D model space.
[0063] FIG. 3 shows a second mapping from the virtual data to a feature map in the 3D model space.
[0064] Specifically, the system 300 shown in FIG. 3 receives as inputs (e.g., generated by the system 200 shown in FIG. 2) virtual camera data including data from a virtual rear camera 310a and a virtual front camera 320b, virtual radar data including data from a virtual rear radar 320 and a virtual front radar 320b, and virtual lidar data including data from a first virtual lidar 330a and a second virtual lidar 330b.
[0065] The camera data 310a, 310b are processed by a pinhole encoder 312 to obtain camera embeddings 314. These embeddings 314 are then processed by a pinhole decoder 316 to obtain sensor virtualization parameters 318 for the camera data. From these, feature maps 340a, 340b for the back and front cameras are obtained.
[0066] Data from the virtual back radar 320a and the virtual front radar 320b are processed in the radar neural network 322 to obtain sensor virtualization parameters 234 associated with the radar data. From these, feature maps 342a, 342b for the back and front radars are derived.
[0067] Similarly, data from the virtual lidars 330a, 330b are processed by a lidar neural network 332 to obtain sensor virtualization parameters 334 for the lidar data. From these, feature maps 344a, 344b for the lidar sensors are obtained.
[0068] The neural networks 322, 332 can be trained using the virtual sensor data and the training data where feature maps are available. Preferably, each feature map corresponds to one virtual sensor. Therefore, they can also be referred to as virtual sensor feature maps.
[0069] The sensor virtualization parameters may include parameters of the placement of the virtual sensors relative to the 3D model space, which are known during training time and can be considered as inputs for constructing the bird's-eye feature map 410.
[0070] Since the first transformation has already been done from the physical sensor to the virtual sensor, information about the physical placement of the physical sensor is no longer needed at this stage. If the physical sensor includes a front camera, the virtual sensor will usually also include a front camera, but the virtual camera may have a completely different field of view and even a different placement. Therefore, the placement details of the physical and virtual front cameras are different, and it is the placement of the virtual camera that is important when generating the feature map.
[0071] For example, during virtualization of the front camera, we know that it is looking forward. Therefore, the corresponding virtual feature map can be inserted at the appropriate place in the final bird's-eye view raster. If the virtual camera is a front camera, the ego vehicle is located at the origin and the feature map starts from the center of the bird's-eye view raster (i.e., from index (BEV raster width / 2)) since the front camera sees the environment in front of the ego vehicle. If it is a rear camera, the corresponding feature map is inserted into the BEV raster starting from index 0 (behind the ego vehicle).
[0072] The system 300 obtains a feature map that can be further processed to obtain an estimated prediction.
[0073] FIG. 4 shows a flow chart for processing a feature map 410, for example a feature map determined by the system of FIG.
[0074] The feature map 410 can be processed to obtain an auxiliary camera prediction 412, an auxiliary radar prediction 414, and / or an auxiliary lidar prediction 416. The feature map 410 is further processed by a fusion neural network 420, which generates a final prediction 422 in the 3D model space.
[0075] The 3D model space can be an overhead raster, specifically a 306 degree overhead raster. Once we have a 360 degree overhead raster with feature maps stored in the model space, we can add prediction heads. For example, one head is responsible for learning whether or not there is an object in a raster cell. At this point, the network can be trained. We implemented heads as convolutional layers whose output channel size corresponds to the dimensionality of the feature being predicted. For example, the presence of an object in a raster cell is a convolutional layer with one output channel where the probability of objectness is predicted for each cell.
[0076] However, it is beneficial to impose a sensor fusion sub-network on top of the 3D model spatial raster feature map to train the network to properly combine features from different sensor modalities. In a preferred implementation, a lightweight CNN is utilized as the sensor fusion sub-network. When a fusion sub-network is used, an auxiliary head can be added to all virtual sensor models to facilitate training, and intermediate predictions can be concatenated to the input of the sensor fusion sub-network (similar to the method described in the previous section). The auxiliary loss calculated for the virtual sensor model is also added to the final loss calculated for the entire perceptual range (i.e., on the overhead raster). The robustness of the perception system can be further improved by randomly turning off some of the sensors during training. This method is similar to dropout regularization, except that instead of masking a set of neurons, the sensor input is masked.
[0077] 3 are trained together with parameters of the fused neural network 420, which may act as a prediction head. Furthermore, the first mapping from physical sensor data to virtual sensor data may also include parameters (e.g., parameters of one or more further neural networks) that may also be obtained in the same training process.
[0078] Thus, end-to-end training can be performed from virtual sensor data (or starting from the original physical sensor data) to final predictions.
Claims
1. A method (100) for training a neural network (312, 322, 332, 420) to predict objects in the vicinity of a vehicle, comprising: Acquiring physical sensor data from a physical sensor having one or more modalities (110); performing a first mapping of the physical sensor data to one or more virtual sensors to obtain virtual sensor data (120); performing a second mapping from the virtual sensor data to a 3D model space (130); and training (140) the neural network (312, 322, 332, 420) based on the virtual sensor data and one or more annotations in the 3D model space.
2. The method of claim 1 , further comprising acquiring training data using a fleet of vehicles, the fleet using different physical sensors.
3. 3. The method (100) of claim 1, wherein performing (120) the first mapping comprises applying a transformation from the physical sensor data to obtain the virtual sensor data, the transformation being based on a difference between actual physical properties of the physical sensor and virtual physical properties of the virtual sensor.
4. The method (100) of any one of claims 1 to 3, wherein the 3D model space comprises an overhead raster.
5. 5. The method (100) of claim 1, wherein performing the second mapping (130) includes embedding using virtual sensor data obtained from a second one of the physical sensors if a failure of a first one of the physical sensors is detected, the method (100) comprising: embedding using virtual sensor data obtained from a second one of the physical sensors, the first and second sensors using different modalities.
6. The method (100) of any one of claims 1 to 5, wherein the virtual sensors consist of one virtual sensor for each virtual modality.
7. 7. The method (100) of any one of claims 1 to 6, further comprising training parameters of an encoder (312) and a decoder (316) of a transformer model, the encoder mapping from the virtual sensor data to a latent space and the decoder mapping from the latent space to the 3D model space.
8. The method (100) of any one of claims 1 to 7, wherein the neural network (312, 322, 332, 420) comprises a feature mapping sub-network that maps from the virtual sensor data to a feature map (340a, 340b, 342a, 342b, 344a, 344b, 410) in the 3D model space, and a processing head that maps from the feature map (340a, 340b, 342a, 342b, 344a, 344b, 410) to an annotation space.
9. 9. The method (100) of claim 8, wherein the one or more modalities include at least two modalities, and the feature map (340a, 340b, 342a, 342b, 344a, 344b, 410) includes feature sub-maps in the 3D model space for each of the at least two modalities, preferably each feature sub-map fed into the processing head.
10. The method (100) of any one of claims 1 to 9, wherein said one or more annotations comprise the presence of an object and / or a label of an object.
11. The method (100) of any one of claims 1 to 10, further comprising performing a fusion between said at least two modalities in said 3D model space.
12. 12. The method (100) of claim 1, wherein training the neural network (312, 322, 332, 420) comprises training the neural network (312, 322, 332, 420) multiple times, and at least one intervening training data from one or more of the virtual sensors and / or the physical sensors is omitted.
13. 13. The method (100) of any one of claims 1 to 12, wherein the first mapping comprises one or more first parameters, the second mapping comprises one or more second parameters, a processing head for obtaining one or more annotations comprises one or more third parameters, and training the neural network (312, 322, 332, 420) comprises end-to-end training to obtain the first, second, and third parameters.
14. A system (200, 300) for training a neural network (312, 322, 332, 420) to predict objects in the vicinity of a vehicle, said system (200, 300) configured to perform the method (100) according to any one of claims 1 to 13.
15. A computer readable storage medium storing program code comprising instructions which, when executed by a processor, performs the method (100) of any one of claims 1 to 13.