Method and apparatus for training a machine learning model for processing multimodal data
The method enhances geometric understanding in multimodal sensor data processing by using a machine learning model with encoders and pose regressive heads, improving feature extraction and pose estimation for applications like autonomous vehicles and SLAM systems.
Patent Information
- Application Number
- DE102024203137
- Authority / Receiving Office
- DE · DE
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-04-05
- Publication Date
- 2025-10-09
AI Technical Summary
Existing pretraining methods for multimodal sensor data lack a deep geometric understanding, which is crucial for applications requiring precise spatial relationships and characteristics, such as in autonomous vehicles and remote sensing.
A method involving a machine learning model with two encoders, a transformer, and pose regressive heads is used to compress and merge sensor data into a common feature embedding space, optimizing pose estimates by minimizing a loss function to enhance geometric awareness.
The method improves the quality of feature extraction across multiple sensor modalities, enabling precise pose estimation and geometric understanding, suitable for applications like autonomous vehicles and SLAM systems.
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
[0001] The invention relates to a method and a device for training a machine learning model for processing multimodal (sensor) data. State of the art
[0002] Image and / or other sensor data-based pretraining methods such as DINO and Masked Image Modeling play a central role in the development of advanced visual processing systems. These methods have enabled advances in generating features useful for a wide range of applications, from image recognition to image segmentation. Despite their successes, however, they face limitations when it comes to generating features with a deep geographic understanding. This is particularly relevant in areas where a precise understanding of spatial relationships and properties is critical, such as autonomous vehicles and / or remote sensing.Particularly given the increasing complexity and diversity of sensors in systems such as autonomous vehicles, including cameras, radar, LiDAR, IMU, GNSS, and more, the ability to leverage precisely aligned sensor data and develop models with a deep understanding of the physical world is becoming increasingly important.
[0003] Most modern pretraining methods rely on cross-modal masked autoencoders and contrastive learning to bridge the gap between different modalities. Such methods aim to develop models that can effectively work with data from different modalities (e.g., text, images, audio data) and are relevant for tasks where understanding and integrating information from multiple sources is crucial. Different approaches are used to learn data representations.
[0004] Cross-modal masked autoencoders are an extension of the autoencoder concept, where input data is partially masked (i.e., certain parts are intentionally hidden) and then reconstructed by the model. In a cross-modal context, data from different modalities is processed together. For example, such a model can receive radar data from a radar sensor and simultaneously acquired image data from a camera, with parts of the radar data or the image data masked. The goal is to correctly reconstruct the masked parts by drawing on information from the other modality.
[0005] Contrastive learning is an approach that aims to position similar (or positive) data points closer together and dissimilar (or negative) data points farther apart in an embedding space. When used in a cross-modal context, pairs or groups of data points from different modalities are used to train the model to recognize the correspondences between the modalities. For example, such a model can learn to pair radar data and image data by learning which radar data and image data belong together (similar) and which do not (dissimilar).
[0006] Although these methods represent progress, they lack the deep geometric understanding that is crucial for many applications. Thus, the challenge of generating geometry-aware features remains unsolved. As previously described, existing methods focus on inpainting (reconstructing or restoring) masked signals, neglecting the model-driven information about spatial alignments, rotations, and orientations that are essential for deep geometric understanding.
[0007] It is an object of the invention to provide an improved method and / or apparatus for training a machine learning model for processing multimodal (sensor) data.
[0008] The object is achieved by a method according to the features of patent claim 1. The object is achieved by a device according to the features of patent claim 10. Disclosure of the invention
[0009] According to a first aspect, there is provided a method for training a machine learning model for processing multimodal sensor data, the machine learning model comprising a first and a second encoder, a transformer, a first pose regressor head and a second pose regressor head, the method comprising the steps of: - Providing first sensor data acquired by a first sensor and second sensor data acquired by a second sensor; - compressing the first sensor data by the first encoder into a feature dimension with features from the first sensor data; - compressing the second sensor data by the second encoder into a feature dimension with features from the second sensor data; - merging the features of the first and second sensor data into a common feature embedding space by the transformer; - decoding the merged features from the joint feature embedding space to output a pose estimate for the first sensor by the first pose regressor head; - decoding the merged features from the joint feature embedding space to output a pose estimate for the second sensor by the second pose regressor head; - Minimizing a loss function to optimize the pose estimation for the first and second sensors; and - Deploying the trained machine learning model to process multimodal sensor data.
[0010] It is understood that the steps according to the invention, as well as other optional steps, do not necessarily have to be performed in the order shown, but can also be performed in a different order. Furthermore, additional intermediate steps can be provided. The individual steps can also comprise one or more substeps without thereby departing from the scope of the method according to the invention.
[0011] According to a second aspect, a device for training a machine learning model for processing multimodal data is provided, comprising the machine learning model, a first and a second encoder, a transformer, a first pose regressor head and a second pose regressor head, the device comprising an evaluation and computing device which is designed to carry out the following steps: - Providing first sensor data acquired by a first sensor and second sensor data acquired by a second sensor; - compressing the first sensor data by the first encoder into a feature dimension with features from the first sensor data; - compressing the second sensor data by the second encoder into a feature dimension with features from the second sensor data; - merging the features of the first and second sensor data into a common feature embedding space by the transformer; - decoding the merged features from the joint feature embedding space to output a pose estimate for the first sensor by the first pose regressor head; - decoding the merged features from the joint feature embedding space to output a pose estimate for the second sensor by the second pose regressor head; - Minimizing a loss function to optimize the pose estimation for the first and second sensors; and - Deploying the trained machine learning model to process multimodal data.
[0012] The statements made for the procedure apply accordingly to the system. It is understood that linguistic modifications of procedurally formulated features can be reformulated for the system according to common linguistic practice, without such formulations having to be explicitly listed here.
[0013] The use of sensor position for pretraining, a promising direction that remains unexplored, could play a key role in the development of pretraining methods that are not only cross-modal but also possess profound geometric awareness.
[0014] Unlike prior art methods, the present method does not require technically complex approaches such as masked image modeling, and can therefore be designed to be technically simpler and less complex. The processing of multimodal data, i.e., data originating from sensor sources of different sensor types, can be optimized by the present method because the training of the machine learning model trained for multimodal data processing is improved.
[0015] The basic concept of the method is to use a single real-world entity that connects all sensors to a single source of information available across all sensor modalities. In this case, this entity is the sensor pose. Estimating the pose as reliably as possible directly from the features encoded by the sensors requires a deep geometric and physical understanding. Furthermore, a semantic understanding of the world is needed. This geometric, physical, and semantic understanding can be trained into the machine learning model. Only the real-world pose is required as shared information between the sensors to scale large data sets for pretraining. The real-world pose information can be provided in a simple manner.
[0016] Each modality, i.e., each sensor, preferably has its own encoder, which compresses the raw information of the respective sensor measurement into a feature dimension. Transformer-based encoders are particularly suitable for this purpose, as they can be easily adapted to different modalities. A further transformer then summarizes the information from the features of each modality into a common feature embedding space. From this common embedding space, preferably n copies (in particular, one copy for each sensor) of a pose regressor head then decode this information and estimate a pose of the respective sensor with respect to a global coordinate system or relative between two sensors.
[0017] The method is used to improve the quality of feature extraction across multiple sensor modalities. In particular, it can optimize the solution of geometric regression problems. The method is preferably designed as a pretraining method for later use in multiple state estimation or perceptual backbones. However, the present method, or a machine learning model trained according to the method, can also be used directly as a pose estimation method.
[0018] The method, or the machine learning model trained using the method, can be used for multimodal foundation models or SLAM systems. The method, or the machine learning model trained using the method, can also be used in the field of parking assistance systems.
[0019] The machine learning model in this case comprises two encoders, one for each sensor data source. Each encoder is preferably specialized in compressing the data from its respective sensor and converting it into a feature dimension. This enables efficient and specific processing of different sensor data types. After initial processing by the encoders, the features from both sensor data sources are merged into a common feature embedding space by a transformer. This step enables the model to recognize and exploit relationships and dependencies between the features of the different sensors. The model has several separate pose regressor heads (corresponding to the number of sensors), one for each sensor.Pose regressor heads in a machine learning model are specialized network components designed to estimate the pose of an object or entity from the processed features. A "pose" refers to the spatial arrangement or orientation of a sensor or other object in space, which can be defined by parameters such as position, orientation, and possibly scale. Accurate pose estimation is important for tasks such as object detection, motion tracking, and interaction between physical and virtual objects. This configuration allows the model to make individual pose estimates for each sensor based on the combined features in the shared embedding space. By minimizing a loss function, the accuracy of the pose estimates for both sensors is improved. These optimization steps are for fine-tuning the model to enable accurate predictions.After training, the model is configured to process multimodal sensor data. It can be used for practical applications that require precise pose estimation from data from different sensors.
[0020] Datasets consisting of multiple sensors, such as nuScenes, can be used as training data. Other datasets can also be used.
[0021] In another aspect, the first encoder is associated with the first sensor and comprises a transformer-based encoder, a vision encoder, a radar encoder, or a lidar encoder. Alternatively or additionally (i.e., "and / or"), the second encoder is associated with the second sensor and comprises a transformer-based encoder, a vision encoder, a radar encoder, or a lidar encoder.
[0022] Encoders transform raw input data into a higher-quality, condensed feature representation (feature embedding), which is used for further processing steps within the model. Each of these encoder types uses specific architectures and techniques tailored to the properties of the respective data. A transformer-based encoder is based on the transformer architecture, which was originally developed for processing sequence-to-sequence tasks in natural language processing (NLP). This architecture uses mechanisms such as self-attention and positional encoding to capture relationships between elements in the input data regardless of their distance within the sequence. In a transformer-based encoder, the input data is converted into a set of feature vectors, which can then be used for tasks such as classification, regression, or other specific analyses.A vision encoder is designed for processing image data or visual information. It typically uses convolutional neural networks (CNNs) or newer architectures such as vision transformers (ViTs) to transform the raw pixel values of an image into a compact, informative feature representation. This representation captures important visual features such as edges, textures, shapes, and object relationships in the image. Vision encoders are fundamental to computer vision tasks such as object detection, image segmentation, and image classification. A radar or lidar encoder specializes in processing radar data, which is usually in the form of signals or point clouds and provides information about the distance, speed, and angular position of objects relative to the radar sensor. These encoders can be based on techniques such as deep learning, specially adapted CNNs, or networks optimized for point clouds such as PointNet.They aim to convert the raw radar signals into a feature representation that can be used for object detection and tracking, collision avoidance, and other radar-based applications.
[0023] In another aspect, the pose estimation for the first sensor and the pose estimation for the second sensor are performed with respect to a global coordinate system or in relative relation between the sensors.
[0024] "Pose estimation with respect to a global coordinate system" means that the position and orientation (pose) of an entity, i.e., from each sensor, is estimated relative to a fixed, common coordinate system. A global coordinate system provides a uniform reference frame against which all measurements and estimations are calibrated. This enables consistent interpretation of pose data across sensors, regardless of their individual position or orientation. "Relative pose estimation between sensors" refers to determining the position and orientation of one entity relative to another without the need for a global reference frame. This method considers the spatial relationships between the sensors and the observed object and can be useful for detecting and interpreting changes or movements of the object relative to the sensors.
[0025] In another aspect, minimizing the loss function to optimize the pose estimation for the first and second sensors comprises comparing with a ground truth value of a real pose of the first sensor and the second sensor.
[0026] Minimizing the loss function is a step that aims to improve the accuracy of the pose estimations performed by the model. This is done by comparing the poses predicted by the model with the actual, real-world poses of the objects as captured by the two sensors. The real-world poses serve as ground truth values. A loss function quantifies the difference, or error, between the model's predicted poses and the actual, real-world poses of the objects. The goal is to minimize this error. Optimizing the loss function is done by adjusting the internal (hyper-)parameters of the machine model (e.g., weights in a neural network) to improve the accuracy of the pose estimates. Ground truth values represent the true, real-world poses of the objects as captured by the sensors.They serve as a reference to evaluate the accuracy of the model's estimates. The comparison includes both the pose estimates from the first and second sensors, meaning the model should learn to estimate accurate poses for the sensors. This optimization method allows the model to detect and account for differences in the data or performance between the two sensors, resulting in improved overall system performance.
[0027] In another aspect, the pose estimation for the first and / or second sensor comprises a rotation estimation and a translation estimation. The rotation estimation involves solving a regression-by-classification problem.
[0028] In this case, pose estimation is preferably divided into a rotation estimation and a translation estimation.
[0029] In another aspect, solving the regression by classification problem comprises partitioning a rotation space into a voxel grid and classifying which voxel optimally represents a rotation; and regressing an actual rotation as an offset from a voxel center onto the rotation estimate.
[0030] Rotation estimation is preferably performed on a specific orthogonal group, so that the rotation estimation problem can be formulated as a two-stage problem. Rotation estimation is formulated as a regression by classification problem, i.e., the entire space of 3D rotations is divided into a voxel grid. It is then classified which voxel optimally represents the rotation. Subsequently, the actual rotation, as an offset from the voxel center, is regressed on the actual fine-grained rotation estimate. Since it may not be possible to define a unique world coordinate system, the problem can be viewed as a problem of relative pose estimation between sensor pairs. Subsequently, a global coordinate system can be reconstructed and unified for the sensors using coordinate transformation. The approach of RelPose++ (https: / / ar-xiv.org / 2019 / 06 / 09 / 13 / the-world-coordinate-system-relationships-and-relations-to-the-sensors) can be used for translation estimation.org / abs / 2305.04926) can be used.
[0031] In a further aspect, the first sensor comprises a lidar sensor and / or a radar sensor and / or an ultrasonic sensor and / or a camera sensor and / or an infrared sensor and / or an acceleration sensor and / or a Global Navigation Satellite System (GNSS) sensor. Alternatively or additionally, the second sensor comprises a lidar sensor and / or a radar sensor and / or an ultrasonic sensor and / or a camera sensor and / or an infrared sensor and / or an acceleration sensor and / or a Global Navigation Satellite System (GNSS) sensor.
[0032] It is understood that other types of sensors are also conceivable and that the list given is not to be understood as restrictive.
[0033] In a further aspect, the first and second sensors are arranged at different positions of a vehicle in order to detect the vehicle and / or a vehicle environment.
[0034] It goes without saying that the vehicle can also have more than two sensors, especially different ones. Alternatively, the first and second sensors can also be arranged at different positions on a robot, a medical device, an industrial machine, and / or a quality control system and / or a safety control system.
[0035] In a further aspect, a computer program with program code is claimed for executing at least parts of the method according to the invention in one of its embodiments when the computer program is executed on a computer. In other words, the invention provides a computer program (product) comprising instructions that, when executed by a computer, cause the computer to execute the method / steps of the method according to the invention in one of its embodiments.
[0036] In a further aspect, a computer-readable data carrier with program code of a computer program is proposed for executing at least parts of the method according to the invention in one of its embodiments when the computer program is executed on a computer. In other words, the invention relates to a computer-readable (storage) medium comprising instructions which, when executed by a computer, cause the computer to execute the method / steps of the method according to the invention in one of its embodiments.
[0037] The described designs and further training courses can be combined as desired.
[0038] Further possible embodiments, developments and implementations of the invention also include combinations of features of the invention described previously or below with regard to the embodiments that are not explicitly mentioned. Short description of the drawings
[0039] The accompanying drawings are intended to provide a further understanding of embodiments of the invention. They illustrate embodiments and, in conjunction with the description, serve to explain principles and concepts of the invention.
[0040] Other embodiments and many of the aforementioned advantages will become apparent upon review of the drawings. The elements illustrated in the drawings are not necessarily drawn to scale.
[0041] They show: Fig. 1 is a schematic flow diagram of an embodiment of the present method; and Fig. 2 is a schematic block diagram of an embodiment of the present device.
[0042] In the figures of the drawings, the same reference symbols designate the same or functionally equivalent elements, parts or components, unless otherwise stated.
[0043] Fig. 1 shows a schematic flow diagram of a method for training a machine learning model for processing multimodal sensor data.
[0044] In any embodiment, the method can be carried out at least partially by a device 100, which for this purpose can comprise several components not shown in detail, for example, one or more provision devices and / or at least one evaluation and computing device. It is understood that the provision device can be designed jointly with the evaluation and computing device or can be different from it. Furthermore, the device 100 can comprise a storage device and / or an output device and / or a display device and / or an input device.
[0045] The computer-implemented method comprises at least the following steps: In a step S1, first sensor data acquired by a first sensor and second sensor data acquired by a second sensor are provided. In a step S2, the first sensor data is compressed by the first encoder into a feature dimension with features from the first sensor data. In a step S3, the second sensor data is compressed by the second encoder into a feature dimension with features from the second sensor data. In a step S4, the features of the first and second sensor data are merged into a common feature embedding space by the transformer. In a step S5, the merged features from the common feature embedding space are decoded to output a pose estimate for the first sensor by the first pose regressor head. In a step S6, the merged features from the common feature embedding space are decoded to output a pose estimate for the second sensor by the second pose regressor head. In step S7, a loss function is minimized to optimize the pose estimation for the first and second sensors. Minimizing the loss function to optimize the pose estimation for the first and second sensors involves comparing it with a ground truth value of a real pose of the first sensor and the second sensor. In a step S8, the trained machine learning model is provided for processing multimodal sensor data.
[0046] In Fig. Figure 2 shows a block diagram of an embodiment of the present method and device. In particular, the schematic structure of the machine learning model 200 is shown.
[0047] The machine learning model 200 includes, for example, a first encoder 202, a second encoder 204, and a third encoder 206. The machine learning model 200 can, of course, also include only two encoders or more than three encoders.
[0048] The first encoder 202 is assigned to a first sensor 208. The first sensor 208 is a camera sensor. The second encoder 202 is assigned to a second sensor 210. The second sensor 210 is a radar sensor. The third encoder 206 is assigned to a third sensor 212. The third sensor 212 is another sensor that is different from the first and second sensors, for example, an acceleration sensor. The first sensor 208 provides first sensor data to the first encoder 202 for compression S2. The second sensor 210 provides second sensor data to the second encoder 204 for compression S3. The third sensor 212 provides third sensor data to the third encoder 206 for compression (equivalent to steps S2 and S3). The features of the first, second and third sensor data are combined into a common feature embedding space 214 by a transformer 216.
[0049] The machine learning model 200 further comprises a first pose regressor head 218, a second pose regressor head 220, and a third pose regressor head 222. The first pose regressor head 218 decodes S5 the merged features from the shared feature embedding space to output a pose estimate 224 for the first sensor 208. The second pose regressor head 220 decodes S6 the merged features from the shared feature embedding space to output a pose estimate 226 for the second sensor 210. The third pose regressor head 222 decodes (similar to S5 and S6) the merged features from the shared feature embedding space to output a pose estimate 228 for the third sensor 212.
Claims
[1] A method for training a machine learning model (200) for processing multimodal data, comprising the machine learning model (200), a first and a second encoder (202, 204), a transformer (216), a first pose regressor head (218) and a second pose regressor head (220), the method comprising the steps: - Providing (S1) first sensor data acquired by a first sensor (208) and second sensor data acquired by a second sensor (210); - compressing (S2) the first sensor data by the first encoder (202) into a feature dimension with features from the first sensor data; - compressing (S3) the second sensor data by the second encoder (204) into a feature dimension with features from the second sensor data; - merging (S4) the features of the first and second sensor data into a common feature embedding space (214) by the transformer (216); - decoding (S5) the merged features from the common feature embedding space (214) to output a pose estimate (224) for the first sensor (208) by the first pose regressor head (218); - decoding (S6) the merged features from the joint feature embedding space to output a pose estimate (226) for the second sensor (210) by the second pose regressor head (220); - minimizing (S7) a loss function for optimizing the pose estimation (224, 226) for the first and second sensors (208, 210); and - Providing (S8) the trained machine learning model (200) for processing multimodal data. [2] The method of claim 1, wherein the first encoder (202) is associated with the first sensor (208) and comprises a transformer-based encoder, a vision encoder, a radar encoder, or a lidar encoder; and / or wherein the second encoder (204) is associated with the second sensor (210) and comprises a transformer-based encoder, a vision encoder, a radar encoder, or a lidar encoder. [3] The method of claim 1 or 2, wherein the pose estimation (224) for the first sensor (208) and the pose estimation (226) for the second sensor (210) are performed with respect to a global coordinate system or in relative relation between the sensors (208, 210). [4] The method of any preceding claim, wherein minimizing the loss function to optimize the pose estimation for the first and second sensors (208, 210) comprises comparing with a ground truth value of a real pose of the first sensor (208) and the second sensor (210). [5] The method of any preceding claim, wherein the pose estimation (224, 226) for the first and / or second sensor (208, 210) comprises a rotation estimation and a translation estimation, wherein the rotation estimation comprises solving a regression by classification problem. [6] The method of claim 5, wherein solving the regression by classification problem comprises: partitioning a rotation space into a voxel grid and classifying which voxel optimally represents a rotation; and regressing an actual rotation as an offset from a voxel center onto the rotation estimate. [7] Method according to one of the preceding claims, wherein the first sensor (208) comprises a lidar sensor and / or a radar sensor and / or an ultrasonic sensor and / or a camera sensor and / or an infrared sensor and / or an acceleration sensor and / or a Global Navigation Satellite System, GNSS, sensor, and / or wherein the second sensor (210) comprises a lidar sensor and / or a radar sensor and / or an ultrasonic sensor and / or a camera sensor and / or an infrared sensor and / or an acceleration sensor and / or a Global Navigation Satellite System, GNSS, sensor. [8] Method according to one of the preceding claims, wherein the first and the second sensor (208, 210) are arranged at different positions of a vehicle in order to detect the vehicle and / or a vehicle environment. [9] Computer program with program code for carrying out at least parts of a method according to one of claims 1 to 8 when the computer program is executed on a computer. [10] Computer-readable data carrier with program code of a computer program for carrying out at least parts of a method according to one of claims 1 to 8 when the computer program is executed on a computer. [11] Device (100) for training a machine learning model (200) for processing multimodal data, comprising the machine learning model (200), a first and a second encoder (202, 204), a transformer (216), a first pose regressor head (218) and a second pose regressor head (220), the device (100) comprising an evaluation and computing device which is designed to carry out the following steps: - Providing (S1) first sensor data acquired by a first sensor (208) and second sensor data acquired by a second sensor (210); - compressing (S2) the first sensor data by the first encoder (202) into a feature dimension with features from the first sensor data; - compressing (S3) the second sensor data by the second encoder (204) into a feature dimension with features from the second sensor data; - merging (S4) the features of the first and second sensor data into a common feature embedding space (214) by the transformer (216); - decoding (S5) the merged features from the common feature embedding space (214) to output a pose estimate (224) for the first sensor (208) by the first pose regressor head (218); - decoding (S6) the merged features from the joint feature embedding space to output a pose estimate (226) for the second sensor (210) by the second pose regressor head (220); - minimizing (S7) a loss function for optimizing the pose estimation (224, 226) for the first and second sensors (208, 210); and - Providing (S8) the trained machine learning model (200) for processing multimodal data.