Method and device for providing training data for training a segmentation model

By adapting training datasets to match the target sensor's specifications and utilizing a labeling model and autoencoder, the method addresses the high effort in providing labeled data for segmentation models, achieving precise and efficient semantic segmentation of lidar and radar sensor data.

DE102024119321A1Pending Publication Date: 2026-01-08BAYERISCHE MOTOREN WERKE AG
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
DE102024119321
Authority / Receiving Office
DE · DE
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-07-08
Publication Date
2026-01-08

AI Technical Summary

Technical Problem

Training a segmentation model for semantic segmentation of detection points from lidar and/or radar sensors is associated with a high level of effort, particularly in providing labeled training data.

Method used

A device and method are employed to determine a target set of labeled training datasets by adapting initial datasets based on the target environmental sensor's specification, including adjustments for detection range, spatial resolution, and measurement accuracy, and using a labeling model and autoencoder to enhance the quality and efficiency of training data provision.

Benefits of technology

The solution enables precise and robust training of a segmentation model by efficiently providing labeled training data that aligns with the sensor's specifications, improving the accuracy and efficiency of semantic segmentation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

The invention relates to a device (200) for determining a target set (203) of labeled training data sets for training a segmentation model (220) configured to semantically segment a frame (130) of detection points (111) of a target environmental sensor (102), wherein the device (200) is configured to determine a target specification of the target environmental sensor (102); to adapt training data sets from an output set (201) of labeled training data sets acquired by at least one output environmental sensor (102) with an output specification, depending on the target specification; and to determine the target set (203) of labeled training data sets for training the segmentation model (220) based on the adapted training data sets.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] The invention relates to a method and a corresponding device designed to provide training data for teaching a segmentation model for the semantic segmentation of detection points of a lidar and / or radar sensor.

[0002] A vehicle capable of at least partial automation has one or more environmental sensors, each configured to acquire sensor data relating to the vehicle's surroundings. This sensor data can be used to operate a driving function that automates longitudinal and / or lateral movement of the vehicle. In particular, the vehicle may include at least one lidar sensor and / or radar sensor as an environmental sensor, configured to acquire a cloud and / or frame of detection points as sensor data during a measurement period.

[0003] The individual frames of detection points (for one or more measurement periods) can be analyzed to detect objects in the vehicle's vicinity and / or to determine object information relating to one or more objects in the vehicle's vicinity. For this purpose, a pre-trained segmentation model can be used, which is configured to assign each detection point of a frame a semantic class from a plurality of predefined classes. The detection points can then be grouped into one or more distinct objects based on their assigned class.

[0004] Training a segmentation model is typically associated with a relatively high level of effort, particularly in terms of providing labeled training data for training the segmentation model.

[0005] This document addresses the technical task of efficiently providing training data for training a precise and robust segmentation model for the semantic segmentation of frames of detection points.

[0006] The problem is solved by each of the independent claims. Advantageous embodiments are described, inter alia, in the dependent claims. It should be noted that additional features of a claim dependent on an independent claim, without the features of the independent claim itself or only in combination with a subset of the features of the independent claim, can constitute a separate invention independent of the combination of all features of the independent claim, which can be made the subject of an independent claim, a divisional application, or a subsequent application. This applies equally to technical teachings described in the description, which can constitute an invention independent of the features of the independent claims.

[0007] According to one aspect, a device for determining a target set of labeled training datasets for training a segmentation model is described. This training is intended to enable the segmentation model to semantically segment a frame of detection points from a target environmental sensor (in particular, a lidar and / or radar sensor). Within the framework of semantic segmentation, each detection point can be assigned to an object class from a multitude of different object classes. Examples of object classes are: a vehicle, a pedestrian, a motorcyclist, a stationary object, a moving object, and / or background or clutter.

[0008] The device is configured to determine a target specification for the target environmental sensor. The target specification can specify one or more parameters relating to the detection of points by the target environmental sensor. Each of these parameters can have a value or range of values ​​(specified in the target specification). The one or more parameters can include the spatial extent of the detection area, the spatial resolution of detection points, and / or the measurement accuracy.

[0009] An initial set of labeled training datasets can be provided, each acquired by at least one output environmental sensor with an output specification. The output specification can define a value or range of values ​​for one or more parameters. For at least one parameter, the target specification can have a range of values ​​or a value that is a subset of the output specification's range or is reduced compared to the output specification's value.

[0010] The device is configured to adapt the training datasets from the initial set of labeled training datasets depending on the target specification. Each training dataset from the set of labeled training data can comprise a training frame with a plurality of detection points. The device can be configured to adapt the training frame, in particular the plurality of detection points within the training frame, depending on the target specification in order to determine the corresponding adapted training dataset. The device can, in particular, be configured to adapt the training frame such that the adapted training frame is technically detectable by the target environmental sensor with the target specification.

[0011] The device is further configured to determine the target set of labeled training datasets for training the segmentation model based on the adapted training datasets. In particular, the adapted training datasets can be included in the target set as labeled training datasets.

[0012] The device can also be configured to train the segmentation model based on the set of labeled training data and using a learning method (e.g., using supervised learning).

[0013] By taking into account the target specification of the target environmental sensor when transforming training datasets, a target set of labeled training datasets can be efficiently provided for training a segmentation model for detection point frames captured by the target environmental sensor.

[0014] As previously explained, the target specification can define the spatial extent of the target environmental sensor's detection range, particularly with respect to radial distance and / or angular range (azimuth and / or elevation). The device can be configured to remove one or more detection points from the training frame that lie outside the target environmental sensor's detection range in order to determine the adapted training frame. Preferably, only the one or more detection points that lie within the target environmental sensor's detection range are retained.

[0015] The device can be configured to determine the object class of a detection point from a set of different object classes based on its label from the training frame. Similarly, the respective object class can be determined for each individual detection point from the training frame.

[0016] The device can further be configured to determine an object-class-specific detection range of the target environmental sensor, wherein the target environmental sensor can have a different detection range for each set of different object classes. In particular, an object-class-specific detection range can be determined for each individual detection point from the training frame.

[0017] The device can be configured to remove one or more detection points from the training frame that are outside the object-class-specific detection range of the target environmental sensor for the object class of the one or more detection points in order to determine the adapted training frame.

[0018] Alternatively or additionally, the target specification can specify the spatial resolution, particularly in relation to the radial direction and / or in relation to the (azimuth and / or height) angle at which detection points are captured by the target environmental sensor.

[0019] The device can be configured to remove one or more detection points from the training frame, depending on the spatial resolution of the target environmental sensor, in particular randomly and / or using a random number generator, in order to determine the adapted training frame. The number of detection points removed from the training frame can be increased as the spatial resolution of the target environmental sensor decreases.

[0020] The device can be configured to determine an object-class-specific spatial resolution of the target environmental sensor, whereby the target environmental sensor can have a different spatial resolution for each set of different object classes. Based on the labels of the detection points from the training frame, an object-class-specific spatial resolution of the target environmental sensor can thus be determined for each individual detection point.

[0021] The device can be configured to remove one or more detection points from the training frame, in particular randomly, depending on the respective object-class-specific spatial resolution of the target environmental sensor, in order to determine the adapted training frame.

[0022] Alternatively or additionally, the target specification can specify the measurement accuracy, in particular the signal-to-noise ratio, of the target environmental sensor. The device can be configured to superimpose detection points from the training frame with a noise signal to determine the adapted training frame. The noise signal, in particular its amplitude, can depend on the measurement accuracy of the target environmental sensor.

[0023] The aforementioned transformation measures enable the efficient provision of a particularly precise target set of labeled training datasets.

[0024] For a training frame, object information may be available for each detection point, indicating whether or not the respective detection point is assigned to a specific object (e.g., a particular vehicle). The object information of a detection point might, for example, specify an identifier for the object to which the detection point is assigned. Different detection points may be assigned to different objects (with different identifiers). The one or more transformation measures can take the object information of the individual detection points into account to further improve the quality of the training frame's transformation.

[0025] The target environmental sensor can be configured such that adjacent detection points exhibit a combined distance, based on the deviation of movement speeds and the spatial separation of the detection points, that is equal to or greater than a predefined distance threshold. Deleting one or more detection points (as part of a transformation process) can be achieved such that the detection points of the adapted training frame, in pairs, exhibit a combined distance that is equal to or greater than the distance threshold. This further improves the accuracy of the adapted training frame.

[0026] The target environmental sensor can be configured such that the frequency of detection points within a frame varies along the radial distance and / or along the (azimuth and / or elevation) angle, particularly the horizontal viewing angle. The frequency of detection points can thus be a function of the radial distance and / or the (azimuth and / or elevation) angle, particularly the horizontal viewing angle (which can be referred to as the frequency function). This frequency function can, if necessary, be derived from the target specification of the target environmental sensor.As part of a transformation process, detection points can be deleted in such a way that the detection points of the adapted training frame correspond to the frequency function of the target environmental sensor, and / or that the detection points of the adapted training frame are closer to the frequency function of the target environmental sensor than the detection points of the original training frame. This can further improve the accuracy of the adapted training frame.

[0027] According to another aspect, a further device for determining an additional set of labeled data sets for training a segmentation model is described. This model is designed to semantically segment a frame of detection points from an environmental sensor (in particular, the target environmental sensor). The device is configured to train a labeling model and a verification unit with an autoencoder using a set of labeled training data sets (in particular, the aforementioned target set of labeled training data sets).

[0028] The labeling model can include a first submodel that is designed to determine a feature vector with a specific number of features for each detection point of a (training) dataset. Each detection point can comprise a detection vector with measured values ​​for one or more measurands. Examples of measurands include: • the (radial) distance of the detection point from the environmental sensor; • the (azimuth and / or altitude) angle of the detection point (relative to the environmental sensor); • the (radial) velocity of the detection point; and / or • the intensity of the detection point.

[0029] The first sub-model can be configured to determine a feature vector with a multitude of features based on the detection vector of the detection point (and based on the detection vectors of the surrounding detection points). Typically, the number of features in the feature vector is higher, e.g., by a factor of 2 or more, in particular 10 or more, than the number of measured quantities in the detection vector.

[0030] The labeling model can further comprise a second sub-model configured to determine, based on the feature vectors for each detection point of the (training) dataset, a label that indicates the (most probable) object class of the respective detection point. Furthermore, the second sub-model can be configured to determine a confidence value for the confidence (e.g., reliability) of the determined label for each detection point.

[0031] The autoencoder of the verification unit can be configured to determine a reconstructed feature vector for each detection point of a data set. The autoencoder can comprise an encoder and a subsequent decoder. The device can be configured to train the autoencoder of the verification unit based on the set of labeled training data sets in such a way that an error function is reduced. This function depends on the deviation of the feature vectors for the individual detection points of the training data sets from the set of labeled training data sets from the feature vectors reconstructed by the autoencoder. As a result, the autoencoder can be made to represent a distribution of the detection points (or the feature vectors of the detection points) of the training data sets from the set of labeled training data sets.

[0032] The device is configured to label a set of unlabeled data records using the labeling model in order to provide an additional set of labeled data records. For each detection point of an unlabeled data record, a label can be determined that specifies the object class of the respective detection point.

[0033] The device can further be configured to determine a confidence value for the label determined by the labeling model for each detection point of an unlabeled dataset (from the set of unlabeled datasets). Furthermore, the device can be configured to include a detection point from an unlabeled dataset, along with its determined label, in the corresponding labeled dataset, or not, depending on the determined confidence value. The detection point may only be included if the confidence value is equal to or greater than a predefined confidence threshold. In this way, an additional set of automatically labeled datasets can be efficiently provided.

[0034] Furthermore, the device is configured to determine a weight for each detection point of the data sets from the additional set, based on the verification unit, during the training of the segmentation model. In particular, the device can be configured to determine a reconstructed feature vector for the feature vector of a detection point from a data set in the additional set using the autoencoder. The weight for the detection point can then be determined based on the deviation of the reconstructed feature vector from the original feature vector.

[0035] The device can be set up to train the segmentation model based on the additional set of labeled data sets, based on the corresponding set of weights for the individual detection points of the labeled data sets from the additional set, and using a learning method (e.g., using supervised learning).

[0036] The weighting can (due to the trained autoencoder) depend on how well the feature vector of a detection point matches the distribution of feature vectors of the detection points from the training datasets of the set of labeled training datasets used to train the labeling model. By considering the weighting during the training of the segmentation model, the accuracy of the segmentation model can be further improved.

[0037] According to another aspect, a method for determining a target set of labeled training datasets for training a segmentation model is described, wherein the segmentation model is trained to semantically segment a frame of detection points from a target environmental sensor (in particular, a lidar and / or radar sensor). The method includes determining a target specification of the target environmental sensor, as well as adapting training datasets from an initial set of labeled training datasets acquired by at least one output environmental sensor with an output specification, depending on the target specification. Furthermore, the method includes determining the target set of labeled training datasets for training the segmentation model based on the adapted training datasets.

[0038] According to another aspect, a method for determining an additional set of labeled datasets for training a segmentation model is described. The segmentation model is trained to semantically segment a frame of detection points from an environmental sensor (in particular, a lidar and / or radar sensor). The method includes training a labeling model and a verification unit with an autoencoder using a set of labeled training datasets. Furthermore, the method includes labeling a set of unlabeled datasets using the labeling model to provide an additional set of labeled datasets.The procedure further includes determining, for each detection point of the individual data sets from the additional set, a weighting of the respective detection point based on the verification unit, whereby the weighting of the respective detection point can be taken into account when training the segmentation model.

[0039] Another aspect described is a software (SW) program. The SW program can be configured to run on a processor (e.g., on a server) and thereby execute at least one of the procedures described in this document.

[0040] According to another aspect, a storage medium is described. The storage medium can include a software program that is configured to run on a processor and thereby execute at least one of the procedures described in this document.

[0041] It should be noted that the methods, devices, and systems described in this document can be used both alone and in combination with other methods, devices, and systems described in this document. Furthermore, any aspect of the methods, devices, and systems described in this document can be combined with one another in a variety of ways. In particular, the features of the claims can be combined with one another in a variety of ways. Features listed in parentheses are to be understood as optional features.

[0042] The invention will now be described in more detail using exemplary embodiments. Fig. 1a an exemplary vehicle with one or more lidar and / or radar sensors; Fig. 1b an exemplary frame of detection points from a lidar and / or radar sensor; Fig. 2a an exemplary device for training a segmentation model; Fig. 2b an exemplary labeling model; Fig. 3. An exemplary review unit for assessing the quality of automatically generated labels; and Fig. 4a and Fig. 4b each a flowchart of an exemplary procedure for providing training data for training a segmentation model for the semantic segmentation of detection point frames.

[0043] As stated at the beginning, this document deals with the efficient provision of training data for teaching a segmentation model for evaluating the detection points of a lidar and / or radar sensor. In this context, it shows Fig. 1a An exemplary vehicle 100 having at least one lidar and / or radar sensor 102 configured to acquire sensor data relating to the environment of the vehicle 100, in particular the environment in front of the vehicle 100. The lidar and / or radar sensor 102 may be configured to transmit a signal (e.g., a laser signal or a radio signal) and to receive a receive signal dependent on the transmitted signal. The receive signal may be based on a reflection of the transmitted signal from an object 105 in the environment of the vehicle 100. The lidar and / or radar sensor 102 may have a specific detection range 120, wherein the detection range 120 has a plurality of different detection directions.

[0044] The lidar and / or radar sensor 102 can be configured to detect a detection point 111 in a specific detection direction (i.e., for a specific azimuth angle and elevation angle) based on the transmitted signal if the transmitted signal is reflected by an object 105 or by the ground 110 on which the vehicle 100 is traveling. Conversely, the lidar and / or radar sensor 102 typically does not detect a detection point 111 if the transmitted signal is not reflected or not reflected strongly enough. Consequently, the lidar and / or radar sensor 102 can provide a cloud or frame 130 of detection points 111 for its detection range 120, indicating one or more objects 105 and, if applicable, the ground 110 in the vicinity of the vehicle 100 (see Fig. 1b).

[0045] The cloud 130 of detection points 111 detected by a lidar and / or radar sensor 102 can be aggregated into a (two-dimensional) grid of detection points 111. The grid can have N x M grid cells (e.g., with N, M equal to or greater than 50, 100, or 500). Each grid cell can contain a maximum of one detection point 111 (e.g., the detection point 111 with the highest intensity). Providing a grid of detection points 111 enables efficient analysis of the cloud or frame 130 of detection points 111 using neural networks, particularly convolutional neural networks (CNNs).

[0046] The aspects described in this document can generally be applied to clouds or frames 130 of detection points 111, regardless of whether they are in a raster (i.e., rasterized) or non-rasterized. For non-rasterized clouds or frames 130, a PointNet++ network can be used, for example.

[0047] A segmentation model with one or more neural networks can be trained for the semantic segmentation of an input raster (i.e., a frame 130) of detection points 111. Based on this input raster, the model generates an output raster that specifies a semantic object class from a defined set of distinct object classes for each raster cell and / or detection point 111. Objects 105 can then be identified by their respective object classes based on the output raster. Examples of object classes include: a passenger car, a truck, a pedestrian, an unknown object, etc. Furthermore, a class for "background" can be included in the classification (typically encompassing stationary objects (such as trees, houses, etc.) and clutter).

[0048] Training a segmentation model typically requires a relatively large amount of labeled training data, in particular a relatively large amount of labeled detection point rasters or detection point frames. Labeling training data (typically manually) is relatively time-consuming. This document describes one or more measures that enable the efficient provision of labeled training data for training a segmentation model for the semantic segmentation of the detection point frames 130 of an environmental sensor 102.

[0049] The environmental sensor 102 installed in a vehicle 100 can be selected from a set of different environmental sensors 102, each with different specifications regarding the detection of detection points 111. These specifications may differ in terms of, • the size of the detection range 120, e.g. in relation to the dimensions: distance, azimuth angle and / or elevation angle; and / or • the spatial resolution of detection points 111, e.g. in relation to the dimensions: distance, azimuth angle and / or elevation angle; • a distance at which different detection points 111 can be recognized as different; and / or • the measurement accuracy of the recorded detection points 111, e.g. in relation to the dimensions: distance, azimuth angle and / or elevation angle.

[0050] Fig. Figure 2a shows a block diagram of an exemplary device 200 for training a segmentation model for the semantic segmentation of the detection point frames 130 acquired by an environmental sensor 102 with a specific target specification. For training the segmentation model, an output set 201 of labeled training data acquired by at least one environmental sensor 102 with a specific output specification that differs from the target specification can be provided. In particular, the size of the detection area 120 for one or more dimensions and / or the spatial resolution for one or more dimensions and / or the accuracy of the acquired detection points of the output specification can be higher than that of the target specification.

[0051] The device 200 comprises a transformation unit 202, which is configured to generate a target set 203 of training data (i.e., training records) based on the output set 201 of labeled training data (i.e., training records). This target set can be used to train the segmentation model 220 for the frames 130 of the environmental sensor 102 with the target specification. For each data record of the output set 201, a corresponding data record of the target set 203 can be generated.

[0052] The transformation unit 202 is configured to take the target specification and / or the output specification into account when transforming the individual data sets. In particular, the size of the detection range 120 and / or the resolution of the target specification can be taken into account in order to generate a target set 203 of training data that could at least technically have been detected by the environmental sensor 102 with the target specification.

[0053] A data set of the initial set 201 typically comprises one frame 130 of detection points 111. The density of the detection points 111 and / or the spatial arrangement of the detection points 111 are based on the initial specification of the environmental sensor 102 with which the frame 130 was acquired. Each individual detection point 111 has an object class label.

[0054] Transformation unit 202 can be configured to perform one or more transformation actions to transform a data set of output quantity 201 into a corresponding data set of target quantity 203. The one or more transformation actions can depend on the target specification. Examples of such actions are: • Deleting one or more detection points 111 of frame 130 of a data set from the output set 201 that lie outside the detection range 120 of the environmental sensor 102 with the target specification. The labels of the detection points 111 can be used to account for class-specific differences in the detection range 120 of the environmental sensor 102; and / or • Deleting one or more detection points 111 of frame 130 of a data set from the initial set 201 to match the spatial resolution of the detection points 111 to the spatial resolution of the target specification. The one or more detection points 111 can be deleted randomly and / or using a random number generator. The number of detection points 111 deleted can depend on the spatial resolution of the target specification and can increase with decreasing resolution. Furthermore, the number of detection points 111 deleted can depend on the labels of the detection points 111 and can differ for different object classes; and / or • Applying additional noise to the individual detection points 111 of frame 130 of a data set from the initial set 201 in order to adapt the value accuracy of the detection points 111 to the measurement accuracy of the target specification. Noise, in particular Gaussian noise, can be added to the individual detection points 111 to apply this noise.

[0055] The detection range 120 of the environmental sensor 102 can differ for different object classes. For example, the detection range 120 can be larger for relatively large and / or metallic objects than for relatively small and / or non-metallic objects. When deleting a detection point 111 that lies outside the detection range 120 of the environmental sensor 102 with the target specification, the object class of detection point 110 can be determined based on the label of detection point 111. Furthermore, the detection range 120 of the environmental sensor 102 with the target specification can be determined for the identified object class.It can then be determined whether the detection point 110 lies within the (object class-dependent) detection range 120 of the environmental sensor 102 with the target specification (and can therefore be retained) or whether the detection point 110 lies outside the detection range 120 of the environmental sensor 102 with the target specification (and is therefore deleted).

[0056] Similarly, the spatial resolution of the environmental sensor 102 may depend on the object class in relation to the target specification. Consequently, the number of detection points 111 that are deleted to adjust their spatial resolution to match the target specification may depend on the labels of the detection points 111 and the object class they each indicate.

[0057] By performing one or more transformation operations, the individual data records of the output set 201 can be precisely transformed into corresponding data records of the target set 203, so that a target set of labeled training data can be provided that conforms to the target specification of the environmental sensor 102. The target set 203 of labeled training data can be used to train the segmentation model 220 (using a supervised learning method).

[0058] The in Fig. The device 200 shown in Figure 2a is further configured to train a labeling model 212 using the target set 203 of labeled training data. This labeling model is trained to automatically label the data records from a set 211 of unlabeled training data in order to provide an additional set 213 of labeled training data, which can then be used for further training of the segmentation model 220. This further improves the accuracy of the segmentation model 220.

[0059] The labeling model 212 may correspond to the segmentation model 220, which was trained (possibly alone) on the basis of the target set 203 of labeled training data. The labeling model 212 may be configured to perform a semantic segmentation of detection point frames 130 that were detected by an environmental sensor 102 configured according to the target specification. Within the framework of the semantic segmentation, each individual detection point of a detection point frame 130 can be assigned an object class (as a label).

[0060] The set 211 of unlabeled training data comprises a multitude of unlabeled datasets 214, where each dataset 214 corresponds to a detection point frame 130 or comprises a detection point frame 130. The labeling model 212 is configured to determine a label 217 for each dataset 214, where the label 217 specifies an object class for each detection point 111 of frame 130 of dataset 214. The labeling model 212 can also be configured to provide confidence values ​​219 for each dataset 214, indicating the confidence of the determined labels 217 for each detection point 111 of frame 130 of the respective dataset 214.

[0061] The device 200 can include a selection unit 215 configured to select, based on the confidence values ​​219 of the individual data sets 214, the detection points 111 of the data sets 214 that have a label 217 with a sufficiently high confidence. In particular, the detection points 111 of the data sets 214 can be selected for which the confidence value 219 is equal to or greater than a predefined confidence threshold.

[0062] Therefore, based on the set 211 of unlabeled training data, an additional set 213 of labeled training data can be determined, wherein the additional set 213 comprises the detection points 111 of the frames 130 of the data sets 214 (together with the respective determined labels 217) from the set 211 of unlabeled training data, for which the determined labels 217 have a sufficiently high confidence.

[0063] Fig. Figure 2b shows further details of an exemplary labeling model 212. The labeling model 212 can comprise a first submodel 231 and a subsequent second submodel 232. The first submodel 231 can be configured to determine the value of a feature vector 224 for each of the individual detection points 111 of the frame 130 of a data set 214. The frame 130 can, for example, comprise Q detection points 111, e.g., with Q ≥ 10, in particular Q ≥ 100. The individual detection points 111 can each have values ​​for one or more measured variables. Exemplary measured variables are: the distance of the detection point 111 from the environmental sensor 102; the (azimuth and / or elevation) angle of the detection point 111; the (radial)

[0064] The speed of movement of detection point 111; and / or the intensity of detection point 111. Each detection point 111 can thus have a detection vector with measured values ​​for one or more measured quantities.

[0065] The first submodel 231 can be configured to determine feature vectors 224 for the Q detection points 111. The individual feature vectors 224 typically have a dimension that is larger than the dimension of the detection vector, e.g., by a factor of 10 or more. The feature vector 224 determined for a specific detection point 111 is typically dependent on the detection points 111 in the vicinity of that specific detection point 111.

[0066] The second submodel 232 can be configured to determine the labels 217 and the confidence values ​​219 for the corresponding frame 130 of detection points 111, based on the frame 130 of feature vectors 224. A label 217 and a confidence value 219 can be determined for each detection point 111.

[0067] The in Fig. The device 200 shown in Figure 2a comprises a verification unit 216, which is configured to check for each of the individual detection points 111 of the data sets 214 from the set 211 and / or from the additional set 213 whether the individual detection points 111 (or the feature vectors 224 determined for the individual detection points 111) of the data sets 214 match the distribution of the detection points 111 (or the feature vectors 224) of the set 203 of labeled training data that was used for training the labeling model 212. If the verification unit 216 detects that a detection point 111 from a data set 214 does not match the distribution of the training data used, it can be concluded that the label 217 determined for this detection point 111 has a relatively high probability of containing errors and should therefore not be used, or at least less strongly, when training the segmentation model 220.

[0068] The verification unit 216 can be set up to determine a weighting or weight 218 for each of the individual detection points 111 of the data sets 214 from the additional set 213, which indicates how strongly the respective detection point 111 should be weighted when training the segmentation model 220.

[0069] The inspection unit 216 can, as in Fig. Figure 3 shows that (as part of an autoencoder) an encoder 311 is configured to determine a feature set 303 for a detection point 111 of frame 130 of a data record 214. Advantageously, the encoder 311 can take the feature vector 224 for detection point 111 as its input. Based on the feature vector 224 for detection point 111, the encoder 311 can then determine a feature set 304 for detection point 111. The feature set 304 typically has a significantly reduced dimension compared to the feature vector 224.

[0070] The verification unit 216 further comprises (as part of the autoencoder) a decoder 312, which is configured to determine a reconstructed feature vector 304 based on the feature set 303. The reconstructed feature vector 304 can be compared with the original feature vector 224, and the weight or weighting 218 can be determined based on this comparison. The verification unit 216 can, in particular, be configured to determine a deviation measure for the deviation of the reconstructed feature vector 304 from the original feature vector 224. The weight or weighting 218 can be increased with decreasing deviation or with decreasing value of the deviation measure.

[0071] The encoder 311 and the decoder 312 of the autoencoder of the verification unit 216 may have been pre-trained using the set 203 of labeled training data, in particular using an error function designed to reduce the deviation measure for the deviation of the respective reconstructed feature vector 304 from the original feature vector 224 for the individual detection points 111 from the individual data records 214 from the set 203. The error function and / or the deviation measure may each depend on the pairwise difference of the individual features of feature vector 224 and the reconstructed feature vector 304. For example, the mean squared difference of the features may be taken into account. The individual feature differences may optionally be weighted with an expected value or a mean value of the feature differences of the respective feature.

[0072] The segmentation model 220 (and, if applicable, the labeling model 212) can be further trained based on the additional set 213 of labeled data records 214 and taking into account the determined weights 218 for the individual data records 214. In this way, a precise and robust segmentation model 220 can be provided for the semantic segmentation of detection point frames 130.

[0073] Fig. Figure 4a shows a flowchart of a (possibly computer-implemented) procedure 400 for determining a target set 203 of labeled training datasets for training a segmentation model 220. The segmentation model 220 can be trained such that the segmentation model 200 is trained to semantically segment a frame 130 of detection points 111 of a target environmental sensor 102 (in particular a lidar and / or radar sensor).

[0074] The procedure 400 comprises determining 401 a target specification of the target environmental sensor 102. The target specification can specify value ranges or values ​​of one or more parameters for the acquisition of detection point frames 130 by the target environmental sensor 102.

[0075] Furthermore, the method 400 comprises the adaptation 402 of training data sets from an initial set 201 of labeled training data sets, acquired by at least one initial environmental sensor 102 with an initial specification, depending on the target specification. The target specification and the initial specification are different. In particular, the values ​​or value ranges of one or more acquisition parameters can differ.

[0076] Procedure 400 further includes determining 403 the target set 203 of labeled training datasets for training the segmentation model 220 based on the adapted training datasets. In particular, the adapted training datasets can each be included in the target set 203 as labeled training datasets.

[0077] Fig.Figure 4b shows a flowchart of another (possibly computer-implemented) method 410 for determining an additional set 213 of labeled data records 214 for training a segmentation model 220. The segmentation model 220 can be trained such that the segmentation model 200 is capable of semantically segmenting a frame 130 of detection points 111 of an environmental sensor 102 (in particular a lidar and / or radar sensor). Methods 400 and 410 can be combined to provide a particularly large set 203 and 213 of labeled data records for training the segmentation model 220.

[0078] The procedure 410 includes the training 411, using a set 203 of labeled training datasets. • a labeling model 212; and • a verification unit 216 with an autoencoder 311, 312.

[0079] The set 203 of labeled training datasets can be determined using procedure 400.

[0080] Method 410 further comprises labeling 412 a set 211 of unlabeled data records 214 using labeling model 212 to provide an additional set 213 of labeled data records 214. Labeling model 212 can, in particular, be designed as a segmentation model for semantic segmentation. The label 217 for a data record 214 can specify an object class for each of the individual detection points 111 of frame 130 of data record 214. In other words, a label 217 can be determined for each of the individual detection points 111 of frame 130 of data record 214, specifying the object class of the respective detection point 111.

[0081] Furthermore, the procedure 410 includes determining 413 for the individual data records 214 from the additional set 213, each based on the verification unit 216, a weighting 218 of the individual detection points 111 of the frame 130 of the respective data record 214 during the training of the segmentation model 220. The weighting 218 of the individual detection point 111 of the frame 130 of the labeled data record 214 can be used to weight the individual detection points 111 more or less strongly during the training of the segmentation model 220.

[0082] The measures described in this document enable the efficient training of a segmentation model 220 for semantic segmentation with high segmentation accuracy.

[0083] The present invention is not limited to the embodiments shown. In particular, it should be noted that the description and the figures are intended only to illustrate the principle of the proposed methods, devices, and systems by way of example.

Claims

[1] Device (200) for determining a target set (203) of labeled training data sets for training a segmentation model (220) trained to semantically segment a frame (130) of detection points (111) of a target environmental sensor (102), wherein the device (200) is configured, - to determine a target specification of the target environmental sensor (102); - To adapt training datasets from an initial set (201) of labeled training datasets acquired by at least one initial environmental sensor (102) with an initial specification, depending on the target specification; and - to determine the target set (203) of labeled training datasets for training the segmentation model (220) based on the adapted training datasets. [2] Device (200) according to claim 1, wherein - a training dataset from the set (201) of labeled training data comprises a training frame (130) with a plurality of detection points (111); and - the device (200) is set up to adapt the training frame (130), in particular the plurality of detection points (111) of the training frame (130), depending on the target specification, in order to determine the corresponding adapted training data set. [3] Device (200) according to claim 2, wherein the device (200) is configured to adapt the training frame (130) such that the adapted training frame (130) can be technically detected by the target environment sensor (102) with the target specification. [4] Device (200) according to one of claims 2 to 3, wherein - the target specification specifies a spatial extent of a detection range (120) of the target environmental sensor (102), in particular with respect to a radial distance and / or with respect to an angular range; and - the device (200) is configured to remove one or more detection points (111) from the training frame (130) that are outside the detection range (120) of the target environmental sensor (102) in order to determine the adapted training frame (130). [5] Device (200) according to claim 4, wherein the device (200) is configured, - to determine an object class of the detection point (111) from a set of different object classes using a label (217) of a detection point (111) from the training frame (130); - to determine an object-class-specific detection range (120) of the target environmental sensor (102); wherein the target environmental sensor (120) has a different detection range (120) for each set of different object classes; and - to remove one or more detection points (111) from the training frame (130) that are outside the object-class-specific detection range (120) of the target environmental sensor (102) for the object class of the one or more detection points (111) in order to determine the adapted training frame (130). [6] Device (200) according to any one of claims 2 to 5, wherein - the target specification specifies a spatial resolution, in particular with respect to a radial direction and / or with respect to an angle, at which detection points (102) are detected by the target environmental sensor (102); and - the device (200) is configured to remove one or more detection points (111) from the training frame (130), in particular randomly, depending on the spatial resolution of the target environment sensor (102), in order to determine the adapted training frame (130). [7] Device (200) according to claim 6, wherein the device (200) is configured to increase the number of detection points (111) removed from the training frame (130) as the spatial resolution of the target environment sensor (102) decreases. [8] Device (200) according to one of claims 6 to 7, wherein the device (200) is configured, - to determine an object class of one or more detection points (111) from a set of different object classes based on labels (217) of the one or more detection points (111) from the training frame (130); - to determine an object-class-specific spatial resolution of the target environmental sensor (102); wherein the target environmental sensor (120) has a different spatial resolution for the set of different object classes; and - depending on the object-class-specific spatial resolution of the target environment sensor (102), in particular randomly, to remove one or more detection points (111) from the training frame (130) in order to determine the adapted training frame (130). [9] Device (200) according to any one of claims 2 to 8, wherein - the target specification specifies a measurement accuracy, in particular a signal-to-noise ratio, of the target environmental sensor (102); - the device (200) is configured to superimpose detection points (111) from the training frame (130) with a noise signal in order to determine the adapted training frame (130); and - the noise signal, in particular an amplitude of the noise signal, depends on the measurement accuracy of the target environmental sensor (102). [10] Device (200) according to one of the preceding claims, wherein - the output specification and the target specification each specify one or more parameters relating to the detection of detection points (111) by the output environmental sensor (102) and by the target environmental sensor (102), respectively; - each of which has a range of values; - which include one or more parameters, in particular a spatial extent of the detection range (120), a spatial resolution of detection points (111) and / or a measurement accuracy; and - at least for one parameter, the target specification has a range of values ​​that is a subrange of the range of values ​​of the output specification. [11] Device (200) according to one of the preceding claims, wherein the device (200) is configured to train the segmentation model (220) on the basis of the set (203) of labeled training data and using a learning method. [12] Device (200) for determining an additional set (213) of labeled data sets (214) for training a segmentation model (220) configured to semantically segment a frame (130) of detection points (111) of an environmental sensor (102), wherein the device (200) is configured - based on a set (203) of labeled training datasets - to train a labeling model (212); and - to teach a verification unit (216) with an autoencoder (311, 312); - to label a set (211) of unlabeled data records (214) using the labeling model (212) in order to provide an additional set (213) of labeled data records (214); and - to determine a weighting (218) of individual detection points (111) of the respective data set (214) from the additional set (213) on the basis of the verification unit (216) when training the segmentation model (220) for each individual data set (214). [13] Device (200) according to claim 12, wherein - the labeling model (212) includes a first sub-model (231) which is designed to determine a feature vector (224) for each of the individual detection points (111) of a data set (214); - the autoencoder (311, 312) of the verification unit (216) is configured to determine a reconstructed feature vector (304) for each feature vector (224) of a detection point (111) of a data set (214); and - the device (200) is set up to train the autoencoder (311, 312) of the verification unit (216) on the basis of the set (203) of labeled training data sets in such a way that an error function which depends on a deviation of the feature vectors (224) for the individual detection points (111) of the training data sets of the set (203) of labeled training data sets from the feature vectors (304) reconstructed by the autoencoder (311, 312) is reduced. [14] Device (200) according to one of claims 12 to 13, wherein - the autoencoder (311, 312) of the verification unit (216) is configured to determine, for each feature vector (224) for a detection point (111) of a data set (214) from the additional set (213), a reconstructed feature vector (304); and - the device (200) is set up, - to determine a reconstructed feature vector (304) for a feature vector (224) for a detection point (111) of a data set (214) from the additional set (213) using the autoencoder (311, 312); and - to determine the weighting (218) for the detection point (111) of the data set (214) based on a deviation of the reconstructed feature vector (304) from the feature vector (224). [15] Device (200) according to any one of claims 12 to 14, wherein the device (200) is configured, - to determine a confidence value (219) for a label (217) determined by the labeling model (212) for a detection point (111) of an unlabeled data set (214) using the labeling model (212); and - to include the detection point (111) of the unlabeled data set (214) together with the determined label (217) depending on the determined confidence value (219) as the detection point (111) of the labeled data set (214) or not. [16] Device (200) according to any one of claims 12 to 15, wherein the device (200) is configured to train the segmentation model (220) based on the additional set (213) of labeled data sets (214), based on the corresponding set of weights (218) and using a learning method. [17] Method (400) for determining a target set (203) of labeled training datasets for training a segmentation model (220) trained to semantically segment a frame (130) of detection points (111) of a target environmental sensor (102), wherein the method (400) comprises, - Determine (401) a target specification of the target environmental sensor (102); - Adapting (402) training datasets from an initial set (201) of labeled training datasets acquired by at least one initial environmental sensor (102) with an initial specification, depending on the target specification; and - Determine (403) the target set (203) of labeled training datasets for training the segmentation model (220) based on the adapted training datasets. [18] Method (410) for determining an additional set (213) of labeled data sets (214) for training a segmentation model (220) trained to semantically segment a frame (130) of detection points (111) of an environment sensor (102), wherein the method (410) comprises, - Training (411) using a set (203) of labeled training datasets - a labeling model (212); and - a verification unit (216) with an autoencoder (311, 312); - Labeling (412) a set (211) of unlabeled records (214) using the labeling model (212) to provide an additional set (213) of labeled records (214); and - Determine (413) for individual detection points (111) of the individual data sets (214) from the additional set (213) each based on the verification unit (216) of a weighting (218) of the respective detection point 111 when training the segmentation model (220).

Citation Information

Patent Citations

  • Method and device for semantic segmentation and clutter detection

    DE102023123114A1