Computer-implemented method for training a machine learning model
Patent Information
- Application Number
- DE102024200355
- Authority / Receiving Office
- DE · DE
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-01-16
- Publication Date
- 2025-07-17
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
State of the art
[0001] In general, machine learning models (also referred to as machine learning models) can be trained using a training dataset with input data elements and an associated reference dataset that has reference labels for the input data elements. This can also be referred to as a data-driven approach. The reference labels of the reference dataset can be generated in a supervised or at least partially supervised manner by subjecting them to quality control. However, this generation of reference labels is complex and costly. To reduce the effort (e.g., cost), an algorithm can be used to generate the reference labels, which determines one or more reference labels for each input data element of the training dataset. This can also be referred to as autolabeling. However, in this case, hardly any statements can be made about the correctness of the reference labels.Furthermore, various such algorithms are generally available, and depending on the application, it is either impossible or almost impossible to evaluate which algorithm is better suited for a particular application (i.e., generates more accurate reference labels). Traditionally, an algorithm and / or a reference dataset generated by the algorithm is selected and used to train the machine learning model. To evaluate the accuracy of the reference labels, additional quality indicators can be introduced, although these are often unreliable. Disclosure of the invention
[0002] The present disclosure relates to a computer-implemented method for training a machine learning model. According to various embodiments, a machine learning model is trained not only with the reference labels from a single reference dataset, but with reference labels from multiple reference datasets.
[0003] This makes it possible to create a machine learning model that has high accuracy and / or high sensitivity with comparatively little effort (e.g., due to the use of auto-labeled reference data sets). It was also recognized that the algorithms for generating the reference labels have different strengths (e.g., the reference labels of some object attributes (e.g., with regard to an object type and / or a regression attribute) are more accurate compared to other algorithms) and weaknesses (e.g., the reference labels of some objects are less accurate compared to other algorithms). According to various embodiments, these strengths and weaknesses are taken into account when training the machine learning model. This can further increase the accuracy of the trained machine learning model.
[0004] Example 1 is a computer-implemented method for training a machine learning model, the method comprising: providing first training data comprising a training data set with a plurality of input data elements and a plurality of (two or more) reference data sets, wherein each reference data set of the plurality of reference data sets (each) comprises a plurality of reference labels (a reference label may also be referred to as ground truth information and / or a ground truth label), each reference label being assigned (exactly) one object of an input data element of the plurality of input data elements, and wherein each object of a plurality of (i.e., two or more than two) objects (of the plurality of input data elements) is assigned two or more than two reference labels,by assigning a respective reference label from two or more than two reference data sets of the plurality of reference data sets to a respective object of the plurality of objects; determining second training data comprising the training data set and reference labels from a plurality of reference data sets of the plurality of reference data sets by selecting one or more than one reference label from the two or more than two reference labels or no reference label for each object of the plurality of objects; and training the machine learning model using the second training data.
[0005] Clearly, reference labels are selected from different reference datasets, thereby increasing the accuracy of the trained machine learning model.
[0006] Example 2 is configured according to Example 1, wherein selecting the one or more reference labels for a respective object of the plurality of objects comprises: selecting exactly one reference label of the two or more reference labels according to a confidence metric; or determining a respective confidence value for each reference label of the two or more reference labels and selecting the reference labels of the two or more reference labels whose respective confidence value is greater than or equal to a confidence threshold.
[0007] In this way, the most trustworthy reference label(s) can be selected, thereby increasing the accuracy of the trained machine learning model.
[0008] Example 3 is configured according to Example 1, wherein selecting the one or more reference labels for a respective object of the plurality of objects comprises: selecting the one or more reference labels by filtering the two or more reference labels using a non-maximum suppression method according to a confidence metric and a similarity metric (e.g., intersection set over union set).
[0009] In this way, overlapping reference labels can be filtered out to select only the most trustworthy reference label from the overlapping reference labels, thereby increasing the accuracy of the trained machine learning model.
[0010] Example 4 is configured according to Example 1, wherein selecting the one or more reference labels for a respective object of the plurality of objects comprises: determining whether a plurality of reference labels of the two or more reference labels have one or more similarity values, each representing a similarity between two of the plurality of reference labels according to a similarity metric, that are greater than or equal to a similarity threshold; if it is determined that a plurality of reference labels have the one or more similarity values that are greater than or equal to the similarity threshold, selecting the plurality of reference labels or (exactly) one of the plurality of reference labels (e.g., according to a confidence metric).
[0011] In this way, reference labels can be selected when multiple reference labels detect the object, which can reduce the probability of detecting ghost objects.
[0012] Example 5 is configured according to Example 1, wherein selecting the one or more reference labels for a respective object of the plurality of objects comprises: determining whether a plurality of reference labels of the two or more reference labels have one or more similarity values, each representing a similarity between two of the plurality of reference labels according to a similarity metric, that are greater than or equal to a similarity threshold; if it is determined that a plurality of reference labels have the one or more similarity values that are greater than or equal to the similarity threshold, determining whether a number of the plurality of reference labels is greater than or equal to a number threshold; and if it is determined that the number of the plurality of reference labels is greater than or equal to the number threshold, selecting the plurality of reference labels or selecting a reference label of the plurality of reference labels (e.g.according to a confidence metric).
[0013] Example 6 is configured according to any one of examples 3 to 5, wherein the similarity metric is a similarity metric of objects with respect to at least one object attribute (e.g., a position and / or an orientation and / or an extent) of the respective object (e.g., an intersection set over a union set).
[0014] Example 7 is set up according to example 2 or 3, wherein the confidence metric takes into account an objectness of the respective object.
[0015] Example 8 is configured according to any one of examples 2, 3, or 7, wherein each object has a plurality of object attributes, and wherein each reference label indicates a respective reference object attribute value for each object attribute of the plurality of object attributes; wherein the (first and) second training data for each reference data set of the plurality of reference data sets comprises respective weighting data comprising a respective weighting factor for each object attribute of the plurality of object attributes;wherein selecting the one or more reference labels for a respective object using the confidence metric comprises: for each reference label of the two or more reference labels, determining a respective confidence value, determining a plurality of weighted confidence values by determining a respective weighted confidence value for each reference label of the two or more reference labels using the (determined) confidence value and the weighting factors of the reference data set having the reference label associated with the object attributes of the respective object, and selecting the one or more reference labels for the respective object using the plurality of weighted confidence values (e.g., the reference label for which the largest weighted confidence value is determined).
[0016] Illustratively, the weighting data of a reference data set can comprise a weighting vector which has a respective weighting factor for each object attribute of the plurality of object attributes.
[0017] By using the weighted confidence values, for example, the respective strengths and weaknesses of the labeling algorithms that are or were used to generate the reference datasets can be taken into account, which can further increase the accuracy of the trained machine learning model.
[0018] Example 9 is configured according to Example 1, wherein determining the second training data comprises selecting, for each object of the plurality of objects, each reference label of the two or more than two reference labels; wherein the machine learning model is configured to, in response to an input of an input data item, output an output data item comprising a prediction regarding at least one object of the input data item;wherein training the machine learning model using the second training data comprises iteratively training the machine learning model using the plurality of input data items of the training data set, wherein an iteration of training for a respective input data item of the plurality of input data items comprises: inputting the respective input data item to the machine learning model to generate an associated output data item, the output data item comprising a plurality of output data item points, a plurality of output data item points of which are associated with the at least one object;wherein, if the at least one object is an object of the plurality of objects to which two or more than two reference labels are associated, training the machine learning model comprises: for each output data element point of the plurality of output data element points, selecting (exactly) one reference label of the two or more than two reference labels according to a predefined criterion, and adapting the machine learning model using the reference label selected for each of the plurality of output data element points;
[0019] In this way, all reference labels (i.e., their union) can be used, and filtering occurs when determining the error function (e.g., in a center-head approach) and / or during post-processing of an output from the machine learning model (e.g., using non-maximum suppression). In this way, the sensitivity of the trained machine learning model can be significantly increased. While this can lead to the detection of ghost objects (i.e., false-positive objects), it can be advantageous when using the trained machine learning model in safety-critical applications, as the number of false-negative detections is also reduced. For example, in autonomous driving, it can be advantageous for the vehicle to stop at ghost objects or avoid them but have a lower number of false-negative object detections (e.g., failing to detect a pedestrian).
[0020] Example 10 is a computer-implemented method for training a machine learning model, wherein the machine learning model is configured to output, in response to an input of an input data item, an output data item comprising a prediction regarding at least one object of the input data item, the method comprising: providing training data comprising a training data set with a plurality of input data items and a plurality of (two or more) reference data sets, wherein each reference data set of the plurality of reference data sets (respectively) comprises a plurality of reference labels (a reference label may also be referred to as ground truth information and / or a ground truth label), each reference label being associated (exactly) with one object of an input data item of the plurality of input data items,and wherein each object of a plurality of (i.e., two or more than two) objects (of the plurality of input data elements) is assigned two or more than two reference labels, in that a respective object of the plurality of objects is assigned a respective reference label from two or more than two reference data sets of the plurality of reference data sets; iteratively training the machine learning model using the plurality of input data elements of the training data set, wherein an iteration of the training for a respective input data element of the plurality of input data elements comprises: inputting the respective input data element into the machine learning model to generate an associated output data element, wherein the output data element has a plurality of output data element points, of which a plurality of output data element points are assigned to the at least one object; wherein, if the at least one object is an object of the plurality of objects,to which two or more than two reference labels are assigned, comprising training the machine learning model: for each output data element point of the plurality of output data element points, selecting (exactly) one reference label of the two or more than two reference labels according to a predefined criterion, and adapting the machine learning model using the reference label selected for each of the plurality of output data element points.
[0021] Example 11 is configured according to example 9 or 10, wherein the predefined criterion indicates that the reference label of the two or more than two reference labels is selected that has the smallest similarity distance to an object attribute vector (e.g., the smallest distance to a reference point) of the at least one object.
[0022] Example 12 is configured according to any one of Examples 9 to 11, wherein each output data element point of the plurality of output data element points each comprises a predicted object attribute vector (e.g., predicted object attributes, such as a predicted anchor point and / or a predicted center point, a predicted object velocity, a predicted object extent, etc.) of the at least one object; and wherein the predefined criterion indicates that for a respective output data element point, the reference label of the two or more than two reference labels is selected that lies within a similarity threshold and that has a smallest similarity distance to the predicted object attribute vector.
[0023] Due to the iterative nature of training machine learning models, it may happen that in some iterations, a reference label has the smallest similarity distance to the predicted reference vector, while in other iterations, a different reference label has the smallest similarity distance to the predicted reference vector. This can clearly induce noise (which can also be considered augmentation), which can lead to improved training.
[0024] According to examples 11 and 12, multiple reference labels can be incorporated into the error function when training the machine learning model, thus reducing the number of false negative detections and increasing the sensitivity of the trained machine learning model.
[0025] Example 13 is configured according to example 9 or 10, wherein the (second) training data for each reference data set of the plurality of reference data sets comprises respective weighting data comprising a respective weighting factor for each object attribute of a plurality of object attributes that an object can have; wherein each output data element point of the plurality of output data element points comprises a respective predicted object attribute vector (e.g. comprising an object type, a predicted anchor point, a predicted center point, etc.) of the at least one object; and wherein selecting the reference label of the two or more than two reference labels for a respective output data element point of the plurality of output data element points according to the predefined criterion comprises: for each reference label of the two or more than two reference labels, determining a respective weighted similarity distance between the reference label (e.g.a reference vector thereof) and the predicted object attribute vector, determining a plurality of weighted similarity distances, wherein the weighted similarity distance is determined using the weighting factor data of the reference data set having the reference label, and selecting the reference label for the respective output data element point using the weighted similarity distances (e.g., the reference label having the smallest weighted similarity distance).
[0026] By using weighting factors, for example, the respective strengths and weaknesses of the labeling algorithms that are or were used to generate the reference datasets can be taken into account, which can further increase the accuracy of the trained machine learning model.
[0027] Optionally, a subsequent normalization can be carried out so that objects to which a comparatively large number of reference labels are assigned are not penalized more severely (within the framework of the error function) than objects to which a comparatively few reference labels are assigned.
[0028] Example 14 is configured according to any one of Examples 9 to 13, wherein the output data item is a grid cell representation (e.g., in bird's eye view or (in three-dimensional space) in voxel representation), and wherein each output data item point of the output data item is a grid cell of the grid cell representation.
[0029] Example 15 is configured according to any one of Examples 9 to 14, wherein each input data item of the plurality of input data items comprises a lidar point cloud and / or a camera image (e.g., an RGB image, an RGBD image, etc.), and / or a radar point cloud and / or ultrasonic sensor data.
[0030] Example 16 is configured according to any one of Examples 9 to 15, wherein, if the at least one object is not an object of the plurality of objects, training the machine learning model comprises: adapting the machine learning model using the reference label associated with the at least one object.
[0031] Example 17 is configured according to any one of Examples 9 to 16, wherein the machine learning model in inference is configured to generate, in response to an input of an input data item, a plurality of predictions regarding at least one object of the input data item, generate one or more than one filtered prediction regarding the at least one object by filtering the plurality of predictions using a non-maximum suppression method according to a confidence metric and a similarity metric (e.g., intersection set over union set), and output an output data item comprising the one or more than one filtered prediction.
[0032] Example 18 is configured according to any one of Examples 1 to 17, wherein providing the first training data comprises: providing the training data set with the plurality of input data elements; and (e.g., unsupervised) generating a respective reference data set of the plurality of reference data sets using an associated algorithm for (automatic, i.e., unsupervised) determination of reference labels (also referred to as a label algorithm).
[0033] Example 19 is configured according to Example 18, wherein the algorithm associated with a respective reference data set is different from the algorithms of the other reference data sets of the plurality of reference data sets.
[0034] Example 20 is configured according to any one of Examples 1 to 19, wherein each reference label indicates a respective reference object attribute value for each object attribute of a plurality of object attributes of the object with which the reference label is associated.
[0035] Example 21 is configured according to any one of Examples 1 to 20, wherein the machine learning model comprises a plurality of submodels, each submodel being configured to output, in response to an input of an input data item, an output data item comprising a respective prediction regarding at least one object of the input data item; wherein the machine learning model is configured to determine a plurality of predictions regarding at least one object of an input data item by inputting the input data item to each submodel of the plurality of submodels, and to make a prediction of the plurality of predictions using a confidence metric and / or a similarity metric (e.g.using a non-maximum suppression method according to the confidence metric and the similarity metric); wherein determining the second training data comprises selecting, for each object of the plurality of objects, each reference label of the two or more than two reference labels; wherein training the machine learning model comprises: for each reference data set of the plurality of reference data sets, training a respective submodel of the plurality of submodels.
[0036] This allows all reference labels to be used during training, and filtering can then be performed in inference (e.g., using non-maximum suppression). This can increase the sensitivity of the trained machine learning model (by reducing the number of false negative detections).
[0037] Example 22 is a method for determining a prediction regarding at least one object of an input data item, the method comprising: providing a plurality of machine learning models, each machine learning model trained using a same training data set and a respective reference data set of a plurality of reference data sets, wherein the respective reference data set is different from the other reference data sets of the plurality of reference data sets, wherein each machine learning model is configured to output, in response to an input of an input data item, an output data item comprising a respective prediction regarding at least one object of the input data item;Determining a plurality of predictions regarding at least one object of an input data item by inputting the input data item to each machine learning model of the plurality of machine learning models; and selecting a prediction of the plurality of predictions using a confidence metric and / or a similarity metric (e.g., using a non-maximum suppression method according to the confidence metric and the similarity metric).
[0038] Example 23 is configured according to example 21 or 22, further comprising: training the plurality of machine learning models (e.g., the plurality of submodels) using the training dataset and the plurality of reference datasets.
[0039] Example 24 is configured according to Example 23, further comprising: providing the training data set; and (e.g., unsupervised) generating a respective reference data set of the plurality of reference data sets using an associated algorithm for (automatic, i.e., unsupervised) determination of reference labels (also referred to as a label algorithm).
[0040] Example 25 is configured according to Example 24, wherein the algorithm associated with a respective reference data set is different from the algorithms of the other reference data sets of the plurality of reference data sets.
[0041] Example 26 is configured according to any one of Examples 21 to 25, wherein selecting the prediction comprises: selecting the prediction of the plurality of predictions having a largest confidence value according to the confidence metric.
[0042] Example 27 is configured according to any one of Examples 21 to 25, wherein selecting the prediction comprises: determining whether a plurality of predictions of the plurality of predictions have one or more similarity values, each representing a similarity between two of the plurality of predictions according to the similarity metric, that are greater than or equal to a similarity threshold; if it is determined that a plurality of predictions have the one or more similarity values that are greater than or equal to the similarity threshold, selecting the prediction of the plurality of predictions that has a greatest confidence value according to the confidence metric.
[0043] Example 28 is configured according to any one of Examples 21 to 25, wherein each machine learning model is associated with respective weighting data having a respective weighting factor for each object attribute of a plurality of object attributes; wherein selecting the prediction comprises: for each prediction of the plurality of predictions, determining a confidence value according to the confidence metric; determining a plurality of weighted confidence values by determining, for each prediction of the plurality of predictions, a respective weighted confidence value of the plurality of weighted confidence values using the confidence value of the prediction and the weighting factors associated with the machine learning model associated with the prediction for the object attributes of the at least one object; selecting the prediction using the plurality of weighted confidence values (e.g.the prediction for which the largest weighted confidence value is determined).
[0044] Example 29 is configured according to any one of Examples 21 to 25, wherein each machine learning model is associated with respective weighting data comprising a respective weighting factor for each object attribute of a plurality of object attributes; wherein selecting the prediction comprises: determining whether a plurality of predictions of the plurality of predictions have one or more similarity values, each representing a similarity between two of the plurality of predictions according to the similarity metric, that are greater than or equal to a similarity threshold; and, if it is determined that a plurality of predictions have the one or more similarity values that are greater than or equal to the similarity threshold: for each prediction of the plurality of predictions, determining a confidence value according to the confidence metric;Determining a plurality of weighted confidence values by determining, for each prediction of the plurality of predictions, a respective weighted confidence value of the plurality of weighted confidence values using the confidence value of the prediction and the weighting factors associated with the machine learning model associated with the prediction for the object attributes of the at least one object; selecting the prediction using the plurality of weighted confidence values (e.g., the prediction for which the largest weighted confidence value is determined).
[0045] By using the weighting factors according to examples 28 and 29, for example, the respective strengths and weaknesses of the labeling algorithms that are or were used to generate the reference datasets can be taken into account with respect to each object attribute, which can further increase the accuracy of the trained machine learning model.
[0046] Example 30 is configured according to any one of Examples 21 to 25, wherein each object type of a plurality of object types that an object of an input data item may have is associated with exactly one machine learning model of the plurality of machine learning models; wherein selecting the prediction comprises: selecting the prediction of the machine learning model to which the object type of the at least one object is associated.
[0047] Example 31 is configured according to any one of Examples 21 to 30, wherein each prediction of the plurality of predictions indicates a respective object type of the at least one object; and wherein the prediction of the plurality of predictions is selected only if a predefined number of machine learning models indicate a same object type of the at least one object.
[0048] Example 32 is configured according to Example 31, wherein each machine learning model is associated with respective weighting data having a respective weighting factor for each object type of a plurality of object types; wherein each prediction of the plurality of predictions indicates a respective object type of the at least one object; wherein the method comprises determining a weighted count value that is a sum of the weighting factors of the plurality of machine learning models for a same object type of the at least one object; and wherein the prediction of the plurality of predictions is selected only if the weighted count value is greater than or equal to a count threshold.
[0049] Example 33 is a computer program including instructions that, when executed by a processor, cause the processor to perform a method according to any one of Examples 1 to 32.
[0050] Example 34 is a computer-readable medium storing instructions that, when executed by a processor, cause the processor to perform a method according to any one of Examples 1 to 32.
[0051] Example 35 is a data processing unit configured to perform the method according to any one of Examples 1 to 32.
[0052] Example 36 is configured according to any one of examples 1 to 35, wherein at least one object is assigned exactly one reference label (i.e., the at least one object is not an object of the plurality of objects), wherein when determining the second training data, exactly one reference label is selected for the object or (e.g., in the case of a majority decision) exactly one reference label is not selected (so that the object is considered a ghost object).
[0053] Example 37 is configured according to any one of Examples 1 to 36, wherein the machine learning model is configured to (after training) filter out a prediction from a plurality of predictions in inference (in post-processing, for example using non-maximum suppression). This may, for example, resolve an ambiguity caused by many reference labels.
[0054] In the drawings, like reference characters generally refer to the same parts throughout the several views. The drawings are not necessarily to scale, emphasis instead generally being placed upon illustrating the principles of the invention. In the following description, various aspects are described with reference to the following drawings. Fig. 1A and Fig. 1B each show an exemplary input data element of a training dataset and associated reference labels from multiple reference datasets according to different aspects. Fig. Figure 2 shows a graphical illustration of the intersection set over the union of two reference labels as an example similarity metric. Fig. 3 shows an illustration of an exemplary grid cell representation and associated reference labels from two reference datasets according to different aspects. Fig. 4 shows an example of a first training data set and a second training data set generated therefrom according to various aspects. Fig. 5 shows a flowchart of a computer-implemented method for training a machine learning model according to various aspects.
[0055] The following detailed description refers to the accompanying drawings, which, by way of illustration, show specific details and aspects of this disclosure in which the invention may be practiced. Other aspects may be utilized, and structural, logical, and electrical changes may be made without departing from the scope of the invention. The various aspects of this disclosure are not necessarily mutually exclusive, as some aspects of this disclosure may be combined with one or more other aspects of this disclosure to form new aspects.
[0056] In summary, according to various embodiments, a computer-implemented method for training a machine learning model is provided as described in Fig. 5 and Example 1 is described. Various aspects of this procedure are described in more detail below.
[0057] The method may be performed by one or more computers having one or more data processing units. The term "data processing unit" may be understood as any type of entity that enables the processing of data or signals. The data or signals may, for example, be processed according to at least one (i.e., one or more than one) specific function performed by the data processing unit. A data processing unit may comprise or be formed from an analog circuit, a digital circuit, a logic circuit, a microprocessor, a microcontroller, a central processing unit (CPU), a graphics processing unit (GPU), a digital signal processor (DSP), a programmable gate array (FPGA) integrated circuit, or any combination thereof.Any other way of implementing the respective functions described in more detail herein may also be understood as a data processing unit or logic circuit arrangement. One or more of the method steps described in detail herein may be performed (e.g., implemented) by a data processing unit through one or more specific functions performed by the data processing unit.
[0058] According to various embodiments, the method is therefore particularly computer-implemented.
[0059] After training, the machine learning model can be applied to sensor data obtained from at least one sensor. The output of the machine learning model thus provides a result about a physical state of an environment of the at least one sensor and / or the at least one sensor itself, or the method can include using the output of the trained machine learning model, which it provides in response to an input of sensor data, as such a result.
[0060] The result about the physical state can in particular contain the information whether the specified event has occurred (e.g. a speaker has started speaking or a lane change has taken place).
[0061] In other words, the method may comprise inferring or predicting the physical state of an existing real object from measurements of physical properties (i.e., sensor data regarding the object) by means of the trained machine learning model (i.e., its output in response to the sensor data).
[0062] For example, after training, the machine learning model is used to generate a control signal for a robotic device by supplying it with sensor data regarding the robotic device and / or its environment. The term "robotic device" can be understood as referring to any technical system (having a mechanical part whose movement is controlled), such as a computer-controlled machine, a vehicle, a household appliance, a power tool, a manufacturing machine, a personal assistant, or an access control system. Thus, the method described herein can comprise generating control instructions for controlling a robotic device (and optionally further controlling the robotic device) using an output of the trained machine learning model. The machine learning model described herein can, for example, be an image classifier and / or perform vehicle environment recognition.
[0063] Various embodiments may receive and use time series of sensor data from various sensors, such as video, radar, LiDAR, ultrasound, motion, thermal imaging, etc., for example, to obtain sensor data regarding demonstrations or states of the system (robot and object or objects), configurations, and scenarios. Sensor data may be measured for periods of time or simulated (corresponding to one or more event times and in one or more predetermined events). The sensor data may be processed. This may include classifying the sensor data or performing semantic segmentation on the sensor data, for example, to detect the presence of objects (in the environment in which the sensor data was obtained). The trained machine learning model described herein may recognize these objects with increased accuracy.
[0064] The machine learning model can be any type of model based on machine learning, such as a reinforcement learning model (e.g., using Q-learning, Temporal Difference (TD), Deep Adversarial Networks, etc.) and / or a classification model (e.g., a linear classifier (e.g., a logistic regression classifier or a Naive Bayes classifier), a support vector machine, a decision tree, a boosted tree classifier, a random forest classifier, a neural network, or a nearest neighbor model, etc.). A neural network can be or include any type of neural network, such as a convolutional neural network (CNN), a variational autoencoder network (VNA), or a multivariate neural network (MNE). autoencoder network, VAE), a thinned autoencoder network.sparse autoencoder network (SAE), a recurrent neural network (RNN), a deconvolutional neural network (DNN), a generative adversarial network (GAN), a forward-thinking neural network, a sum-product neural network, transformer-based architectures, etc.
[0065] As in Fig. 4, according to various aspects, first training data 402 may be provided. The first training data 402 may include a training data set 404 having a plurality of input data elements. The first training data 402 may include a plurality of (N) reference data sets 406 (where N may be any integer greater than or equal to two). Each reference data set of the plurality of reference data sets may each include a plurality of reference labels (a reference label may also be referred to as ground truth information and / or a ground truth label). Here, each reference label may be associated with exactly one object of an input data element of the plurality of input data elements of the training data set.
[0066] It is understood that an input data item, as described herein, represents an input to the machine learning model and may comprise one or more than one data item. In some embodiments, the input data item may comprise exactly one data item (e.g., a camera image). In other embodiments, the input data item may comprise more than one data item (e.g., a lidar point cloud, a radar point cloud, and a camera image). The expression that one element, parameter, etc. "represents" another element, parameter, etc. may be understood to mean that these are related to one another, e.g., the element and / or parameter is a (e.g., unique, e.g., one-to-one) function of the other element and / or parameter.
[0067] In one illustrative example, the input data item may comprise or be a trace, and the object may be a peak of the trace. In another illustrative example, as in Fig. 1A, the input data item 102 may comprise or be an image, and the object 104 may be an object shown in the image. In the following, various examples refer to object detection (e.g., from camera images and / or lidar point clouds and / or radar point clouds). It is understood that this is for illustrative purposes and its description is exemplary of all other input data items.
[0068] At least one (e.g., each) reference data set 406(n) of the plurality of reference data sets 406 can be generated using an associated label algorithm (i.e., an algorithm for automatically generating reference labels). In this case, several different label algorithms (e.g., a video-based label algorithm, a lidar-based label algorithm, etc.) can be used to generate a corresponding number of reference data sets. As explained herein, the use of label algorithms (as part of autolabeling) can significantly reduce the effort (e.g., time expenditure, personnel expenditure, cost expenditure, etc.). In this case, two or more reference labels from different reference data sets can be assigned to an object. As described in Fig. 1A, the object 104 may be assigned a first reference label 106(1) from a first reference data set 406(1), a second reference label 106(2) from a second reference data set 406(2), and a third reference label 106(3) from a third reference data set 406(1). In this regard, it is understood that it may happen that a labeling algorithm does not recognize an object (e.g., false negative) and therefore does not generate a reference label for this object. Likewise, it may happen that a labeling algorithm generates a reference label for a (false positive) recognized object (also referred to as a ghost object), even though the input data element does not contain this object. As shown in Fig. As shown in Figure 1B, it may also happen that some reference labels overlap and others do not. No statements can be made regarding the accuracy of the reference labels.
[0069] According to various embodiments, as in Fig. 4, second training data 408 is generated, which comprises the training data set 404 and reference labels 410 from a plurality of reference data sets of the plurality of reference data sets 406. Fig. 5 shows a flowchart of an associated method 500 for computer-implemented training of the machine learning model. The method 500 includes (in 502) providing the first training data 402. The method 500 includes (in 504) determining the second training data 408. The method 500 includes (in 506) training the machine learning model using the second training data 408.
[0070] When generating the second training data 408, one or more reference labels are selected for each object to which more than one reference label is assigned. In some aspects, only a portion of the reference labels assigned to the object are selected. This may also be referred to as pre-filtering. In other aspects, all reference labels assigned to the object are selected, and filtering occurs elsewhere. This may also be referred to as post-filtering. If all reference labels are selected, the second training data corresponds to the first training data. These two options are described in more detail below.
[0071] Each reference label can be assigned a confidence value according to a confidence metric. For example, the labeling algorithm can output the confidence value when generating the reference label.
[0072] In a first embodiment, for an object to which two or more reference labels are assigned, exactly one reference label can be selected. This selection can be made based on the confidence value. For example, the reference label assigned the highest confidence value can be selected. The confidence value can indicate the object's object-like nature. Alternatively, the reference label can be selected arbitrarily or deterministically.
[0073] An object described herein may be characterized by a variety of object attributes. An object attribute may, for example, be a regression attribute (e.g., a bounding box) or a classification attribute (e.g., an object class / type).
[0074] A reference label can specify a reference object attribute vector, which has a respective reference object attribute value for each object attribute of the plurality of object attributes. Various aspects relate to a similarity distance between a reference label and an object attribute vector of an object. This similarity distance can be a similarity distance between the reference object attribute vector of the reference label and the object attribute vector of the object. The similarity distance can be determined according to a similarity metric. The similarity distance can take into account, for each object attribute of the plurality of object attributes, a similarity between an object attribute value of the object attribute specified by the object attribute vector of the object and a reference object attribute value of the object attribute specified by the reference object attribute vector.
[0075] As described herein, a labeling algorithm may have certain strengths such that some object attributes are detected with higher accuracy compared to other labeling algorithms (e.g., a lidar-based labeling algorithm may provide a more accurate distance estimate than other labeling algorithms and / or a video-based labeling algorithm may detect pedestrians better than other labeling algorithms). For example, the input data item may be a camera image representing a vehicle's surroundings; for example, a first labeling algorithm may be able to detect small objects (e.g., object type: pedestrians) with higher accuracy than a second labeling algorithm, but the second labeling algorithm may be able to detect larger objects (e.g., object type: trucks) with higher accuracy than the first labeling algorithm.According to various embodiments, these strengths and weaknesses of the labeling algorithms are taken into account by assigning to each reference data set for each object attribute of the plurality of object attributes a respective weighting factor which represents the strengths / weaknesses of the labeling algorithm by means of which the reference data set was generated.
[0076] In various aspects, the selection is described as an object attribute by way of example for the object type. It is understood that this is for illustrative purposes and that the selection may, for example, be made additionally or alternatively based on one or more other object attributes. For example, some object attributes may be detected by some sensors and not by others. In an illustrative example, the object attribute may be a speed: some sensors, such as radar sensors, can detect the radial object speed, whereas other sensors, such as lidar sensors and video sensors, cannot detect the speed directly (via the measuring principle). This object attribute may therefore be associated with a reference data set generated using a radar-based algorithm.This assignment of object attributes can also be used when two objects can be associated with each other. Various weighting factors are also described here. It is understood that, with respect to such an object attribute, different weighting factors can be assigned to the reference data sets (for example, a reference data set generated using the radar-based algorithm can be assigned a larger weighting factor with respect to the object attribute speed than a reference data set generated using a lidar-based algorithm and / or a reference data set generated using a video-based algorithm).
[0077] In one example, for each reference label, a respective weighted confidence value may be determined using the confidence value and the weighting factors associated with the object attributes of the object of the reference data set having the reference label.
[0078] In another example, the reference label of the reference data set that has the greatest weighting factor for a specific object attribute (e.g., the object type) of the object can be selected, regardless of the confidence value. For example, each object type can be assigned exactly one reference data set, and the reference label of the assigned reference data set is selected depending on the object type of the object. If the assigned reference data set does not have a reference label for an object, either a different reference label or no reference label can be selected. If no reference label is selected, the number of ghost objects can be reduced. It is understood that the object type is an example of an object attribute.
[0079] In a second embodiment, the two or more reference labels associated with an object can be filtered using a non-maximum suppression method. With non-maximum suppression, the reference label with the largest (e.g., weighted) confidence value can be selected iteratively. Then, for all other reference labels of the two or more reference labels, a respective similarity value representing a similarity to the selected reference label can be determined according to a similarity metric. This similarity metric can be any type of similarity metric that can describe the similarity of two reference labels. For example, for the exemplary example that the input data element 102 comprises the image with the object 104, the similarity metric can be a similarity with respect to an object attribute specified by the respective reference label.The object attribute may include one or more of the following attributes: a position of the object 104, an orientation of the object 104, an extent of the object 104, etc. In an illustrative example, the similarity metric may be an intersection set over a union of two reference labels. This is illustrated graphically in . Fig. 2 for a similarity value of 200 determined in this way for the first reference label 106(1) and the second reference label 106(2). The intersection set over the union (IoU) for the sets A and B can be described as IoU(A,B)=|A∩B||A∪B|.
[0080] With non-maximum suppression, all reference labels whose similarity value to the selected reference label is greater than or equal to a similarity threshold can then be filtered out (i.e., not selected). If reference labels of the two or more reference labels are then neither selected nor filtered out, the next reference label with the largest (e.g., weighted) confidence value can be selected from the remaining reference labels in a subsequent iteration. Then, for all other remaining reference labels, a similarity value for the similarity to the reference label selected in this iteration can be determined and then filtered accordingly. It is understood that non-maximum suppression is exemplary and any other similarity metric (e.g., an L1 norm, an L2 norm, etc.) can be used.
[0081] In a third embodiment, for an object to which two or more than two reference labels are assigned, a respective similarity value can be determined for each pair of reference labels for all reference labels of the two or more than two reference labels, and the reference labels can be selected that have similarity values to one another that are greater than or equal to a similarity threshold. Optionally, it can be determined whether a number of reference labels that have a similarity to one another that is greater than or equal to the similarity threshold is greater than a number threshold, and then the reference labels or one of these reference labels (e.g., the reference label with the largest (e.g., weighted) confidence value) can be selected if their number is greater than or equal to the number threshold. Illustratively, one or more than one reference labels can only be selected if a minimum number of reference labels are similar (e.g.,B. overlap each other). As an example, N can be 10, so that the plurality of reference data sets comprises ten reference data sets. The number threshold can be seven, for example, meaning that one or more reference labels are only selected if the number of reference labels with a similarity to each other is greater than or equal to the similarity threshold, which is greater than or equal to seven. Optionally, the weighting factor can be taken into account, so that the reference label is included in a weighted number of reference labels not with the value 1, but with the weighting factor.
[0082] As an alternative to the above pre-filtering, all reference labels of the multiple reference labels for the second training data can be selected and post-filtering can be performed during training and / or the algorithm by non-maximum suppression post-processing.
[0083] In general, the machine learning model may be configured to output, in response to an input of an input data item 102 of the training data set 404, an output data item comprising a prediction regarding at least one object 104 of the input data item 102. During training, in each iteration, an input data item of the plurality of input data items of the training data set 404 may be input into the machine learning model to generate a respective output data item.
[0084] An output data element may comprise one or more than one data element, as explained above for the input data element. A data element of the output data element may comprise a plurality of output data element points, of which a plurality of output data element points are associated with the at least one object 104. In one example, the output data element may be an image (e.g., a segmentation image), and an output data element point may be a pixel, wherein a plurality of pixels may be associated with the object 104. In another example, the output data element may be a measurement curve, and an output data element point may be a measurement point of the measurement curve, wherein a plurality of measurement points may be associated with the object 104 (e.g., a measurement peak). In yet another example, the output data element may be a grid cell representation (e.g., an occupancy grid) (e.g.,in bird's eye view) and an output data element point may be a grid cell of the grid cell representation, wherein multiple grid cells may be associated with the object 104.
[0085] For illustrative purposes, various aspects described above are described with respect to an object. It is understood that a reference label may be present even if the object specified by the reference label does not exist (i.e., is a (false positive) ghost object).
[0086] During the training of the machine learning model, several reference labels can then be included in the error function according to various aspects. For example, a reference label can be selected for each output data element point of the plurality of output data element points assigned to the object. The reference label for an output data element point can be selected according to a predefined criterion. The predefined criterion can, for example, specify that the reference label (of the two or more than two reference labels) is selected that has the smallest distance to a reference point of the object. According to various aspects, each output data element point can specify a predicted reference point (e.g., a predicted anchor point, a predicted center point, etc.) of the object.In this case, the predefined criterion may, for example, specify that the reference label (of the two or more reference labels) with the smallest distance to the predicted reference point be selected. Optionally, the weighting factors of the reference labels may be taken into account. For example, the respective distance of the reference labels may be weighted according to the weighting factor of the respective reference label, and the reference label (of the two or more reference labels) with the smallest weighted distance to the predicted reference point may be selected.
[0087] Fig. 3 shows an exemplary output data element as a grid cell representation and associated reference labels from two reference data sets according to various aspects. In this example, grid cells 3,2; 4,2; 3,3; 4,3; 5,3; 4,4; and 5,4 can each detect the object and each predict a reference point (e.g., the anchor point) of the object. All reference labels that lie within a predefined distance to the predicted reference point can be determined, and then the reference label whose (e.g., weighted) distance to the predicted reference point is the smallest can be selected. For example, in Fig.3 for the grid cells 3,2; 4,2; and 3,3 it can be determined that only the first reference label 106(1) lies within the predefined distance to the predicted reference point, and this first reference label 106(1) can be selected. Correspondingly, for the grid cells 5,3; 4,4; and 5,4 it can be determined that only the second reference label 106(2) lies within the predefined distance to the predicted reference point, and this second reference label 106(2) can be selected. For the grid cell 3,4 it can be determined that both the first reference label 106(1) and the second reference label 106(2) lie within the predefined distance to the predicted reference point, and the reference label of these two whose (e.g. weighted) distance (e.g. the distance of its reference point) to the predicted reference point is the smallest can be selected.In this way, the prediction of false-negative object detections can be significantly reduced and thus the sensitivity of the machine learning model can be increased.
[0088] In the inference of the machine learning model, ambiguities can be resolved by predicting multiple overlapping references using a non-maximum suppression method and / or a center-head method.
[0089] During training, for example, the following cases can occur: In case 1, overlapping references are automatically filtered out during inference using non-maximum suppression. During training, overlapping references also flow more into the error function due to an increased spatial distribution through grid cell reference points. In total, more reference points are therefore included in the error function during training than without the use of multiple reference labels. In case 2, there is one reference label from the first reference data set and no reference label from the second reference data set. In this case, the object can be a real object or a ghost object. In this case, the object is simply included in the error function using the reference label from the first reference data set. In case 3, the reference labels of the reference data sets do not overlap, so they are included as two objects in training.As a result, the trained machine learning model has fewer false negative detections (i.e., missing detections).
[0090] In a variant of the post-filtering, the machine learning model can have a plurality of submodels, each submodel being configured to output an output data item in response to an input of an input data item, which output data item has a respective prediction regarding at least one object of the input data item. For each reference data set of the plurality of reference data sets, exactly one submodel of the plurality of submodels can be trained. Thus, the machine learning model can be configured to determine a plurality of predictions regarding at least one object of an input data item by inputting the input data item to each submodel of the plurality of submodels. The machine learning model can then select one of the plurality of predictions using the confidence metric and / or the similarity metric. This can be done according to any suitable ensemble decision (e.g., aweighted) majority decision). For example, the prediction can be selected using non-maximum suppression, as explained above for pre-filtering. The weighting factors can also be taken into account here. For example, each submodel can have an associated weighting factor, and the prediction output by this submodel can be weighted according to the associated weighting factor. As explained above, each object attribute can be assigned a respective weighting factor. Consequently, each submodel can be assigned a respective weighting vector that has a weighting factor for each object attribute.
[0091] Illustratively, the weighting described here can, in some aspects, be performed on an object vector-by-attribute basis. The network output can also be viewed as a kind of probability over the number of detections of the N submodels through an ensemble prediction (i.e., different submodels provide different predictions).
[0092] As explained above, when multiple overlapping reference labels (i.e., reference labels whose similarity is greater than the predefined similarity threshold) exist, either exactly one reference label can be selected or several (e.g., all) of the overlapping reference labels can be used. This also leads to fewer missed detections (i.e., a lower false negative rate) of the trained machine learning model. Optionally, only overlapping reference labels can be used here, and / or a confidence threshold can be offered to reduce sensitivity. Depending on various aspects, the overlapping reference labels can be used in a weighted manner.
[0093] As clearly described herein, according to various embodiments, reference labels from multiple reference datasets can be used when training a machine learning model to increase the sensitivity of the (trained) machine learning model. For this purpose, the reference labels of all (available) reference datasets can be pre-filtered or post-filtered. In this case, the following can be done: - for a low (but still higher) sensitivity, only reference labels that are similar to each other are used in order to have a low false-positive rate; - for high sensitivity, take all reference labels and filter similar reference labels (e.g. using non-maximum suppression); - for a very high sensitivity, all references are used (without pre-filtering) and filtering of references is then done in inference.
[0094] As explained above, for very high sensitivity, either the reference labels can be selected during training or the machine learning model can have a large number of submodels and the predictions of all submodels can be filtered in inference.
[0095] As described herein, in each of these variants, weighting (according to the weighting factors) can be applied to account for strengths and weaknesses of the labeling algorithms used to generate or have generated the reference datasets.
[0096] According to various aspects, using the outputs of the N submodels, a statistic on estimates can be generated. This allows, for example, to predict distributions that may represent a probability or a certainty (also referred to as confidence). The selection of one or more reference labels can then be performed using this information.
Claims
[1] A computer-implemented method (500) for training a machine learning model, the method (500) comprising: • Providing (502) first training data (402) comprising a training data set (404) with a plurality of input data elements and a plurality of reference data sets (406), wherein each reference data set of the plurality of reference data sets (406) comprises a plurality of reference labels, each reference label being assigned to an object (104) of an input data element (102) of the plurality of input data elements, and wherein each object of a plurality of objects is assigned two or more than two reference labels in each case, by assigning a respective reference label from two or more than two reference data sets of the plurality of reference data sets (406) to a respective object of the plurality of objects; • Determining (504) second training data (408) comprising the training data set (404) and reference labels (410) from a plurality of reference data sets of the plurality of reference data sets (406) by selecting one or more than one reference label of the two or more than two reference labels or no reference label for each object of the plurality of objects; and • Training (506) the machine learning model using the second training data (408). [2] The method (500) of claim 1, wherein selecting the one or more reference labels for a respective one of the plurality of objects comprises: Selecting exactly one of the two or more than two reference labels according to a confidence metric; or Determining a respective confidence value for each reference label of the two or more than two reference labels and selecting the reference labels of the two or more than two reference labels whose respective confidence value is greater than or equal to a confidence threshold; or Selecting the one or more reference labels by filtering the two or more reference labels using a non-maximum suppression method according to a confidence metric and a similarity metric; or Determining whether a plurality of reference labels of the two or more than two reference labels have one or more similarity values, each of which represents a similarity between two of the plurality of reference labels according to a similarity metric, that are greater than or equal to a similarity threshold; if it is determined that a plurality of reference labels have the one or more similarity values that are greater than or equal to the similarity threshold, selecting the plurality of reference labels or selecting exactly one reference label of the plurality of reference labels; or Determining whether a plurality of reference labels of the two or more than two reference labels have one or more similarity values, each representing a similarity between two of the plurality of reference labels according to a similarity metric, that are greater than or equal to a similarity threshold; if it is determined that a plurality of reference labels have the one or more similarity values that are greater than or equal to the similarity threshold, determining whether a number of the plurality of reference labels is greater than or equal to a number threshold; and if it is determined that the number of the plurality of reference labels is greater than or equal to the number threshold, selecting the plurality of reference labels or selecting a reference label of the plurality of reference labels. [3] Method (500) according to claim 2, wherein each object has a plurality of object attributes, and wherein each reference label indicates a respective reference object attribute value for each object attribute of the plurality of object attributes; wherein the second training data for each reference data set of the plurality of reference data sets (406) comprises respective weighting data comprising a respective weighting factor for each object attribute of the plurality of object attributes; wherein selecting the one or more reference labels for a respective object using the confidence metric comprises: • for each of the two or more reference labels, determining a respective confidence value, • Determining a plurality of weighted confidence values by determining a respective weighted confidence value for each reference label of the two or more than two reference labels using the confidence value and the weighting factors associated with the object attributes of the respective object of the reference data set having the reference label, and • Selecting one or more reference labels for the respective object using the plurality of weighted confidence values. [4] Method (500) according to claim 1, wherein determining the second training data (408) comprises selecting, for each object of the plurality of objects, each reference label of the two or more than two reference labels; wherein the machine learning model is configured to output, in response to an input of an input data item, an output data item comprising a prediction regarding at least one object of the input data item; wherein training (504) the machine learning model using the second training data (408) comprises iteratively training the machine learning model using the plurality of input data elements of the training data set (404), wherein an iteration of the training for a respective input data element of the plurality of input data elements comprises: • inputting the respective input data item into the machine learning model to generate an associated output data item, wherein the output data item comprises a plurality of output data item points, of which a plurality of output data item points are associated with the at least one object; • wherein, if the at least one object is an object of the plurality of objects to which two or more than two reference labels are associated, training the machine learning model comprises: ◯ for each output data element point of the plurality of output data element points, selecting one of the two or more than two reference labels according to a predefined criterion, and ◯ Adapt the machine learning model using the reference label selected for each of the multiple output data element points. [5] Method (500) according to claim 4, wherein each output data element point of the plurality of output data element points each comprises a predicted object attribute vector of the at least one object; and wherein the predefined criterion specifies that for a respective output data element point, the reference label of the two or more than two reference labels is selected which lies within a similarity threshold and which has a smallest similarity distance to the predicted object attribute vector. [6] Method (500) according to claim 4, wherein the second training data for each reference data set of the plurality of reference data sets (406) comprises respective weighting data comprising a respective weighting factor for each object attribute of a plurality of object attributes that an object may have; wherein each output data element point of the plurality of output data element points comprises a respective predicted object attribute vector comprising a respective predicted object attribute value for each object attribute of the plurality of object attributes; and wherein selecting the reference label of the two or more than two reference labels for a respective output data element point of the plurality of output data element points according to the predefined criterion comprises: • for each reference label of the two or more than two reference labels, determining a respective weighted similarity distance between the reference label and the predicted object attribute vector, wherein the weighted similarity distance is determined using the weight data of the reference data set having the reference label, and • Selecting the reference label for each output data element point using the weighted similarity distances. [7] Method (500) according to claim 1, wherein the machine learning model comprises a plurality of submodels, each submodel being configured to output, in response to an input of an input data item, an output data item comprising a respective prediction relating to at least one object of the input data item; wherein the machine learning model is configured to determine a plurality of predictions regarding at least one object of an input data item by inputting the input data item to each submodel of the plurality of submodels, and to select a prediction of the plurality of predictions using a confidence metric and / or a similarity metric; wherein determining the second training data comprises selecting each reference label of the two or more than two reference labels for each object of the plurality of objects; and wherein training the machine learning model comprises: for each reference data set of the plurality of reference data sets (406), training a respective submodel of the plurality of submodels. [8] Data processing unit configured to carry out a method (500) according to any one of claims 1 to 7. [9] A computer program comprising instructions which, when executed by a processor, cause the processor to perform a method (500) according to any one of claims 1 to 7. [10] A computer-readable medium storing instructions that, when executed by a processor, cause the processor to perform a method (500) according to any one of claims 1 to 7.