Vehicle, device, computer program, machine learning-based model, method for training a machine learning-based model, and method for adjusting a sample for training a machine learning-based model

The method addresses the challenge of insufficient training data in machine learning models by using concept activation vectors to adjust samples for improved training and debugging, enhancing prediction accuracy and network transparency.

DE102023212519A1Pending Publication Date: 2025-07-10CONTINENTAL AUTOMOTIVE TECHNOLOGIES GMBH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
DE102023212519
Authority / Receiving Office
DE · DE
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-11-16
Filing Date
2023-12-12
Publication Date
2025-07-10

AI Technical Summary

Technical Problem

Existing machine learning-based models suffer from insufficient training data, leading to errors and a lack of human interpretable concepts for debugging, with current methods being labor-intensive, unsupervised, or unable to visualize concept representations effectively.

Method used

A method for adjusting samples by obtaining concept activation vectors (CAVs) and ground truth to reduce deviations, allowing for human interpretable corrections and improved training of machine learning models.

Benefits of technology

Enables accurate and reliable predictions by generating corrected input samples based on human interpretable concepts, reducing labeling costs and making networks transparent through multi-layer analysis.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

Embodiments of the present disclosure relate to a vehicle, a device, a computer program, a machine learning-based model, a method for training a machine learning-based model, and a method for adjusting a sample for training a machine learning-based model. The method for adjusting the sample includes obtaining a concept activation factor (CAV) of the segmentation model for the sample to obtain a ground truth for the CAV. The method also includes adjusting the sample based on a comparison of the CAV to the ground truth such that a deviation between the CAV and the ground truth is reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Embodiments of the present disclosure relate to a vehicle, an apparatus, a computer program, a machine learning-based model, a method for training a machine learning-based model, and a method for adjusting a sample for training a machine learning-based model.Fault Detection in ML ModelsSome object detection tracking programs use an original network to assess the correctness of the object tracking and train an auxiliary network to predict the correctness of the tracking using the activation functions from the intermediate layer. In addition, a context (temporal information, activations and raw pixel / optical flow) is fed into the fault detection network. The fault detection network predicts the correctness of the fault of the object tracking from the activations and the context.Another object tracking architecture provides disturbance prediction using two modules, where the first module (correlation filter) takes an original image as well as the bounding box of the object as input and outputs a correlation map reflecting similarity of the target based on the object and context. The input to the second module, called anomaly tracking, is the correlation map from the previous module, and the output is a probability of predicting the accuracy of the tracking ranging from 0 to 1.In another approach, a concept-based object detection program is used to detect complex objects using simple ones that are structural parts of the complex. The goal is to analyze user queries (speech, text, etc.) and select the appropriate object recognition program for the given task: excellent object recognition, concept-based recognition models with known classes and unknown classes. The framework is not designed to identify cases of failure.Another framework is designed to process queries to known and unknown classes in the case of a known object class, and uses a known recognition model to recognize the query object. Otherwise, a neural network may segment potential objects and generate an object vector for each object, generate a query object vector, generate a correlation ratio between the vectors, and associate the object based on the correlation ratio. The query may include a text string or a selection in an image.Another approach provides a method for overseem monitoring for detecting an abnormality occurring in the area, preventing false detection caused by changes in an environmental condition, and including means for photographing a monitoring area on a road having one or more viewpoints, and object extracting means for extracting an area of an object occurring in the monitoring area and a pixel value from a video detected by the photographing means. The abnormality detection system includes object detection means for identifying a type of the object from a set of local features by dividing the range of the object and the pixel value detected by the object extraction means into blocks based on the angle of view and a set of determination criteria for positions in the video, and abnormality detection means for detecting the presence or absence of an obstacle in the video from information on the type of the object detected by the object detection means.Another approach proposes a mechanism for determining a quota indicative of success of segmenting a 3D image, i.e., a success quota. The mechanism proposes to obtain one or more 2D images of different target views of a target object in the 3D image by processing a segmentation result of the 3D image. (A view) of each 2D image is classified using an automated classifier. The classification results are used to determine a success rate, which indicates, for example, whether or how close to the 3D segmentation result represents a segmentation result with sufficient accuracy on the basis of checked data (ground truth), for example for a medical decision making.Another approach proposes to transform an opaque neural network into an explanatory neural network by, for example, extracting rules for explaining decisions of the network.Machine learning-based models may suffer from insufficient training data, in particular, that introduces errors such as distorted data into the model.The targeted application of the present disclosure thus enables human experts to trace back an error that has occurred in a sample to their actual cause in order to be able to make informed decisions regarding the correction of the error. 1. existing methods for debugging a TNN are unable to correct input samples at the level of human interpretable concepts; however, this is highly desirable because human error understanding is necessary for the informed development of debugging strategies. For example, a human skilled artisan should be able to observe which changes to the input must be made to alter the output of the TNN, in alignment with human interpretable concepts (alias features). 2. implementation of such a recovery paradigm is hampered by the following disadvantages of existing methods for retrieving and processing learned concepts of TNN: a. Trade-off between label cost and concept integrity: i. TCAV, NetDis, Net2Ve: These and similar supervised methods require the labor intensive collection of labels such as per pixel segmentation or pixel samples. ii. ICE, ACE, ECLAD: These and similar unsupervised methods may not retrieve concepts and are not applicable to concept integrity analysis because they "answer" which (not necessarily human interpretable) concepts learned the model to make decisions, and not, "Whether the model has learned a particular set of concepts defined by humans", e.g., CBM, CBM-AUC: In these and similar concept Bottleneck approaches, only a single layer is interpretable; moreover, Bottleneck neurons are only proxies that are soft trained to conform to the meaning of a human-accessible concept, but may not do so. c. TCAV, CBM, CBM-AUC: In some methods, concept representations may be impossible or difficult to visualize directly and / or locally; e.g., locating the concepts used by the model in a particular single sample. However, such visualization is detrimental to the purposes of prescribed security control and debugging, d. Net2Ve, TCAV: The methods generate a single vector representation of the concept that may be distributed in different regions of the feature space, which may have a complex fragmented representation structure. Therefore, the concept may be encoded differently due to various factors.Thus, there may be a need for a sample adjustment approach for training a machine learning-based model.This need may be met by the subject matter of the appended independent claims. The appended independent claims disclose advantageous embodiments thereof.The embodiments of the present disclosure provide a method of adjusting a sample to train a machine learning-based model. The method includes obtaining a concept activation vector (CAV) of the model for segmenting the sample to obtain a ground truth for the CAV. The CAV thus represents the features in this one sample which fall within the scope of the concept class. Within the scope of the present disclosure, the CAV is also understood as a numerical representation of the concept (CNR) or as a vector representation of the concept (CVR).The method also includes adjusting the sample based on a comparison of the CAV to the ground truth such that a deviation between the CAV and the ground truth is reduced.In this way, for example, irrevocable or detrimental features in the sample can be removed and / or features can be added. Training a machine learning-based model using the adjusted sample may thus result in a more trained machine learning-based model, e.g., one that makes more accurate and / or reliable predictions.In some embodiments, the method further comprises obtaining CAV and a corresponding ground truth for one or more further samples, obtaining at least one cluster of the CAV and at least one cluster of the ground truth based on their similarity, and obtaining a projection function for mapping the CAV cluster to the ground truth cluster. In this case, adjusting the sample may include adjusting the sample using the projection function.In practice, the projection function may be applicable to or suitable not only for the aforementioned sample, but also for / for other samples. For the adaptation of further random samples, therefore, a further corresponding CAV does not have to be obtained. Therefore, the projection function makes it possible to save additional effort in the adaptation of further random samples.In practice, the CAV and ground truth can be obtained on a single sample basis. Accordingly, obtaining the CAV may include obtaining a first CAV and a corresponding ground truth for a first sample and a second CAV and a corresponding ground truth for a second sample.This has the following main advantages: 1) the exact distribution of information about the respective concept specific to this sample is obtained. This allows obtaining local explanations, e.g. for comparing the response of different nets to a single sample. In FIG. 2 ), the differences between the samples regarding the distribution of the concept information can be analyzed.In some applications, the method further comprises providing the adjusted sample to train a machine learning-based model.In this way, training the machine learning-based model may be improved such that predictions of the trained machine learning-based model may be more accurate and / or reliable.Further embodiments provide a method for training a machine learning-based model. The method comprises obtaining an adjusted sample obtainable by the proposed method for adjusting the sample. Further, the method of training the machine learning based model includes training the machine learning based model using training data including the adjusted sample.Other embodiments of the present disclosure provide a machine learning-based model obtainable by the proposed training method.Further embodiments of the present disclosure provide a computer program comprising instructions which, when the computer program is executed by a computer, cause the computer to perform one of the proposed methods and / or to apply the proposed machine learning-based model.Other embodiments provide an apparatus comprising one or more interfaces for communication and a data processing circuit configured to perform any of the proposed methods.Further embodiments provide a vehicle comprising the proposed device.Further, embodiments will now be described with reference to the accompanying drawings. Note that the embodiments illustrated by the mentioned drawings show only optional embodiments as an example, and the scope of the present disclosure is by no means limited to the presented embodiments.Brief Description of the DrawingsFIG. 1 illustrates a flow diagram schematically illustrating an embodiment of a method for adjusting a sample to train a machine learning-based model; FIG. 2 illustrates a flow diagram schematically illustrating an embodiment of a method for training a machine learning-based model; FIG. 3 illustrates, by way of example, concept-based classification; FIG. 4 schematically illustrates various dimensionality of CAV; FIG. 5 schematically illustrates how CAVs may be optimized; FIG. 6 schematically illustrates how single sample CAVs can be optimized for multiple samples; FIG. 7 schematically illustrates how a sample may be adjusted for training a machine learning based model according to the proposed approach; FIG. 8 schematically illustrates how predicted segmentation may be obtained; FIG. 9 schematically illustrates a process for retrieving CAV for predicted segmentations; FIG. 10 schematically illustrates how projection functions may be obtained; FIG. 11 schematically illustrates how the projection functions may be applied for sample matching to train a machine learning-based model; and FIG. 12 illustrates a block diagram schematically illustrating an embodiment of a device according to the proposed approach.The present disclosure addresses the above-mentioned problems. In particular, the proposed approach provides the following advantages: 1. generating corrected input samples, e.g., images, based on human interpretable learned concepts. 2. optional automated labeling: Reducing the labeling cost for the monitored approach by alternative strategies of automated generation of segmentations. 3. the proposed method uses internal representations in several layers and thus makes the network as a whole transparent, not just in a single Bottleneck. 4. the proposed approach allows the creation of a Salizz explanation map to enable manual quality controls by the human. 5. the proposed method allows searching for all possible areas per sample in connection with the concept by single sample analysis of inputs and not by generalized optimization of sample sets.Hereinafter, embodiments thereof will be described in detail with reference to the drawings.FIG. 1 illustrates a flow diagram schematically illustrating an embodiment of a method 100 for adjusting a sample to train a machine learning-based model.As can be seen from the flow chart, the method 100 comprises obtaining 110 a concept activation vector (CAV) (also referred to as "concept vector") of the model for segmenting the sample.The sample contains, for example, the entire correspondence of an image. In practice, the image may be a camera image.In enlightening AI, a concept refers to an abstraction or high-level representation of the feature or pattern learned by the model in training. In other words, a concept (also referred to as a "semantic concept") is a scalar or vector in the deep neural network (TNN) feature space that corresponds to a portion of the input space (e.g., a portion of the image). A semantic concept corresponds to a natural language concept that can be described with a word or phrase (e.g., "yellow", "pedestrian head", "person", etc.). Further examples are: "car", "truck", "bus", "bicycle" and / or "motorcycle". A concept scalar represents the concept assignment (strength, meaning) of a single concept for the current prediction.In the context of the present disclosure, the CAV may be understood as a vector in the direction of the activation values of the set of examples of the respective concept. In other words, concept activation vectors (also referred to as "concept vectors") represent the direction or "center of mass" (centroid) of concepts in the feature space, depending on the implementation.Dimensionality of the concept vector also depends on the implementation.Further, the method 100 includes obtaining 120 a ground truth for the CAV. Ground truth can be understood as ideal segmentation for the samples that can be reproduced by an "ideal" CAV, e.g. an ideal segmentation of an image for the concept "bicycle" or "motorcycle". In practice, the ideal segmentation may be determined manually, (semi) automatically, and / or from a sample database that provides not only the samples but also corresponding examples of ideal segmentation for one or more concepts.Further, the method 100 comprises adjusting 130 the sample based on a comparison of the CAV and its ground truth such that a deviation between the CAV and the ground truth is reduced. In this way, irrevocable features can be removed from the sample, thereby improving training of a machine learning-based model using the sample as training data.For adapting 130 the sample, the sample can be optimized iteratively, for example. For this purpose, for example, a suitable loss function may be minimized, as will be explained in more detail later.The adjusted sample may then be used in the respective method for training a machine learning-based model, as set forth in more detail below with reference to FIG. 2.FIG. 2 illustrates a flow diagram schematically illustrating an embodiment of a method 200 for training a machine learning-based model. The method 200 comprises obtaining 210 an adjusted sample obtainable by means of the proposed method for adjusting a sample, and training 220 the machine learning-based model using training data, including the adjusted sample.Those skilled in the art having benefit of the present disclosure will appreciate that the adjusted sample may be used in various machine learning techniques, including supervised learning or unsupervised learning.In other words, to summarize the proposed approach, the present disclosure provides a method that helps human skilled artisans analyze and repair machine learning (ML) based models, particularly, but not exclusively, for the computer vision domain. The method can be used to obtain a "corrected" version for an input with a wrong or undesired prediction, which leads to correct predictions of the TNN (deep neural network). Herein, the applied human interpretable changes are more specifically directed to predefined (e.g., user selected) human interpretable semantic concepts arising from natural language. By comparing the original and (interpretable) "corrected" images, the human skilled person can (1) obtain insights into and understand conceptual disturbance reasons and consequently (2) generate solutions and recommendations for error correction in the network.The proposed method can be used to generate and represent the corrected sample obtained from a fault, resulting in a correct prediction of the network. Thus, one skilled in the art can understand which changes to the input must be made for the model to function properly and acquires an understanding of the underlying faults in the model. The proposed method allows the generation of corrected samples for a single input or the generalization of the information for different inputs and thus the obtaining of information about possible fault cases in the network. Thus, a person skilled in the art can understand a single sample and a general model. After many samples have been checked, the following corrective measures are possible:One skilled in the art may derive constraints for additional data records to optimize data acquisition and model training processes; additional constraints may be added to the optimization criteria. In summary, the proposed method provides a human expert with possibilities for analyzing an ML model, identifying current errors and proposing changes to the data processing and training processes.Optimization of CAVAn example approach:Purely Global CAVGlobal CAVs are typically generated for a large amount of data and / or an entire data set. In contrast, the present invention proposes to obtain the CAV only for a few selected samples or the same sample (local CAV).Conceptually, it is proposed to find neural entities that contain the information relevant to a particular semantic concept, i.e., a concept that is writable in natural language, for example, objects or object attributes. This information is (can) (be) encoded in different units in each layer. It is proposed to use a standard representation of the information distribution over TNN units: concept activation vectors (CAV). In the present case, a CAV is a vector in (a subspace) the latent space of the TNN, i.e. the space spanned by (a subset) of the hidden layer output units of the TNN. The CAV entries represent weights for weighting the outputs of the units, e.g. the outputs of neurons or convolutional filters. The optimal CAV for a particular concept is characterized by the following property: Assume that an input sample is provided along with a ground truth about where / whether the concept is present in the input. When the output of the set of (hidden) units corresponding to the CAV is multiplied by the CAV (vector product), the resulting scalar matches the ground truth of the concept. A CAV corresponding to convolution filters in one layer is applicable to each activation pixel, for example. Internal multiplication with the pixels results in a segmentation mask that indicates where a concept is present in the image (see FIG. 5 ). A standard process for obtaining an approximately optimal CAV for a concept (defined by an input and label record) is to fix the TNN weights and optimize the CAV entries so that the multiplication outputs match the ground truth (see FIG. 5 ). The process is as follows: 1. input: a. a trained TNN b. a concept data set which consists of input images and the associated ground truth designations for whether / where the selected concept is present. 2. selecting the layer in which to search for the CAV. 3. initializing the CAV weights randomly. 4. iterative optimization of the CAV weights as follows: a. obtaining intermediate outputs: an input image is passed to the ML model to be tested, which makes predictions. For the chosen layer, activations are stored, e.g., obtaining a CAV prediction: multiplying the activations by the CAV; in the case of a convolutional TNN and a CAV for application to activation map pixels, applying the CAV to all activation map pixels and obtaining a segmentation mask. c. updating the CAV weights to improve the quality of the obtained CAV prediction in terms of ground truth, e.g., optimization by means of cube loss.The proposed approach: local CAVA major difference between existing approaches for obtaining CAV and the proposed approach is as follows: It is proposed to obtain CAV on a single sample basis. This is done by optimizing the CAV not on a sample set but for a single sample (always using the same sample for the optimization). This has the following main advantages: 1) the exact distribution of the information about the respective concept specific to this sample is obtained. This allows obtaining local explanations, e.g. for comparing the response of different nets to a single sample or for comparing the response of a single net to different samples containing the same concept. In FIG. 2 ), the differences between the samples regarding the distribution of the concept information can be analyzed. 3) single sample-based CAVs for the same concept and from different samples can still be combined with information about the global distribution of information about the concept within the TNN layer. In the present case, it is proposed to obtain single sample-based CAV for a plurality of samples that may belong to different classes (see FIG. 6 ). The resulting set CAV (ground truth CAV) can be used in future analysis to generalize the findings obtained from each of the multiple samples.Sample Error Correction (Single Sample, Local Clarity)The following setting is considered: An input image is provided which causes a false prediction regarding the presence of a concept, e.g. the localization of a car, by an object recognition program, accepting, e.g. a false-negative. It could now be assumed that the input image contains all the information necessary to make a correct prediction, but the TNN loses the information provided at a certain point or interprets it incorrectly, resulting in the wrong prediction. The proposed approach allows to determine: 1) In which layer information is available that allows to generate the correct results? → With the aid of the layers containing this information, a single sample CAV that reproduces the ground truth well should also be found. Now, the success of finding a CAV that produces masks that match the desired output is considered. 2) What features in the image would have to be added / removed / changed to obtain correct (re)for the concept in a layer (especially those where loss of information was found)? → Iterative optimization of the image to reduce the difference between a single sample CAV matching the prediction and the single sample CAV matching the desired ground truth. Specifically, the proposed approach for 1) proposes iteratively optimizing the input sample to cause the resulting TNN activations to match the difference between the single sample CAV obtained for the GT segmentation mask (GT) and those obtained for the predicted segmentation mask (see FIG. 7 ). To this end, it is proposed to minimize the distance between the predicted segmentation mask CAV (calculated after each forward propagation) and the desired GT segmentation mask CAV. After each iteration of the optimization of the image, the "corrected" output must obtain more distinct features that reduce the distance between the GT-CAV and the predicted CAV and thus reduce the difference between the GT segmentation and the predicted segmentation.In the case of the object recognition task, the predicted segmentation mask is defined as a Jacquard index (loU) of the GT segmentation and the masks obtained from prediction envelopes (see FIG. 8 ). One skilled in the art will appreciate that other loss functions may also be used, such as cube loss, mean absolute error (MAF), or mean squared error (MSE).The general process of retrieving CAV for predicted segmentations is illustrated in Figure 9. The process is similar to that in FIG. 6, but predictions are additionally used (see FIG. 6 ).Sample Error Correction (Multiple Samples, Global Clarity)Error correction can be generalized by projection functions that can be learned by clustering GT and predicted CAV (see FIG. 10 ). The projection function generalizes the mapping of predicted CAV to GT-CAV (or vice versa). Using these findings, information about typical errors that may occur in models (cluster groups and projection functions) can be obtained, and errors can be reconstructed using projected CAV. Specifically, the use case is considered to be a set of fault cases, i.e., samples that cause erroneous predictions. It is now desired to identify 1) the fault case, i.e. clusters of fault cases where errors are (could) caused by the same semantic error in a layer representation of the TNN; this also includes a function which can classify new errors by known fault cases; 2) identify the common features of samples of a fault case which cause the errors; 3) for each fault case, find a globally applicable input transform (the projection function) which corrects the input in such a way that the TNN generates the expected correct outputs; this function should also be applicable to new fault cases. The proposed approach: For 1), it is proposed to group each GT and predicted CAV using standard clustering algorithms. A cluster of predicted CAV then contains samples in which the TNN used similar features in its latent space to generate the (erroneous) output and is interpreted as a failure case. Fig. 2) deals with understanding of the features that lead to the failure case. For this, it is proposed to associate each cluster of predicted CAV with a cluster of ground truth CAV (with / without considering that CAV should be associated on the same sample). It is proposed to identify or learn a projection function on the input that moves the predicted CAV (i.e., the features used by the TNN in the prediction layer) into the clusters of its corresponding correct GT-CAV (i.e., the features it should use to obtain ground truth). The projection functions are therefore defined in this case by the transformations in the feature space that they should cause. These functions may be linear or non-linear mapping functions and have been learned from and with respect to clusters of GT and predicted CAV. For example, the linear transformation in latent space that moves the predicted CAV clusters under consideration to the GT-CAV clusters is determined; then the projection function is defined to iteratively alter the image in a manner such that the predicted CAV follows this linear transformation.The process of generalized error correction using a projection function (see FIG. 11 ) is similar to a process of single sample optimization (see FIG. 7 ). Here, however, instead of a continuous re-estimation of the GT-CAV, it is proposed to use the projection obtained with the known projection function during the first iteration. Such an approach allows generation and reproduction of any type of (known, projectable) error and can be used for error recovery, analysis and generation of recommendations for the correction of NN.The approach allows for a more natural explanation of ML models so that a human skilled artisan can obtain a visual indication of potential errors in the ML model.The approach of the method according to the invention can also be implemented in a device as explained in more detail below with reference to FIG. 12.FIG. 12 illustrates a block diagram schematically illustrating an embodiment of such a device 1200. The apparatus comprises one or more interfaces 1210 for the communication and a data processing circuit 1220 configured to carry out the proposed method.In embodiments, the one or more interfaces 1210 may include wired and / or wireless interfaces for transmitting and / or receiving communication signals in connection with the execution of the proposed concept. In practice, the interfaces include, for example, pins, wires, antennas, and / or the like. The interfaces can also comprise means for (analog and / or digital) signal or data processing in connection with communication, e.g. filters, random samples, analog-to-digital converters, signal acquisition and / or reconstruction means as well as signal amplifiers, compressors and / or any encoding / decoding means.Data processing circuit 1220 may correspond to or include any type of programmable hardware. For example, examples of the data processing circuit 1220 include a memory, a microcontroller, field programmable gate arrays, one or more central and / or graphics processing units. To execute the proposed method, the data processing circuit 1220 may be configured to access or retrieve an appropriate computer program for executing the proposed method from a memory of the data processing circuit 1220 or a separate memory communicatively coupled to the data processing circuit 1220.In practice, the proposed device may be installed in a vehicle. Embodiments may thus provide a vehicle comprising the proposed device. In implementations, the device is, for example, a part or component of the ADS.In the foregoing description, it will be appreciated that various features are incorporated in examples for the purpose of streamlining the disclosure. This method of the invention is not to be construed as reflecting the intention that the claimed examples require more features than are expressly recited in each claim. Rather, the subject matter may lie in less than all features of a single disclosed example, as reflected in the following claims. Thus, the following claims are hereby incorporated into the specification, with each claim standing on its own as a separate example. Although each claim may stand upon itself as a separate example, it is noted that although a dependent claim may refer to a specific combination with one or more other claims in the claims, other examples may also include a combination of the dependent claim with the subject matter of any other dependent claim or a combination of each feature with other dependent or independent claims. Such combinations are proposed herein unless it is indicated that a specific combination is not intended. Moreover, it is intended that features of a claim may also be included in any other independent claim, even if that claim is not made directly dependent on the independent claim.Although specific embodiments have been illustrated and described herein, it will be understood by those of ordinary skill in the art that a variety of alternative and / or equivalent implementations may be substituted for the specific embodiments shown and described without departing from the scope of the present embodiments. This application is intended to cover any adaptations or variations of the specific embodiments discussed herein. Therefore, it is intended that the embodiments be limited only by the claims and their equivalents.

Claims

A method (100) of adjusting a sample for training a machine learning based model, the method (100) comprising: obtaining (110) a concept activation vector CAV of the model for segmentation for the sample; obtaining (120) a ground truth for the CAV; and adjusting (130) the sample based on a comparison of the CAV and the ground truth such that a deviation between the CAV and the ground truth is reduced.The method (100) of claim 1, wherein the method further comprises: obtaining CAV and a corresponding ground truth for one or more further samples, obtaining at least one cluster of the CAV and at least one cluster of the ground truth based on their similarity, and obtaining a projection function for mapping the cluster of CAV to the cluster of the ground truth, and wherein adjusting the sample comprises adjusting the sample using the projection function.The method (100) of claim 1 or 2, wherein obtaining the CAV comprises obtaining a first CAV and a corresponding ground truth for a first sample and a second CAV and a corresponding ground truth for a second sample.The method of any preceding claim, wherein the method further comprises providing the adjusted sample to train a machine learning based model.A method (200) for training a machine learning-based model, the method comprising: obtaining (210) an adjusted sample obtainable by the method of any one of the preceding claims; and training (220) the machine learning-based model using training data including the adjusted sample.A machine learning-based model obtainable by the method of claim 5.A computer program comprising instructions which, when the computer program is executed by a computer, cause the computer to perform any of the methods (100, 200) of any of claims 1 to 5 and / or to apply the machine learning-based model of claim 6.An apparatus (1200) comprising: one or more interfaces (1210) for communication; and a data processing circuit (1220) configured to perform any of the methods (100, 200) of any of claims 1 to 5 and / or to apply the machine learning based model of claim 6.A vehicle comprising the apparatus (1200) of claim 8.