Vehicle, device, computer program and method for embedding and evaluating visual concepts within the latent feature space of object detectors aimed at improving the safety and reliability of autonomous systems
The Siamese neural network training method embeds visual concepts into the latent feature space of object detectors, enabling interpretable object detection and improving safety in autonomous driving by ensuring detections are explainable and reliable.
Patent Information
- Application Number
- DE102024200029
- Authority / Receiving Office
- DE · DE
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-01-03
- Publication Date
- 2025-07-03
AI Technical Summary
Existing AI-based object detectors in autonomous driving systems are difficult to interpret, making it challenging to understand and rectify errors, which poses a safety risk in safety-critical applications.
A Siamese neural network training method is employed to embed predefined visual concepts into the latent feature space of object detectors, allowing for the classification of objects and ensuring that detections can be explained by these concepts, with a metric to assess concept embedding quality.
The method enables interpretable object detection, allowing for reliable recognition of objects and identification of unexplainable detections, thereby enhancing safety in autonomous driving systems by providing a mechanism to verify and improve the detection models.
Smart Images

Figure 00000013_0000 
Figure 00000014_0000 
Figure 00000015_0000
Abstract
Description
[0001] The present disclosure relates to a vehicle, a device, a computer program, and a method for training interpretable object detectors, which include embedding visual concepts into the latent feature space of the object classification module of the object detectors. This allows for checking whether the object detections align with previously defined visual concepts. Detections that do not exhibit an affinity to defined concepts are referred to as "unexplainable."
[0002] Computer vision plays a crucial role in various applications, including automotive, robotics, and the like.
[0003] Many autonomous driving systems are based on the "Sense-Plan-Act" architecture. One of the prerequisites for this architecture is the extraction of an environmental model by analyzing sensor data. Artificial intelligence (AI)-based object detectors are playing an increasingly crucial role for this purpose. For example, they are often used to analyze images from cameras in autonomous vehicles.
[0004] A notable concern related to the use of AI-based object detectors is safety: modern detectors deliver impressive results, but their internal state representations are difficult for humans to understand. Error cases are therefore difficult to analyze and resolve. For example, if an object detector fails to detect a pedestrian, it may be unclear what caused the error. This poses a major problem for any safety-critical application such as autonomous driving.
[0005] In recent years, explainable AI (EAI) has become an emerging trend in AI research. The goal is to make the internal workings of AI models transparent and their decisions verifiable. However, most of this work focuses on simple object classification tasks. Therefore, they cannot be easily applied to other tasks such as object detection.
[0006] In addition, many methods focus exclusively on post-hoc model analyses. This is helpful for finding vulnerabilities in previously trained AI models, but does not allow for the integration of security-relevant information into the training process itself.
[0007] Thus, there may be a need for an improved concept for AI object detectors, especially to make the processes of object detectors more explainable.
[0008] The present disclosure proposes a solution to the challenges mentioned.
[0009] The use of visual concepts is proposed. A key difference lies in the training scheme used, which ensures the consistent and plausible acquisition of these concepts. Furthermore, a novel assessment methodology is introduced that allows the measurement of the ability to learn these concepts.
[0010] The embodiments presented in the present disclosure allow the establishment of domain-specific prior knowledge through the use of visual concept images. These images represent object properties such as color, size, texture, orientation, occlusion, and / or other attributes. An interpretable object detector can be trained to embed visual concepts into the latent feature space of the AI model. This has the advantage that objects can be reliably recognized and it can be verified whether the observed detections can be "explained" by the visual concepts. Unexplainable detections may indicate unidentified problems in the perception stack. In such a case, all automated driving functions can be interrupted, and the human can assume control instead. It is also proposed to record sensor data from these events in a database.Such data can be helpful for improving detection models by incorporating them into the training loop. Need for Explainable AI (EKI):
[0011] In recent years, deep neural networks (TNNs) have become ubiquitous in many computer vision tasks, such as image classification and object detection. These models are known to be able to learn rich and diverse feature representations when trained on large image datasets. This typically involves end-to-end optimization of model parameters without imposing any prior constraints or conditions on the learned features. Such an approach has been shown to far outperform traditional hand-crafted features. However, the internal feature representations learned by most deep learning models are opaque and difficult for humans to understand. This has significant implications for their application in real-world scenarios. For example, if an AI model produces incorrect results, it is difficult to understand the underlying cause and how to fix the problem.This limits their potential use in safety-critical tasks such as autonomous driving. EKI: Local explanations (feature salience):
[0012] To mitigate this problem, methods can be applied to explain the behavior of AI models. Some of these methods are based on salience maps: here, the goal is to find the crucial input features used by an AI model for its predictions. The Grad-CAM (Gradient-weighted Class Activation Mapping) method, for example, finds regions in images that were often responsible for the predictions of image classification models. However, this method is limited to the analysis of individual images and cannot explain how objects are fundamentally recognized. Furthermore, the salience maps themselves can be ambiguous and therefore difficult to interpret. EKI: Global explanations using concepts:
[0013] Other methods are based on feature space analysis. One example is the TCAV (Testing with Concept Activation Vectors) method, which attempts to find the visual concepts used by a TNN (deep neural network) to classify images. One might expect that, for example, an image classification model capable of recognizing images of zebras would also recognize striped patterns. To test such assumptions, TCAV uses a separate dataset of concept images to find feature activations correlated with a concept of interest (also defined as concept activation vectors). Visual concepts represent object properties such as color, texture, orientation, and size. In contrast to salience-based methods such as Grad-CAM, concept-based methods are more general in their explanations and provide insights into how a model recognizes entire classes of objects (also called global explanations).
[0014] Concepts could play a crucial role in security, as they are well-suited to explaining how an AI model classifies an object. For example, if an AI model has been trained to recognize people, the concepts could be used to check whether certain body parts, such as the head, shoulders, arms, or legs, are also recognized.
[0015] Some methods focus on the automated discovery of learned concepts. However, this may not be as relevant for safety-critical applications. For example, an AI model trained to use medical X-ray images would have to base its decisions on a specific set of visual concepts, such as bone spurs in the case of evidence of arthritis. Here, the problem is not to determine which concepts the model learned during training, but to ensure that the model uses a specific set of previously defined visual concepts for its decisions. Since post-hoc analysis methods do not intervene in the model training, they can only verify whether the model has learned a specific concept.
[0016] Many approaches used for visual concept analysis are post-hoc, meaning they cannot influence model training. Post-hoc approaches can be helpful for finding inconsistencies and errors in TNN feature representations, but these are not easily remedied. This is because post-hoc approaches have no means of reinforcing the presence of a particular concept in a TNN's feature space. For example, if a TCAV reveals that the model is missing an important concept, it remains unclear how to correct the problem. Knowledge contribution using concept bottleneck models
[0017] Recently, several methods have demonstrated how visual concepts can be used in the training phase of AI models. For example, Concept Bottleneck Networks (CBNs) are being introduced, in which an image classification model is divided into two information processing parts.
[0018] A first part maps the input image for a set of concept activations, while the second part uses this information to classify the image. The concept activations serve as a "bottleneck" for the information, forcing the classification to base its predictions solely on the provided concepts. By manipulating the concept activations, it is possible to verify whether a classification result was caused by the presence or absence of a prescribed concept.
[0019] However, CBN's feature representations are still not fully interpretable. Only a specific part, the concept activation layer, is understandable to humans. For example, there is no mechanism that explains how the concepts themselves are recognized in the input images. Furthermore, the number of concepts required to properly perform the classification task may be unknown.
[0020] Concept bottleneck models (CBM) modify the TNN architecture by introducing a specific layer, the concept bottleneck layer, in which each neural or convolutional filter corresponds to a visual concept. However, if the set of previously selected concepts is incomplete, recognition performance suffers significantly. Another problem is that if the concept layer does not use strictly binary activations, there is a significant risk that other information will be implicitly encoded in the bottleneck layer in an entangled and uninterpretable manner. Another problem with CBM is that the concepts must be orthogonal (each concept is represented by a unit vector): this does not allow for computational similarities between different concepts and prevents the learning of efficient feature representations.
[0021] The completeness of the previously selected concepts is usually assumed to be known in advance. However, this only applies to simple recognition tasks. Some methods provide a measure of completeness for the concepts used in CBM. This can be helpful in finding deficiencies in the selection of visual concepts. However, if the actual understanding of the recognition task itself is lacking, 100% completeness may not be achieved because it is still unknown which concepts are missing. Therefore, "residual" dimensions are introduced that are not aligned with a previous concept and therefore do not allow for 100% completeness of the concepts. However, the number of these residual dimensions is not fixed, which means that the number of missing concepts would have to be estimated in advance. Conceptual alignment of the feature space
[0022] Other methods provide concept whitening modules (CWM) that make the TNN feature space interpretable. The idea is to teach a transformation matrix to disentangle the latent features by aligning them to specific concepts. The method is based on conventional principal component analysis (PCA), which is widely used to map high-dimensional feature spaces into lower dimensions. The major advantage of CWM is that the TNN features of a layer become easily interpretable, since each dimension in the transformed feature space directly aligns with a visual concept.
[0023] One of the main problems with CWM is that model training cannot jointly optimize the image classification task and the CWM. Instead, an alternating training scheme is required, in which the CWM parameters must be frozen during classification training, and vice versa. Such training typically takes considerably longer, and proper convergence may also be difficult.
[0024] Another problem with many methods is that the use of visual concepts is limited to a single layer or block of a TNN. This means that any efforts to make the feature representations interpretable are limited to small parts of the models. However, TNNs typically learn hierarchical representations spanning many layers: some visual concepts are embedded in the early layers (e.g., colors and texture) and others in the intermediate or final layers (e.g., shapes, object parts).
[0025] An underlying goal of the present disclosure is to make the feature space of TNNs human-interpretable. To this end, it is proposed to ensure that a TNN learns a predefined set of visual concepts relevant to a particular task. For example, concepts used for a pedestrian detector could be body parts such as the head, arms, legs, and / or the like. Other concepts used could be orientations relative to the camera, for example, when a person is viewed from the front or back. All concepts must be clearly distinguishable in the feature space representations of a TNN.
[0026] Furthermore, the present disclosure relates to a useful way to measure how well visual concepts are embedded in the feature space of the TNN. For example, concept whitening modules align the feature space of a TNN layer with visual concepts. However, concept whitening modules do not include a metric that explains how well the concepts are represented: for example, whether they are well separated in the feature space of a TNN.
[0027] For a better understanding, it is helpful to find relationships between visual concepts. Therefore, the way concepts are represented in the feature space should allow for the measurement of similarities between them (e.g., "arm" is more similar to "leg" than "head"). Many concept analysis methods do not allow for computational similarities between visual concepts.
[0028] It should be noted that introducing visual concepts into the feature space of a TNN should not prevent the learning of good feature representations during training, especially when the previously defined concepts are incomplete and cannot efficiently solve the classification task, e.g., due to selection bias or insufficient understanding of the task (e.g., "skirt" may be relevant in pedestrian detection along with "leg," but may not be part of the previously defined concept). Concept bottleneck models, for example, lack sufficient recognition accuracy if the previously selected concepts were poorly chosen.
[0029] It is also important to be able to verify whether the output of a TNN can be trusted at inference time. Some methods achieve this through uncertainty modeling. For example, object detections with high uncertainty scores would be considered less reliable than those with low uncertainty scores. While uncertainty scores have been used successfully in the past, the use of visual concepts for this purpose has not been considered. Testing whether the feature activations of a TNN are consistent with a set of visual concepts at runtime could also be an important criterion for safety and reliability. Therefore, it should be possible to extract the concept information for an input sample on the fly, e.g., for further plausibility checks.
[0030] Embodiments of the present disclosure introduce a methodology for training interpretable object detectors. Within this framework, an interpretable object classification model is trained to evaluate region proposals generated by conventional box detection models. The proposal proposes the use of a Siamese training setup to learn a multi-objective function. This function is designed to simultaneously classify box proposals and incorporate visual concepts into the latent feature space of the box assessment model. The method involves acquiring training data for two specific objectives: the training data for the first objective comprises one or more images depicting objects along with previously defined class labels assigned to those objects.The second training dataset includes one or more concept images representing different concepts, each related to the previously defined concept labels.
[0031] Embodiments of the present disclosure provide a method for training an object detector using a Siamese neural network. The method includes obtaining first training data and second training data. The first training data includes one or more images of one or more objects and predefined class labels for the objects. The second training data includes one or more concept images of one or more concepts and predefined concept labels for the concepts.
[0032] Furthermore, the method involves training the box assessment model using the first training dataset to classify objects and simultaneously training it on the second training dataset to minimize the distance between feature activations related to that concept within the latent space.
[0033] Furthermore, the method includes training the Siamese neural network based on the first training data to classify the objects and training the Siamese neural network based on the second training data so that a distance of the feature activations of this concept in the latent space is reduced.
[0034] As one skilled in the art will appreciate, this approach provides a simple yet efficient method for reinforcing the learning of visual semantic concepts during a training phase of a TNN. Unlike other methods, the proposed approach is not limited to a single layer.
[0035] In practice, training the object detector or its box assessment model may involve obtaining feature maps from one or more layers and unwrapping them to obtain feature vectors for the concept images. Further, some embodiments include adapting the model based on a loss function applied to the feature vectors such that, in a latent space of the model, a distance between feature vectors of this concept is reduced and a distance between feature vectors of other concepts is increased.
[0036] In certain embodiments, the method additionally includes group assignment of the feature vectors of visual concepts within the latent feature space. The cluster centers and concept labels from this method serve as tools for assessing clustering quality. Successful concept embeddings are likely to produce clear, well-defined clusters, while suboptimal concept embeddings may lead to erroneous cluster assignments.
[0037] In some further embodiments, the method further comprises determining clusters for the feature vectors of that concept in the latent feature space.
[0038] Optionally, the method further comprises assessing an ability of the Siamese neural network to recognize the concepts based on the clusters.
[0039] In some embodiments, assessing the ability of the Siamese neural network to recognize the concepts includes obtaining the V-measure for the clusters and assessing the ability of the Siamese neural network to recognize the concepts based on the V-measure.
[0040] In some embodiments, the assessment of concept embedding quality through group assignment may be based on the V-measure, which measures the accuracy of the resulting clusters. This may additionally involve calculating pairwise V-measures between certain visual concepts to accurately identify poorly selected or conflicting concepts.
[0041] Obtaining the V-measure for the clusters may include obtaining a pairwise V-measure for a first and a second cluster of different concepts.
[0042] The method may also include refining each cluster center using the expectation-maximization (EM) algorithm to derive a density function for the clusters. Additionally, the method may facilitate the application of a mixture distribution model to the density function, examining whether features from the concept images align with any of the clusters to assess the ability to recognize the respective concepts.
[0043] This allows for the simultaneous classification of box proposals from object detectors and the calculation of affinity scores for the feature activations of the box classification model and the visual concept clusters. Detections without significant affinity to a cluster are identified and labeled as "unexplainable."
[0044] The method may further comprise specifying a respective center of the clusters using the EM (Expectation-Maximization) algorithm to obtain a density function of the clusters. Furthermore, the method may provide for applying a mixture distribution model to the density function to test whether features from the concept images correspond to at least one of the clusters in order to assess the ability to recognize a corresponding concept from the concept images.
[0045] In some embodiments, the method further comprises adapting the neural network for classification based on the ability to recognize the concepts.
[0046] Further embodiments provide a neural network that can be obtained by the method proposed herein.
[0047] Alternative embodiments provide a method for evaluating an object detector. The method involves applying the neural network generated by the previously described training process to test the object detector. In particular, the trained neural network is used to evaluate whether the object detector classifies objects in an understandable manner. In certain cases, the method allows for the verification of the explanatory capabilities of the object detector by examining whether the classification of objects is aligned with recognized concepts—for example, by confirming whether wheels are identified when an object is classified as a car. In this way, a means is provided to ensure the reliability of the object detector.
[0048] Other embodiments provide a method for testing an object detector. The method comprises applying the neural network, which can be obtained by the aforementioned training method, to test the object detector. The trained neural network is trained, for example, to verify whether the object detector classifies objects understandably. In some embodiments, the method allows, for example, to verify whether the classification of objects by the object detector is explainable based on recognized concepts, e.g., whether wheels are detected when an object is classified as a car. This allows verification of whether the object detector functions reliably.
[0049] Those skilled in the art will recognize that the proposed methods can be implemented as a computer-implemented method and / or in a computer program.
[0050] Accordingly, further embodiments provide a computer program comprising instructions that, when executed by a computer, cause the computer to perform the proposed method.
[0051] Other embodiments provide an apparatus comprising one or more interfaces for communication and a data processing circuit configured to carry out the method proposed herein.
[0052] In practice, the proposed method can be applied in a vehicle. Accordingly, some embodiments of the present disclosure provide a vehicle including the proposed device.
[0053] Further, embodiments will now be described with reference to the accompanying drawings. It should be noted that the embodiments illustrated by the aforementioned drawings merely show optional embodiments by way of example, and the scope of the present disclosure is by no means limited to the presented embodiments. Short description of the drawings Fig. 1 is a flowchart schematically illustrating one embodiment of a method for training a Siamese neural network; Fig. 2 is a flowchart schematically illustrating another embodiment of a method for training a Siamese neural network; Fig. 3 is a flowchart schematically illustrating a method for testing the trained neural network; Fig. 4 is a flowchart schematically illustrating a use case of a trained neural network; Fig. Figure 5 is a block diagram schematically illustrating an embodiment of a device according to the proposed method. Interpretable object detection
[0054] The present disclosure specifically considers the problem of interpretable object detection. Object detectors are critical components of automated driving systems. For example, 2D and 3D object detectors are commonly used to analyze images from cameras in autonomous vehicles. Object detection is a well-studied task, and many methods have been proposed over the years. Object detectors such as EfficientDet, CenterNet, YOLOv7, and MaskRCNN demonstrate excellent performance in real-world applications after training on large image datasets. However, as with many classification models, their latent feature spaces are difficult for humans to interpret.
[0055] However, the field of explainable AI (EAI) focuses on image classification rather than object detection. Unfortunately, extending it to object detection is difficult because detection models are more complex and use many different architectures.
[0056] It is therefore an object of the present disclosure to provide an interpretable object detector. Embodiments of the present disclosure will be described in more detail below with reference to the accompanying drawings.
[0057] In particular, the present disclosure proposes a two-step method for efficiently solving the problems: first, an object detector is run over a plurality of input images to obtain bounding rectangles for detected objects. The detected bounding rectangles are not considered the final output, but merely as box proposals for a second step. In the second step, the individual box proposals are cropped and re-enlarged, resulting in a set of equally sized images (one image per detected rectangle). These images are then processed by an interpretable neural classification network. The idea is to discard all detections that cannot be explained by a set of previously defined visual concepts.
[0058] In practice, a single convolutional deep neural network (TNN) (with feedforward architecture) can be trained using images and ground-truth labels with a suitable loss function, e.g., the categorical cross-entropy loss function. This allows, for example, the training of classification models, and, assuming the use of a large and diverse dataset, high classification accuracy can be expected. One idea of the present disclosure is the embedding of visual concepts during training. However, a single convolutional neural network does not allow control of its latent feature space and the concepts learned by the TNN.
[0059] The embodiments of the present disclosure are based on the recognition that a Siamese neural network can be used for this purpose.
[0060] Fig. 1 accordingly represents a flowchart schematically illustrating one embodiment of a method 100 for training a Siamese neural network.
[0061] A Siamese neural network (sometimes referred to as a dual neural network) is a type of neural network architecture that, when working hand in hand on two different input tensors, uses the same weights to compute comparable output tensors. One idea of the present disclosure is to use the two-part architecture for training / learning concept embeddings. This allows the Siamese neural network to be trained for different capabilities. For this purpose, the use of different training data for subnetworks of the Siamese neural network is proposed.
[0062] Accordingly, the method 100 includes obtaining 110 first training data and second training data. The first training data includes one or more images of one or more objects and predefined class labels for the objects, and the second training data includes one or more concept images of one or more concepts and predefined concept labels for the concepts.
[0063] For automotive applications, the first and second training data may indicate objects that may appear as a vehicle, e.g., infrastructure objects, traffic signs, road users, and / or the like. Accordingly, the first training data includes, e.g., images of traffic scenes and the class labels, e.g., classes such as "vehicle," "pedestrian," "traffic light," and / or other object classes.
[0064] In the context of the present disclosure, concepts may be understood as visual concepts that refer to identifiable elements or patterns within visual data (e.g., images or videos). In particular, the concepts may refer to visual properties of an object class. In practice, an object class includes, for example, certain visual features that make it distinguishable. For example, the object class "vehicle" may be characterized by the following concepts: "wheels" and "headlights." Accordingly, the second training data includes, for example, images of wheels, headlights, and / or other concepts and the corresponding concept labels.
[0065] The proposed method 100 proposes training 120 the Siamese neural network based on the first training data to classify the objects. For example, a first subnetwork is fed with the first training data, and the parameters of the first subnetwork are adjusted so that it classifies the objects better, i.e., so that the classes predicted by the first subnetwork correspond to the respective class labels.
[0066] Furthermore, the method 100 proposes training 130 the Siamese neural network based on the second training data such that the distance between the feature activations of this concept in the latent space is reduced. For this purpose, for example, a second subnetwork that shares parameters with the first subnetwork is fed with the second training data, and parameters of the second subnetwork are adjusted accordingly, as explained in more detail later.
[0067] In this way, the resulting neural network can not only classify objects but also identify concepts. In practice, the resulting neural network allows, for example, to understand why and how—that is, based on which of the concepts—it came to a particular conclusion.
[0068] In particular, the proposed method provides a simple yet efficient method for improving the learning of visual semantic concepts during the training phase of a neural network. Unlike other methods, the proposed method is not limited to a single layer and does not require a "complete" set of predefined concepts to reliably classify images. The use of concepts also does not prevent the learning of good feature representations to efficiently solve the classification task.
[0069] Further aspects and features are described below with reference to Fig. 2 described.
[0070] Fig. Figure 2 is a flowchart schematically illustrating another embodiment of the proposed method for training a Siamese neural network, including a first subnetwork 218 and a second subnetwork 219 used in parallel. As can be seen from Fig. As can be seen in Figure 2, the subnetworks 218 and 219 each comprise an input 212 and 222, respectively, and convolutional blocks 213. In this context, a convolutional block is a building block used in a convolutional neural network (ANN) for image recognition. It comprises one or more convolutional layers designed to extract features from input images and / or videos. In accordance with the idea of a Siamese network, the subnetworks share the same configuration, i.e., they have the same architecture, weights, and hyperparameters. This parallel structure is only required for training. For inference, a single TNN initialized with the common parameters is used. Siamese networks have been widely used in the past, for example, in object tracking applications.
[0071] The proposed training optimizes / improves the joint parameters for two different tasks simultaneously: object classification and concept embedding in the feature space.
[0072] To train the first task, an object classification dataset (first training data), including images 211 and annotations (class labels) 217, is used to train the object detector to classify objects in the images 211. In the present use case, for example, the task is to distinguish between objects (e.g., objects relevant to the desired use case) and their background. In automotive applications, the objects may include, for example, road users, obstacles, and / or traffic infrastructure, while the background may include landscape objects.
[0073] To train the classification task, the images 211 are processed by the convolutional blocks of the first subnetwork 218 and classified, for example, as a (foreground) object or as a background (see step 214). Subsequently, the parameters of the neural network are adjusted so that the predictions regarding the classes are more appropriate, i.e., the predicted classes more often match the ground truth labels 217 of the training data for classification.
[0074] The second task, for example, provides the embedding of concepts into the neural network. To train this second task, the use of a dataset (second training data) containing visual concept images 221 is proposed. The concept images 221 represent, for example, object properties such as color, texture, orientation, and size. For a pedestrian detector, for example, detecting pedestrians on the road is essential in autonomous driving applications. An important concept is orientation, as it indicates the direction in which pedestrians are moving. It is therefore proposed that the pedestrian detector understands orientations and can distinguish whether a person is seen from the front, the side, or the back. For this purpose, it is considered to capture many images of pedestrians and group them according to their orientation.For this purpose, a corresponding concept data set with concept labels for different concepts is created.
[0075] Preferably, all feature activations of the same concept should be similar or at least close to each other in the latent space. However, they should be different from those of other concepts. For example, a pedestrian can be seen either from the front or from the back. One idea of the present disclosure is to group the visual concepts in the latent feature space of the TNN. To provide a uniform group assignment, i.e., clusters that have high homogeneity and completeness, it is proposed to unroll or "flatten" feature maps of the TNN layers into a one-dimensional vector (see step 224) and to use a loss function, e.g., a semi-hard triplet loss function, for any discrepancy between predicted concept labels for the concept images 221 and concept labels 227 of the training data, to learn concept embeddings.In this way, it can be ensured that Euclidean distances between feature vectors of the same concept class are reduced or even minimized, while distances between different concept classes are increased or even maximized. To achieve this, the proposed method proposes obtaining feature maps from one or more layers of the Siamese neural network and unrolling the feature maps to obtain one-dimensional feature vectors for the concept images. Various methods can be used to unroll the feature maps into a vector, such as flattening and concatenation (see step 225). It is also possible to reduce the dimensionality of the feature vectors before flattening, for example, by using multiple pooling layers.
[0076] The proposed method is relatively efficient because it only relies on calculating distances between vectors. No complex matrix-vector multiplications are required. Thus, the proposed method allows the use of features from multiple layers of the TNN.
[0077] Apart from that, the present disclosure provides a method for assessing a capability of the Siamese neural network to recognize concepts based on the clusters, i.e., how well the concepts are embedded in the feature space of the neural network. This is described below with reference to Fig. 3 explained in more detail.
[0078] Many different metrics have been proposed to evaluate the performance of image recognition models. For example, classification models can be evaluated by computing their accuracy, and object detectors are often evaluated using the mean average accuracy. Good evaluation metrics are also important for measuring how well visual concepts are embedded in the latent feature space of a TNN. However, recently introduced EKI methods do not provide a solution to this problem.
[0079] The proposed method of testing how well the concepts are grouped in the latent feature space is described in Fig. 3 schematically visualized: As can be seen, the second subnet 219 is applied to the concept images 221 together with the previously mentioned flattening, concatenation and normalization steps (see steps 224 and 225).
[0080] For assessment, it is proposed to cluster the feature vectors that indicate the concepts. For this purpose, the Kmeans++ clustering algorithm is applied to the concept feature vectors. Alternatively, other clustering algorithms can also be used.
[0081] Furthermore, the proposed method proposes that the clustering result be assessed, i.e., it is checked how reliably feature vectors of the same concept are ultimately located in the same cluster. In practice, for example, the V-measure (see step 229) is applied to determine how many feature vectors of the same concept are ultimately located in the same cluster. For example, the number of cluster centers, which are used as hyperparameters in Kmeans++ clustering, can be considered as the number of different concept classes. The V-measure is a numerical value between zero and one, where one means that all data points could be correctly grouped, and zero means that the data are randomly distributed.
[0082] Optionally, pairwise V-measures can be calculated for the different visual concepts. This allows the generation of a confusion matrix that indicates whether some of the concepts overlap and are therefore indistinguishable from one another. For example, it is possible that some concepts were poorly selected or are semantically inconsistent. Analyzing the V-measure is a useful way to detect such problems. In the case of indistinguishable concept groups, the neural network (e.g., subnetwork 218) can be further trained, e.g., using special training data that indicates the indistinguishable concept groups to increase their distance from each other in the latent space.
[0083] During training, the proposed model learns to embed visual concepts into its latent feature space, and a clustering algorithm, e.g., the Kmeans++ algorithm, is applied for their group assignment.
[0084] The clusters can then be used to check whether the classification of images (e.g., by an object detector) can be explained based on the embedded concepts, as explained in more detail below.
[0085] For this purpose, a cluster center can be obtained, which, for example, indicates a mean value. The obtained cluster centers can be further refined using an EZ (Expectation Maximization) algorithm. This EM algorithm also calculates covariance matrices for each cluster. In this way, each cluster is modeled as a multivariate Gaussian function, which allows the use of a Gaussian mixture model (GMM) at runtime to test whether the features extracted from an input region correspond to a concept cluster in the GMM.
[0086] This process is in Fig. 4 is schematically visualized.
[0087] In the exemplary use case according to the invention, the resulting interpretable classification model is run, for example, on cropped and re-enlarged image areas obtained by an object detector 230. In practice, the object detector 230 is fed, for example, with the images 211 and determines bounding rectangles for objects detected in the images 211; the images are then cropped from the images 211 and re-enlarged, for example, to obtain the original width and / or height, thus obtaining cropped and re-enlarged images 211'.
[0088] The images 211' are then classified to correspond to or contain an object of interest or a background area. Based on this classification, areas considered background can be discarded. Furthermore, it is tested whether detected (foreground) objects (of interest) can be explained by the visual concepts. As in the training phase, the feature space of the TNN layers is unrolled into a feature vector (see steps 244 and 245), and the GMM is applied to check whether this vector can be assigned to a visual concept. To this end, the GMM provides a probability that the vector is located in an area in the feature space that also contains latent space vectors of samples of a concept. If the probability is low, the vector is not assigned to the concept and must be treated as an outlier / unknown concept.If it is high (e.g., > 60%), it is assigned to the concept. If the vector does not appear to have a sufficiently strong affinity to a concept cluster in the GMM, the proposed method suggests marking the input region as "unexplainable."
[0089] In summary, the proposed method (using the V-measure), unlike other methods, therefore provides a metric that indicates how well the visual concepts are embedded in the feature representation of the TNN, i.e., previously defined similarities of concepts are taken into account by the TNN in its latent space(s).
[0090] Furthermore, the proposed method allows for verifying whether a detection / classification can be explained by the visual concepts at runtime. This allows the identification and / or marking of object detections that cannot be explained using visual concepts. Subsequently, the object detector can be retrained based on the determined unexplainable classifications. In practice, for example, objects that cannot be classified in an explainable manner, i.e., by a specific combination of concepts, can be marked and used as training data along with the corresponding labels for retraining or further training of the object detector (based on machine learning).
[0091] The proposed method also allows for the calculation of similarities between different visual concepts by calculating the Euclidean distances between the concept cluster centers. Therefore, it provides further insights into how the concepts are related to each other. This allows, for example, more thorough testing of an object detector. In practice, this allows, for example, to verify whether the object detector can even distinguish between similar concepts.
[0092] As the person skilled in the art will recognize, the proposed concept can be implemented not only in a computer program, but also on a device, as described below with reference to Fig. 5 is explained in more detail.
[0093] Fig.Figure 5 is a block diagram schematically illustrating one embodiment of such a device 500. The device comprises one or more interfaces 510 for communication and a data processing circuit 520 configured to carry out the proposed method.
[0094] In embodiments, the one or more interfaces 510 may comprise wired and / or wireless interfaces for transmitting and / or receiving communication signals in connection with the implementation of the proposed concept. In practice, the interfaces comprise, for example, pins, wires, antennas, and / or the like. The interfaces may also comprise means for (analog and / or digital) signal or data processing in connection with the communication, e.g., filters, samples, analog-to-digital converters, signal acquisition and / or reconstruction means, as well as signal amplifiers, compressors, and / or any encryption / decryption means.
[0095] The data processing circuit 520 may correspond to or include any type of programmable hardware. Examples of the data processing circuit 520 include, for example, a memory, a microcontroller, field-programmable gate arrays, one or more central and / or graphics processing units. To carry out the proposed method, the data processing circuit 520 may be configured to access or retrieve a suitable computer program for executing the proposed method from a memory of the data processing circuit 520 or a separate memory communicatively coupled to the data processing circuit 520.
[0096] The proposed method can also be used in automotive applications. For example, the proposed method can be implemented in a vehicle, where it allows for verifying whether predictions of an algorithm applied to the vehicle are plausible and / or consistent. For example, explainable predictions are considered valid, while others are not. This can be used, for example, to detect errors to prevent a vehicle from inappropriate driving behavior based on an erroneous algorithm result.
[0097] In the foregoing description, it will be appreciated that various features are grouped together in examples for the purpose of streamlining the disclosure. This method of disclosure should not be construed as reflecting an intent that the claimed examples require more features than are expressly recited in each claim. Rather, the subject matter may lie in fewer than all of the features of a single disclosed example, as the following claims reflect. Thus, the following claims are hereby incorporated into the specification, with each claim capable of standing on its own as a separate example.While each claim may stand on its own as a separate example, it should be noted that although a dependent claim in the claims may refer to a specific combination with one or more other claims, other examples may also include a combination of the dependent claim with the subject matter of any other dependent claim, or a combination of each feature with other dependent or independent claims. Such combinations are suggested herein unless it is stated that a specific combination is not intended. Furthermore, it is intended that features of one claim may also be included in any other independent claim, even if that claim is not directly made dependent on the independent claim.
[0098] Although specific embodiments have been illustrated and described herein, those of ordinary skill in the art will recognize that the specific embodiments shown and described may be replaced with a variety of alternative and / or equivalent implementations without departing from the scope of the inventive embodiments. This application is intended to cover any adaptations or variations of the embodiments discussed herein. Therefore, the embodiments are intended to be limited only by the claims and their equivalents.
Claims
[1] A method (100) for training an object detector using a Siamese neural network, the method (100) comprising: Obtaining (110) first training data and second training data, wherein the first training data includes one or more images of one or more objects and previously defined class labels for the objects, and wherein the second training data includes one or more concept images of one or more concepts and previously defined concept labels for the concepts; Training (120) the Siamese neural network based on the first training data to classify the objects; and Training (130) the Siamese neural network based on the second training data such that a distance of the feature activations of the same concept in the latent space is reduced. [2] The method (100) of claim 1, wherein training the Siamese neural network comprises: Obtaining feature maps from one or more layers of the Siamese neural network; Unrolling the feature maps to obtain one-dimensional feature vectors for the concept images; and Adapting the Siamese neural network based on a loss function applied to the one-dimensional feature vectors such that, in a latent space of the Siamese neural network, a distance between feature vectors of the same concept is reduced and a distance between feature vectors of other concepts is increased. [3] The method (100) of claim 2, wherein the method (100) further comprises: Determining clusters for the feature vectors of the same concept in the latent feature space. [4] The method (100) of claim 3, wherein the method (100) further comprises assessing an ability of the Siamese neural network to recognize the concepts based on the clusters. [5] The method (100) of claim 4, wherein assessing the ability of the Siamese neural network to recognize the concepts comprises: Obtaining the V-measure for the clusters; and Evaluate the ability of the Siamese neural network to recognize the concepts based on the V-measure. [6] The method (100) of claim 5, wherein obtaining the V-measure for the clusters comprises obtaining a pairwise V-measure for a first and a second cluster of different concepts. [7] The method (100) of any one of claims 3 to 6, wherein the method (100) further comprises: Precisely specifying a corresponding center of the clusters using an Expectation-Maximization (EM) algorithm to obtain a density function of the clusters; and Applying a mixture distribution model to the density function to test whether features from the concept images correspond to at least one of the clusters to assess the ability to recognize a corresponding concept from the concept images. [8] The method (100) of any one of claims 4 to 7, wherein the method (100) further comprises adapting the Siamese neural network for classification based on the ability to recognize the concepts. [9] Neural network obtainable by means of the method (100) according to any one of the preceding claims. [10] A method for testing an object detector, the method comprising applying the neural network of claim 9 for error detection. [11] A computer program comprising instructions which, when executed by a computer, cause the computer to perform the method (100) according to any one of claims 1 to 9. [12] Facility (500) comprising: one or more interfaces (510) for communication; and a data processing circuit (520) configured to carry out the method (100) according to any one of claims 1 to 9. [13] A vehicle comprising the device (500) according to claim 12.