Training neural network to more reliably detect object even if object is of unknown type

By introducing an object-specific contribution into the loss function, independent of category information, and optimizing the neural network parameters, the reliability problem of OOD object detection is solved, and the safety and real-time performance of the autonomous system are improved.

CN120877236APending Publication Date: 2025-10-31ROBERT BOSCH GMBH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510544920.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2024-04-30
Filing Date
2025-04-28
Publication Date
2025-10-31

AI Technical Summary

Technical Problem

Existing object detectors struggle to reliably detect objects that were not seen during training or that are significantly different from known categories (out of distribution, OOD), especially in autonomous driving and robotic applications, potentially leading to erroneous emergency braking or avoidance maneuvers.

Method used

By introducing an object-specific contribution into the loss function, independent of category information, and utilizing a feature extractor network and a classifier head, combined with an object-specific head or a regressor head, the parameters of the neural network are optimized to detect the presence or absence of objects. Furthermore, the detection reliability is improved through cross-entropy and bounding box intersection consistency.

Benefits of technology

It improves the reliability of detecting unknown objects, avoids false detections, ensures the security and real-time performance of autonomous systems, does not affect the detection performance of known categories, and is suitable for applications with limited hardware resources.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120877236A_ABST
    Figure CN120877236A_ABST
Patent Text Reader

Abstract

A method for training a neural network configured to extract features from an image by means of a feature extractor network and to determine classification scores relative to one or more classes in a given set of classes from these features by means of a classifier head, the method comprises the steps of: providing a training image and a corresponding reference truth classification score; processing the training images or regions thereof into classification scores using a neural network; calculating a value of a loss function dependent at least on the deviation of the o classification score from the reference truth classification score, and on the presence or absence of the object, but independent of the object nature contribution of the category information; and optimizing the parameters characterizing the behavior of the neural network towards the goal of improving the value of the loss function.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to image classification, and more particularly to the detection and / or localization of objects of known and unknown types in images. Such detection is crucial for safety-related applications, such as the autonomous movement of vehicles or robots. Background Technology

[0002] Autonomous operation of vehicles or robots in company premises or even on public roads requires constant monitoring of their surroundings. Acquiring and analyzing images of these environments is a crucial part of this monitoring. Of particular importance is the detection of objects that the vehicle and / or robot may collide with.

[0003] Object detectors locate objects of interest and assign classification scores relative to one or more categories of a given classification to detected object instances. That is, they can classify portions of an image into a specific type of object. However, they struggle to reliably detect out-of-distribution (OOD) objects, which were not seen during training or are significantly different from known categories. An example of such an object is cargo lost by another vehicle on the road, such as skis or a piece of furniture. Summary of the Invention

[0004] This invention provides a method for training a neural network. The neural network is configured to extract features from an image using a feature extractor network. A classifier head of the neural network then uses these features to determine a classification score relative to one or more categories in a given set of categories. For example, the feature extractor may include one or more convolutional layers, and the classifier head may include one or more fully connected layers. Above the classifier head, the neural network may include other heads that utilize the features, such as one or more regressor heads that determine the value of any sought quantity.

[0005] In this method, training images and corresponding ground truth classification scores are provided. These training images or their regions are then processed into classification scores using a neural network. Specifically, the object detector first detects regions of interest (ROIs) in the image that indicate the presence of object instances, and then the neural network maps each ROI to a classification score. To evaluate the quality of the neural network's output, the value of a loss function is calculated. This loss function includes at least two contributions:

[0006] • The deviation between the classification score and the baseline true classification score, and

[0007] • The objectness contribution depends on the presence or absence of objects, but is independent of category information. In particular, this objectness contribution may be related to a specific region of interest in the image, rather than the entire image.

[0008] For example, the objectivity contribution could be determined based on the classification score, but this is not required. Instead, the objectivity contribution can be determined based on other outputs of the neural network, such as the output of a separate head dedicated to objectivity, or even directly on the features.

[0009] The parameters characterizing the behavior of the neural network are optimized with the goal of improving the loss function value.

[0010] The inventors have discovered that including the object's contribution in the loss improves the reliability of object detection. Specifically, reliability can include ensuring that if an image shows the presence of an actual object at a particular location, pixels or other portions of the image belonging to that object are correctly identified. Furthermore, reliability can include ensuring that "ghost" objects are not detected if the image shows no actual object. The relative importance of these objectives depends on the application at hand. For example, in autonomous driving applications, detecting the presence of an object is extremely important to prevent collisions between automated vehicles or robots and objects. However, in public road traffic, avoiding false detections of objects that do not actually exist is also crucial. False detections can trigger unwarranted emergency braking or evasive maneuvers. Such maneuvers can startle other road users, potentially leading to rear-end collisions.

[0011] Improved detection (also for out-of-distribution, unseen objects) is particularly important for rapid response to completely unexpected situations in road traffic. For example, encountering a cow, sofa, or ski on a highway is highly unusual. However, if the sofa or ski is indeed present (because someone lost it during transport), or if the cow is present (because it escaped from the farm), triggering emergency braking and / or swerving is crucial.

[0012] Furthermore, the independence of objectivity contributions from category information differs first and foremost from using category-agnostic cues alone. Objectivity contributions, independent of category information, can fully utilize features that indicate categories in specific combinations. For example, when representing a person, a specific combination of values ​​for the features “height,” “body shape,” and “facial expression” can indicate the category “investment banker,” while other combinations of these feature values ​​can indicate the category “stubborn criminal.” Objectivity analysis can determine exactly where a person exists in an image using the features “height,” “body shape,” and “facial expression,” independent of that person’s latent category. In contrast, using category-agnostic cues alone would mean excluding the features “height,” “body shape,” and “facial expression” from the analysis of whether a person is actually present. This would be a bad thing, because, for example, a person walking on a road between two parked cars might be partially obscured by the parked cars in the image, leaving only their face visible. Then, the feature “facial expression” might be the only cue that allows timely detection of the person before they walk onto the road directly in front of an automated vehicle.

[0013] In particular, the independence of objectivity contributions from category information ensures that the neural network will truly learn the concept of objectivity in a general sense, rather than simply expanding its known set of categories by adding more categories. This cannot be guaranteed merely by exposing the neural network to outliers that do not belong to the original category set; such exposure to outliers can also trigger the learning of objectivity concepts. Furthermore, inference by the trained network is not slowed down, whereas inference might be slowed down, for example, when simply adding anomaly detection to an existing neural network. Therefore, the real-time performance important for autonomous driving and other time-critical applications is not hampered.

[0014] When we talk about objectivity in a general sense, objects typically contain well-connected surfaces and have specific geometric structures. This kind of cue is common across different category categories. Learning this cue to detect the presence of objects allows for generalization from known categories to unknown categories during reasoning. This would be impossible if the cue were category-specific, because objectivity would only be biased towards known categories.

[0015] Learning object properties during training requires some effort, but the ability to detect objects does not necessarily impose a significant computational burden during inference. This is particularly important for automotive and other mobile applications where hardware resources are limited, but rapid decisions about the presence or absence of objects are needed. That is, during inference, only a limited amount of additional resources can be dedicated to the additional capability of detecting unknown objects.

[0016] Furthermore, learning object properties does not degrade the performance of a neural network on known categories within a given set of categories. For example, if contributions related to classification scores and object properties are added to the loss function, this cannot offset the large value of the contribution related to classification scores, even if the object property contribution is too low for a particular training image. That is, a neural network cannot avoid the burden of becoming good at determining classification scores by becoming good at determining object properties.

[0017] In a particularly advantageous embodiment, the objectivity contribution depends on the output of a separate objectivity head of the neural network, which at least predicts occupancy in a class-agnostic manner. Occupancy is a measure of whether a feature indicates the presence of an object. This further translates to whether a specific region in an image indicates the presence of an actual object. In one example, the objectivity head may include several convolutional layers with non-linear activation functions interspersed therebetween. Such an objectivity head can predict a single log-odds value. For example, a sigmoid mapping can then map this single log-odds value to a value between 0 and 1. In particular, the presence of a dedicated objectivity head provides an additional possibility for maintaining objectivity training without degrading performance on known classes. For example, objectivity training can be restricted to optimizing parameters characterizing the behavior of the objectivity head, while parameters characterizing the behavior of the classifier head and feature extractor remain frozen. Furthermore, the presence of a separate objectivity head ensures that the classifier head will output its classification score relative to a given known class without additional delay. That is, the determination of objectivity is performed purely on top of the determination of the classification score.

[0018] Instead of using a dedicated object-oriented head to determine occupancy or in combination with it, occupancy can be determined based on the output of a regressor in a neural network. For example, the YOLOX architecture may include...

[0019] • Feature Pyramid Network (FPN), used as a feature extractor for the output feature map set;

[0020] • The regressor head maps features from the feature map to the bounding box coordinates of the detected object instances;

[0021] • A classifier head that maps features from the feature map to the class labels of detected object instances; and

[0022] • The objectivity head uses intermediate working products from the regressor head to predict class-agnostic objectivity, and may optionally also use features from the classifier head. That is, for example, features from bounding box regression can be fused with features from the classifier head.

[0023] Therefore, in another particularly advantageous embodiment, the neural network is further configured to predict the bounding boxes of objects. These bounding boxes corresponding to object instances provide another objectivity concept that can be rationalized in reference to the output of the objectivity head.

[0024] Therefore, in another particularly advantageous embodiment, the objectivity contribution depends on the degree to which the occupancy rate aligns with one or more intersections between the predicted bounding box and the ground truth bounding box. That is, the predicted bounding box localization can be used to generate the ground truth loss for occupancy rate prediction: if the predicted bounding box has high overlap with the ground truth bounding box annotation, this indicates that the occupancy rate o should be high; otherwise, the occupancy rate o should be low. For example, the occupancy rate o can be compared with the following expression.

[0025]

[0026] Where |·| represents the area in pixels, b p It predicts the bounding box, and This represents the i-th bounding box among a total of n bounding boxes, regardless of its category.

[0027] In particular, one advantage of this method is that the detection of unknown objects does not depend on portions of the image that indicate the object's presence, portions sufficient to definitively identify the object by matching it with a benchmark ground truth for a specific category. That is, even if an object is partially occluded or otherwise difficult to identify, it is still possible to detect that at least one object exists. For example, if a vehicle is partially occluded, it may be difficult to distinguish the specific type of vehicle (e.g., bus, van, or brand and model). Furthermore, if a person is partially visible between parked cars, it is irrelevant if the specific type of person cannot be determined. What matters is detecting a person, or more abstractly, detecting that at least one object that should not be run over.

[0028] One way to measure the consistency between this expression on one side and the occupancy rate o on the other is through cross-entropy, such as binary cross-entropy (BCE). Therefore, the objectivity contribution L of the loss... occ It can be done in form

[0029]

[0030] In information theory, cross-entropy is a measure of the quality of a (probabilistic) distribution model. Therefore, given a distribution, optimizing model parameters toward minimizing cross-entropy is equivalent to maximizing the model's log-likelihood.

[0031] A particular advantage of having an occupancy score o is that this score o can be determined for all images. That is, it can be trained using all available training images (not just training images outside the known categories of a given set of categories). In contrast, training for determining the classification score relative to a newly introduced OOD category is very likely to overfit to the available OOD examples, as these available OOD examples are far fewer in number than the in-distribution training examples.

[0032] In response to L occ The expression can be further simplified by the following approximation for intersection and union calculation.

[0033]

[0034] That is, b can be calculated. p Boundary boxes of each reference truth value Instead of calculating the predicted bounding box b, smaller and simpler intersections are used. p bounding boxes with all baseline truth values The intersections between the unions are quite complex. For most of these intersections, it will be quickly determined that they are empty without requiring extensive computation, thus resulting in a net saving in computational time.

[0035] Therefore, in another particularly advantageous embodiment, the intersection between the union of the predicted bounding box and the reference ground truth bounding box is approximately the sum of the intersections between the predicted bounding box and each reference ground truth bounding box.

[0036] In another particularly advantageous embodiment, the given set of categories is expanded by additional categories of objects that do not belong to any category in the given set of categories. In this way, the classifier head of the neural network gains the opportunity to express the discovery that a detected object is an unseen object. That is, the output of the classifier head can distinguish between "no object" on one hand and "object present but unseen" on the other. Without additional OOD categories, the classifier head would have to express both "no object" and "object present but unseen" with low scores for all given categories, or it might even be tempted to output high classification scores for any category in the given set, all of which would be erroneous. Furthermore, additional classification scores for unseen objects in additional categories will be obtained during inference with only a small (if any) additional computational burden.

[0037] In another particularly advantageous embodiment, the set of training images is expanded using training images that do not belong to any category in the given set of categories. In this way, the neural network gains an improved opportunity to detect unseen objects. The additional training images used for this expansion can be derived from any suitable source. For example, multiple images from different datasets can be combined into a single image using any well-known augmentation technique (such as Mosaic or Mixup) to expose the neural network to unseen objects. Exposure to diverse objects enhances the acquisition of a more general understanding of objectivity and can be performed in any suitable manner. For example, training images from other datasets can be used, and they can be further modified using any suitable data augmentation technique. In one example, a dataset with training images for traffic conditions for automated driving can be expanded using additional training images from: the general MS COCO (Microsoft Contextual Common Objects) large-scale object detection, segmentation, and captioning dataset, and / or the LVIS dataset for large-scale lexical instance segmentation. In particular, this enhances the tendency of objectivity scores to respond to objects from both known and unknown categories, while remaining silent on "filler" categories such as roads and sky.

[0038] In another particularly advantageous embodiment, images acquired by at least one sensor are processed by a trained machine learning model into classification scores, and optionally also into occupancy and / or objectivity scores. The improved training then has the effect of making object detection more reliable, regardless of whether the objects belong to categories within or outside the original given set of categories.

[0039] In another particularly advantageous embodiment, depth information is used to verify the presence of an object in response to a classification score and / or occupancy indicating its presence. For this purpose, depth information of the image region associated with the object is acquired. It is then determined whether this depth information indicates a depth change that could be expected given the presence of the object. If this determination is negative, i.e., if the expected depth change does not exist, the object detection is determined to be a false detection. Specifically, if the image exhibits features that somehow resemble an object but are not actually objects, these features will not be falsely detected as objects. An example of such features is shadows. Although they are produced by the presence of actual objects, they appear in other locations where no objects are present. Examples of depth changes indicating the presence of an object include small local depth changes within the bounding box associated with the object. In contrast, for example, the flat surface of a road only exhibits continuous local depth changes.

[0040] In another particularly advantageous embodiment, the classification score and occupancy rate are evaluated together to verify the presence of an object. To this end, if the classification score and / or occupancy rate indicate the presence of an object, the product of the maximum classification score and occupancy rate associated with the detected object is calculated. If this product is below a predetermined threshold, the detection of the object is determined to be a false detection. As discussed earlier, this approach works best if there are additional categories for objects that do not belong to any category in the given category set. The presence of an object can then be confirmed using two separate heads (i.e., the classifier head and the object attribute head) before determining that the object actually exists.

[0041] In another particularly advantageous embodiment, an actuation signal is calculated at least in part based on classification scores and / or occupancy rates output by a trained machine learning model, and / or based on object detection. The actuation signal is then used to actuate vehicles, robots, driver assistance systems, quality control systems, surveillance systems, and / or medical imaging systems. In this way, the probability that the response of the corresponding actuated system to the actuation signal is appropriate given the conditions characterized by the acquired images is improved. In particular, fewer responses to the actual presence of an object are missed, and fewer responses are performed to object detection that does not correspond to an actual present object. For example, in an automated driving system, emergency braking or evasive maneuvers will be triggered more reliably if an object is indeed present in the vehicle's path, but there will be no "sudden" emergency braking or evasive maneuvers without apparent cause if no object is actually present in the vehicle's path.

[0042] This method can be implemented entirely or partially by a computer and embodied in software. Therefore, the present invention also relates to a computer program having machine-readable instructions that, when executed by one or more computers and / or computing instances, cause those computers and / or computing instances to perform the method described above. In this document, control units for vehicles or robots, as well as other embedded systems capable of executing machine-readable instructions, should also be considered as computers. Computing instances include virtual machines, containers, or other execution environments that allow the execution of machine-readable instructions in the cloud.

[0043] Non-transitory storage media and / or downloadable products may include computer programs. Downloadable products are electronic products that can be sold online and transmitted over a network for immediate fulfillment. One or more computers and / or computing instances may be equipped with the computer programs and / or the non-transitory storage media and / or downloadable products.

[0044] The invention will be described below using accompanying drawings, without any intention to limit the scope of the invention. Attached Figure Description

[0045] Figure 1 Exemplary embodiments of the method 100 for training neural network 1;

[0046] Figure 2 Visualization of an exemplary processing pipeline for training and inference according to method 100;

[0047] Figure 3 : An example of an image that causes false object detection. Detailed Implementation

[0048] Figure 1 This is a schematic flowchart of an exemplary embodiment of a method 100 for training a neural network 1, wherein the neural network 1 is configured as follows:

[0049] • Feature 3 is extracted from image 2 using feature extractor network 4, and

[0050] • Using the classifier head 6, a classification score 5 is determined relative to one or more categories in a given set of categories based on these features 3.

[0051] According to box 105, neural network 1 can be further configured to predict bounding boxes 10 for objects detected in image 2.

[0052] In step 110, training image 2a and corresponding benchmark ground truth classification score 5a are provided.

[0053] According to box 111, the set of training images 2a can be expanded using training images 2a* that do not belong to any category in the given set of categories.

[0054] In step 120, the training image 2a is processed into a classification score 5 using neural network 1.

[0055] In step 130, the value 7a of the given loss function 7 is calculated. The loss function 7 depends at least on

[0056] • The deviation between classification score 5 and the baseline true classification score 5a, and

[0057] • The objectivity contribution depends on the presence or absence of the object, but is independent of category information.

[0058] According to box 131, the objectivity contribution can depend on the output of an additional objectivity head 8 of neural network 1, which predicts occupancy 9 in a class-agnostic manner. This occupancy 9 is a measure of whether feature 3 indicates the presence of an object.

[0059] According to box 132, the objectivity contribution may depend on the degree to which the occupancy 9 is consistent with one or more intersections between the predicted bounding box 10 and the baseline ground truth bounding box 10a.

[0060] According to box 132a, the intersection between the union of the predicted bounding box 10 and the reference ground truth bounding box 10a can be approximated as the sum of the intersections between the predicted bounding box 10 and each reference ground truth bounding box 10a.

[0061] According to box 132b, the consistency between occupancy 9 and the intersection can be measured by cross-entropy.

[0062] According to box 133, a given set of categories can be expanded by additional categories of objects that do not belong to any category in the given set of categories.

[0063] In step 140, the parameters 1a that characterize the behavior of the neural network 1 are optimized with the goal of improving the value 7a of the loss function 7. The final optimized state of the parameters is labeled with reference numeral 1a*, and this final optimized state characterizes the training state 1* of the neural network 1.

[0064] exist Figure 1 In the example shown, in step 150, the image 2 acquired by at least one sensor 11 is processed by a trained machine learning model 1* into a classification score 5, and optionally also into an occupancy score 9 and / or an objectivity score.

[0065] In step 160, it is checked whether the classification score 5 and / or occupancy rate 9 and / or objectivity score indicate the presence of an object. If this is the case (truth value 1), false positive detections can be eliminated by one or two of the following methods, which can also be repeatedly performed on multiple instances of the detected object.

[0066] According to the first method, in step 170, depth information 12 of the image region associated with the object (labeled O here) is acquired. Then, in step 180, it is determined whether the depth information 12 indicates a depth change that can be expected given the presence of the object. If this is not the case (true value 0), then in step 190, it is determined that the detection of object O is a false detection.

[0067] According to the second method, in step 200, the product 13 of the maximum classification score 5 and the occupancy rate 9 associated with the detected object O is calculated. In step 210, it is checked whether the product 13 is higher than a predetermined threshold 14. If this is not the case (true value 0), then in step 220, it is determined that the detection of object O is a false detection.

[0068] In step 230, an actuation signal 230a is calculated, at least in part, based on the classification score 5 and / or occupancy rate 9 output by the trained machine learning model 1 and / or the detection of object O. In step 240, the actuation signal 230a is used to actuate the vehicle 50, the driver assistance system 51, the robot 60, the quality inspection system 70, the monitoring system 80, and / or the medical imaging system 90.

[0069] Figure 2 An exemplary processing pipeline for training and for inference according to the method 100 described above is illustrated.

[0070] Figure 2 a illustrates the pipeline used for training. According to box 111, training images 2a from the applicable domain (here: autonomous driving) are combined with additional training images 2a* from outside that domain. The combined set of training images is fed to the feature extractor 4 of the neural network 1. This produces extracted features 3.

[0071] The extracted feature 3 is fed to the classifier head 6, which produces a classification score 5 relative to the in-distribution class (ID) in the given class set and relative to the new out-of-distribution class (OOD) associated with the object of the unknown class.

[0072] The extracted feature 3 is also fed into the object attribute head 8. This object attribute head 8 predicts at least occupancy 9 in a class-agnostic manner. Occupancy 9 is a measure of whether feature 3 indicates the presence of an object. Figure 2 In the example shown, the objectality head 8 is further configured to determine an "obj" score, which is another concept related to objectality. The objectality head 8 is also used to predict the bounding box 10 of the object.

[0073] exist Figure 2 In the example shown, the training of neural network 1 is directed towards the following objective:

[0074] ·For example, by the loss component L box The measured predicted bounding box 10 corresponds to the baseline true bounding box 10a;

[0075] ·For example, by the loss component L occ The measured occupancy rate of 9 is consistent with that of bounding boxes 10 and 10a.

[0076] ·For example, by the loss component L obj The object fraction “obj”, measured in any suitable manner, is consistent with the available benchmark truth; and

[0077] • For example, the classification loss component L cls The measured classification score 5 corresponds to the baseline true value classification score 5a.

[0078] Figure 2 b begins with the following assumption: Neural Network 1 has been based on... Figure 2 The pipeline shown in Figure a is trained. Figure 2 b illustrates an exemplary inference pipeline. Image 2 is fed to a trained neural network 1*. The neural network 1* then outputs both an occupancy map 9 and a classification score 5 for image 2. For objects within the distribution of known categories, both the occupancy map 9 and the classification score 5 are visualized using bounding boxes 10 (ID), and for objects outside the distribution of unknown categories, both the occupancy map 9 and the classification score 5 are visualized using bounding boxes 10 (OOD).

[0079] exist Figure 2 In the example shown in b, for an exemplary object O whose presence is not indicated by classification score 5, occupancy map 9 has a high occupancy score. Furthermore, classification score 5 indicates the presence of another object O' that is not readily apparent from occupancy map 9. In step 200, the product 13 of classification score 5 and occupancy map 9 is calculated, and if the product is above a threshold 14, object O is determined to exist. Figure 2 In the example shown, the total set of detected objects O is the union of the objects detected using both classification score 5 and occupancy graph 9.

[0080] Figure 2 c illustrates how these objects O are further filtered using depth information 12 according to steps 180 and 190 of method 100. Figure 2 In the example shown, for object O, there is matching depth information 12, so object O is retained. In contrast, for object O', no matching depth information 12 is available. Therefore, object O' is discarded as a false detection.

[0081] Figure 3 Image 2 shows some examples of road scenes that caused erroneous object detections. The bounding boxes for these erroneous detections are drawn with dashed lines.

[0082] From the Fishyscapes dataset Figure 3 In case a, manholes flush with the road surface and graffiti painted on the road surface cause false detections.

[0083] From the Cityscapes dataset Figure 3 In b, subtle changes in road texture caused by road surface repair lead to false detections.

[0084] From the BDD100K dataset Figure 3 In C, road markings cause error detection.

Claims

1. A method (100) for training a neural network (1), the neural network (1) being configured to extract features (3) from an image (2) by means of a feature extractor network (4), and to determine a classification score (5) relative to one or more categories in a given set of categories by means of a classifier head (6), the method (100) comprising the steps of: • Provide (110) training images (2a) and corresponding baseline ground truth classification scores (5a); • Using the neural network (1), these training images (2a) or regions of these training images (2a) are processed (120) into classification scores (5); • Calculate the value (7a) of the loss function (7), which depends at least on o The deviation between the classification score (5) and the baseline true classification score (5a), and o The objectivity contribution depends on whether the object exists or not, but is independent of category information; as well as • Optimize (140) the parameters (1a) that characterize the behavior of the neural network (1) with the goal of improving the value (7a) of the loss function (7).

2. The method (100) according to claim 1, wherein the objectivity contribution depends (131) on the output of an additional objectivity head (8) of the neural network (1), the additional objectivity head (8) predicting at least occupancy (9) in a class-agnostic manner, the occupancy (9) being a measure of whether the feature (3) indicates the presence of an object.

3. The method (100) according to any one of claims 1 to 2, wherein the neural network (1) is further configured (105) to predict the bounding box (10) of an object.

4. The method (100) according to claims 2 and 3, wherein the objectivity contribution depends (132) on the degree to which the occupancy rate (9) is consistent with one or more intersections between the predicted bounding box (10) and the baseline ground truth bounding box (10a).

5. The method (100) according to claim 4, wherein the intersection between the union of the predicted bounding box (10) and the reference ground truth bounding box (10a) is approximately (132a) the sum of the intersections between the predicted bounding box (10) and each reference ground truth bounding box (10a).

6. The method (100) according to any one of claims 4 to 5, wherein the consistency between the occupancy rate (9) and the intersection is measured by cross-entropy (132b).

7. The method (100) according to any one of claims 1 to 6, wherein the given set of categories is expanded (133) by additional categories of objects that do not belong to any category of the given set of categories.

8. The method (100) according to claim 7, wherein the set of training images (2a) is expanded (111) by using training images (2a*) that do not belong to any category in the given set of categories.

9. The method (100) according to any one of claims 1 to 8, further comprising: The image (2) acquired by at least one sensor (11) is processed (150) into a classification score (5) by a trained machine learning model (1*), and optionally also into an occupancy rate (9) and / or an objectivity score.

10. The method (10) according to claim 9, further comprising: In response to the classification score (5) and / or the occupancy rate (9) and / or the objectivity score indication (160) of the presence of object (O): • Obtain (170) depth information (12) of the image region associated with the object (O); • Determine (180) whether the depth information (12) indicates a depth change that can be expected given the presence of the object (O); as well as • If the determination is negative, then the detection of the object (O) is determined to be an erroneous detection. (190) 11. The method (100) according to any one of claims 9 to 10, further comprising: In response to the classification score (5) and / or the occupancy rate (9) and / or the objectivity score indication (160) of the presence of object (O): • Calculate (200) the product (13) of the maximum classification score (5) associated with the detected object (O) and the occupancy rate (9); and • If the product (13) is (210) below a predetermined threshold (14), then (220) the detection of the object (O) is determined to be an erroneous detection.

12. The method (100) according to any one of claims 9 to 11, further comprising: • The actuation signal (230a) is calculated (230) based at least in part on the classification score (5) and / or occupancy rate (9) output by the trained machine learning model (1), and / or on the detection of the object (O); and • The actuation signal (230a) is used to actuate (240) the vehicle (50), the driver assistance system (51), the robot (60), the quality inspection system (70), the monitoring system (80), and / or the medical imaging system (90).

13. A computer program comprising machine-readable instructions that, when executed by one or more computers and / or computing instances, cause the one or more computers and / or computing instances to perform the method (100) according to any one of claims 1 to 12.

14. A non-transitory machine-readable storage medium and / or downloadable product having a computer program according to claim 13.

15. One or more computers and / or computing instances having a computer program as claimed in claim 13, and / or having a non-transitory machine-readable storage medium and / or downloadable product as claimed in claim 14.