Method for detecting objects from temporally resolved sensor data
By leveraging temporal consistency and masking techniques, the method reduces resource requirements for DNN-based object detection, enabling reliable real-time outlier detection in resource-constrained environments.
Patent Information
- Application Number
- DE102023212875
- Authority / Receiving Office
- DE · DE
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2023-12-18
- Publication Date
- 2025-06-18
AI Technical Summary
Existing object detection methods using deep neural networks (DNNs) require significant computational resources and memory, making real-time outlier detection challenging, especially in resource-constrained environments like autonomous vehicles.
A method that utilizes temporal consistency in sensor data to reduce the computational load by pre-calculating outlier information for DNNs, estimating missing information based on changes in successive frames, and applying masking techniques to focus on relevant areas for object detection, thereby minimizing resource requirements while maintaining accuracy.
This approach allows for accurate real-time outlier detection with reduced computational and memory consumption, enhancing the reliability of DNN-based object detection systems without compromising performance.
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
TECHNICAL FIELDThe invention relates to a (computer-implemented) method for detecting objects from time-resolved sensor data by means of an artificial deep neural network which is monitored while it carries out the object detection. The invention further relates to a method for controlling an autonomous movement device and corresponding control units or devices.BACKGROUNDIn monitoring deep neural networks (DNNs) for detecting objects from sensor data at runtime, the goal is to identify input-output combinations for which the DNN outputs are poorly trusted and may require error handling (in real-time). One possible error handling is, for example, the consideration of further alternative evaluation algorithms for object detection. To identify such exceptional cases (outliers) in practice, different reasons for the low trustworthiness and different information sources for the estimation are differentiated. Reasons may be that inputs were too rarely / not represented in the training dataset (out-of-distribution inputs), or that the DNN behavior conflicts with given boundary conditions (implausible outputs).Information sources for monitoring can be both the inputs and the DNN behavior, i.e., intermediate / final outputs with or without temporal evolution, as well as outputs of surrogate models.EP 4099210 A1 shows a method for training a neural network for semantic image segmentation and use in a vehicle control unit or a backend system. In a first step, image data of a sequence of individual images are received. Subsequently, a single image-based evaluation of the semantic segmentation predictions of one or more objects in individual images takes place. Furthermore, a sequence-based evaluation of the temporal characteristics of semantic segmentation predictions of the one or more objects is carried out in at least two individual images. The results of the frame-based evaluation and the sequence-based evaluation are combined.EP 3896651 A1 discloses a method for evaluating temporal characteristics of semantic image segmentation, in particular for detecting perception instabilities in time-sequential images of a video sequence. In a first step, image data of two successive individual images are received. Semantic segmentation predictions are then determined for each of the two sequential frames. The displacements between the image data of the two sequential frames are estimated. The estimated displacements are applied to the semantic segmentation prediction of the first of the two sequential frames to generate an expected semantic segmentation prediction for the second of the two sequential frames. The semantic segmentation prediction of the second of the two sequential individual images is evaluated on the basis of the expected semantic segmentation prediction for the second of the two sequential individual images.EP 0680026 A2 shows a monitoring system in road traffic, in which exteriors are detected and processed in the form of pixels or points from sensor image data, which differ significantly in brightness or intensity from the pixels surrounding them.WO 2022016011 A1 shows anomaly detection by means of artificial neural networks from statistical data acquired by a plurality of sensors for quality control of manufacturing and processing methods for electronic devices. The method includes, among other things, processing a plurality of outlier values using a detector neural network to generate an anomaly value indicative of a probability of an anomaly associated with the manufacturing process.DE 10 2021 211 503 B3 shows a computer-implemented method for monitoring the logical consistency of an artificial neural network. First, activation data (e.g., in the form of activation values or activation maps) of the artificial neural network are read in, which result from input data. The activation data are passed to at least one trained concept model that is trained for the recognition and, if appropriate, localization of a partial feature of the features contained in the input data and for the output of a calibrated partial feature mask.The final output data are linked to the partial feature truth values by means of a fuzzy logic unit in such a way that a continuous logical consistency truth value results therefrom. The logical consistency truth value is evaluated by means of an evaluation unit, wherein a logical inconsistency of the final output data in an inconsistency range is determined if the consistency truth value falls below a predefined threshold value.If the predefined threshold value is undershot, different measures can be initiated. An alarm can be triggered, for example. Furthermore, it is conceivable that an uncertainty measure (possibly locally) is increased.With regard to autonomous driving, it is preferred to switch on a redundant (possibly more expensive, i.e. more computationally intensive) evaluation of ambient sensor data in order to add these later estimates. Further (local, possibly more computationally intensive) checks of the output of the artificial neural network can also be activated. Another possibility is that the system assumes a safe state (in particular as long as the safety is not otherwise confirmed). Finally, it is also conceivable to prompt a driver of an autonomously driving vehicle by means of a display or other information for intervention in the vehicle control.SUMMARY OF THE INVENTIONIt is an object of the invention to minimize the resource requirement of object detection methods with integrated monitoring methods (outlier detection), so that they can be reliably carried out in real time even with limited hardware conditions.The object is achieved by the subject matter of the independent claims. Preferred developments are the subject matter of the dependent claims.A first aspect of the invention relates to a (computer-implemented) method for detecting objects from temporally resolved sensor data by a trained neural network having a plurality of layers (layers). The method comprises the following steps:receiving a set of time-resolved sensor data as input data for the trained neural network. The set includes first sensor data acquired at a first time t and second sensor data acquired at a subsequent second time t+1. The sensor data can be, in particular, data of an environment detection sensor, a camera, a camera system having a plurality of cameras, a lidar sensor, a radar sensor. The sensor data may be image data or correspond to image data. The sensor data is generally subject to temporal consistency.determining first object detection data as output data for the first sensor data as input data by the trained neural network.monitoring (or validation or plausibility checking) the object detection by ascertaining a first outlier score (for input-output combinations of the trained neural network) in the following manner: a) precalculating a plurality of first outlier information items from a plurality of intermediate outputs of the trained neural network. The precalculation may require a high demand for resources, e.g. computing capacity or storage capacity. b) calculation of the first outlier score on the basis of the (all) first outlier information from the precalculation.determining second object detection data as output data for the second sensor data as input data by the trained neural network and determining a second outlier score by:c) Determination of changes (F t+1) between first and second sensor data (It, I t+1). These may already be present as a result of object recognition provided these changes take into account in successive sensor data.d) precalculating second outlier information from at least one intermediate output of the trained neural network, wherein fewer second outlier information are calculated than first outlier information.e) estimating the missing second outlier information from corresponding first outlier information taking into account the determined changes.f) calculating the second outlier score from the second outlier information, andoutputting the determined (first and second) object detection data and (first and second) outlier scores.Alternatively to the complete output of object detection data and / or outlier scores, first or second object detection data / outlier scores and information for modification with respect to the first or second object detection data / outlier scores can be output in each case. The output can be effected to subsequent processing steps / units.The outlier score is, for example, a numerical value and specifies the extent to which the ascertained output deviates from a typical case.Advantages of the Invention:Simultaneous use of DNN-internal information for the most accurate possible prediction of outliers without unacceptable efficiency losses; in particular: 1. improvement of real-time capability by targeted reduction of the necessary additional evaluations without appreciable loss of accuracy 2. reduced memory consumptionutilizing the natural temporal consistency of sensor recordings, especially. Camera Images of Road ScenariosSimple implementation as an addition for existing complicated outlier score calculations.This allows general outlier detection for DNN-based perception methods under real conditions. A prerequisite is merely the condition that objects are so slow that they can be recognized over a plurality of sensor data (frames) following one another in time (this could no longer be fulfilled in the case of a rifleball as object).Preferably, in step c), spatial changes of features (e.g. object components) between first and second sensor data are considered as changes, and in step e) the estimation comprises a spatial transformation of the first outlier information(s).Preferably, in steps b) and optionally also d), the outlier information relating to a plurality of layers (or relating to each layer) of the trained neural network each comprise an intermediate output as outlier information. These can be referred to as layer-by-layer outlier information.In one embodiment, each layer of the trained neural network comprises at least one artificial neuron. In the course of the precalculation in steps b) and optionally also d), activation map data are determined by the trained neural network. The activation map data indicates which activation value that neuron has for each neuron. The following steps are carried out:aa) generating masking data from the activation map data and / or the sensor data and / or the object detection data;masking the activation map data from the masking data to obtain masked activation map data, the masked activation map data including non-masked activation values and masked activation values; andcc) determining layer-by-layer outlier information for at least one network layer for the object detection data on the basis of the masked activation map data, wherein only the non-masked values are taken into account when determining the layer-by-layer outlier information. Layer-by-layer outlier information can be, in particular, a numerical value.By using masking, it is possible to focus on areas which are actually of interest for object detection. Furthermore, it can be avoided that activations from the background or from adjacent objects are included in the calculation of the outlier score. In this case, it is advantageous to choose masking on the basis of temporal consistency: all (relevant) points are intended to be covered over time. However, not all relevant points need to be covered for this at any time, but only the integral over time should provide good coverage.Relevance thus does not need to be viewed as a property of sensor data at a single point in time. Masking also allows for a massive reduction in computational and / or storage requirements and may therefore enable the real-time capability of the ideas described herein.It is preferred that in step aa) the masking data are generated by means of a masking threshold value in such a way that those activation values which are below the masking threshold value are masked.In other words, an outlier score is calculated on the basis of neural activations. For better efficiency, outlier information per layer is calculated, which is merged into an overall score. Thus, no or fewer intermediate storage(s) of the neural activations are necessary.Preferably, in step aa), the masking data is generated from the granularity data such that those activation values are masked that spatially conform to regions in the input data that were irrelevant to the output data. Such masking data are also referred to as a Salicy mask.For example, masking is effected by means of Salicy maps which are generated with the aid of backpropagation.An advantage here is: as a result, no neural activations from the forward propagation have to be temporarily stored until the mask calculation by the back propagation is ended. Instead, just after the activations of each layer are calculated, the activations can be masked and the outlier information for the layer evaluated.It is preferred that the masking takes place at the level of the artificial neurons.Alternatively, the masking can be carried out at the level of the sensor data.It is preferred that the temporally resolved sensor data correspond to a sequence of camera images.In this case, changes between first and second sensor data can preferably be ascertained as ascertainment of the optical flow between first and second camera images. This can be done in the form of an (optical) flow card.In one embodiment, a measure is initiated if the first or the second outlier score exceeds a (monitoring) threshold value. One measure can be, for example: switching on further redundant perception or object detection algorithms; outputting a warning; in the case of a control of an autonomous movement device or of an autonomous vehicle: taking a safe state; feedback to safety drivers; reporting to manufacturers, Incident Reporting.It is preferred that in step aa) the masking data is generated from the output data such that those activation values are masked that do not match a predetermined reference class, bounding box, or combination of both.It is preferable that in step aa), the mask data is inverted after being generated or invertedly generated.It is preferred that in step aa) the masking data is generated from the output data by placing a grid of input boxes over the input data and masking those activation values at which the associated input boxes overlap a bounding box in the output data at least partially, preferably more than half by area.It is preferred that in step aa) at least a part of the masking data is generated on the basis of a different criterion or in a different manner than at least a different part of the masking data.It is preferred that steps aa) and bb) are performed multiple times, each for a set of mask data, each set of mask data having been generated on the basis of a different criterion or in a different manner than the preceding set or all preceding sets of mask data.It is preferred that in step cc) the outlier score is determined by determining a neuron coverage for the activation map data and outputting it as the outlier score; or by determining an average neuron coverage when training the DNN, wherein the neuron coverage determined in operation is compared with the average neuron coverage and the difference is output as the outlier score.It is preferable that in step cc), the outlier score is obtained by obtaining a correlation value for the activation map data and outputting it as the outlier score.It is preferred that in step cc) the outlier score is determined by clustering the activation map data obtained during training of the DNN, wherein a distance to the nearest cluster is determined for the activation map data determined during operation and the distance is output as the outlier score.A further aspect of the invention relates to a computer-implemented perception method, in particular for object detection (object detection), the method comprising: a) providing image data by means of a storage medium or a sensor; b) processing the image data by a deep neural network trained for the object detection in order to obtain object detection data which indicate where an object of a specific semantic class is located in the image data; c) monitoring the deep neural network by a method described above in order to obtain an outlier score for the result from step b); d) assigning the outlier score to the object detection data in order to process the outlier score together with the object detection data in order to generate a control signal.A further aspect of the invention relates to a computer-implemented method for controlling an autonomous movement device by a control unit, the method comprising: a) carrying out a method described above in order to obtain object detection data which are indicative of an object type and object position of an object contained in the temporally resolved sensor data, wherein the object detection data additionally contain an outlier score for at least two points in time; b) generating a control signal for the control unit, wherein the control signal causes the control unit to control the autonomous movement device in accordance with the control signal.Examples of autonomous movement devices are, for example: autonomous vehicles, ships, drones or also robots.A further aspect of the invention relates to a data processing device which has units which are configured to carry out one, more or all steps of a method described above.A further aspect of the invention relates to an autonomous movement device which comprises at least one sensor and a data processing device designed as a control unit, wherein the control unit is connected to the sensor in order to process the image data thereof.A further aspect of the invention relates to a computer program comprising instructions for a data processing device, wherein the instructions cause the data processing device to perform one, more or all steps of a method described above.A further aspect of the invention relates to a machine-readable data carrier medium or data carrier signal which comprises or contains the computer program described above.A further aspect of the invention relates to a computer-implemented method for monitoring an artificial deep neural network (DNN) having a plurality of network layers each with at least one artificial neuron, the method comprising: a) supplying input data to a deep neural network to be monitored trained on a task to obtain output data and activation map data therefrom, the output data being generated according to the trained task, the activation map data indicating for each neuron which activation value this neuron has; b) supplying the input data, the output data and the activation map data to a computer-implemented network observer performing the following steps:aa) generating mask data from the activation map data and / or the input data and / or the output data;masking the activation map data from the masking data to obtain masked activation map data, the masked activation map data including non-masked activation values and masked activation values;cc) determining an outlier score for the output data based on the masked activation map data, wherein only the non-masked values are taken into account when determining the outlier score, wherein the outlier score is a numerical value;dd) assigning the outlier score to the respective output data and optionally to the input data for further joint processing.Furthermore, no outliers are required in the training dataset. By considering the intermediate outputs, i.e. the activations, the chances of processing errors or inconsistencies being detected are improved. Further, the distinction between examples that are "in-distribution" but still difficult to process for the DNN (e.g., model-specific decision boundaries) and examples that are out-of-distribution from the DNN perspective can be improved.One reason for undesirable behavior of the neural network may be erroneous or incomplete training data. In order to exclude such errors, it is an approach for the entire data set to check neuron activation. It can thus be detected whether there are dead neurons or whether certain neurons are activated in an undesirable manner in boundary cases. The aim is to detect the occurrence of these situations, in order to prevent them from subsequently being able to do so.The so-called neuron coverage specifies the ratio of activated neurons to the total number of neurons in a monitored neural network. A neuron is considered to be activated when it exceeds a certain threshold value (threshold). The selected threshold value generally depends on the activation interval available. If the activation of the neuron can assume values between 0 and 1, for example, then the threshold value chosen is typically 0.5.Neuron Coverage can be calculated for a single picture or by summing the activations per neuron. In a large part of the literature, the neuron coverage is applied to the entire network, i.e. all layers and all neurons, to the entire input, i.e. the entire input image, and for all inputs, i.e. the entire data record.However, this approach can be awkward, because very large output data are usually produced, which cannot be created or processed at the runtime of the monitored neural network or in real time, or only with great difficulties. For this purpose, there is the possibility of being limited to individual layers of the neural network or to individual image sections.In one embodiment, the masking data may be configured such that the masking is applied to the outputs of selected layers of the network. In one embodiment, it is possible that the masking data is associated with specific areas of the input data, for example specific image areas in the case of image data as input data.In one embodiment, the masking data may be generated based on different masking criteria. In other words, the masking data can be generated by combining different masking criteria and applied once.In one embodiment, the mask data may be generated separately according to mask criteria. The masking data are then preferably applied successively.In the approach presented here, the neuron coverage is considered as an outlier score, which is determined for masked regions of an image or of an image section. However, another conventional method may be used for the outlier score.One concept is to cut relevant regions from the input data by means of a mask. The relevant areas are associated with object detection.For object detection, the relevant areas are for example those which contribute to the final decision, e.g. the paw and the head of a dog for the recognition of a dog, but not the ground or an adjacent table.The masking of the bounding box is preferably based only on activations that lie on the object. However, this is only useful if the respective neurons can be assigned to a spatial region of the input data. This is the case in particular when one of the network architectures explained below is involved. It can thus be avoided that activations are taken into account in the calculation of the neuron coverage, which activations are only located in background regions or regions of adjacent objects.An idea used here is the selective masking of activation maps (NNs) in neural networks (NNs), such as convolutional neural networks (CNNs for short), transformer networks and recurrent neural networks (RNNs for short), in particular gated recurrent units (GRU for short). Thus, only a portion of the activation information is forwarded to the network observer (network observer) for processing. Neuron coverage methods are preferably used for abnormality detection.Another idea is to use an automatic masking criterion. One possibility for realizing this is the use of Salicy Maps (Salicy Maps). The granularity map may use the spatial mapping of image regions in the input to neural activations. This is particularly applicable to CNNs and similar architectures.In one embodiment, the granularity map is generated perturbation-based. In this case, random features in the input are removed or changed and the effect on the DNN output is measured. From this, the importance of the input features for the output can be determined. Preferably, multiple evaluations of the monitored DNN are performed on altered versions of the input. Therefore, DNN-internal information can be omitted. Examples of these are SLAP (SHapley Additive exPlans) disclosed in Lundberg et al. 2017 "A Unified Approach to Interpreting Model Predictions." In Advances in Neural Information Processing Systems 30, 4765-74, http: / / papers.nips.cc / paper / 7062-a-unified Approach-to-interpreting-model-predictive ns.pdf., and LIME (Local Interpretable Model-agnostic Explantations) disclosed in Riiero et al. 2016 "Why Should I Trust You?": Explaining the Predictions of Any Classifier." In Proc. 22nd ACM SIGKDD Int. Conf. Knowledge Discovery and Data Mining 1135-44. KDD '16th https: / / doi.org / 10.1145 / 2939672.2939778., the contents of which are incorporated herein by reference.In one embodiment, the granularity maps may be generated based on backpropagation. For this purpose, the activations of the output are followed back until the input according to given rules. A special case is gradient backpropagation. This is advantageous because this information is generally obtained anyway during the training of the DNN and can thus be used further. Examples are LRP (Layerwise Relevance Propagation) disclosed in Bach et al. 2015 "On Pixel-Wise Illustrations for Non-Linear Classifier Explanations by Layer-Wise Relevance Propagation." PLOS ONE 10 (7): e0330140. https: / / doi.org / 10.1371 / journal.pon.0130140., and Sensitivity Analysis disclosed in Baehrens et al. 2010, "How to Explain Individual Classification Decisions." Journal of Machine Learning Research 11 (August): 1803-31, the contents of which are incorporated herein by reference. A further advantage is that no spatial assignment of the activations to input image areas is required for the LPR. Furthermore, in the case of backpropagation-based methods-such as LRP-the granularity maps can also be created directly on the activation maps. A calculation step from the input back to the activations can thus be omitted.In one embodiment, the granularity maps may be generated based on activation maps of convolutional networks. For CNNs, the neural activations in the intermediate layers may be associated with spatial regions in the input image. The resolution capability is generally limited by potentially lower resolution of the activation maps and the receptive field, disclosed in Luo et al. 2016, "Understanding the Effective Receptive Field in Deep Convolutional Neural Networks." In Advances in Neural Information Processing Systems, 29:4898-4906. Barcelona, Spain: Curran Associates, Inc. https: / / proceedings.neurips.cc / paper / 2016 / hasch / c8067ad193 7f7 28f51288b3eb986 afaa-Abstract.html, of the respective neurons. This assignment can be used to assign higher relevance for the output to image areas with high intermediate activation. No additional DNN evaluations are needed, but no class specific information is obtained except in combination with backpropagation techniques.Examples thereof are CAM (class activation mapping) disclosed in Zhou et al. 2016, "Learning Deep Features for Distinctive Localization." In Proc. 2016 IEEE Conf. Comput. Vision and Pattern Recognition, 2921-29. Las Vegas NV, USA: IEEE Computer Society https: / / doi.org / 10.1109 / CVPR.2016.319., and mixed forms such as Grad-CAM (Gradient Class Activation Mapping) disclosed in Selvaraju et al. 2017 "Grad-CAM: Visual Explants from Deep Networks via Gradient-Based Localization." In Proc. 2017 IEEE Int. Conf. Computer Vision, 618-26. Venice: IEEE. https: / / doi.org / 10.1109 / ICCV.2017,74, the contents of which are incorporated herein by reference.In one embodiment, masking using Salicy Maps may include selecting one or more outputs. Then, the granularity map is calculated with respect to the outputs based on the activations of the input neurons, i.e. the original image input, or on the activation maps. Once the granularity map has been calculated for the input, the granularity map values are preferably assigned to the corresponding neural activations. It is preferred that the masks are binarized by applying a threshold value to the Granularity Map values.Generally, the method described herein comprises at least the following steps.A masking criterion is evaluated on the image or the activation maps. In other words, a suitable mask for neuron activation is generated. The activation maps can be masked, in particular with the aim of focusing on the activation patterns and / or simplifying the activation patterns. The mask can also result in a reduction in the activation map.For each detection or each relevant image area, an outlier criterion is evaluated and an outlier score is determined.The Outlier score is used further in the respective application. The outlier score is indicative of an amount of deviation of the monitored neural network from a similar rule rule. The outlier score can be used, for example, to increase the uncertainty in a sense plan act pipeline and cause it to draw (dynamic) additional information for the evaluation. Examples of such information are temporal or spatial information suitable for verifying or discarding the detection associated with a high outlier score. It should be taken into account that additional information required for verification can often be "expensive" in the evaluation. It can therefore be advantageous to evaluate the additional information only in the case of high uncertainty or high probability of error.The outlier criterion can be a function which accepts as input the activation of the neurons from one or more layers unenlighted by the mask and uses this to determine an outlier score, for example as a number. In this case, the function can be selected such that a high external score indicates that the input examined is an external for the monitored NN or deep neural network (DNN). In other words, the outlier score may indicate that the result of the NN is based on a test input that is more likely to belong to the edge region of the training data distribution.The test input can be, for example, a rare case (for example, costumed persons), an Adversary Example (i.e. an outlier with respect to non-semantic features) or a case generally not covered by the training data (for example a night image if only daylight training data were present).The determination of the outlier score can be based on a post-hoc analysis of the activations within the network. The weights of the network are trained only for the application. The outlier score is only determined when the activations are established therefrom after a feedforward of test data in each case for an example.Various approaches may be used for determination based on one or more given activation maps.In one embodiment, a neuron coverage criterion may be used. In this case, a proportion of activated neurons, i.e. those which exceed a previously defined threshold value, can be determined. In this configuration, the monitored NN can be checked, in particular, for false negatives. This is attributed to the fact that in such a case, above-average high activations indicate an overlooked object.Another possibility is the use of a neural correlation criterion. The neuron coverage criterion can be extended in such a way that not only the ratio of activated neurons to total neurons but also their activation patterns are evaluated. Thus, correlations of different neurons within one or more layers can be considered. It is thus possible to take into account functions for the outlier score that are mapped only by the combination of a plurality of layers or activations.In one embodiment, statistical values of the activations can be used as a basis for the outlier score.In another embodiment, clustering of the activation patterns may be used. In this case, the activation pattern that has been caused by an input can be compared with the closest known cluster of activation patterns. In particular, the distance between the example and the next known cluster is determined in this abstract space.These basic possibilities can be modified even further in order to reduce or minimize the size of the activation patterns which simultaneously result in an outlier score. The required computing resources, in particular computing power / time and working memory, can thus be reduced and, surprisingly, the accuracy of the outlier score can be improved at the same time. This is believed to be due to the fact that the computation of the score in this case is based on more relevant regions and the impact of background regions or regions of other objects is reduced.In one embodiment, this is achieved by means of masking on the basis of the previously determined mask. The neurons can be masked in this case by their activation for the determination of the outlier score being simply set to 0. In other words, masked neurons are considered not activated (0-valued).In one embodiment, region-of-interest (Rol) pooling may be used. In order to determine the outlier score for a detection or an image region, the activation pattern can be restricted to an area of predefined or predicted size around the selected image region. Additional measures may be helpful to obtain normalized size activation patterns. Object size can be restricted by additional pooling (rol pooling as in legacy standard detectors, e.g., Faster R-CNN).Anchors having multiple standard sizes per anchor may be used. In this case, a separate outlier criterion can be used for each size category. In other words, the outlier criterion can be evaluated per anchor and per size category. Preferably, only one preselected size category can be evaluated for an input and an anchor, e.g., that best matches the bounding box size. One embodiment provides that different outlier criteria are applied for different reference classes, similar to the class-by-class neuron coverage. Reference classes may be obtained from objectness, object classes / super classes or object subclasses.The objectivity indicates a value for a binary classification, whether it is an "object" or a "no object". This value is of particular interest for the evaluation of false negatives.The object classes / superclasses may, for example, denote vehicle, pedestrian, etc.Object subclasses can be used if there are corresponding DNN outputs (e.g., van, truck, adult, child). The reference class can already be included as a reference class during masking, for example, by only leaving all pedestrians unmasked.The masking preferably uses a function which accepts an input, the activation patterns of a DNN, and the DNN output and returns a binary mask for the activation patterns from this, for example. It is also conceivable that the mask is "fuzzy" in design, so that certain input data points are at least partially passed through by the mask.Masking makes it possible to simplify the activation patterns for further processing. On the one hand, this allows the computational and storage outlay to be reduced. On the other hand, the adjusted evaluation of the outlier criterion can be limited to the relevant regions in the activation patterns. It is thus not only possible to determine the presence of an outlier; instead, this allows localization where the outlier is located in the input.The masking function may be given or limited by one or more criteria applied in succession. The criteria help determine which activations are relevant. If a reference class has been selected, the criteria can also be specific for the respective reference class.One possibility is the use of preselection by means of criticality criteria, e.g. plausibility or (in)constituency (strange position, temporal inconsistency such as tracking: person disappears / appears, dreck on the recording optics, etc.); uncertainty of the DNN output; and criticality of the image area (e.g. edge areas).In one embodiment, the masking criteria aid in identifying false positives, for example, Salinity Maps and DNN outputs.Salinity maps allow to use only areas of the activation maps that spatially match areas in the input that were relevant for the output. The DNN outputs can be mapped back into the image. For example, in the case of semantic segmentation or instance segregation for objects, a reduction to the segmented surfaces can take place, which preferably match the selected reference class and / or the selected object. Bounding boxes that preferably match the selected reference class can be considered. It is also conceivable to connect the projections to one another.For the identification of false negatives, it is preferred to invert the masking in comparison with false positives. Thus, the regions are used for calculating the outlier score that could contain false negatives. One way of reacting is to invert the regions that would be marked for false positives (e.g., regions with predicted detections, regions that were important for detections, or the like). As a result, those regions can be selected which probably do not contain false positives and correspondingly possibly false negatives. A high outlier score in these areas (especially when the reference class "no object" is selected) can then indicate a false negative.In one embodiment, a sliding window (sliding window) is used. In the sliding window, a parallel evaluation per image area is carried out for a grid of (partially) overlapping image areas of the same size. This is preferably restricted to image areas which do not belong to detection. It can thus be answered whether an image region in which no detection is present contains an anomaly and thus possibly a false negative.Usually, no uniform image areas are specified, e.g. by detections. In other words, there are normally no bounding boxes for "no detection". The image areas can therefore be explicitly selected. For this purpose, a sliding window approach is proposed, in which a grid of overlapping boxes is preferably placed over the image area of the input. Bounding boxes belonging to detection or largely overlapping detection are preferably totally excluded. An outlier score is determined for each of the remaining boxes on the neural activations that spatially belong to the remaining bounding boxes by means of the same outlier criterion.The outlier score in this case may be an indicator of whether the activations that did not result in detection at a location are abnormal (and thus may be a false negative). Abnormal activations may be, for example, excessively strong activations. This is because in general, it can be assumed well that regions without objects generally have low activations.The measures disclosed herein may be used for DNN-based perception applications, e.g., in an advanced driver assistance system (ADAS). An application in the field of perception for mobile robots, interior monitoring, intelligent infrastructure, camera-based quality checking in production or automated (pre)engineering is also conceivable.In quality control for DNN-based perception applications, the quality of detections can be checked using the Outlier Score. For a good detector, a stable behavior of the activations with respect to disturbances such as noise, adventrial attacks, etc. is desired. If a detector fluctuates greatly in the activations in the event of small changes to the input images, this can indicate a certain susceptibility to faults of the monitored DNN. The outlier score can be used as a measure to measure and appropriately respond to the robustness of this DNN, for example, by discarding the test result and re-quality control. It is also possible to supply the quality-controlled object to a manual control and, if appropriate, to use the result of the manual control for the further training of the monitored DNN.In the (pre)training of training and test data sets, these can be improved by a continuous extension of the samples. Here, the Outlier Score can be used to identify critical samples that should be subject to manual testing, if necessary. Furthermore, a verification or quality control of the monitored DNN can be carried out on the basis of the Outlier Score.It is likewise possible to perform operation-time controls, i.e. control at the runtime of the DNN. This is advantageous, for example, when using autonomous mobile systems in a traffic area. In this case, the Outlier Score can be used to identify critical detections at runtime, in the sense of detections that are not sufficiently reliable. These identified critical detections can then be checked by means of an additional detector or scene information.BRIEF DESCRIPTION OF THE DRAWINGSExemplary embodiments of the invention are explained in more detail below with reference to the attached schematic drawings. The following shows: FIG. 1 shows an embodiment of an autonomous vehicle; FIG. 2 shows an exemplary embodiment of a monitoring method; FIG. 3 is a diagram illustrating the location of an abnormality; FIG. 4 shows an exemplary embodiment of an object detection method with monitoring and error handling; FIG. 5 shows two variants of a selection method: a) permanently predefined and b) heuristically; FIG. 6 shows a decision flow diagram for an exemplary embodiment in the case of temporally successive sensor data; FIG. 7 shows data flow diagrams at successive points in time; and FIG. 8 shows an embodiment for the efficient spatial and temporal calculation of Salicy masks.DETAILED DESCRIPTION OF THE EMBODIMENTSFIG. 1 shows an autonomous vehicle 10. The vehicle 10 has a vehicle control unit 12 which is configured to control the vehicle 10, that is to say, for example, to accelerate, decelerate or steer the vehicle 10. The vehicle control unit 12 is an example of a control unit. It is also possible that the vehicle control unit 12 controls display for the driver. The vehicle control unit 12 is preferably designed as a data processing device.The vehicle 10 comprises a sensor system 14. The sensor system 14 can be configured to sense the environment of the vehicle 10. The sensor system 14 comprises at least one sensor 16, for example an imaging sensor, such as a camera. Preferably, the vehicle has so many sensors 16 arranged on it that the entire 360° environment of the vehicle 10 can be detected. Each sensor 16 has its own sensor region 18 which is detected by the sensor 16.An object 20 can be located in the sensor region 18. This can be a traffic sign, for example. Other examples of the object 20 include obstacles (guardrail, bollard, and the like), traffic lights, pedestrians, cyclists, and generally any other class of object useful in autonomous driving in the traffic area.FIG. 2 shows that the vehicle control unit 12 includes a deep neural network (DNN for short) 22 trained for object detection, especially in road traffic. For example, the DNN 22 may be trained to recognize traffic signs. The DNN 22 may be a convolutional neural network (CNN). Such networks for object detection are known and will therefore not be explained in more detail.The vehicle controller 12 further includes a computer-implemented network observer 24 that monitors the DNN 22. The network observer 24 is configured to determine an outlier score 26 that indicates whether, and to what extent if appropriate, output data of the DNN 22 is due to anomalous behavior. Preferably, the output data may be assumed to be based on anomalous behavior of the DNN 22 when the outlier score 26 is outside a predefined range or exceeds a predefined threshold. The outlier score 26 is preferably a numerical value.With reference to FIGS. 2 and 3, a method for controlling the autonomous vehicle 10 will be explained in more detail. As an example, driving the autonomous vehicle 10 on a road with traffic signs is provided.The vehicle control unit 12 causes the sensor system 14 to capture image data 28 of the environment of the autonomous vehicle 10 by means of the sensors 16. The image data 28 includes object data indicative of the object 20. Unlike in the training examples, however, the object 20 can be modified in an unfavorable manner, for example by contamination of the sensors 16 or of the object 20, unfavorable light incidence and the like.The image data 28 is supplied as input data 30 to the DNN 22. The DNN 22 performs object detection in a manner known per se in the feed-forward and generates object detection data 32 as output data 34. the object detection data 32 preferably includes a bounding box for each detected object, a semantic classification of the detected object, and optionally a probability indication indicating which probability the classification of the detected object matches the true class of the object 20.In the present example, the object detection data 32 may include a bounding box. In other words, the DNN 22 was able to determine the presence of the object 20 and its position in the image data 28. Furthermore, the DNN 22 can generally also still recognize that it is a traffic sign, in particular on account of the external shape. However, it may be that the speed could not be detected, for example because of dirt on the sensors 16, partial dirt or occlusion of the road sign or light incidence which only highlights the contours of the object 20 against the background, but makes the print "invisible" (for example a counter light situation with a low sun).In the feed-forward of the input data 30, activation map data 36 is generated. The activation map data 36 includes, for each layer of the DNN 22, an activation map 38 that includes the activation values of the individual neurons of the respective layer of the DNN 22.The input data 30, output data 34, and activation map data 36 are provided to and processed by the network observer 24.The network observer 24 generates masking data 40 therefrom. The masking data 40 specifies which activation map data 36 is taken into account in the further course and which is not. For example, the mask data 40 may include a mask map 42 for each activation map 38. The masking data 40 can be generated, for example, by means of a threshold value. The network observer 24 may mask the activation map data 36, for example, by setting all data that is no longer to be considered according to the masking data 40 to a zero value. However, the zero value may not necessarily correspond to the number zero.Each mask map 42 may be generated from the associated enable map 38, for example, by setting all entries that are equal to or above a predefined threshold to "1" while setting all other entries to "0". Here, "1" preferably means that the corresponding entry in the activation map 38 to be masked is to be taken into account further, while in the case of "0", the corresponding entry of the activation map 38 is not to be taken into account further.The network observer 24 generates masked activation map data 44 from the masking data 40 and the activation map data 36 in this variant, the dimensions of which remain unchanged. The masked activation map data 44 then typically forms a sparse matrix.In one variation, the masked activation map data 44 may be stored and processed more efficiently as a smaller matrix along with a location indication, depending on the masking criterion selected. This is because, if, in an intermediate step, for example, a sliding window is used to generate masks, the non-zero values of the masked activation map data 44 are of the same size for all positions of the sliding window. Thus, the masked activation map data 44 can be represented as activations of only the size of the sliding window, together with an indicator at which location the sliding window was in the input data 30, for example image data.The network observer 24 can determine the outlier score 26 on the basis of the masked activation map data 44. The network observer 24 can determine the outlier score 26 on the basis of the so-called neuron coverage. Neuron Coverage corresponds to the ratio of activated neurons to the total number of all neurons.The network observer 24 therefore preferably determines the current neuron coverage for the object 20 that has just been processed. It is also possible for the network observer 24 to compare this result with the average neuron coverage of the DNN 22 determined during training. The difference between the current neuron Coverage and the average neuron Coverage may alternatively be output as the outlier score 26. This variant is suitable in particular for false negatives, because these are usually associated with an above-average current neuron coverage.In the case of the road sign, the current neuron coverage is significantly increased because not only the activations associated with a particular road sign are above the threshold value, but all activations associated with a road sign having at least a similar external shape.The network observer 24 then assigns the output data 34 to the outlier score 26 and transmits it to the vehicle control unit 12.The vehicle control unit 12 processes the output data 34 and the outlier score 26. If the vehicle control unit 12 determines that the outlier score 26 is above a predefined threshold value, the output data 34 can be considered to be comparatively unreliable. In response, the vehicle control unit 12 may cause, for example, the repetition of the image acquisition and subsequent evaluation by the DNN 22. The vehicle controller 12 may also activate other sensors of the sensor system 14 to obtain supplemental input data that augments the previous input data 30 to collectively analyze that data. The vehicle control unit 12 may cause a database query. The vehicle control unit 12 may cause a warning to be issued to the driver that a traffic sign could not be recognized so that the driver can intervene in a corrective manner.In the example of the road sign, the vehicle control unit 12 can ascertain the position of the autonomous vehicle 10 by means of a satellite navigation system, for example, and then ascertain the speed restriction on the basis of corresponding stored road map data. The vehicle control unit 12 may then cause the autonomous vehicle 10 to accelerate or decelerate.In one variation, the network observer 24 reduces the input data 30, the output data 34, and / or the activation map data 36 before being used to generate the mask data 40. For example, the resolution can be reduced. A MaxPool function can be used.Other embodiments of the invention will be described below only insofar as they are different from the embodiment described so far.Further variants of the determination of the outlier score 26 are explained below. It should be noted that the different types of determination of the outlier score 26 (including the embodiment described above) may be combined with each of the different types of determination of the mask data 40 described later.In another embodiment, the network observer 24 may determine the outlier score 26 based on the neural correlation. At this time, the number of activated neurons is not simply counted, but the activation patterns included in the activation map data 36 are evaluated. In particular, at least one correlation value may be determined within the activation maps 38 and / or activation maps 38 for the different layers of the DNN 22. The more the correlation value(s) for the input data 30 deviate from the average correlation values for regular test data, the more likely it is to be an outlier. Accordingly, the resulting outlier score 26 is higher.In another embodiment, the network observer 24 may determine the outlier score 26 by clustering. For this purpose, the activation patterns are first recorded in the activation maps 38 during the training of the DNN 22 and clustered by means of a clustering method known per se, such as the k-means algorithm. For the activation maps 38 generated by the input data 30 at the feed-forward, more specifically the activation patterns contained therein, the distance to the nearest cluster is determined and output as an outlier score 26. Consequently, if the activation patterns deviate significantly from the trained situations, the outlier score 26 is higher and the output data 34 is thus marked as more unreliable.Further variants of the determination and use of the masking data 40 are explained below. It should be appreciated that these embodiments are compatible with and can be used with any way of determining the outlier score 26.In another embodiment, the network observer 24 may generate the mask data 40 based on the salinity data. The granularity data indicates which areas of the activation map data 36 spatially match those areas of the input data 30 that were relevant to the output data 34. In other words, each masking map 42 may be based on a Salicy map (Salicy Mask) determined for each layer of the DNN 22. The methods by which Salicy maps can be determined are known per se and are not explained in more detail here. Reference is made in particular to the following publications:Lundberg et al. 2017 "A Unified Approach to Interpreting Model Predictions." In Advances in Neural Information Processing Systems 30, 4765-74.Ribeiro et al. 2016 "Why Should I Trust You?": Explaining the Predictions of Any Classifier." In Proc. 22nd ACM SIGKDD Int. Conf. Knowledge Discovery and Data Mining 1135-44. KDD '16.Bach et al. 2015 "On Pixel-Wise Explantations for Non-Linear Classifier Decisions by Layer-Wise Relevance Propagation." PLOS ONE 10 (7): e01130.Baehrens et al. 2010, "How to Explain Individual Classification Decisions." Journal of Machine Learning Research 11 (August): 1803-31.Luo et al. 2016 "Understanding the Effective Receptive Field in Deep Convolutional Neural Networks." In Advances in Neural Information Processing Systems, 29:4898-4906. Barcelona, Spain: Curran Associates, Inc.Zhou et al. 2016, "Learning Deep Features for Distinctive Localization." In Proc. 2016 IEEE Conf. Comput. Vision and Pattern Recognition, 2921-29. Las Vegas, NV, USA: IEEE Computer Society.Selvaraju et al. 2017 "Grad-CAM: Visual Explantations from deep networks via gradient-based localization." In Proc. 2017 IEEE Int. Conf. Computer Vision, 618-26. Venice: IEEE. https: / / doi.org / 10.1109 / ICCV.2017,74.In one embodiment, the network observer 24 may use the output data 34 to generate the mask data 40 that may be mapped back into the input data 30. In semantic segmentation or instance segmentation for objects, this includes the reduction to those segmented areas that belong to a selected reference class or match the selected object. For the object 20, for example, those regions of the output data 34 which correspond to the reference class "traffic signs" in the input data 30 can be used to generate the masking data 40. It is also conceivable to use the bounding boxes contained in the object detection data 32 to generate the masking data 40. Also, both approaches may be combined.In a further embodiment, the uncertainties determined by the semantic segmentation can serve as masking data 40. Thus, areas with uncertainty values that exceed or fall below a predetermined threshold value can be examined for atypical activations.The embodiments explained above are suitable in particular for identifying false positives. If false negative exit is to be detected, the network observer 24 can invert or create the masking data 40 in an inverted manner. In other words, the entries "1" are replaced with "0", and vice versa. The resulting outlier score 26 in this masking is then an indicator of a false negative.In another embodiment, particularly suitable for false negatives, the network observer 24 may superimpose a grid of overlapping input boxes on the input data 30. Based on the output data 34, in particular the object detection data 32, such as the bounding boxes, those input boxes which belong to a bounding box or have an area overlap with such a bounding box which is more than half can be excluded. Using the non-excluded input boxes, the network observer 24 then determines the masking data 40.In contrast to other methods, the measures presented here can not only be used to determine whether the output data 34 represents an outlier; rather, it is additionally possible, on the basis of the masking, to determine where the output data 34 contains the cause of the high outlier score 26. In other words, the network observer 24 is not only able to recognize that an outlier is present, but can also determine the position of the outlier / anomaly in the output data 34 and thus also in the input data 30.The object detection and its monitoring in the case of time-resolved sensor data will be discussed in more detail below.RequirementsA starting point or a sequence of a corresponding method is schematically illustrated in FIG. 4. The starting point of an embodiment comprises:A DNN for object detection on the basis of sensor recordings (in particular. camera images) of road scenes for ADAS functions in a vehicle;an additional method for outlier score calculation for input-output combinations of the DNN, comprising: 1. an "expensive" precalculation (e.g., compute or memory intensive) 2. a final outlier score calculation based on the results of the precalculation; anda subsequent evaluation of the determined outlier score, which adapts the system reaction at a high score (e.g. switching on further redundant perception algorithms; taking a safe state; feedback to safety drivers; reporting to manufacturers, cf. Incident Reporting according to the European Union AI Act: https: / / artificialtelligenceact.eu / or comparable requirements of further future regulations in different states or regions).The illustrated starting point ("base approach") may be extended by embodiments. Examples of embodiments are explained in more detail below and illustrated with reference to FIGS. 6 and 7.General idea: Perform the precalculation of the information ("outlier information") necessary for the final calculation of the outlier score not in each frame (individual image of an image sequence of, for example, camera images), but instead replace the calculation of selected substeps by approximated values which are estimated on the basis of the already calculated information from previous frames.For the estimation, for example, the optical flow of the sensor data may be determined and used to estimate future values of both the input and (intermediate) output of the DNN by transforming them along the optical flow.Specifically, the approach includes the following principles:a selection of the substeps / calculations which are not performed but rather are approximatedapproximating the results of these sub-computationsSelection of Sub-ComputationsHere, the following forms of partial calculations are considered, which are partially replaced (visual comparison in):Calculation of masks for the activation mapscalculations with intermediate outputs / scores per layerIn both cases mentioned, it is possible to decide, according to fixed (simple case) or heuristic criteria per frame, whether the result is calculated anew and exactly or approximated.This illustrates the representation in FIGS. 5 aand 5 b In the simple case (FIG. 5 a: Recalculation for a layer L 1, L 2, L 3, L 4 every 2 frames), a frequency at which recalculations take place can be determined in a simple manner. In step t (S t) the first and third layers L 1, L 3 are recomputed (white rectangle), the second and fourth layers L 2, L 4 are approximated (hatched rectangle). In step t+1 (S t+1) the first and third layers L 1, L 3 are approximated, the second and fourth layers are recomputed. Step t+2 (S t+2) see Step t (S t), Step t+3 (S t+3) see Step t+1 (S t+1). In this case, it is / should be ensured during calculations per layer that not all layers are approximated simultaneously: by alternating recalculation of the layers it can be ensured that a certain percentage of the layer-specific calculations is always exact.In the heuristic case (FIG. 5 b), in each step S 1, S 2, S 3, S 4 the calculations that are approximated (hatched rectangle) are selected randomly. In FIG. 5 b, for example, only one layer is recomputed per step (white rectangle): in step t (S t) the third layer L 3, in step t+1 (S t+1) the fourth layer L 4, in step t+2 (S t+2) the first layer L 1 and in step t+3 (S t+3) the fourth layer L 4. Here, the probability with which a recalculation is started should increase with the duration since the last recalculation. In FIG. 5 b, the second layer L 2 is more likely to be recomputed in the following step (not shown), since it has not undergone any recomputing since at least four steps. An advantage over the simple case is that no systematic errors can occur due to the fixed selection. One drawback is less control over the actual availability of accurate precalculation results.Approximation of Partial Calculation ResultsTwo possibilities are provided here:Simple case: Simple replacement by the results from the last such partial calculation. This is possible under the assumptions that partial results hardly change between frames (temporal coexistence) and smaller changes have little influence on the final outlier score.Optical Flow: If even small changes in the partial results may have an influence on the final outlier score, certain forms of intermediate results may be estimated based on the optical flow in the inputs. The sub-calculations to be approximated must satisfy the following condition: their results must have spatial regions (e.g. pixels) which can be spatially unambiguously assigned to regions in the input (e.g. image sections). The following steps can then be applied for the approximation: 1. from the time t- 1, the following information is stored: a. the input frame and / or system information (e.g. ego motion) which can serve for calculating the optical flow, e.g. the intermediate result used there from the partial calculation (exactly calculated or approximated) 2. for the input frame in step t, the optical flow between the current and the previous frame is calculated and stored on the basis of the information from the current frame and step 1.1. 3. approximation:For each intermediate output to be approximated: The optical flow of 2. is applied as a transformation to the respective intermediate output from time t-1 in step 1.This is schematically illustrated in a region of FIG. 6, FIG. 6 overall shows a decision flow diagram for the proposed addition of the estimated outer score precalculation. In comparison with the basic approach (see FIG. 4 ), it is decided whether an exact precalculation of the individual outlier information ("outlier-score precalculation") is carried out. The outlier information that is not exactly precomputed is estimated. The outlier score is calculated from the present exactly calculated and estimated outlier information ("final outlier score calculation"). As has been the case, in the case of a high exterior score, error handling is carried out, e.g. countermeasures can be initiated at the time t.As shown in FIG. 6, a change in input data for the DNN is calculated or estimated based on the sensor value at time t and a sensor value at previous time t- 1 and / or other system information such as ego motion ("ego motion") information. This change at time t can, as shown, be included in the estimation of the outlier information that is not calculated exactly. Additionally or alternatively, the change at time t may be stored ("input change of the last n times"). In a comparable manner, the intermediate information ascertained at the n preceding points in time can be stored in a second memory (dashed arrow). The information from both memories can be taken into account in the decision whether an exact calculation is performed. The latter enables a heuristic decision, in which the number of frames is taken into account, since the last exact precalculation of this outlier information.FIG. 7 shows a schematic data flow diagram of an added exemplary embodiment (cf. again with the basic approach in FIG. 4 ). At time t, the sensor value is ready at time t at the input of the DNN performing the object detection. An outlier score is determined to check the results of the object detection. Here, at time t, the "DNN (Intermediate) outputs" are used to calculate all outlier information exactly ("(expensive) outlier-score precalculation"). The outlier information ("intermediate information") can be present, for example, in the form of a heat map. The "final outlier score calculation" is performed on the basis of the pre-calculated outlier information and provides the "outlier score". If the outlier score exceeds a threshold value ("error treatment at high score"), at least one countermeasure is initiated ("countermeasures at time t").At time t+1, the sensor value is ready at the input of the DNN at time t+1. In the "estimated outlier score precalculation", not all outlier information is calculated exactly again: on the basis of "DNN (intermediate) outputs" (at the time t+1) and intermediate information at the time t, the current outlier information ("estimated intermediate info") is completely determined, just estimated and possibly also calculated exactly in parts. The "final outlier score calculation" (at time t+1) occurs based on this and provides the "outlier score" (at time t+1) from which it is checked whether a "high score fault treatment" is required. If yes, "countermeasures are initiated at time t+1".At time t+2, the same procedure as at time t+1 is followed. However, it is quite possible and reasonable for outlier information to be estimated other than at the time t+1. As a basis for the "estimated outlier score precalculation" at the time t+2, the "DNN (intermediate) outputs" (at the time t+2) and intermediate information at the time t+1. After j-1 times, j is a natural number greater than or equal to 2 at which outlier information was estimated, all outlier information is calculated exactly again at the following time t+j ("(expensive) outlier score precalculation") on the basis of the sensor value at the time t+j. This ensures that, that the uncertainty from the estimates at the j-1 times since the last complete exact calculation (time t) of the outlier information cannot become arbitrarily large.Example 1:The given outlier score calculation evaluates the neural activations of the DNN to obtain the outlier score. The method is multistage: 1. (expensive precalculation) Anomaly detection is applied per layer in order to obtain a layer-specific outlier score. 2. (final calculation) The individual scores are merged into an overall score, e.g. by a simple trained linear classifier.One embodiment now provides for not calculating an anomaly score for each layer in each frame, but rather only for a part of the layers in each case. The scores for the non-computed layers are estimated from the result of the last pre-computation for that layer. The supplement routine may provide, for example, to simply replace the missing values by their predecessors from previous calculations.At the beginning, values are calculated once for all layers.Thereafter, each layer is no longer calculated, but rather some values are occupied by previously calculated values, according to one of the following example schemes: 1. simple division: An even layer is calculated (as shown in FIG. 5 a) always alternately a. odd layer is supplemented from the previous frame b. odd layer is calculated, even supplemented from the previous frame 2. heuristic division with each X layers (cf. FIG. 5 b), X is a natural number which is smaller than the total number of layers: Each layer has a certain probability of being recomputed. This probability increases with the duration since the last recalculation. At each pass, X layers are randomly selected from the probabilities calculated.Example 2:The requirements are fundamentally the same as in Example 1, except that the Outlier Score ( 204) is additionally masked with the aid of Salicy masks (At, ß t+1, ß t+2,..., A t+j). Although masking reduces the storage and computing effort for the outlier score precalculation by using less data, the Salicy mask (At, ß t+1, ß t+2,..., A t+j) must be calculated for this purpose. This shows which areas of the input were relevant for the output of the DNN and normally requires at least one additional (backward) evaluation of the DNN. FIG. 8 shows an embodiment of an efficient spatial and temporal calculation of Salicy masks (At, ß t+1, ß t+2,..., A t+j). The time axis t is plotted to the right, the processing flow p to the bottom. In order to obtain a first outlier score ( 204) from the incoming first sensor data I t at a time t, a granularity mask At is calculated by applying a granularity method, e.g. layer-by-layer relevance propagation ( 200). The Salicy mask At is used to calculate the Outlier score using a monitoring model (203). For the determination of a second outlier score ( 204) at time t+1, a flow map F t+1 is calculated on the basis of second sensor data I t+1 and the first sensor data I t e.g. by means of optical flow ( 201). By using the previous Salicy mask At and the current flow map F t+1 a calculation of an approximated Salicy mask ß t+1(202). is performed. The procedure is analogous for the determination of a third outlier score ( 204) at the time t+2. The approximated Salicy mask is determined based on the current flowchart F t+2 and the approximated Salicy mask ß t+1 from the previous time step (202). At a later time t+j, the Salicy mask A t+j is calculated only on the basis of the sensor value I t+j at the time t+j. Using the monitoring model, the Outlier score ( 203) is calculated using the (exactly) computed Saliency Mask A t+j. Thus, the Outlier score is obtained at time t+j (204).Instead of recomputing the Salicy mask (At) for each frame (It, I t+1,...) the following is provided within the scope of the embodiment according to example 2 (cf. FIG. 8 ): 1. store input information at the time t (and hold for the next i steps in the memory, i is a natural number). 2. Nimm, that the mask At for the frame t has been determined (estimated or recomputed). 3. for frame I t+1 at time t+1, determine the optical flow F t+1, i.e., an approximation of the change that occurred per image region between the current frame at time t and the frame(s) through time t-i. 4. for frame t+1, not recalculate the Salicy mask, but instead flip the current value of the optical flow F t+1 as a transform to the Salicy mask Atto estimate the true Salicy mask for frame t+1. The result of this estimation is referred to as estimated Salicy mask A t+1.An advantage here is: as a result, no neural activations from the forward propagation have to be temporarily stored until the mask calculation by the back propagation is ended. Instead, immediately after the calculation of the activations of each layer (L 1, L 2, L 3, L 4) the activations can be masked and the outlier information, i.e. the outlier score for the layer, can be evaluated.Example 3: Example 3:Again, the requirements are the same as in Example 1, except that no single outlier score per layer is calculated, but the precalculation per layer uses an additional CNN operation and produces memory intensive output. This must be stored in the memory until the final calculation of the outlier score, i.e. especially until the evaluation of the last relevant layer.In the third embodiment, the following approach is provided to reduce the memory requirement:Turn on the steps of example 2 per layer, except that in 2nd no mask is determined, but rather the result of the layer-specific precalculation is determined, and that in 4th the optical flow is applied to this precalculation result. This is possible because the spatial regions can be assigned to the input with respective regions in the pre-calculation result, since it was a CNN operation.Additionally, for better accuracy, the procedure of Example 1 may be used to heuristically choose the layers from which the activation maps are not re-stored / evaluated.REFERENCE NUMERALS10 Autonomous vehicle 12 Vehicle control unit 14 Sensor system 16 Sensor 18 Sensor region 20 Object 22 Deep Neural Network (DNN) 24 Network observer 26 Outlier Score 28 Image data 30 Input data 32 Object detection data 34 Output data 36 Activation map data 38 Activation map 40 Masking data 42 Masking map 44 Masked Activation map data 200 Application of a Salicy method 201 Calculation of the flow map 202 Calculation of an approximated Salicy mask by using the previous Salicy mask and the current flow map 203 Use of the Salience mask to calculate an outlier score using a monitoring model 204 Obtained Outlier Score t Time; Time p Processing Flow St, S t+1,... Step t, Step t+1,... L 1, L 2, L 3, L 4 Layer 1, Layer 2, Layer 3, Layer 4I t, I t+1,... Picture t, Picture t+1,... AtAlien mask t F t+1, F t+2,... Flow Map t+1, Flow Map t+2,... ß t+1, ß t+2,... approximated Salicymask t+1, approximated Salicymask t+2,...References included in the specificationThis list of documents cited by the applicant has been produced in an automated manner and is only included for the better information of the reader. The list is not part of the German patent application or utility model application. The DPMA does not take any adhesion for any faults or omissions.Patent Literature citedEP 4099210 A1
[0004] EP 3896651 A1
[0005] EP 0680026 A2
[0006] WO 2022016011 A1
[0007] DE 10 2021 211 503 B3
[0008] Cited Non-Patent LiteratureLundberg et al. 2017 "A Unified Approach to Interpreting Model Predictions." In Advances in Neural Information Processing Systems 30, 4765-74, http: / / papers.nips.cc / paper / 7062-a-unified Approach-to-interpreting-model-predicate ns.pdf.
[0063] Ribeiro et al. 2016 "Why Should I Trust You?": Explaining the Predictions of Any Classifier." In Proc. 22nd ACM SIGKDD Int. Conf. Knowledge Discovery and Data Mining 1135-44. KDD '16th https: / / doi.org / 10.1145 / 2939672.2939778
[0063] Bach et al. 2015 "On Pixel-Wise Explantations for Non-Linear Classifier Decisions by Layer-Wise Relevance Propagation." PLOS ONE 10 (7): e0113140. https: / / doi.org / 10.1371 / journal.pone.0130140
[0064] Baehrens et al. 2010, "How to Explain Individual Classification Decisions." Journal of Machine Learning Research 11 (August): 1803-31 [0064, 0129]Luo et al. 2016 "Understanding the Effective Receptive Field in Deep Convolutional Neural Networks." In Advances in Neural Information Processing Systems, 29:4898-4906. Barcelona, Spain: Curran Associates, Inc. https: / / proceedings.neurips.cc / paper / 2016 / hashi / c8067d193 7f7 28f51288b3eb986
[0065] Zhou et al. 2016, "Learning Deep Features for Distinctive Localization." In Proc. 2016 IEEE Conf. Comput. Vision and Pattern Recognition, 2921-29. Las Vegas, NV, USA: IEEE Computer Society. https: / / doi.org / 10.1109 / CVPR.2016.319
[0066] Selvaraju et al. 2017 "Grad-CAM: Visual Explantations from deep networks via gradient-based localization." In Proc. 2017 IEEE Int. Conf. Computer Vision, 618-26. Venice: IEEE. https: / / doi.org / 10.1109 / ICCV.2017.74 [0066, 0129]Lundberg et al. 2017 "A Unified Approach to Interpreting Model Predictions." In Advances in Neural Information Processing Systems 30, 4765-74
[0129] Ribeiro et al. 2016 "Why Should I Trust You?": Explaining the Predictions of Any Classifier." In Proc. 22nd ACM SIGKDD Int. Conf. Knowledge Discovery and Data Mining 1135-44. KDD '16
[0129] Bach et al. 2015 "On Pixel-Wise Explantations for Non-Linear Classifier Decisions by Layer-Wise Relevance Propagation." PLOS ONE 10 (7): e0330140
[0129] Luo et al. 2016 "Understanding the Effective Receptive Field in Deep Convolutional Neural Networks." In Advances in Neural Information Processing Systems, 29:4898-4906. Barcelona, Spain: Curran Associates, Inc.
[0129] Zhou et al. 2016, "Learning Deep Features for Distinctive Localization." In Proc. 2016 IEEE Conf. Comput. Vision and Pattern Recognition, 2921-29. Las Vegas, NV, USA: IEEE Computer Society
[0129]
Claims
Method for detecting objects from time-resolved sensor data by a trained neural network (22; DNN) having a plurality of layers (L 1, L 2,...) having the steps of - receiving a set of time-resolved sensor data (It, I t+1,...) as input data for the trained neural network comprising first sensor data (I t), which have been acquired at a first point in time t, and second sensor data (I t+1), which have been acquired at a subsequent second point in time t+1, - determining first object detection data (32) as output data for the first sensor data (I t) as input data by the trained neural network (22; DNN), - monitoring object detection by determining a first outlier score (26; 204) in the following manner: a) precalculating a plurality of first outlier information from a plurality of intermediate outputs of the trained neural network (22; DNN), b) calculating the first outlier score (26; 204) on the basis of the first outlier information from the precalculation - determining second object detection data (32) as output data for the second sensor data (I t+1) as input data by the trained neural network (22; DNN) and determining a second outlier score (26; 204) by: c) determining changes (F t+1) between first and second sensor data (It, I t+1); d) precalculating second outlier information, wherein fewer second outlier information are calculated than first outlier information, e) estimating the missing second outlier information from corresponding first outlier information taking into account the determined changes (F t+1), f), calculating the second outlier score (26; 204) on the basis of the second outlier information, and - outputting the determined object detection data (32) and outlier scores (26; 204).Method according to claim 1, wherein in step c) spatial changes of features between first and second sensor data (It, I t+1) are considered changes, and in step e) the estimating comprises a spatial transformation of the first outlier information(s).Method according to claim 1 or 2, wherein in steps b) and optionally also d) the outlier information relating to a plurality of layers (L 1, L 2,...) of the trained neural network (22; DNN) each comprise an intermediate output as outlier information.Method according to any one of the preceding claims, wherein each layer (L 1, L 2,...) of the trained neural network comprises at least one artificial neuron, wherein, in the course of the precalculation in steps b) and optionally also d), activation map data are determined by the trained neural network, wherein the activation map data indicate for each neuron which activation value this neuron has, aa) generating masking data from the activation map data and / or the sensor data and / or the object detection data; bb) masking the activation map data on the basis of the masking data in order to obtain masked activation map data, wherein the masked activation map data contain non-masked activation values and masked activation values; cc) determining layer-wise outlier information for at least one network layer for the object detection data on the basis of the masked activation map data, wherein only the non-masked values are taken into account when determining the layer-wise outlier information.The method of claim 4, wherein in step aa), the masking data (40) is generated from the granularity data such that those activation values are masked that spatially conform to areas in the input data (30) that were irrelevant to the output data (34).Method according to Claim 5, wherein the masking is effected at the level of the artificial neurons.Method according to Claim 5, wherein the masking is carried out at the level of the sensor data (It, I t+1,...).Method according to one of the preceding claims, wherein temporally resolved sensor data (It, I t+1,...) correspond to a sequence of camera images.Method according to Claim 8, wherein the ascertainment of changes between first and second sensor data (It, I t+1) takes place as ascertainment of the optical flow (F t+1) between first and second camera images.The method of any preceding claim, wherein an action is initiated if the first or second outlier score (26; 204) exceeds a threshold.A computer implemented method for controlling an autonomous moving device by a control unit, the method comprising: a) performing a method according to any one of the preceding claims to obtain object detection data indicative of an object type and object position of an object included in the time resolved sensor data, the object detection data additionally including an outlier score for at least two times; b) generating a control signal for the control unit, the control signal causing the control unit to control the autonomous moving device according to the control signal.A data processing device comprising units arranged to perform one, more or all steps of a method according to any of claims 1 to 11.An autonomous movement device comprising at least one sensor and a data processing device configured as a control unit according to claim 12, wherein the control unit is connected to the sensor in order to process its sensor data.A computer program comprising instructions for a data processing device, the instructions causing the data processing device to perform one, more or all steps of a method according to any one of claims 1 to 11.A machine readable medium or data carrier signal comprising the computer program of claim 14.
Citation Information
Patent Citations
Method for monitoring logical consistency in a machine learning model and associated monitoring device
DE102021211503B3
Vehicular traffic monitoring system
EP0680026A2
Method and apparatus for evaluating temporal characteristics of semantic image segmentation
EP3896651A1
Method for training a neural network for semantic image segmentation
EP4099210A1
Anomaly detection from aggregate statistics using neural networks
WO2022016011A1