Method for evaluating a trained deep neural network
Patent Information
- Application Number
- EP2023744735
- Authority / Receiving Office
- EP · EP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2022-08-16
- Filing Date
- 2023-07-18
- Publication Date
- 2025-06-25
AI Technical Summary
Existing methods for evaluating the performance and robustness of trained deep neural networks in vehicle environments are inadequate, as they do not provide a clear assessment of how the network learns to represent features and can lead to overfitting, and there is a lack of evaluation for already trained networks.
A method that involves determining reference values using evaluation data sets, generating additional training data sets for varying scenarios, and assessing the network's performance and robustness through additional training steps without altering its operational state, allowing for the evaluation of an already trained deep neural network by resetting parameters if unsuitable.
Enables effective evaluation of a deep neural network's suitability for object recognition tasks, preventing overfitting and ensuring the network's ability to generalize, while maintaining its original training state for practical operation.
Smart Images

Figure 1.1
Abstract
Description
[0001] Description
[0002] METHOD FOR EVALUATING A TRAINED DEEP NEURAL
[0003] NETWORK
[0004] Technical area
[0005] The invention relates to a method for evaluating a trained deep neural network. The invention further relates to a computer program and a computer program product for implementing the method.
[0006] State of the art
[0007] The perception or modeling of a vehicle environment represents a major challenge in the development of automated driving functions or advanced driver assistance systems (ADAS). Due to their outstanding performance, deep neural networks (DNNs) play a crucial role in object detection, i.e., classification and localization of sensor-detected objects.
[0008] Using a training algorithm and a number of training iterations, a feature space is learned in the internal structures of a deep neural network, which can be used to represent the objects to be recognized. The performance of a DNN can generally be improved by expanding a corresponding training dataset. After successful basic training, training datasets can be provided whose objects are difficult for the DNN to detect, although correct detection is still expected based on the requirements. Such training datasets can contain images with rare situations (corner cases) or images with limited image quality (e.g., with regard to contrast, brightness, etc.).
[0009] Performance as a criterion says nothing about how a DNN learns its feature space based on given training data over a finite number of training iterations. The internal structures of a DNN, which can also be referred to as hidden or latent structures, are essentially incomprehensible from the outside. Thus, during or after training, it is also unclear from the outside whether learned representations stabilize in one area of the feature space or whether completely different areas are used to represent what has been learned from one training step to the next.
[0010] In addition to performance, robustness is also a particularly important criterion for assessing sufficient training. The goal of training is not to feed a DNN so many data sets that every relevant situation has been considered once during training ("training by memorizing all situations"), as such an approach carries the risk of overfitting. Instead, the DNN should acquire the ability to generalize in object recognition through training, which can be assessed by sufficient robustness.
[0011] In the German patent application 10 2021 207 505.3, a method for training a deep neural network for object recognition in the environment of a motor vehicle was presented, which enables an evaluation of a current learning state of a training.
[0012] However, a subsequent evaluation of an already trained deep neural network is not possible with the described method.
[0013] Brief description of the invention
[0014] Against this background, the object of the invention is to provide a method for evaluating a deep neural network whose training for object recognition in the environment of a motor vehicle has already been completed.
[0015] Accordingly, a method according to the main claim, as well as a computer program and a computer program product according to the dependent claims, are proposed. Further embodiments are the subject of the respective dependent claims.
[0016] According to a first aspect of the invention, the object is achieved by a method for evaluating a deep neural network that has been trained for object recognition in the environment of a motor vehicle. At least one reference variable is determined based on a predetermined number of evaluation data sets and with the aid of the deep neural network. Furthermore, a number of test data sets are provided, from which a number of additional training data sets are generated using a selected method for varying data sets. The deep neural network is trained in an additional training step using the generated additional training data sets. After completing the additional training step, at least one evaluation variable is determined based on the predetermined number of evaluation data sets and with the aid of the deep neural network.The deep neural network is considered unsuitable for the selected method if a difference between the least one reference value and a corresponding evaluation value exceeds a predetermined threshold.
[0017] Accordingly, the deep neural network can be considered suitable for the selected method if the difference between the least one reference value and a corresponding evaluation value complies with the predetermined threshold.
[0018] One idea behind the present invention is to use additional training steps not for an extended training of a deep neural network, but to use them for the evaluation of an already trained deep neural network.
[0019] According to a further development of the method, it can therefore be provided that originally learned parameters of the trained deep neural network are first read out and stored before a first additional training step is carried out.
[0020] By freezing the training state of the deep neural network, the learned parameters of the deep neural network can be reset to the original learning state after evaluation. The stored parameters can include, for example, weighting parameters and prototypes, or concepts of feature extraction layers and perception layers for the internal structures of a deep neural network.
[0021] According to a further development of the method, all parameters of the deep neural network can be reset to the stored originally learned parameters after at least one evaluation variable has been determined.
[0022] Since the additional training steps are not intended to result in any additional training, for example, to avoid the risk of overfitting, the original read-out and saved training status is reset as soon as the deep neural network is evaluated as suitable for the selected method. The additional training performed thus only affects an evaluation step based on the selected method for varying data sets. An additional training step therefore has no influence on the practical operation of an evaluated deep neural network.
[0023] According to a further development of the method, a number of different methods for varying data sets can be provided, with a number of additional training data sets being provided for each of the different methods for varying data sets, so that a number of consecutive additional training steps corresponding to the number of methods are run through. Thus, the deep neural network can be evaluated successively based on different critical limit cases.
[0024] In this context, parameters of the deep neural network can be reset to the stored originally learned parameters after each additional training step, as soon as the corresponding evaluation variables have been determined.
[0025] This means that each additional training step for a particular method for varying data sets can begin at the same initial training level of the deep neural network. A previous additional training step therefore has no influence on the evaluation of a currently selected method.
[0026] According to a further development, a corresponding latent representation dataset can be generated by the deep neural network for each evaluation dataset, whereby relevant latent representations for the annotated data of each evaluation dataset are selected from the corresponding latent representation dataset. A distance mean in the latent space can be determined from the selected relevant latent representations of the evaluation datasets, whereby the determined distance mean is used as a reference value and / or as an evaluation value.
[0027] Alternatively or additionally, output data can be generated for each evaluation data set by the deep neural network, whereby a performance value is determined by comparing the generated output data with the reference data of the respective evaluation data set, whereby the performance value is used as a reference value and / or as an evaluation value.
[0028] According to a further development, the method for varying data sets can include a parameterized change of all data in a data set.
[0029] Alternatively or additionally, the method for varying data sets may involve adding objects to a respective data set.
[0030] According to a further aspect of the invention, the object is achieved by a computer program which, when executed on a computing unit within a maneuver planning device, instructs the respective computing unit to execute the method.
[0031] According to a further aspect of the invention, the object is achieved by a computer program product with a program code for carrying out the method, which is stored on a computer-readable medium. Brief description of the drawing figures
[0032] Further features and details emerge from the following description, in which at least one embodiment is described in detail—possibly with reference to the drawing. Described and / or illustrated features constitute the subject matter, either individually or in any meaningful combination, possibly independently of the claims, and may, in particular, also be the subject of one or more separate applications. Identical, similar, and / or functionally equivalent parts are provided with the same reference numerals. Here, the following are shown:
[0033] Figure 1 shows a block diagram of a deep neural network;
[0034] Figure 2 shows an image input dataset with an annotated bounding box and a latent representation dataset Z;
[0035] Figure 3 shows a flowchart of an evaluation method according to the invention.
[0036] Description of the execution types
[0037] Figure 1 shows a block diagram of a deep neural network 10 (DNN) that has been trained to detect objects, for example pedestrians, in camera-based 2D image data from a vehicle's surroundings. The DNN 10 has a number of feature extraction layers 11 and a number of perception layers 12. The feature extraction layers 11 are designed to generate a latent representation data set Z for a current image input data set 1 and to pass it on to the perception layers 12. The perception layers 12 are designed to compare latent representations from a current latent representation data set Z with a number of learned prototypes for different classes of objects for similarity, so that objects in the 2D image data can be detected, i.e., classified and localized, based thereon.
[0038] The image input data set 1 shown in Figure 2 contains or describes a 2D image of a vehicle's surroundings, which may have been captured, for example, with a vehicle's front-end camera. The image input data set 1 has a number of pixels arranged in rows and columns according to the height (H) and a width (W) of the image. Each pixel of the image input data set 1 is described, for example, by a vector in the three-dimensional RGB color space. The DNN 10 can, for example, be designed as a convolutional neural network (CNN), the architecture of which provides for special convolution and bundling structures in the feature extraction layers 11.
[0039] Using the feature extraction layers 11, a latent representation dataset Z is generated for each current image input dataset 1, thereby achieving a data reduction with respect to the number of data points. For example, from an image input dataset 1 with 2048 x 1024 pixels, a latent representation dataset Z with 512 x 256 latent representations Zjj is generated.
[0040] While each pixel is defined by a vector of three color values in the RGB color space, each latent representation contains a vector of features in an n-dimensional latent space. Each latent representation encodes, in the n features (e.g., 256 features), semantic relationships between image input data from a receptive field that has been incorporated into the respective latent representation.
[0041] How a respective image input data set 1 is mapped to the latent representation data set Z is determined on the one hand by the network architecture of the feature extraction levels 11 of the DNN 10, and on the other hand by a number of associated weighting parameters that have been learned by machine learning from training data.
[0042] The perception layers 12 of the DNN 10 have a number of prototypes that can be represented by a vector in the same latent space as the latent representations of the latent representation dataset Z. Thus, all latent representations can be compared with a number of prototypes that have also been learned using machine learning from the training data. From the comparison, the DNN 10 obtains output datasets 3 that contain classes and positions of objects that have been recognized by the perception layers 12 in the latent representation dataset Z.
[0043] The internal structures of DNN 10, which can also be referred to as hidden or latent structures, are fundamentally incomprehensible from the outside. However, the latent representations Zjj can be assigned to a spatially corresponding group of pixels in the input image data set 1.
[0044] For training and evaluation purposes, the DNN 10 is connected to a training algorithm 20. With the training algorithm 20, the weighting parameters and prototypes can be gradually optimized using a sequence of training iteration steps and by minimizing an internal cost function. For each training iteration step, a training dataset is provided, which comprises an image input dataset 1 and an associated reference dataset 2. With each image input dataset 1, a reference dataset 2 is provided, which is represented in Figure 2 as an annotated bounding box. The training algorithm 20 is provided with classes and positions of objects to be recognized that are present in an image of the corresponding image input dataset 1 and that are to be recognized by the DNN 10.The DNN 10 further feeds an output data set 3 to the training algorithm 20, which contains classes and positions of objects that have been recognized by the perception layers 12 from the latent representation data set Z. The weighting parameters and prototypes of the DNN 10 are adjusted by the training algorithm 10 in each training iteration step such that the classes and positions of objects in the output data set 3 match those in the reference data set 2 as closely as possible.
[0045] Figure 3 shows a sequence of an evaluation method 100 according to the invention, with which a trained deep neural network can be evaluated with regard to its robustness and / or performance.
[0046] In a first step 101, a DNN 10 and an associated training algorithm 20 are provided, which can be supplied as a fully trained hardware and / or software component, for example, by a supplier. The previously learned weighting parameters and prototypes of the feature extraction layers 11 and the perception layers 12 are frozen and saved as a data set.
[0047] In a second step 102, a predetermined number of evaluation data sets are provided, each evaluation data set comprising an image input data set 1 and an associated reference data set 2. The evaluation data sets are fed to the DNN 10 and the associated training algorithm 20 in a defined order. For each evaluation data set, the DNN 10 generates a latent representation data set Z and an output data set 3 with data about classes and positions of detected objects.
[0048] The data from reference dataset 2 are compared with those from output dataset 3. Based on the number of objects that were not detected or were falsely positive by DNN 10 compared to the reference data, an error rate F is determined and stored for each evaluation dataset. An average and / or maximum error rate is determined and stored for all evaluation datasets as a reference value for the performance of DNN 10.
[0049] Furthermore, based on the reference dataset 2, all latent representations Zjj that match the positions of the annotated bounding boxes are determined and stored. A distance average is calculated from the stored latent representations Zjj of all evaluation datasets, which is stored as a reference value for the robustness of the DNN 10.
[0050] In a third step 103, a number of test data sets and a number of methods for varying data sets are provided. For each method for varying data sets, a number of additional training data sets are generated.
[0051] The test data sets each contain an image input data set 1 and a corresponding reference data set 2, with the classes and positions of the objects to be recognized, which are annotated or specified as usual.
[0052] Methods for varying data sets can include global changes to the image, such as changes in contrast or brightness. However, the addition of additional objects to the image (e.g., people, pedestrians, etc.) is also possible. Furthermore, an image can be overlaid with snowfall, rain, or other image noise.
[0053] For example, if the contrast change method is selected, an additional training data set is generated based on each test data set, the pixels of which in the image input data set 1 have been manipulated or adjusted with regard to the contrast values by a specified value.
[0054] In a fourth step 104, the DNN 10 is placed into training mode and trained with a number of additional training data sets in an additional training step that were generated for a first selected method for varying data sets. Once a corresponding additional training step has been completed with the number of additional training data sets, the training mode of the DNN 10 is terminated.
[0055] In a fifth step 105, the evaluation data sets are fed to the DNN 10 and the associated training algorithm 20 in the defined order. For each evaluation data set, the DNN 10 again generates a latent representation data set Z and an output data set 3 with data about classes and positions of detected objects.
[0056] In a sixth step 106, for each evaluation dataset, all latent representations Zjj that match the positions of the annotated bounding boxes are determined based on the reference dataset 2. From the determined latent representations Zjj of all evaluation datasets, a distance mean is again determined, which is stored as an evaluation parameter for the robustness of the DNN 10.
[0057] A difference is calculated between the determined robustness evaluation value and the reference value for the robustness of the DNN 10 stored in the second step 102 and compared with a threshold value for a robustness change. If the threshold value is not met, i.e., if the robustness has deteriorated beyond a certain level after the additional training, the deep neural network 10 is assessed as unsuitable for the selected method. The process for evaluating the DNN 10 can be terminated with step 110.
[0058] If the threshold for the robustness change has been met, in a seventh step 107, an error rate F is again determined and stored for each evaluation data set by comparing the reference data sets 2 with the output data sets 3. As previously described in the second step 102, the error rate F results from the number of objects that were not detected or were falsely detected by the DNN 10 compared to the reference data. Accordingly, an average and / or maximum error rate is again determined and stored for all evaluation data sets, which serves as an evaluation parameter for the performance of the DNN 10.
[0059] In the seventh step 107, a difference is determined between the determined performance evaluation value and the reference value for the performance of the DNN 10 stored in the second step 102 and compared with a predetermined threshold for a change in performance. If the threshold is not met, i.e., if the performance has deteriorated beyond a certain level after the additional training, the deep neural network 10 is assessed as unsuitable for the selected method. The method for evaluating the DNN 10 can be terminated accordingly 110.
[0060] If the threshold for the change in performance has been met, the DNN 10 is assessed as suitable for the selected method. In a subsequent eighth step 108, the DNN 10 can be reset by playing back the originally learned weighting parameters and prototypes of the feature extraction levels 11 and the perception levels 12, respectively, stored as a data set. If it is further determined in the eighth step 108 that further methods for varying data sets are available, i.e. methods not yet used for additional training, the DNN 10 is again placed in training mode in the fourth step 104 and trained with a number of additional training data sets in an additional training session, which were generated for a further selected method for varying data sets.
[0061] If it is determined in the eighth step 108 that all additional training data sets have been used for all methods for varying data sets, the DNN 10 is evaluated as suitable for all methods used. The process is terminated with a positive evaluation of the DNN 10 in step 109.
[0062] If in the sixth step 106 and / or seventh step 107 the deep neural network 10 is assessed as unsuitable for a selected method, alternatively to terminating 110 the method, it can be provided to carry out further additional training for the methods not yet used for varying data sets in order to obtain a complete list of the methods for which the deep neural network is suitable and for which it is not.
[0063] Although the subject matter has been illustrated and explained in detail by means of exemplary embodiments, the invention is not limited by the disclosed examples, and other variations may be derived therefrom by those skilled in the art. It is therefore clear that a multitude of possible variations exist. It is also clear that the exemplary embodiments mentioned are merely examples and should not be construed as limiting the scope, possible applications, or configuration of the invention in any way.Rather, the foregoing description and the description of the figures enable the person skilled in the art to implement the exemplary embodiments in concrete terms. With knowledge of the disclosed inventive concept, the person skilled in the art can make various changes, for example, with regard to the function or arrangement of individual elements mentioned in an exemplary embodiment, without departing from the scope of protection defined by the claims and their legal equivalents, such as further explanations in the description. List of reference symbols.
[0064] 1 image input data set
[0065] 2 Reference data set
[0066] 3 Output data set
[0067] 10 deep neural network (DNN)
[0068] 11 feature extraction levels
[0069] 12 levels of perception
[0070] 20 Training algorithm
[0071] F error rate
[0072] Z latent representation dataset
[0073] Zi,j latent representation
[0074] 100 evaluation procedures
[0075] 101 first step
[0076] 102 second step
[0077] 103 third step
[0078] 104 fourth step
[0079] 105 fifth step
[0080] 106 sixth step
[0081] 107 seventh step
[0082] 108 eighth step
[0083] 109 End; rated as suitable for all methods used
[0084] 110 End; rated as unsuitable for a selected method
Claims
Claims 1. A method for evaluating a deep neural network (10) that has been trained for object recognition in the environment of a motor vehicle, wherein at least one reference value is determined based on a predetermined number of evaluation data sets and with the aid of the deep neural network (10), and wherein a number of test data sets are provided, from which a number of additional training data sets are generated by means of a selected method for varying data sets, wherein the deep neural network (10) is trained in an additional training step with the generated additional training data sets, wherein after completing the additional training step, at least one evaluation value is determined based on the predetermined number of evaluation data sets and with the aid of the deep neural network (10), and wherein the deep neural network (10) is assessed as unsuitable for the selected method,if a difference between the least one reference value and a corresponding evaluation value exceeds a predetermined limit.
2. Method according to the preceding claim 1, wherein originally learned parameters of the deep neural network (10) are read out and stored before a first additional training step is carried out.
3. Method according to the preceding claim 2, wherein all parameters of the deep neural network (10) are reset to the stored originally learned parameters after the at least one evaluation variable has been determined.
4. Method according to one of the preceding claims 1 to 3, wherein a number of different methods for varying data sets are provided, wherein for each of the different methods for varying data sets a number of additional training data sets are generated from the test data sets, so that a number of successive additional training steps corresponding to the number of methods are carried out.
5. Method according to one of the preceding claims 1 to 4, wherein for each evaluation data set a corresponding latent representation data set (Z) is generated by the deep neural network (10), wherein latent representations (Zjj) relevant to annotated data (2) of each evaluation data set are selected from the corresponding latent representation data set (Z), wherein a distance mean value in the latent space is determined from the selected relevant latent representations of the evaluation data sets, and wherein the determined distance mean value is used as a reference value and / or as an evaluation value.
6. Method according to one of the preceding claims 1 to 5, wherein output data (3) are generated for each evaluation data set by the deep neural network (10), wherein a performance is determined by means of a comparison between the generated output data (3) and the reference data (2) of the respective evaluation data set, wherein the performance is used as a reference variable and / or as an evaluation variable.
7. The method according to any one of the preceding claims, wherein the method for varying data sets comprises a parameterized modification of all data of a data set.
8. The method according to any one of the preceding claims, wherein the method for varying data sets comprises adding objects to a respective data set.
9. A computer program which, when executed on a computing unit, instructs the respective computing unit to carry out a method according to one of claims 1 to 8.
10. A computer program product comprising a program code stored on a computer-readable medium for carrying out the method according to any one of claims 1 to 9.