Method for making a neural network robust against adversarial perturbations
Patent Information
- Application Number
- DE502020011109
- Authority / Receiving Office
- DE · DE
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2019-10-21
- Filing Date
- 2020-10-07
- Publication Date
- 2025-06-12
- Estimated Expiration
- 2040-10-07
AI Technical Summary
Convolutional neural networks are vulnerable to adversarial perturbations in sensor data, leading to misclassification or incorrect semantic segmentation despite semantically unchanged content, which poses a risk in applications like autonomous driving.
A method and device that modify input data of neural networks during application phases to align with statistical properties of training data, using a statistical model to make the data more probable, thereby reducing the effect of adversarial disturbances.
Enhances the robustness of neural networks against adversarial attacks by making input data statistically more probable, ensuring that adversarial disturbances are mitigated without altering the neural network's design or requiring retraining.
Description
[0001] The invention relates to a method for robustifying a neural network against adversarial disturbances and a device for providing a neural network robust against adversarial disturbances. Furthermore, the invention relates to a computer program and a data carrier signal.
[0002] Machine learning, for example based on neural networks, has great potential for application in modern driver assistance systems and automated vehicles. Functions based on deep neural networks process sensor data (e.g., from cameras, radar, or lidar sensors) to derive relevant information. This information includes, for example, the type and position of objects in the vehicle's surroundings, the behavior of the objects, or the road geometry or topology.
[0003] Among neural networks, convolutional neural networks (CNNs) have proven particularly suitable for applications in image processing. Convolutional networks gradually extract various high-quality features from input data (e.g., image data) in an unsupervised manner. During a training phase, the convolutional network independently develops feature maps based on filter channels that locally process the input data to derive local properties. These feature maps are then processed again by additional filter channels, which derive higher-quality feature maps. Based on this information condensed from the input data, the deep neural network ultimately derives its decision and provides this as output data.
[0004] While convolutional networks outperform traditional approaches in terms of functional accuracy, they also have drawbacks. For example, attacks based on adversarial perturbations in the sensor data / input data can lead to misclassification or incorrect semantic segmentation despite the semantically unchanged content of the acquired sensor data.
[0005] Techniques for training deep neural networks are known from DE 10 2018 115 440 A1, for example, those using an iterative approach. A training system for a deep neural network (DNN) is described, which generates a hardened DNN by iteratively training DNNs with images that were misclassified by previous iterations of the DNN. In one or more variants, the logic can, for example, generate an adversary image that is misclassified by a first DNN that was previously trained with a set of sample images. In some variants, the logic can determine a second training set that contains the adversary image that was misclassified by the first DNN and the first training set of one or more sample images. The second training set can be used to train a second DNN.In various variants, the above process can be repeated for a predetermined number of iterations to generate a hardened DNN.
[0006] A method for defending a neural network against adversarial perturbations is known from Tejas Borkar et al., Defending against Adversarial Attacks through Resilient Feature Regeneration, ARXIV.ORG, CORNELL UNIVERSITY LIBRARY, 201 OLIN LIBRARY CORNELL UNIVERSITY ITHACA, NY 14853, June 8, 2019.
[0007] Examples of physical designs of adversarial disturbances from the field of autonomous driving are known from Adith Boloor et al., Simple Physical Adversarial Examples against End-to-End Autonomous Driving Models, 2019 IEEE INTERNATIONAL CONFERENCE ON EMBEDDED SOFTWARE AND SYSTEMS (ICESS), IEEE, June 2, 2019, pages 1-7, DOI: 10.1109 / ICESS.2019.8782514.
[0008] The invention is based on the object of creating an improved method for robustifying a neural network against adversarial disturbances and an improved device for providing a neural network robustified against adversarial disturbances.
[0009] The object is achieved according to the invention by a method having the features of patent claim 1 and a device having the features of patent claim 11. Advantageous embodiments of the invention emerge from the subclaims.
[0010] In particular, a method for robustifying a neural network against adversarial disturbances is provided, wherein by means of at least one manipulator device during an application phase of the neural network, input data of at least one layer of the neural network are at least partially modified before they are fed to the at least one layer in such a way that the modified input data are statistically more probable according to a statistical model than the unchanged input data of the at least one layer, wherein the statistical model respectively maps statistical properties of input data of the at least one layer in the event that training data of a training data set with which the neural network was trained is fed to the neural network.
[0011] Furthermore, in particular, a device for providing a neural network robustified against adversarial disturbances is created, comprising a data processing device for providing the neural network, and at least one manipulator device, wherein the at least one manipulator device is configured to at least partially change input data of at least one layer of the neural network during an application phase of the neural network, before said data is fed to the at least one layer, in such a way that the changed input data are statistically more probable according to a statistical model than the unchanged input data of the at least one layer, wherein the statistical model respectively maps statistical properties of input data of the at least one layer for the case that training data of a training data set with which the neural network was trained is fed to the neural network.
[0012] The method and the invention make it possible to make a neural network more robust against adversarial disturbances. This is achieved by modifying input data from at least one layer of the neural network during an application phase of the neural network, i.e., when applying the trained neural network, such that the modified input data are statistically more probable according to a statistical model than the unchanged input data of at least one layer. The statistical model thereby represents statistical properties of input data of at least one layer for the case in which the neural network is fed training data from a training data set with which the neural network was trained.In other words, based on a statistical distribution of properties in the input data of at least one layer, as they exist when training data from the training data set is fed in, the attempt is made in the application phase—that is, when real data is fed into the neural network—to modify the input data of at least one layer in such a way that it more closely resembles the statistical distribution. This can ensure that adversarial disturbances in the real data lose their effect, meaning that the input data of at least one layer is cleansed of the adversarial disturbances.
[0013] One advantage of the method and device is that the neural network can be made more robust regardless of the specific adversarial disturbance. Another advantage is that the manipulator device and the neural network are independent of each other and can therefore be further developed independently of each other. In particular, the statistical model can be modified, especially refined and improved, without having to modify the neural network itself.
[0014] The method is carried out as a computer-implemented method.
[0015] In particular, the method is carried out using a data processing device. The data processing device comprises, in particular, at least one computing device and at least one memory device. The data processing device provides, in particular, the neural network and the at least one manipulator device.
[0016] In particular, a computer program is also created, comprising instructions which, when the computer program is executed by a computer, cause the computer to carry out the method steps of the method according to the invention.
[0017] In addition, in particular, a data carrier signal is also created, comprising instructions which, when executed by a computer, cause the computer to carry out the method steps of the method according to the invention.
[0018] An adversarial perturbation is, in particular, a deliberate perturbation of the input data of a neural network, in which a semantic content in the input data is not changed, but the perturbation leads to the neural network inferring an incorrect result, i.e., for example, a misclassification or an incorrect semantic segmentation of the input data.
[0019] A neural network is, in particular, a deep neural network, in particular a convolutional neural network (CNN). The neural network is or has been trained, for example, for a specific perceptual function, such as the perception of pedestrians or other objects in captured camera images. The invention is, in particular, independent of the specific design of the neural network, but relates only to the input data of the neural network or the input data of layers of the neural network.
[0020] The input data of the neural network and its individual layers can be one-dimensional or multi-dimensional. In particular, the input data of the neural network is two-dimensional data, for example, image data captured by a camera. The neural network provides, in particular, an inferred result as output data, which is provided, in particular, as an output data signal, for example, in the form of a digital data packet.
[0021] A layer of the neural network can be either an input layer of the neural network or an inner layer of the neural network. A layer can also be an output layer of the neural network, with the output layer serving only to forward input data from the output layer without further modification or processing by the output layer.
[0022] The statistical model, in particular, represents the statistical properties of input data as a function of given training data. If multiple layers are considered or their input data is changed, the statistical model includes, in particular, a separate statistical model for the input data of each of these layers. Furthermore, the statistical model can include additional submodels for input data of elements or groups of elements within layers.
[0023] In one embodiment, the at least one layer comprises an input layer of the neural network, wherein the statistical model associated with the input layer maps statistical properties of the training data of the training data set. The training data in this case particularly comprises sensor data from a sensor system. The sensor data are understood as statistical processes, wherein values in the sensor data can be assigned a probability of occurrence. Statistical distributions can then be modeled for various properties of the training data, which can be used to statistically describe the properties of the training data.If the input data of the neural network, i.e. the input data that is fed to the input layer of the neural network, is camera images, for example, the properties can be concrete values of image elements (pixels) of the camera data or higher-order features, such as spatial frequencies, edge distributions and / or feature maps using learned filters, in their spatial distribution. The statistical model is then adapted (fitted) accordingly for these properties, so that the statistical model subsequently models and maps the statistical distributions of these properties. The individual properties, i.e. the individual features, are generally not statistically independent of one another, but rather mutually dependent. The statistical dependencies can, for example, also be taken into account in the statistical distributions.A camera image, for example, can be modeled as a Markov Random Field. If input data is then fed to the input layer of the neural network in an application phase—in this example, captured camera images—the input data is modified in such a way that the modified input data, particularly with regard to the characteristics of the respective properties, are more probable than the unchanged input data according to the respective statistical distribution in the statistical model.
[0024] In one embodiment, it is provided that the at least one layer comprises at least one inner layer of the neural network, wherein associated input data of the at least one inner layer are activations of a respective preceding layer of the neural network, and wherein, in order to change the input data of the at least one inner layer, the activations of the respective preceding layer are at least partially changed, wherein the statistical model associated with the at least one inner layer maps statistical properties of the activations of the respective preceding layer for the case that the neural network is supplied with the training data of the training data set with which the neural network was trained. In this way, input data of inner layers of the neural network can be changed according to the statistical model.For this purpose, the statistical model includes, for given training data of the neural network, statistical distributions of the properties of the activations of a layer preceding a considered inner layer. By changing the input data of at least one inner layer or the activations of a corresponding preceding layer, the effect of adversarial disturbances can be reduced or even completely eliminated. For this purpose, the activations are modified in such a way that the modified activations are more probable according to the statistical distributions contained in the statistical model.
[0025] In particular, the method can be applied both to input data of the neural network, i.e., to input data of an input layer of the neural network, and to input data of at least one inner layer of the neural network, or to activations of a respective preceding layer. This allows for the removal of adversarial interference both upstream of the neural network and within the neural network itself. Robustness can thus be further improved.
[0026] In one embodiment, the statistical model is generated during or after a training phase of the neural network, starting from the training data set, using a backend server and made available to the at least one manipulator device. This allows the method to be used in the field, for example in a motor vehicle, while the creation or adaptation of the statistical model can be carried out on a powerful backend server with a large training data set. The adapted statistical model is then fed to the at least one manipulator device. Furthermore, after an initial adaptation of the statistical model, it is possible to further improve and refine the statistical model on the backend server (e.g., by including statistical distributions for additional properties), regardless of its use in the field.An improved statistical model is then fed back to the at least one manipulator device, which replaces a previous version of the statistical model with the improved model. This can enable continuous improvement and / or refinement of the statistical model.
[0027] In a further development, it is provided that a ground truth of the training data is taken into account when generating the statistical model. This allows the ground truth of the training data to be represented in the statistical model as well. This also enables, in particular, the modification of the input data of an output layer of the neural network. This can be used, for example, to detect presumably adversarial situations; only with a very high confidence level would the adaptation of the input data of the output layer be assumed to be correct and further processed. After detecting an adversarial situation, a more defensive driving strategy or the use of redundant information could then be preferred.
[0028] In one embodiment, the statistical model is selected during the application phase depending on the current context of the input data of the neural network. This allows the statistical model to be provided in a context-dependent or situation-dependent manner. If the input data of the neural network is, for example, captured camera images of an environment in which the neural network is to recognize objects, a statistical model can be selected and provided depending on the season (summer, winter) and / or the time of day (day / night) of the captured camera images, etc. A specialized statistical model is then provided for each context. This allows the statistical model to map a statistical distribution of properties of the input data in an improved, context-dependent manner. The robustness of the neural network is thereby further improved.
[0029] In one embodiment, the degree of modification of the input data is specified using at least one adaptation parameter. This allows the degree of modification to be specifically adjusted. In particular, this allows the degree of modification to be limited so that any semantic content present in the input data is not affected by the modification.
[0030] In a further embodiment, it is provided that the at least one adaptation parameter specifies a threshold value for a distance measure determined from the unchanged input data and the associated modified input data. If the input data are camera images, for example, a Euclidean distance between the camera images can be defined. The Euclidean distance results in particular from the camera images understood as vectors by forming the square root of the sum of the squared pixel differences of the two camera images. A modification is then carried out, for example, by replacing values for individual pixels in such a way that a modified pixel is more probable according to the statistical model, but a certain distance measure is smaller than the specified adaptation parameter.
[0031] In one embodiment, it is provided that the at least one adaptation parameter is or is predetermined depending on the at least one layer and / or an element of the at least one layer. This allows individual adaptation parameters to be specified for individual elements (e.g., for individual neurons or individual filter kernels) of layers and / or for individual layers. This further improves the robustness of the neural network, since the strength of the change can be individually specified for different areas of the neural network, and thus structural properties and peculiarities of the neural network, which often cause a strong effect of adversarial disturbances, can be taken into account.
[0032] In one embodiment, the neural network provides a function for automated driving of a vehicle and / or for driver assistance of the vehicle and / or for environmental detection. A vehicle is, in particular, a motor vehicle. In principle, however, the vehicle can also be another land, air, water, rail, or space vehicle.
[0033] Features for the design of the device are derived from the description of embodiments of the method. The advantages of the device are the same as those of the embodiments of the method.
[0034] Furthermore, in particular, a motor vehicle is also created, comprising at least one device according to one of the described embodiments.
[0035] The invention will be explained in more detail below using preferred embodiments with reference to the figures. Fig. 1 is a schematic representation of an embodiment of the device for robustifying a neural network against adversarial disturbances; Fig. 2 is a schematic representation of a flow chart to illustrate an embodiment of the method for robustifying a neural network against adversarial disturbances; Fig. 3 is a schematic representation of a flow chart to illustrate the adaptation of the statistical model for the Fig. 2 shown embodiment of the method; Fig. 4 a schematic representation of a flow chart to illustrate a further embodiment of the method for robustifying a neural network against adversarial disturbances; Fig. 5 a schematic representation of a flow chart to illustrate the adaptation of the statistical model for the Fig. 4 shown further embodiment of the method.
[0036] In Fig. 1A schematic representation of an embodiment of the device 1 for providing a neural network robust against adversarial disturbances is shown. The device 1 is arranged in a motor vehicle 50 and serves, for example, to provide a function for automated driving of the motor vehicle 50 and / or for driver assistance of the motor vehicle 50 and / or for environmental detection or environmental recognition.
[0037] The device 1 comprises a data processing device 2. The data processing device 2 comprises a computing device 3, for example a microprocessor, and a memory device 4, which the computing device 3 can access. The data processing device 2 provides the neural network, in particular in the form of a deep neural network or convolutional neural network (CNN), and a manipulator device in the form of an additional program module. To provide the data, the computing device 3 executes instructions stored in the memory device 4.
[0038] Sensor data 10 from a sensor 51 of the motor vehicle 50, for example, camera images from a camera, are fed to the device 1. The neural network performs, for example, object recognition and / or semantic segmentation on the sensor data 10. A result 20 inferred by the neural network is then fed, for example, to a planning device 52 of the motor vehicle 50, which, based thereon, generates a planned trajectory for the motor vehicle 50.
[0039] In order to robustify the neural network against adversarial disturbances, the at least one manipulator device is configured to, during an application phase of the neural network, at least partially modify input data of at least one layer of the neural network before they are fed to the at least one layer in such a way that the modified input data are statistically more probable according to a statistical model than the unchanged input data of the at least one layer. For this purpose, the statistical model maps statistical properties of input data of the at least one layer for the case that the neural network is fed training data from a training data set with which the neural network was trained.
[0040] In Fig. 2A schematic representation of a flowchart is shown to illustrate an embodiment of the method for robustifying a neural network 5 against adversarial disturbances. The embodiment shows, by way of example, the modification of input data 11 of an input layer 5-0 of the neural network 5. In other words, sensor data 10 (input data 11) acquired by a sensor 51 is modified before being fed to the input layer 5-0 of the neural network 5.
[0041] The change is carried out by means of a manipulator device 6. The manipulator device 6 is provided with a statistical model 8 which maps statistical properties of training data of a training data set with which the neural network 5 was trained.
[0042] By means of the manipulator device 6, the input data 11 are at least partially changed during an application phase of the neural network 5 in such a way that changed input data 12 are statistically more probable than the unchanged input data 11 according to the statistical model 7.
[0043] For example, if x denotes a date in the input data 11 and x' denotes a changed date in the changed input data 12, the change is carried out, for example, by the following measures: Replace x by x', so that 1. P(x') = maximal and 2. d(x,x') < ε where P(x') denotes a probability for the occurrence of the data item x' derived from the statistical model, d(x,x') is a distance measure between the data items x and x', and ε is an adjustment parameter. The adjustment parameter ε is determined and specified, for example, based on empirical test runs with test data ("in" and "out of sample"), taking adversarial disturbances into account.
[0044] The manipulator device 6 replaces the data x piecewise, for example, using greedy algorithms (e.g., "replace individual input data values, such as pixel values, as optimally as possible until a budget of permitted manipulations is exhausted"), divide-and-conquer algorithms (e.g., "divide the input data into subsets and manipulate these subsets with a correspondingly adjusted manipulation budget"), and / or gradient-based algorithms (e.g., "calculate the derivative of P(x) and modify the input data in the direction of the resulting gradient"). The adaptation parameter ε represents a heuristic hyperparameter, which is chosen in particular such that the modification remains subsymbolic and does not result in any relevant semantic change.
[0045] The modified input data 12 are then fed to the input layer 5-0 of the neural network 5.
[0046] In Fig. 3is a schematic representation of a flow chart to illustrate the fitting or determination of the statistical model for the Fig. 2 The embodiment of the method shown is shown. The adaptation takes place in particular on a backend server or by means of a backend server.
[0047] The neural network is trained in a method step 100 using training data 30 and a ground truth 31 associated with the training data 30. For example, the neural network can be trained through supervised learning to perform semantic segmentation on camera images.
[0048] The statistical model 7 is then adapted based on the training data 30. For this purpose, statistical distributions of properties of the training data 30 are determined. The properties can relate to both raw information of the training data 30, e.g., pixel values in camera images used for training, as well as higher-level features (e.g., spatial frequencies, edge distributions, etc.). The set of specific statistical distributions forms the statistical model adapted or fitted to the training data 30. After the adaptation (fitting) of the statistical model 7, it is transferred to the manipulator device 6 ( Fig. 2 ).
[0049] Fitting a date refers in particular to determining a probability for the occurrence of the date in the training data set 30: x → P(x)
[0050] It can additionally be provided that a ground truth 31 of the training data 30 is taken into account when generating the statistical model. For this purpose, the statistical model is also adapted based on the ground truth 31 in a method step 102. This makes it possible, in particular, to also change input data of an output layer of the neural network.
[0051] In Fig. 4 A schematic representation of a flowchart is shown to illustrate another embodiment of the method for robustifying a neural network 5 against adversarial disturbances. The embodiment shows, by way of example, the modification of input data 13 of an inner layer 5-i of the neural network 5. In other words, activations 14 inferred by a layer 5-x of the neural network 5 are modified before being fed to the inner layer 5-i of the neural network 5.
[0052] The change is carried out by means of a further manipulator device 8. For this purpose, the further manipulator device 8 changes a data flow within the neural network 5. A statistical model 7 is provided to the further manipulator device 8, wherein the statistical model 7 belonging to the inner layer 5-i maps statistical properties of the activations 14 of the preceding layer 5-x for the case that the neural network 5 is supplied with the training data of the training data set with which the neural network 5 was trained.
[0053] By means of the further manipulator device 8, the input data 13, i.e. the activations 14, are at least partially changed during an application phase of the neural network 5 in such a way that changed input data 15, i.e. changed activations 16, are statistically more probable according to the statistical model 7 than the unchanged input data 13 or the unchanged activations 14.
[0054] In addition, in this embodiment it is provided that the input data 11 of the input layer 5-0 of the neural network 5 - as already described with reference to the Fig. 2 described - be changed.
[0055] For example, if x denotes a date in the input data 11 of the input layer 5-0, x' a changed date, a(x') the activation given the changed date x' of the input layer and a' the changed activation, the change is carried out, for example, by the following measures: Replace a(x') by a', so that 1. P(a', x') = maximum and 2. d(a(x'),a') < ε where P(a', x') denotes a probability for the occurrence of a' given the (changed) data x', d(a(x'),a') is a distance measure between the data a and a' and ε is an adjustment parameter.
[0056] The further manipulation device 8 replaces the activations a piecewise, for example, using greedy algorithms, divide-and-conquer algorithms, and / or gradient-based algorithms (see above). The adaptation parameter ε represents a heuristic hyperparameter and is determined, for example, empirically based on test runs using test data ("in" and "out of sample") and, for fine-tuning the adaptation parameter ε, using adversarial perturbations contained in the test data.
[0057] The input data 15 or modified activations 16 thus modified are then fed to the inner layer 5-i of the neural network 5.
[0058] In Fig. 5 is a schematic representation of a flow chart to illustrate the fitting of the statistical model for the Fig. 4 The embodiment of the method shown is shown. The adaptation takes place, for example, on the backend server or via the backend server. Shown is the adaptation of the statistical model 7 with respect to activations 13 of layer 5-x ( Fig. 4 ). The adaptation of the statistical model 7 with respect to the input data 11 of the input layer 5-0 has already been described in Fig. 3 illustrated by example.
[0059] The neural network is applied multiple times to training data 30 in a method step 200. After each application, activations 14 of layer 5-x ( Fig. 4 ) extracted.
[0060] The activations 14 are statistically evaluated in a method step 201. For this purpose, the statistical model is adapted (fitted) to a distribution of the activations 14 in relation to the training data 30 used in each application. In other words, statistical distributions in the properties of the activations 14 are determined in relation to the training data 30 used in each application of the neural network. The properties include, in particular, both raw information in the activations 14 (for example, in the feature maps provided by individual filters in layer 5-x) and higher-quality features (e.g., spatial frequencies, edge distributions in the feature maps of individual filters in layer 5-x, etc.).
[0061] The set of specific statistical distributions forms the adapted statistical model 7 or at least the part thereof that concerns the activations 14 of layer 5-x. The adapted statistical model 7 is then passed to the further manipulator device 8 ( Fig. 4 ). The statistical model 7 then includes in particular both the statistical distributions for the input data 11 of the input layer 5-0 and the statistical distributions of the activations 14 of the layer 5-x.
[0062] Fitting an activation a(x) to the date x refers in particular to determining a probability for the occurrence of the activation a(x) given the date x in the training data set 30: x , a x → P x a x = P a x x
[0063] The exemplary embodiments only provide for a change of the input data 11 of the input layer 5-0 or the input data 13 of a single inner layer 5-i of the neural network 5 (cf. Fig. 2 and Fig. 4 ). In principle, however, it is possible to consider additional inner layers or just one or more inner layers. The procedure is essentially analogous, with the statistical model always including and providing statistical distributions for properties of input data from the respective layers. List of reference symbols
[0064] 1 Device 2 Data processing device 3 Computing device 4 Storage device 5 Neural network 5-0 Input layer 5-x Layer 5-i Inner layer 6 Manipulator device 7 Statistical model 8 Further manipulator device 10 Sensor data 11 Input data (of the neural network) 12 Modified input data 13 Input data (of the inner layer) 14 Activations (of layer 5-x) 15 Modified input data 16 Modified activations 20 Inferred result 30 Training data 31 Ground truth 50 Motor vehicle 51 Sensor 100-102 Process steps ε Adaptation parameters
Claims
1. Computer-implemented method for making a neural network (5) more robust against adversarial disruptions, wherein the neural network (5) is supplied with sensor data (10) from a sensor (51) of a motor vehicle (50), wherein, by means of at least one manipulator device (6, 8), during a use phase of the neural network (5), input data (11, 13) of at least one layer (5-0, 5-x, 5-i) of the neural network (5) are at least partially modified, before being supplied to the at least one layer (5-0, 5-x, 5-i), in such a way that the modified input data (12, 15) are statistically more probable than the unmodified input data (11, 13) of the at least one layer (5-0, 5-x, 5-i), according to a statistical model (7) provided to the at least one manipulator device (6, 8), wherein the statistical model (7) in each case maps statistical properties of input data (11, 13) of the at least one layer (5-0, 5-x, 5-i) in the event that the neural network (5) is supplied with training data (30) of a training data set on which the neural network (5) was trained.
2. Method according to claim 1, characterized in that the at least one layer (5-0, 5-x, 5-i) comprises an input layer (5-0) of the neural network (5), the statistical model (7) assigned to the input layer (5-0) mapping statistical properties of the training data (30) of the training data set.
3. Method according to claim 1 or 2, characterized in that the at least one layer (5-0, 5-x, 5-i) comprises at least one inner layer (5-i) of the neural network (5), corresponding input data (13) of the at least one inner layer (5-i) being activations (14) of a relevant preceding layer (5-x) of the neural network (5), and the activations (14) of each preceding layer (5-x) being at least partially modified in order to modify the input data (13) of the at least one inner layer (5-i), the statistical model (7) assigned to the at least one inner layer (5-i) mapping statistical properties of the activations (14) of the relevant preceding layer (5-x) in the event that the neural network (5) is supplied with the training data (30) of the training data set on which the neural network (5) was trained.
4. Method according to any of the preceding claims, characterized in that the statistical model (7) is generated during or after a training phase of the neural network (5), based on the training data set, by means of a backend server, and is provided to the at least one manipulator device (6, 8).
5. Method according to claim 4, characterized in that a ground truth (31) of the training data (30) is considered when generating the statistical model (7).
6. Method according to any of the preceding claims, characterized in that the statistical model (7) is selected during the use phase, depending on a current context of the input data (11) of the neural network (5).
7. Method according to any of the preceding claims, characterized in that a strength of the modification of the input data (11, 13) is specified by means of at least one adjustment parameter (ε).
8. Method according to claim 7, characterized in that the at least one adjustment parameter (ε) specifies a threshold value for a distance measure determined from the unmodified input data (11, 13) and corresponding modified input data (12, 15).
9. Method according to claim 7 or 8, characterized in that the at least one adjustment parameter (ε) has been predetermined or is predetermined depending on the at least one layer (5-0, 5-x, 5-i) and / or an element of the at least one layer (5-0, 5-x, 5-i).
10. Method according to any of the preceding claims, characterized in that the neural network (5) provides a function for automated driving of a motor vehicle (50) and / or for driver assistance of the motor vehicle (50) and / or for detecting surroundings.
11. Apparatus for providing a neural network (5) which has been made more robust against adversarial disruptions, wherein the neural network (5) is supplied with sensor data (10) from a sensor (51) of a motor vehicle (50), comprising: a data processing device (2) for providing the neural network (5), and at least one manipulator device (6, 8), wherein the at least one manipulator device (6, 8) is designed toto at least partially modify input data (11, 13) of at least one layer (5-0, 5-x, 5-i) of the neural network (5) during a use phase of the neural network (5), before said data are supplied to the at least one layer (5-0, 5-x, 5-i), in such a way that the modified input data (12, 15) are statistically more probable than the unmodified input data (11, 13) of the at least one layer (5-0, 5-x, 5-i), according to a statistical model (7) provided to the at least one manipulator device (6, 8), wherein the statistical model (7) in each case maps statistical properties of input data (11, 13) of the at least one layer (5-0, 5-x, 5-i) in the event that the neural network (5) is supplied with training data (30) of a training data set on which the neural network (5) was trained.
12. Motor vehicle (50), comprising at least one apparatus according to claim 11.
13. Computer program comprising commands which, when the computer program is executed by a computer, cause the computer to execute the method according to any of claims 1 to 10.
14. Data carrier signal which transmits the computer program according to claim 13.