Method for training artificial neural network and / or verifying robustness of artificial neural network
Patent Information
- Application Number
- JP2022124566
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2021-08-05
- Filing Date
- 2022-08-04
- Publication Date
- 2025-08-07
- Estimated Expiration
- 2042-08-04
AI Technical Summary
Existing systems for training and testing artificial neural networks lack comprehensive integration, leading to incomplete training and testing processes, which can compromise the reliability and robustness of autonomous systems.
A method for training and testing artificial neural networks that involves determining input and output variable limits, using linear layers without activation functions, and applying disturbance models to assess robustness, ensuring the network operates within specified bounds.
Enhances the robustness and reliability of neural networks by training them to withstand disturbances within defined limits, improving their performance in autonomous systems.
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
[Technical Field]
[0001] Background technology Artificial neural networks are used for autonomous driving and for other autonomous operations that machines can perform, and the training and robustness testing of such artificial neural networks is essential for the reliability and reliability of such technical systems. [Background technology]
[0002] When integrating multiple systems, system integrators do not always have access to all aspects of the supplier's training and / or testing processes. These systems use artificial neural networks that are already pre-trained and, in some cases, further trained for specific tasks. [Prior art documents] [Non-patent literature]
[0003] [Non-Patent Document 1] Publication “Scaling provable adversarial defenses” by Eric Wong, Frank R. Schmidt, Jan Hendrik Metzen, and J. Zico Kolter (https: / / arxiv.org / abs / 1805.12514) [Non-patent document 2] Publication “On the Effectiveness of Interval Bound Propagation for Training Verifiably Robust Models” by Sven Gowal, Krishnamurthy Dvijotham, Robert Stanforth, Rudy Bunel, Chongli Qin, Jonathan Uesato, Relja Arandjelovic, Timothy Mann, Pushmeet Kohli (https: / / arxiv.org / abs / 1810.12715) Summary of the Invention [Problem to be solved by the invention]
[0004] It is therefore desirable to be able to train and / or test the robustness of such artificial neural networks.
[0005] Means for doing this are disclosed, for example, in the publication "Scaling provable adversarial defenses" by Eric Wong, Frank R. Schmidt, Jan Hendrik Metzen, and J. Zico Kolter, which is available at https: / / arxiv.org / abs / 1805.12514.
[0006] Means for doing this are disclosed, for example, in the publication "On the Effectiveness of Interval Bound Propagation for Training Verifiably Robust Models" by Sven Gowal, Krishnamurthy Dvijotham, Robert Stanforth, Rudy Bunel, Chongli Qin, Jonathan Uesato, Relja Arandjelovic, Timothy Mann, and Pushmeet Kohli, which is available at https: / / arxiv.org / abs / 1810.12715. [Means for solving the problem]
[0007] Disclosure of the Invention The methods and devices and computer programs according to the independent claims are used to train a robust artificial neural network or to test the robustness of an artificial neural network.
[0008] A method for training an artificial neural network and / or testing the robustness of an artificial neural network configured to determine output variables is envisioned, the method comprising setting input variables to the network having a plurality of dimensions, the method comprising, for each dimension of the input variables or for each dimension of an output side of a linear layer of the artificial neural network without an activation function to which the input variables are mapped by the artificial neural network, determining an upper input variable limit and a lower input variable limit, where at the upper input variable limit, a disturbance variable model that can map the input variable to a disturbed input variable has the largest possible value in this dimension, and at the lower input variable limit, the disturbance variable model has the smallest possible value in this dimension. The method also comprises, for each dimension of the output variable, determining a lower output variable limit for the output variable of the network with a value from a range of values bounded by the lower input variable limit and the upper input variable limit, and determining an upper output variable limit for the output variable with a value from a range of values bounded by the upper input variable limit and the lower input variable limit. The method includes determining the smallest possible value of a function, particularly of real values, from a range of values bounded by a lower output variable limit and an upper output variable limit, and determining the largest possible value of a function, particularly of real values, from a range of values bounded by a lower output variable limit and an upper output variable limit, whereby an output is determined that demonstrates the robustness of the network if the smallest possible value and the largest possible value are within a set interval of acceptable values.
[0009] The input variables preferably represent digital images, in particular video images, radar images, LiDAR images, ultrasound images or infrared images.
[0010] The output variables preferably represent signals for the drive control of a physical system, in particular a computer-controlled machine or information transmission system, preferably a robot, a vehicle, a domestic device, a drive tool of a manufacturing machine, a personal assistance system, an access control system, a surveillance system or in particular a medical imaging system.
[0011] It may be envisaged that the output variables classify the sensor data, in particular for identifying objects in the sensor data or for semantic segmentation, or that the output variables provide a regression of the sensor data, in particular for identifying objects in the sensor data or for semantic segmentation, where the objects are preferably markings or objects on the road surface or objects representing markings or objects on the road surface, in particular objects representing road signs, pedestrians or vehicles.
[0012] For each dimension, it may be assumed that a lower input variable bound is determined, in which the scalar product of the input variable with the unit vector of this dimension is as small as possible.
[0013] For each dimension of the input variable, it may be assumed that a lower input variable limit for this dimension is determined, at which a first scalar product of the input variable and the negative unit vector of this dimension has the largest possible value, and further it may be assumed that for each dimension of the input variable, an upper input variable limit for this dimension is determined, at which a second scalar product of the input variable and the unit vector of this dimension has the largest possible value.
[0014] It may be assumed that the output side of a linear layer of an artificial neural network without an activation function is defined by a matrix, and that for each dimension of the output side a first product determinable by multiplying the negative transpose of the matrix by a unit vector of this dimension is defined, and that a lower input variable limit for this dimension is determined, at which a first scalar product of the input variable and the first product has a value that is as large as possible; and that for each dimension of the output side a second product determinable by multiplying the transpose of the matrix by a unit vector of this dimension is defined, and that an upper input variable limit for this dimension is determined, at which a second scalar product of the input variable and the second product has a value that is as large as possible.
[0015] It may be assumed that the artificial neural network expects a bias for the output, in which case the lower and upper input variable limits are modified relative to the bias.
[0016] It may be assumed that the linear layer of the artificial neural network without an activation function is the input layer of the artificial neural network, or that several linear layers without an activation function are arranged between the input side of the artificial neural network and the output side of the linear layer of the artificial neural network without an activation function.
[0017] It may be assumed that the input variables have been or will be determined in relation to the measured signals, or that the input variables have been or will be selected from a range of values.
[0018] If the smallest possible value or the largest possible value is outside or on the border of a set interval of allowed values, it may be assumed that a value having the largest possible distance from this interval between the smallest possible value and the largest possible value is determined, and the network is trained to reduce this distance.
[0019] The apparatus includes a computing device configured to perform the method.
[0020] The device may include at least one memory for input variables, the input variables representing sensor data or variables measurable by sensors that may predict the sensor data, the output device configured to output a signal, and the computing device configured to determine the signal in relation to the output variables to which the artificial neural network maps the input variables. For the sensor data, for example, a drive control variable or a pre-processed sensor signal is output. For measurable, particularly physical, variables, for example, a sensor signal is output instead of a sensor for another, particularly physical, variable.
[0021] The device may include an input device configured to communicate with the sensor to detect sensor data, and at least one computing device configured to determine input variables in relation to the sensor data.
[0022] The computer program includes computer readable instructions that, when executed, cause the computer to perform the method.
[0023] Further advantageous embodiments will become apparent from the following description and the drawings. [Brief explanation of the drawings]
[0024] [Figure 1] FIG. 1 shows a schematic diagram of an apparatus for training an artificial neural network and / or testing the robustness of an artificial neural network. [Figure 2] FIG. 1 shows a flowchart of a method for training an artificial neural network and / or testing the robustness of an artificial neural network. DETAILED DESCRIPTION OF THE INVENTION
[0025] 1, a device 100 is shown schematically. The device 100 includes at least one computing device, such as at least one processor 102, and at least one memory 104.
[0026] At least one memory 104, in this example, contains a computer program including computer-readable instructions that, when executed by at least one processor 102, perform the methods described below.
[0027] At least one memory 104, in this example, contains an artificial neural network, e.g., the memory stores parameters defining the weights of the individual layers of the network and / or hyperparameters defining the network's architecture and / or activation function.
[0028] The apparatus 100 includes, in this example, an input device 106 configured to receive sensor data from a sensor 108. The input device 106 may be configured to select a value from a range of values that represents a value representative of the sensor data from the sensor 108. It may be assumed that at least one processor 102 is configured to select this value from the range of values. The sensor 108 and the input device 106 are connected, in this example, via a line 110 for communicating the sensor data. The apparatus includes, in this example, a data link 112 that connects the at least one processor 102, the at least one memory 104, and the input device 106 for data transmission.
[0029] In this example, the input variables x∈R of the artificial neural network are used for training or for inference by the artificial neural network on these values. nIt is assumed that the input variable x is determined. It may be assumed that the input variable x is determined from training data or test data including a plurality of sensor data previously measured by the sensors 108. It may be assumed that the training data or test data includes values for the input variable x that replicate the sensor data measurable by the sensors 108. It may be assumed that the training data or test data includes values for the input variable x that can be used by the generative artificial neural network to generate the sensor data measurable by the sensors 108. The sensor data, the input variable x, the training data or the test data may be stored in at least one memory 104.
[0030] The sensor data and / or input variables x may represent a digital image, in particular a video image, a radar image, a LiDAR image, an ultrasound image or an infrared image.
[0031] The artificial neural network may be configured to classify the sensor data, particularly to identify objects in the sensor data or for semantic segmentation.
[0032] The artificial neural network may be configured to provide regression of the sensor data, particularly for identifying objects in the sensor data or for semantic segmentation.
[0033] The artificial neural network is configured to output an output variable y, which represents a signal for driving and controlling a physical system.
[0034] The apparatus 100, in this example, includes an output device 114 configured to output a signal for driving and controlling a physical system in relation to the output variable y. The output device 114, in this example, is connected to at least one processor 102 via a data link 112.
[0035] The output device 114 is configured to drive and control a computer-controlled machine 116, for example, in relation to the output variable y. The output device 114 and the computer-controlled machine 116 are connected in this example via a control line 118 for signal transmission. The computer-controlled machine 116 is, for example, a robot, a vehicle, a household appliance, a driving tool, a manufacturing machine, a personal assistance system, an access control system, a surveillance system, or a medical imaging system, among others.
[0036] Additionally or alternatively, it may be envisaged that the output device 114 is configured to drive and control an information transmission system in relation to the output variable y.
[0037] For autonomous driving or other applications, the output variable y may be envisaged to classify the sensor data, in particular to identify objects in the sensor data or for semantic segmentation.
[0038] For autonomous driving or other applications, it may be envisioned that the output variable y provides a regression of the sensor data, in particular for identifying objects in the sensor data or for semantic segmentation.
[0039] Objects in this context preferably represent markings or objects on or in the road surface, for example road signs, pedestrians or vehicles.
[0040] The artificial neural network in this example includes a first layer, hereinafter referred to as the input layer, and a last layer, hereinafter referred to as the output layer.
[0041] Between the input layer and the output layer, in this example, at least one hidden layer is located in the network. The input layer is a layer without an activation function in this example. The output layer is a layer with an activation function in this example. It is assumed that the hidden layers in the network may be formed with or without an activation function. A layer without an activation function is also referred to as an affine layer in the following.
[0042] The weights of layer j are, in this example, the matrix A j is stored in
[0043] The matrix A for the input layer maps the input variable x of the input layer to the output side of the input layer. The input layer has a number of dimensions, n, in this example.
[0044] The weights of adjacent affine layers may be combined into a single matrix A. In this case, the output is the output of the last layer of the combined layers. In the example of two adjacent affine layers A1 and A2, these affine layers A1 and A2 are combined into a matrix A = A2 · A1.
[0045] A method for training an artificial neural network and / or testing the robustness of an artificial neural network is described below, the steps of which are illustrated diagrammatically in FIG.
[0046] The artificial neural network has an output variable y∈R k is configured to determine
[0047] In what follows, two cases are distinguished. Case 1: y=f w (x) Case 2: y=f A,b,w (x)
[0048] In Case 1, a linear layer of the artificial neural network without an activation function is the input layer of the artificial neural network.
[0049] In case 2, multiple linear layers without activation functions are placed between the input side of the artificial neural network and the output side of the linear layer of the artificial neural network without activation functions. In case 2, at least the input layer is affine. This is because the artificial neural network
number
number
[0050] In both cases, the disturbance variable model T(x) ⊂ R n is assumed, which means that the input variable x is a perturbed input variable
number
[0051] For the disturbance variable model T(x), in this example, the function
number
number
[0052] For the case of \(1 < p < \infty\),
Number
Number
[0053] When \(1\leq i\leq N\),
Number
Number
Number
[0054] L p For the disturbance variable model with the norm, the upper - limit input variable limit is
Number
Number
[0055] In contrast, according to this method, it is possible to provide a solution without determining the maximization part \(T1(x)\).
[0056] This method includes step 202.
[0057] In step 202, the input variable \(x\in R\) having a plurality of \(n\) dimensions nIn step 202, in this example, a disturbance variable model T(x) is set. In step 202, in this example, a real-valued function l:R k In step 202, an interval I ⊂ R of values is set that, in this example, results in a certified artificial neural network.
[0058] It may be assumed that the input variable x has been determined in relation to a measured signal or that the input variable x has been selected from a range of values.
[0059] Then, step 204 is executed.
[0060] Step 204 includes substeps that, in case 1, are performed for each dimension 1≦i≦n of the input variable x. In case 2, these substeps are performed for each dimension 1≦i≦N of the output side of the linear layer of the artificial neural network without an activation function, onto which the input variable x is mapped by the artificial neural network.
[0061] Substep 1: Upper input variable limit u i Decision In this example, the upper input variable limit u i is determined, and this upper input variable limit u i In this case, the disturbance variable model T(x) ⊂ R n has as large a value as possible in dimension i.
[0062] Substep 2: Lower input variable bounds i Decision In this example, the lower input variable limit l i is determined, and this lower bound on the input variable l i In this case, the disturbance variable model T(x) ⊂ R n has the smallest possible value in dimension i.
[0063] Case 1: In this method, for each dimension 1≦i≦n of the input variable x, the lower input variable limit l of this dimension i is i It can be assumed that the lower bound of the input variable l i In this case, the input variable x and the negative unit vector -e of this dimension i i The first scalar product with 〈-e i ;x〉 has the largest possible value. Preferably, these limits are determined for output variables l≦y≦u, where l,u∈R k In this example, for 1≦i≦n, the lower input variable limit l i =-opt(-e i ;x) is determined.
[0064] In this method, for each dimension 1≦i≦n of the input variable x, the upper limit of the input variable u of this dimension i is i It may be assumed that the upper input variable limit u i In this case, the input variable x and the unit vector e of this dimension i are i The second scalar product with i ;x〉 has as large a value as possible. Preferably, these limits are set by the output variable y=f w (x). Preferably, these bounds are determined for output variables l≦y≦u, where l,u∈R. k In this example, for 1≦i≦n, the upper input variable limit u i =opt(e i ;x) is determined.
[0065] It may be assumed that the artificial neural network has a bias b for the input variable x, and in this method, a lower input variable limit l i and the upper input variable limit u i is assumed to be corrected relative to the bias b.
[0066] Case 2: The output of a linear layer of an artificial neural network without an activation function is in this example defined by the matrix A of the affine layer, or is integrated by the matrix A into successive affine layers, starting from the first layer of the artificial neural network.
[0067] For each output dimension 1≦i≦N defined by this matrix A, a first product is defined, which is the negative translation −A of the matrix A. T and the unit vector e of this dimension i i It can be determined by multiplication with
[0068] In this method, for each dimension i, a lower input variable limit l i can be assumed to be determined, and this lower bound on the input variable l i In the first product, the first scalar product 〈-A T e i ;x〉 has the largest possible value.
[0069] For each output dimension 1≦i≦N defined by this matrix A, a second product is defined, which is the transformation A of matrix A. T and the unit vector e of this dimension i i It can be determined by multiplication with
[0070] In this method, the upper input variable limit u of this dimension i i can be assumed to be determined, and this upper input variable limit u i In the second scalar product of the input variable x and the second product, T e i ;x〉 has the largest possible value.
[0071] These limits are, in this example, the output variable y=f A,b,w (x). Preferably, these bounds are determined for output variables l≦y≦u, where l,u∈R. k is.
[0072] It may be assumed that the artificial neural network has a bias b for the output. In this case, the method is based on a lower input limit l i and the upper input variable limit u i is modified relative to the bias b. If the artificial neural network has two successive affine layers, for example, then the matrix A = A2 · A1 and the bias b = A2b1 + b2.
[0073] Then, step 206 is executed.
[0074] In step 206, for each dimension 1≦j≦k of the output variable y, a lower output variable limit is calculated for the output variable y.
number
number
[0075] In this example, the upper output variable limit u y and the lower output variable limit l y are the upper input variable limits u1,…,u n and lower input variable limits l1,…,l n The value is determined by a value from a range of values bounded by
[0076] In the following, as an example, in case 1, the upper output variable limit u y and the lower output variable limit l y Three methods for determining this are described.
[0077] Method 1: Element x of input variable x i For each, the artificial neural network f w (x) is assumed to contain an activation layer φ:R→R, and the element x i One upper bound for each input variable x i ≦u iis stipulated.
[0078] In this example,
number
number
number
number
number
number
[0079] Method 2: Artificial neural network f for input variables x w (x) is assumed to contain a scalar product 〈c,x〉+β with a bias β, and the element x i One upper bound for each input variable x i ≦u i and one lower input variable limit l i ≦x i It is stipulated that:
[0080] In this example, the vector c + and c - is determined, where:
number
number
[0081] upper output variable limit u y and the lower output variable limit l y is the artificial neural network f w It is determined by the mapping resulting from the scalar product of the remainder of (x) and the bias, 〈c,x〉+β.
[0082] Method 3: Artificial neural network f for input variables x w It is assumed that (x) = A·x+b contains a linear layer A with bias b.
[0083] In this example, the element y of the output variable y=A·x+b j For each scalar product with column j, a scalar product is determined from matrix A, where for each scalar product, the calculation is performed as described in method 2.
[0084] If no bias is set, the calculation is performed with a bias b=0 in this example.
[0085] In Case 2, the calculation is performed as described above for Case 1, where the input variable x is replaced by the output defined by matrix A.
[0086] Then, step 208 is executed.
[0087] In step 208, the real-valued function l:R k →The smallest possible value of R
number
[0088] Then, step 210 is executed.
[0089] In step 210, in particular, a real-valued function l:R k →The largest possible value of R
number
[0090] Then, step 212 is executed.
[0091] In step 212, the smallest possible value of y - and the largest possible value of y + If and are within a set interval I ⊂ R of allowed values, the output is determined, which proves the robustness of the network.
[0092] Optionally, training may be assumed, and in step 214, the smallest possible value of y - Or the largest possible value y + is checked whether it is outside or on the boundary of a set interval I⊂R of allowed values.
[0093] The smallest possible value of y - Or the largest possible value y + If lies outside or on the boundary of a set interval I⊂R of allowed values, step 216 is executed. Otherwise, training ends in this example.
[0094] Optionally, it may be envisaged that after training, input variables are detected and mapped to output variables, where in association with each output variable, signals or sensor data for driving and controlling a physical system are specifically determined and / or output.
[0095] In step 216, the smallest possible value of y - and the largest possible value of y + A value is determined from the interval between
number
[0096] Then, step 218 is executed.
[0097] In step 218, the network is trained to reduce this interval.
[0098] Then, step 202 is executed.
Claims
1. 1. A method for training an artificial neural network and / or testing the robustness of an artificial neural network, the artificial neural network being configured to determine output variables, the method comprising: The method includes setting input variables for the network (202) having multiple dimensions; The method comprises determining (204) an upper input variable bound and a lower input variable bound for each dimension of the input variables or for each dimension of an output of a linear layer of the artificial neural network without an activation function to which the input variables are mapped by the artificial neural network, wherein at the upper input variable bound a disturbance variable model that can map the input variables to disturbed input variables has as large a value as possible in the dimension, and at the lower input variable bound the disturbance variable model has as small a value as possible in the dimension; The method includes, for each dimension of the output variable, determining (206) a lower output variable limit for the output variable with a value from a range of values bounded by the lower input variable limit and the upper input variable limit, and determining (206) an upper output variable limit for the output variable with the value from the range of values bounded by the upper input variable limit and the lower input variable limit; The method comprises determining (208) the smallest possible value of the function, in particular of real values, from a range of values bounded by the lower output variable limit and the upper output variable limit, the method comprises determining (210) the largest possible value of the in particular real-valued function by a value from the range of values bounded by the lower output variable limit and the upper output variable limit, determining (212) an output that proves the robustness of the network if the smallest possible value and the largest possible value are within a set interval of acceptable values; A method characterized by:
2. The method of claim 1 , wherein the input variables represent a digital image, in particular a video image, a radar image, a LiDAR image, an ultrasound image or an infrared image.
3. 2. The method according to claim 1, wherein the output variables represent signals for the drive control of a physical system, in particular a computer-controlled machine or an information transmission system, preferably a robot, a vehicle, a domestic device, a drive tool of a manufacturing machine, a personal assistance system, an access control system, a surveillance system, or in particular a medical imaging system.
4. 2. The method of claim 1, wherein the output variables classify the sensor data, in particular for identifying objects in the sensor data or for semantic segmentation, or the output variables provide a regression of the sensor data, in particular for identifying objects in the sensor data or for semantic segmentation, the objects being preferably markings or objects on or representing markings or objects of the road surface, in particular objects representing road signs, pedestrians or vehicles.
5. determining, for each dimension of the input variables, the lower input variable limit for the dimension, wherein at the lower input variable limit, a first scalar product of the input variable and a negative unit vector for the dimension has as large a value as possible; 2. The method of claim 1, further comprising: determining, for each dimension of the input variables, the upper input variable limit for the dimension, at which a second scalar product of the input variable and a unit vector for the dimension has as large a value as possible.
6. the output of the linear layer of the artificial neural network without an activation function is defined by a matrix; defining, for each dimension of the output side, a first product determinable by multiplying the negative transpose of the matrix by a unit vector of the dimension; The method contemplates determining the lower input variable limit for the dimension, where a first scalar product of the input variable and the first product at the lower input variable limit has a value that is as large as possible; defining, for each dimension of the output side, a second product determinable by multiplying the transpose of the matrix by the unit vector of the dimension; 2. The method of claim 1, wherein the method further comprises determining the upper input variable limit for the dimension, wherein at the upper input variable limit a second scalar product of the input variable and the second product has a value that is as large as possible.
7. 7. The method of claim 5 or 6, wherein the artificial neural network predetermines a bias for the input variable or the output, and modifies the lower input variable limit and the upper input variable limit in relation to the bias.
8. 7. The method of claim 6, wherein the linear layer of the artificial neural network without an activation function is an input layer of the artificial neural network, or wherein a plurality of linear layers without an activation function are arranged between the input of the artificial neural network and the output of the linear layer of the artificial neural network without an activation function.
9. 2. The method of claim 1, wherein the input variables have been or are determined in relation to a measured signal, or the input variables have been or are selected from a range of values.
10. If the smallest possible value or the largest possible value is outside or on the border of the set interval of allowed values (214), determine from the interval between the smallest possible value and the largest possible value a value that has the largest possible distance relative to the interval (216); The method of claim 1 , further comprising training the network to reduce the spacing (218).
11. 10. An apparatus (100) comprising at least one computing device (102) configured to perform the method of claim 1.
12. The device (100) includes at least one memory (104) for the input variables and an output device (114); the input variables represent sensor data or represent variables measurable by sensors that may predict sensor data; 12. The apparatus of claim 11, wherein the output device is configured to output a signal, and the computing device is configured to determine the signal in relation to the output variables to which the artificial neural network maps the input variables.
13. The apparatus (100) includes an input device (106), the input device (106) configured to communicate with a sensor (108) to detect sensor data; The apparatus (100) of claim 11, wherein the at least one computing device (102) is configured to determine the input variable (x) in relation to the sensor data.
14. 10. A computer program comprising computer readable instructions that, when executed, cause a computer to perform the method of claim 1.