Method and system for validating a trained artificial neural network (ANN) based on a test data set
The method addresses the challenge of high redundancy in ANNs by validating them through cell-based partitioning and simulation-driven data completion, enhancing the ANN's performance and reliability.
Patent Information
- Application Number
- DE102023211711
- Authority / Receiving Office
- DE · DE
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2023-11-23
- Publication Date
- 2025-05-28
AI Technical Summary
Trained artificial neural networks (ANNs) often have high redundancy in their weights, leading to increased computing power requirements for both execution and training, while also necessitating a method to validate the quality and performance of ANNs.
A computer-implemented method for validating a trained ANN by partitioning the input space into cells using the network architecture, checking for data points in each cell, and generating new data points via a simulation model if necessary to complete the test dataset.
This method effectively validates the ANN by ensuring data points cover all cells, allowing for the determination of ANN quality and performance, and iteratively improving the ANN's performance through retraining.
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
[0001] The invention relates to a method and system for validating a trained artificial neural network (ANN) based on a test data set. State of the art
[0002] Trained artificial neural networks (ANNs) typically exhibit a high degree of redundancy in their weights. This means that many weights are not necessarily necessary to enable the neural network to function efficiently, especially in the case of inference. Since the number of weights strongly influences the complexity of the network architecture underlying the artificial neural network, and as the number of weights increases, so does the computing power required to run the artificial neural network, as well as to train it. Therefore, there is a need to keep the number of weights per network layer as low as possible while still ensuring the ANN's high performance.
[0003] In order to enable an evaluation of the quality and / or performance of ANNs, validation procedures are desirable. Subject of the invention
[0004] An object of the invention is to provide a method and a system for validating a trained artificial neural network (ANN) on the basis of a test data set.
[0005] The object is achieved by a computer-implemented method for validating a trained artificial neural network (ANN) on the basis of a test data set according to the features of patent claim 1. The object is achieved by a system for validating a trained artificial neural network (ANN) on the basis of a test data set according to the features of patent claim 10. Disclosure of the invention
[0006] This invention provides a method for validating a trained artificial neural network (ANN) based on a test data set; the method comprises the steps: - Providing the trained ANN with a network architecture representing a, in particular non-linear, chained, function by which a d-dimensional input space can be mapped to an N-dimensional output space, and further comprising a plurality of weights; - Providing the test data set; - Partitioning the d-dimensional input space into a plurality of cells based on the network architecture of the ANN, wherein the individual cells are each separable from one another by at least one weight-specific hyperplane and / or hyperset; - Check whether at least one data point of the test data set is present in each of the cells in order to validate the ANN; and - if no data point is present in at least one of the cells, generating (S5) at least one new data point by a simulation model based on cell parameters of the at least one of the cells in order to complete the test data set.
[0007] The invention generates a decomposition or partitioning of the input space into cells. In particular, the predicted class is constant in each cell. The invention checks whether a data point is present in each cell. If this is the case, the algorithm or the validation method preferably ends. If this is not the case, the simulation model is given the parameters of the cells without a data point. The simulation model then preferably generates at least one new data point in the relevant cell(s). If necessary, the ANN is iteratively retrained with the newly generated data or data points. Preferably, new test data is generated, on the basis of which the retrained ANN can be evaluated or validated again.
[0008] It is understood that the steps according to the invention, as well as other optional steps, do not necessarily have to be performed in the order shown, but can also be performed in a different order. Furthermore, additional intermediate steps can be provided. The individual steps can also comprise one or more substeps without thereby departing from the scope of the method according to the invention.
[0009] In other words, the present method decomposes or partitions the d-dimensional input space of the ANN into a finite number of cells (“faces”). In or within each of these cells, the output of the ANN is preferably constant if the function of the ANN is to solve a classification problem. In or within each of these cells, the output of the ANN is preferably linear if the function of the ANN is to solve a regression problem. The method pruns the ANN.
[0010] A hyperplane is a mathematical term defined in a multidimensional space. In a two-dimensional space, a hyperplane is a straight line; in a three-dimensional space, it is a plane. Generally speaking, in an n-dimensional space, a hyperplane is an (n-1)-dimensional surface that divides the space into two parts. In ANNs, hyperplanes are used to define classes or decision boundaries by dividing the space into regions corresponding to different categories or states.
[0011] A hyperset is an extension of the concept of a hyperplane. Rather than a single hyperplane, a hyperset is a collection of hyperplanes in space. These hyperplanes, taken together, can be used to define complex spatial structures. Hypersets are used in ANNs to model more complex decision boundaries that cannot be easily described by individual hyperplanes.
[0012] The input space (also called feature space) refers to the space in which the input data of a neural network exists. Each dimension of the input space corresponds to a specific feature or property of the input data. For example, the pixel values of an image in an image classification network could constitute the input space.
[0013] The output space refers to the space of output values of a neural network. This space depends on the type of network. In a multi-class classification network, the output space might contain vectors representing the probabilities for each class. In a regression neural network, the output space might represent numerical values.
[0014] In a preferred embodiment, the ANN has a d-dimensional number of input variables. Such input variables can include sensor signals, for example. The input variables can be pre-processed, for example normalized and / or transformed from a time domain to a frequency domain or from a frequency domain to a time domain and / or filtered and / or smoothed or similar. The term “d-dimensional” means that preferably d different signals are available as input variables. This spans an input space of dimension d. The network architecture further comprises a plurality of bias terms. In an ANN, the bias term is preferably a constant that is added to the weighted inputs of each neuron before the preferably non-linear activation function is applied. It can be interpreted as a type of shift that influences the activation of the neuron.The bias helps control the activation range and the network's overall ability to adapt to the data. The bias term allows a neuron to be activated even if all weighted inputs are zero. This allows the network to respond more flexibly to different patterns and relationships in the data. In a multilayer neural network, each neuron typically has its own bias value. Partitioning the d-dimensional input space into a plurality of cells based on the network architecture comprises at least the following steps: determining hyperplanes and / or hypersets depending on the input variables, at least a subset of the plurality of weights, and a subset of the plurality of bias terms.How exactly the partitioning of the input space as well as the subsequent spaces of the network layers is carried out is described in detail in the figure description, to which explicit reference is made here.
[0015] In a preferred embodiment, if at least one data point of the test data set is present in each of the cells, a validation result of the ANN is provided, on the basis of which a quality and / or goodness and / or performance of the ANN can be determined and / or evaluated. Based on the validation result, it can preferably be decided whether the ANN can be used in the inference, for example by comparing the validation result with a predetermined threshold value. This increases the performance and reliability of the ANN. In a preferred embodiment, the ANN is retrained based on the completed test data set. The retraining sustainably improves the performance of the ANN. This results in a reduced error rate in the inference of the ANN.
[0016] In a preferred embodiment, an expanded and / or new test dataset is provided for the retrained ANN, on the basis of which a renewed validation of the retrained ANN is performed. In this way, the performance of the ANN can be continuously and / or iteratively improved.
[0017] In a preferred embodiment, the retraining is carried out iteratively.
[0018] In a preferred embodiment, the following applies to the d-dimensional input space: d = 10, preferably d < 10, particularly preferably d < 6, where d is an element of the positive natural number set. Therefore, the ANN is preferably a small-dimensional network. For ANNs of higher dimensions, partitioning would only be possible with a significantly higher computational effort due to the large number of input variables.
[0019] In a preferred embodiment, the ANN is part of a complex machine learning model. This allows a part or component of the complex model to be optimized. This leads to an improvement of the overall model.
[0020] In a preferred embodiment, the simulation model is designed to generate at least one data point, in particular through data augmentation, based on the cell parameters, in particular based on the hyperplane and / or hyperset information. The simulation model can, for example, be a data augmentation model.
[0021] In a second aspect, a system for validating a trained artificial neural network (ANN) based on a test data set is specified. The system comprises a provision device configured to perform the following steps: providing the trained ANN with a network architecture representing a, in particular nonlinear, concatenated function by which a d-dimensional input space can be mapped to an N-dimensional output space, and further comprising a plurality of weights; providing the test data set; and an evaluation and / or computing device configured to perform the following steps: partitioning the d-dimensional input space into a plurality of cells based on the network architecture of the ANN, wherein the individual cells are each separable from one another by at least one weight-specific hyperplane and / or hyperset;Checking whether at least one data point of the test data set is present in each of the cells in order to validate the ANN; and if no data point is present in at least one of the cells, generating at least one new data point by a simulation model based on cell parameters of the at least one of the cells in order to complete the test data set.
[0022] The embodiments mentioned for the method apply in the same or similar way to the system.
[0023] The present invention also claims a control and / or computing device that is configured to perform a regression task and / or a classification task by executing the trained, validated, and / or optimized ANN provided according to the present method. The control and / or computing device can be included, purely by way of example, in an autonomous vehicle and / or a robotics system and / or an industrial machine. Such a control and / or computing device can be used, for example, as an embedded device in a plant or system. In general, the present method can be used particularly for small-dimensional artificial neural networks. For example, the present validated and / or optimized ANN can be used as a virtual sensor with up to ten, preferably up to 5-dimensional inputs.The ANN provided here can also be used for feature detection in larger ANNs, for example, for autonomous edge detection. Other examples of virtual sensors in vehicles include virtual temperature sensors (e.g., on the stator / rotor of an electric motor or on an injector magnet), mass estimators for an injection system, coking models, and / or virtual sensors for exhaust gas concentrations. In principle, the present method and the ANN provided by it are applicable to all virtual sensors used in situations where assembly is impossible, where cost pressure is high, or where the measurement setup is complex for series production.
[0024] Furthermore, the presently validated and / or optimized ANN can be used to determine a correction factor for calculating the current consumption of an engine in a vehicle. The present ANN offers an alternative to known methods that require "running in an injector" and may not always work. This behavior can be parameterized in a characteristic map. Furthermore, increased EU requirements necessitate the use of alternative methods. The compactness and low complexity of the presently validated and / or optimized ANN provide such an alternative.
[0025] Furthermore, the presently validated and / or optimized ANN can be used to detect a change point from a discrete gyroscope signal, for example, for fall detection. It is advantageous that the input vectors are small-dimensional, thus allowing the present optimization method to be optimally applied to provide an optimized ANN.
[0026] In addition, the ANN provided here can be used as an edge filter in image processing. This can be illustrated using an excavator bucket as an example, where the ANN can be used as part of the excavator bucket's control loop. The transfer function of the excavator joystick to the bucket's cylinder position can be modeled using such an optimized ANN. Small-dimensional networks with a simple network architecture, for example, of size 3-15-15-1 and / or 2-15-7-1, are preferably considered. With such small ANNs, the decompositions can be determined in less than 5 seconds, for example.
[0027] In particular, due to its low dimensionality, the present ANN can be used as a surrogate function in which the ANN represents a solution of a partial differential equation, which cannot be computed independently in real time.
[0028] The invention also claims a computer program with program code for executing at least parts of the method according to the invention in one of its embodiments when the computer program is executed on a computer. In other words, a computer program (product) comprising instructions that, when executed by a computer, cause the computer to execute the method / steps of the method according to the invention in one of its embodiments.
[0029] The invention also proposes a computer-readable data carrier with program code of a computer program for executing at least parts of the method according to the invention in one of its embodiments when the computer program is executed on a computer. In other words, the invention relates to a computer-readable (storage) medium comprising instructions which, when executed by a computer, cause the computer to execute the method / steps of the method according to the invention in one of its embodiments.
[0030] The described designs and further training courses can be combined as desired.
[0031] Further possible embodiments, developments and implementations of the invention also include combinations of features of the invention described previously or below with regard to the exemplary embodiments that are not explicitly mentioned. Short description of the drawings
[0032] The accompanying drawings are intended to provide a further understanding of embodiments of the invention. They illustrate embodiments and, in conjunction with the description, serve to explain principles and concepts of the invention.
[0033] Other embodiments and many of the aforementioned advantages will become apparent upon review of the drawings. The elements illustrated in the drawings are not necessarily drawn to scale.
[0034] They show: Fig. 1 exemplary hyperplanes for partitioning an input space; Fig. 2 the representation from Fig. 1 with additional coding information; Fig. 3 an exemplary representation of a partitioning with two-dimensional open sets; Fig. 4 an exemplary representation of open 1-dimensional sets; Fig. 5 an exemplary representation of three 0-dimensional sets; Fig. 6 an exemplary representation of a partitioning of a 10-dimensional input space of a first layer of an ANN; Fig. 7 an exemplary representation of a partitioning of a 10-dimensional input space of a second layer of an ANN; Fig. 8 is an exemplary flowchart of an embodiment of the present method; Fig. 9 an exemplary representation of an input space of the first layer of an ANN; Fig. 10 an exemplary representation of an input space of the first layer of an ANN; Fig. 11 an exemplary representation of an input space of the second or a further layer of an ANN; Fig. 12 an exemplary representation of an input space of the second or a further layer of an ANN.
[0035] In the figures of the drawings, the same reference symbols designate the same or functionally identical elements, parts or components, unless otherwise stated.
[0036] The partitioning of the input space in terms of its mathematical foundations with reference to the Fig. 1 to 7. The present invention will be described with reference to the Fig. 8 to 12 explained.
[0037] The partitioning according to the present method preferably considers a multi-layer neural network (KNN) with ReLu nonlinearities as the activation function. This means that the KNN has a function F:Rm1→Rmk+1 defined.
[0038] The following applies: F(x)=Akϕ(Ak−1ϕ(Ak−2...ϕ(A2(ϕ(A1x+b1))+b2)+...)+bk with weight matrices A 1 , ... , A k , and bias terms b 1 , ... , b k . Here, A 1Matrices ie linear mappings between real vector spaces A 1 : R m1 → R m2 , A2:Rm2→Rm3,...,Ak:Rmk→Rmk+1 Likewise, b 1 ∈ R m2 , ..., bk∈Rmk+1 Vectors of a real vector space.
[0039] Overall, F is preferably a concatenated map F:Rm1→Rmk+1 of the m1-dimensional Euclidean space R m1 into the mk+1-dimensional Euclidean space Rmk+1.
[0040] To reduce the notation, d = m 1 and N = m k+1 set, that is, the network F: R examined here d → R N is a function of the d-dimensional Euclidean space R d into the N-dimensional Euclidean space R N .
[0041] When modeling a regression, the description of the function F(x) is preferably complete. The ANN then preferably returns the value F(x) as the output value or prediction value.
[0042] In case of a classification of the input signal x in m k+1 -classes, the output F(x) is preferably followed by a softmax function. This results in the components of F(x) = (f 1 (x), f 2 (x), ... ,f mk+1 (x)) is preferably positive, and normalized to one, ie, for all input variables x, the following applies: f 1 (x), ... ,f mk+1 (x) ≥ 0 for all x and sum f 1 (x) + ... + f mk+1 (x) = 1.
[0043] The classification is preferably carried out with an Argmax function that takes the index of the vector g 1 (x), ..., g mk+1 (x) where g(x) is maximal.
[0044] After explaining the basics, the partitioning of the input space will be described in more detail below. The layers of the ANN preferably partition the input space into cells. In the case of a classification network, it is preferable to distinguish between the first layer (input layer), middle layers (also subsequent layers or layers following the input layer), and the last layer (also output layer).
[0045] In the first layer of the network, consisting of a pair of weight matrix A 1 and bias terms b 1 , and a ReLu nonlinearity ϕ (activation function), an input vector x with the weight matrix A 1 multiplied, the bias term b 1 is added, the resulting vector A 1 X + b 1 a component-wise ReLu ϕ(A 1 x + b 1 ). At the component level, this means: y1=ϕ(A1x+b1)=ϕ([a111a121..a1d1a211a221..a2d1……..…am11am21..amd1][x1x2…xd]+[b11b21…bm1])=[ϕ(a111x1+a121x2+⋯+a1d1xd+b11)ϕ(a211x1+a221x2+⋯+a2d1xd+b21)…ϕ(am11x1+am21x2+⋯+amd1xd+bm1)]
[0046] For ease of reading we have d = m 1 and m = m 2 set.
[0047] In the following, the column vectors of A 1 with the vectors h 1 , ... , hm included. h1=[a111a121 …a1d1],...,hm=[am11am21 ⋯amd1]
[0048] These vectors and the bias term preferably each define open half-spaces. H1−:={x∈Rm1|〈h1,x〉+b11<0},...,Hm−:={x∈Rm1|〈hm,x〉+bm1<0} H1+:={x∈Rm1|〈h1,x〉+b11>0},...,Hm+:={x∈Rm1|〈hm,x〉+bm1>0}
[0049] Furthermore, hyperplanes are introduced, preferably for the purpose of partitioning. H10:={x∈Rm1|〈h1,x〉+b11=0},...,Hm0:={x∈Rm1|〈hm,x〉+bm1=0}
[0050] This system of half-spaces and hyperplanes partitions the input space R n preferably in cells. This connection is described in the Fig. 1 and Fig. 2 is shown as an example. Fig. 1 and Fig. 2 three hyperplanes H 0 , H 1 , H 2 generated according to the above equations. In Fig. 2 the respective hyperplanes H 0 , H 1 , H 2 together with their open half-spaces. Furthermore, Fig. 2 shows three exemplary points (vertices) with their respective coding.
[0051] Fig. 3 shows the different cells defined by the hyperplanes H 0 , H 1 , H 2 are generated. Fig. 4 shows a total of 7 cells, each defining a two-dimensional open set.
[0052] In Fig. 4 shows one-dimensional cells that define one-dimensional relative-open sets.
[0053] In Fig. 5 shows 0-cells that define zero-dimensional sets, i.e., vertices or points.
[0054] In other words, each element x ∈ R m1 a coding code(x) ∈ {-1,0,1}min is assigned such that a -1 is inserted at the k-th position in case that 〈h 1,x 〉 +b 1 1 < 0, a 0 if (h 1 ,x) +b 1 1 = 0 and a +1 if h 1 ,x 〉 +b 1 1 > 0. In this way, preferably each element of the input space is uniquely defined by a code.
[0055] The hyperplanes H 1 0 , H 2 0 ,..., H m 0The first layer generates n-dimensional cells (or n-faces). The n-cells are the connected components of the complement of the union of the hyperplanes. Referring to Fig. 3, this means that the three lines represent the three hyperplanes, where a union of these lines defines a closed set (preferably as sets of points). The complement of this closed set is preferably an open set that divides the input space into connected components (the blue surfaces F 1 2 ,..., F 7 2 .
[0056] The (relatively open connected components) of the edge of the n-faces are the (n-1) cells (see Fig. 3 and Fig. 4). The triangle F 4 2preferably has a boundary consisting of three sides, each of which is preferably defined by bounding hyperplanes or hypersets. These boundaries are preferably the three 1-cells.
[0057] Recursively, the k-cells are defined as the boundary of k+1-cells until 0-cells are formed at the end, which are preferably points of the input space. Connection to neural networks:
[0058] In each n-face, the output of the first layer of the neural ReLu-ANN can be considered as a linear mapping. From k-layer to (k+1)-layer
[0059] Each subsequent layer of a ReLu ANN preferably further partitions the n-cells generated by the previous layers. Each n-cell should preferably be considered individually. Initial layer in a regression:
[0060] In the case of regression, the partitioning process is complete. The last layer, or output layer, is preferably a linear mapping of the outputs of the penultimate layer. Preferably, no further partitioning of the input space occurs. Initial layer in a classification:
[0061] In the previous steps, a partitioning P k of the input space into n-cells (in the input layer), (n-1)-cells (in the next layer), etc. The n-cells are also partitioned in the last layer, the output layer. Preferably, all n-cells are traversed, and the following steps can be performed: 1. A matrix-vector pair L, l is determined that describes a linear relationship in an n-cell. This is done, for example, by: a. An interior point x ̌ of the n-cell is determined, e.g. by forming a common center of the vertices of the n-cells. b. For each layer, the active features are preferably determined. Furthermore, the linear mapping and bias term are preferably determined interactively by: i. it will be A 1 x + b 1 formed. ii. The rows are determined where the vector A 1 x + b 1 is negative. iii. A new matrix-vector pair A1˜,b1˜ formed, whereby A1˜ the matrix A 1 corresponds, preferably with the only difference that at the rows of the vector A 1 x + b 1 is negative, and the rows of A1˜ be set to zero. In this way, for the selected point x ϕ(A1x⌣+b1)=A1˜x⌣+b1˜. iv. Now this is done analogously for the other layers of the ANN: v. It is preferable to use an A 2 (ϕ(A 1 x + b 1 )) + b 2 formed, determining the rows in which the vector A is negative 2 (ϕ(A 1 x + b 1 )) + b 2 is. vi. A new matrix-vector pair A2˜,b2˜ formed, whereby A2˜ A 2 the matrix A 2 , except that at the rows of the vector A 2 (ϕA 1 x + b 1 )) + b 2 is negative, and the rows of A2˜ be set to zero. vii. In this way, L and I can be introduced so that in the cell under investigation: Lx+l=Akϕ(Ak−1ϕ(Ak−2…ϕ(A2(ϕ(A1x+b1))+b2)+…)+bk. 2. Now a new matrix-vector pair (H, d) is determined from the matrix-vector pair L, l in the following way: a. A zero-initialized matrix H of size (0.5 * Number of classes * (Number of classes m - 1)) × d, and a zero-initialized vector d of length 0.5 * Number of classes * (Number of classes - 1)) is formed. b. All combinations of two elements of the possible output classes are traversed. i. for two classes 1,2 there is only one possible combination (1,2) ii. for three classes 1,2,3 there are three possible combinations (1,2), (1,3), (2,3) iii. For four classes 1, 2, 3, 4 there are six possible combinations (1, 2), (1, 3), (1, 4), (2, 3), (2, 4), (3, 4) etc. c. For each combination i, j, the following differences of the i-th column vector from the j-th column vector of L are formed, and analogously the difference of the i-th entry of the bias from the j-th entry of the bias. 3. The cell is now further partitioned by this matrix H and bias term b. 4. Through this construction, the class of the network determined by the Argmax function becomes constant in every cell. The matrix H and the bias term identify those points in the input space where F(x) has two components of equal size; these are the class boundaries of the neural network.
[0062] Thus, the partitioning or decomposition of the input space into cells or faces was described above, whereby the KNN is constant in each cell (i.e., each n-face) of the partitioning or decomposition.
[0063] In the Fig. 6 and Fig. Figure 7 shows further examples of the partitioning described above. The example partitionings of the input space are illustrated using a 2->10->5->4 classification ANN. Fig. Figure 6 shows a partitioning of the input space by the first layer. Ten hyperplanes are shown, which are generated by the weight matrices and bias terms. Fig. Figure 7 shows the partitioning of the input space after the second layer. Each 2-cell partitioning in Fig. 7 is further partitioned by five hyperplanes of the second layer. It is also possible, in principle, that the second layer has no (partitioning) effect in a cell.
[0064] Fig. 8 shows a flowchart of an embodiment of the present method. In any embodiment, the method can be carried out at least partially by a system 1, which for this purpose can comprise several components not shown in detail, for example one or more provision devices and / or at least one evaluation and computing device. It is understood that the provision device can be designed jointly with the evaluation and computing device or can be different from it. Furthermore, the system can comprise a memory device and / or an output device and / or a display device and / or an input device.
[0065] According to the invention, the computer-implemented method comprises at least the following steps: - Providing (S1) the trained ANN with a network architecture representing a, in particular non-linear, chained, function by which a d-dimensional input space can be mapped to an N-dimensional output space, and further comprising a plurality of weights; - Providing (S2) the test data set; - partitioning (S3) the d-dimensional input space into a plurality of cells based on the network architecture of the ANN, wherein the individual cells are each separable from one another by at least one weight-specific hyperplane and / or hyperset; - Checking (S4) whether at least one data point of the test data set is present in each of the cells in order to validate the ANN; and if no data point is present in at least one of the cells, generating (S5) at least one new data point by a simulation model based on cell parameters of the at least one of the cells in order to complete the test data set.
[0066] The Fig. 9 to 12 show the present method as an example for a partitioning of an input space by an ANN with two hidden layers, and for four output classes (network architecture: 2->10->10-4).
[0067] Fig. 9 shows an ANN with a representation of a partitioned input space. The Fig.Figures 10 to 12 show modified ANNs compared to the ANN. A penultimate layer of the ANN preferably divides the input space into 2-faces, 1-faces, etc. To distinguish the test data, it is preferable to also consider the last layer. The circles preferably represent the ANN's test data. The constant classes of the look-up table are visible. Furthermore, it is clear how these classes fit the data.
Claims
[1] A method for validating a trained artificial neural network (ANN) based on a test data set; the method comprising the steps: - Providing (S1) the trained ANN with a network architecture representing a, in particular non-linear, chained, function by which a d-dimensional input space can be mapped to an N-dimensional output space, and further comprising a plurality of weights; - Providing (S2) the test data set; - partitioning (S3) the d-dimensional input space into a plurality of cells based on the network architecture of the ANN, wherein the individual cells are each separable from one another by at least one weight-specific hyperplane and / or hyperset; - Checking (S4) whether at least one data point of the test data set is present in each of the cells in order to validate the ANN; and - if no data point is present in at least one of the cells, generating (S5) at least one new data point by a simulation model based on cell parameters of the at least one of the cells in order to complete the test data set. [2] The method of claim 1, wherein the ANN has a d-dimensional number of input variables, wherein the network architecture further comprises a plurality of bias terms, and wherein partitioning the d-dimensional input space into a plurality of cells based on the network architecture comprises at least the following steps: - Determining hyperplanes and / or hypersets depending on the input variables, at least a subset of the plurality of weights and a subset of the plurality of bias terms. [3] Method according to one of the preceding claims, wherein, if at least one data point of the test data set is present in each of the cells, a validation result of the ANN is provided, on the basis of which a quality of the ANN can be determined. [4] Method according to one of the preceding claims, wherein the ANN is retrained on the basis of the completed test data set. [5] Method according to claim 4, wherein an extended and / or new test data set is provided for the retrained ANN, on the basis of which a validation of the retrained ANN is carried out. [6] Method according to claim 4 or 5, wherein the retraining is carried out iteratively. [7] Method according to one of the preceding claims, wherein the following applies to the d-dimensional input space: d = 10, preferably d < 10, particularly preferably d < 6, where d is an element of the positive natural number set. [8] Method according to one of the preceding claims, wherein the ANN is part of a complex machine learning model [9] Method according to one of claims 1 to 8, wherein the simulation model is designed to generate at least one data point, in particular by data augmentation, on the basis of the cell parameters, in particular on the basis of the hyperplane and / or hyper-set information. [10] System (1) for validating a trained artificial neural network (ANNs) based on a test data set; the system comprising a provision device configured to perform the following steps: providing (S1) the trained ANN with a network architecture representing a, in particular non-linear, concatenated function by which a d-dimensional input space can be mapped to an N-dimensional output space, and further comprising a plurality of weights; providing (S2) the test data set; and an evaluation and / or computing device configured to perform the following steps: partitioning (S3) the d-dimensional input space into a plurality of cells based on the network architecture of the ANN, wherein the individual cells are each separable from one another by at least one weight-specific hyperplane and / or hyperset; checking (S4) whether at least one data point of the test data set is present in each of the cells in order to thus validate the ANN; and if no data point is present in at least one of the cells, generating (S5) at least one new data point by a simulation model based on cell parameters of the at least one of the cells in order to thus complete the test data set. [11] Computer program with program code to carry out at least parts of a method according to one of claims 1 to 9 when the computer program is executed on a computer. [12] Computer-readable data carrier with program code of a computer program for carrying out at least parts of a method according to one of claims 1 to 9 when the computer program is executed on a computer.
Citation Information
Patent Citations
Method and apparatus for testing the robustness of an artificial neural network
DE102019209228A1
Comparing a first KNN with a second KNN
DE102020215430A1