Robust and better trainable artificial neural network

By introducing a standardizer into the processing layer of deep neural networks and using the standardized function switching processing mechanism, the problem of huge changes in output variables when processing data is solved, and the learning ability and robustness of the network are improved.

CN114341887BActive Publication Date: 2025-06-10ROBERT BOSCH GMBH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202080063529.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2019-09-11
Filing Date
2020-07-28
Publication Date
2025-06-10
Estimated Expiration
2040-07-28

AI Technical Summary

Technical Problem

When deep neural networks process data, small changes in input variables may lead to huge changes in output variables, resulting in the network "not learning", that is, the recognition hit rate does not significantly exceed the random guessed hit rate.

Method used

The standardizer is introduced into the processing layer of artificial neural networks. Through translation, standardization and back-translation elements, the input variables are converted into standardized output vectors, and the standardization function switches different processing mechanisms according to the norm of the input vector to reduce the statistical requirements for input variables.

Benefits of technology

Through standardized processing, the network's sensitivity to input variables is reduced, the rounding error and noise influence is reduced, and the learning ability of the network and the robustness of the adversarial examples are improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114341887B_ABST
    Figure CN114341887B_ABST
Patent Text Reader

Abstract

The artificial neural network KNN (1) has processing layers (21-23), each processing layer being configured to process input variables (21a-23a) into output variables (21b-23b) according to the trainable parameters (20) of the KNN (1), wherein at least one normalizer (3) is connected in at least one processing layer (21-23) and / or between at least two processing layers (21-23), wherein the normalizer (3) - comprises a translation element (3a) configured to translate an input variable (31) introduced into the normalizer (3) into one or more input vectors (32) using a predefined transformation (3a'), wherein each input variable (31) exactly enters one input vector (32); - comprises a normalization element (3b) configured to normalize the one or more input vectors (32) into one or more output vectors (34) based on a normalization function (33), wherein the normalization function (33) has at least two different mechanisms (33a, 33b) and switches between the mechanisms (33a, 33b) according to the norm (32a) of the input vector (32) at a point and / or in a region, the position of the point and / or the region depending on a predefined parameter ρ; and - comprises a back-translation element (3c) configured to translate the output vector (34) into an output variable (35) using the inverse (3a'') of the predefined transformation (3a'), the output variable having the same dimension as the input variable (31) fed to the normalizer (3).
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to artificial neural networks, in particular for determining classification, regression, and / or semantic segmentation of physical measurement data. Background Art

[0002] For at least partial autonomous driving of a vehicle in road traffic, it is necessary to observe the vehicle's environment and identify the objects contained in this environment, and, if necessary, determine the position of these objects relative to one's own vehicle. Based on this, a decision can then be made as to whether the presence and / or the identified movement of these objects requires a change in the behavior of one's own vehicle.

[0003] For example, since the optical imaging of a vehicle's environment by a camera is affected by a large number of influencing factors, two images of the same scene will not be exactly the same. Therefore, artificial neural networks KNN with generalization capabilities of an ideal size are typically used to identify objects. These KNNs are trained such that they map the learning input data well to the learning output data according to a cost function. It is then expected that the KNNs will also correctly identify objects in situations that are not the subject of this training.

[0004] In the case of deep networks with many layers, the problem is that it is not possible to control by how many orders of magnitude the numerical values of the data processed by the network move. For example, there may be numbers in the range from 0 to 1 in the first layer of the network, while in deeper layers, numerical values of the order of 1000 can be reached. Then, a small change in the input variable may cause a large change in the output variable. This can lead to the network "not learning", i.e., the recognition hit rate does not significantly exceed the hit rate of random guessing.

[0005] (S. Ioffe, C. Szegedy, “Batch Normalization: Accelerating DeepNetwork Training by Reducing Internal Covariate Shift”, arXiv: 1502.03167v3[cs.LG] (2015)) discloses normalizing the numerical values of the data generated in the KNN for each processed mini-batch of training data to a unified order of magnitude.

[0006] (D. -A. Clevert, T. Unterthirner, S. Hochreiter, "Fast and Accurate Deep Network Learning by Exponential Linear Units (ELUs)", arXiv:1511.07289 [cs.LG] (2016)) discloses activating neurons using a novel activation function that can alleviate the above problems. Summary of the Invention

[0007] In the scope of the present invention, an artificial neural network is developed. The network includes a large number of processing layers connected in series. Each processing layer is configured to process an input variable into an output variable according to the trainable parameters of KNN. Here, in particular, the output variable of each layer can be introduced as an input variable into at least the next layer.

[0008] A novel normalizer is connected in at least one processing layer and / or between at least two processing layers.

[0009] The normalizer includes a translation element. The translation element is configured to translate the input variable introduced into the normalizer into one or more input vectors using a pre-given transformation. Here, each input variable exactly enters one input vector. Thus, a single input vector or a set of input vectors is produced that overall has as much information (i.e., for example, as many values) as the information conveyed in the input variables fed to the normalizer.

[0010] The normalizer further includes a normalization element. The normalization element is configured to normalize the input vector into one or more output vectors based on a normalization function. In the context of the present invention, the normalization of a vector is particularly understood as an arithmetic operation that allows the number of components of the vector and the direction of the vector in a multi-dimensional space to remain unchanged but can change the norm interpreted by the vector in the multi-dimensional space. The norm can, for example, correspond to the length of the vector in the multi-dimensional space. In particular, the normalization function can be designed such that it can map vectors with very different norms to vectors with similar or the same norm.

[0011] The normalization function has at least two different mechanisms and switches between these mechanisms depending on the norm of the input vector at a point and / or in a region, the position of the point and / or the region depending on a pre-given parameter ρ. This means that input vectors with a norm to the left (i.e., slightly smaller) of the point or the region are processed by the normalization function in a different way than input vectors with a norm to the right (i.e., slightly larger) of the point or the region. In particular, one mechanism can, for example, include that the norm of the input vector changes less strongly, absolutely and / or relatively, than prescribed by another mechanism when forming the output vector. One of the mechanisms can, for example, also include not changing the input vector at all, but accepting the input vector unchanged as the output vector.

[0012] The normalizer further includes a backtranslation element. The backtranslation element is configured to translate the output vector into output variables using the inverse of a pre-given transformation. These output variables have the same dimension as the input variables fed to the normalizer. Thereby, the normalizer can be used anywhere between two processing steps of the KNN. Thus, in further processing by the KNN, the output variables of the normalizer can replace those variables that have previously been processed in the KNN and fed to the normalizer as input variables.

[0013] It has been recognized that the numerical stability of the normalization function can be improved precisely by switching the mechanisms depending on the norm of the input vector and the pre-given parameter ρ. In particular, it will counteract the tendency of the normalization function to amplify rounding errors that are inevitable during the mechanical processing of the input variables and the noise that is always contained in physical measurement data.

[0014] Rounding errors and noise create small non-zero numerical values where ideally there should be zero within the KNN. In contrast, the numerical values representing the useful signals contained in the physical measurement data or the conclusions drawn therefrom are significantly larger. If now, between two processing steps in the KNN, the numerical values representing the intermediate results that already exist are combined into vectors and these vectors are normalized, this will cause the original existing distance between the useful signal and its processing product on the one hand and the noise or rounding error on the other hand to be partially or even completely flattened.

[0015] By switching between the mechanisms, it is now possible, for example, to set that all input vectors whose norm does not reach a certain minimum measure will not be changed or will only be slightly changed in terms of their norm. For example, if input vectors with a larger norm are simultaneously mapped to output vectors with the same or a similar norm, there will still be a sufficiently large norm distance to the output vectors traced back to noise or rounding errors.

[0016] This also reduces the statistical requirements for the input variables fed to the normalizer. In principle, it is not necessary to access the following input variables, which can be traced back to different input variable samples fed to the KNN. Instead, if only the numerical values of the intermediate result related to the unique input variable sample fed to the KNN are fed to the normalizer, the main conclusions contained in the intermediate result of the KNN are retained.

[0017] Thus, the advantages that could hitherto be achieved by means of "batch normalization" can be achieved to the same extent or to a greater extent without making the normalization related to the small batch training data processed during KNN training. Thus, the effect of the normalization in particular no longer depends on the size of the small batch selected during training.

[0018] This in turn makes it possible to freely choose the size of the small batch, for example from the perspective of the data throughput during KNN training. To obtain the maximum throughput, it is particularly advantageous to choose the size of the small batch such that the small batch exactly fits the available working memory (e.g., the video RAM of the graphics processor used, GPU) and can be processed in parallel. This is not always the same as the size of the small batch that is also optimal for "batch normalization" in terms of the maximum performance (e.g., classification accuracy) of this network. On the contrary, smaller or larger small batch sizes may be beneficial for "batch normalization", where it is questionable whether the optimal "batch normalization" (and thus the optimal accuracy in the sense of the task) typically takes precedence over the best data throughput during training. In addition, "batch normalization" works very poorly for small batch sizes because the statistics of the small batch can then only very inadequately approximate the statistics of the entire training data.

[0019] In addition, unlike the batch size of "batch normalization", the parameter ρ used by the normalization element is a continuous parameter rather than a discrete parameter. Therefore, this parameter ρ is significantly easier to optimize. For example, this parameter can be trained together with the trainable parameters of the KNN. In contrast, optimizing the batch size of "batch normalization" may require a complete retraining of the KNN for each tested candidate batch size, which correspondingly increases the training effort.

[0020] All in all, the KNN can be effectively trained as a whole while also being robust against attempts at manipulation using so-called "adversarial examples". These attempts aim to deliberately cause misclassification of the artificial neural network, for example, by making minor, unobtrusive changes to the data fed to the artificial neural network. The effect of such changes within the artificial neural network is reduced by the said normalization. Therefore, in order to achieve the desired misclassification, a correspondingly greater manipulation must be carried out at the input of the artificial neural network, and then this manipulation will be noticed with a greater probability.

[0021] In a particularly advantageous design, at least one normalization function is constructed such that input vectors with a norm less than a parameter ρ remain unchanged, and input vectors with a norm greater than the parameter ρ are normalized to a uniform norm while maintaining their direction. Vectors in any multi-dimensional space An example of such a normalization function, as explained above, is:

[0022]

[0023] If the vector norm is less than ρ, the vector remains unchanged. This is the first mechanism of the normalization function Conversely, if is at least equal to ρ, then projects the vector onto the surface of a sphere with radius ρ. This means that the normalized vector then points in the same direction as before but ends on the sphere's surface. This is the second mechanism of the normalization function In the case of a switch occurs between the two mechanisms.

[0024] In another particularly advantageous design, the switching between at least one normalization function's different mechanisms is controlled by a Softplus function, which has a zero-crossing argument when the norm of the input vector equals the parameter ρ. An example of such a function is

[0025] . The Softplus function here is given by

[0026] .

[0027] The advantage of this function is that it is differentiable with respect to ρ. Vectors with a below ρ are now no longer unchanged, but these vectors change significantly less compared to vectors with a larger norm . If tends to 0, then regardless of the value of ρ, the norm of the vector in multi-dimensional space decreases by approximately 25%. There is no norm such that results in an increased norm. Thus, not only are effects such as rounding errors and noise avoided from being amplified, but this effect is even further reduced by further decreasing rather than increasing too-small norms to a uniform level.

[0028] In another particularly advantageous design, at least one predefined transformation from the input variables of the normalizer to the input vectors includes converting the tensor of the input variables into one or more input vectors. The tensor contains f feature maps, and each feature map assigns feature information to n different positions. For example, the tensor can be written as . Then, the normalizer only needs to use the feature information that appears in the unique samples of the input variables from the input to the KNN. Mini-batch samples can still be used, but they are optional.

[0029] In another particularly advantageous design, for each of the f feature maps, at least one predefined transformation includes combining the feature information regarding all positions contained in the feature map into the input vector assigned to the feature map. Thus, for i = 1,..., f, the complete i-th feature map is read out, and the values contained in the feature map are successively written into the input vector : .

[0030] In this way, the tensor X is continuously converted into the input vector , where i = 1,..., f. Thus, a norm is formed over all the feature maps , and the greater the overall specificity of a particular feature in the input variables, the greater this norm.

[0031] In another particularly advantageous design, for each of the n positions, at least one predefined transformation includes combining the feature information assigned to the position by all the feature maps into the input vector assigned to the position. Thus, for j = 1,..., n, the feature information values respectively marked exactly for the position in all the feature maps are read out for the j-th position, and the values thus obtained are successively written into the input vector : .

[0032] In this way, the tensor X is continuously converted into the input vector . The norm is thus formed over the set of features respectively assigned to the individual positions, and the richer the features of the input variables related to the specific position, the greater this norm.

[0033] In another particularly advantageous design, at least one predefined transformation includes combining all the feature information from the tensor X into a single input vector. Then, the richer the features of the overall samples of the input variables fed to the KNN, the greater the norm of this input vector .

[0034] In each of the designs mentioned, the tensor X or the vector , and , further preprocessing can be performed before applying the normalization function. Specifically,

[0035] • All feature information can be separately subtracted from the arithmetic mean formed by all feature information ("overall sample mean" = the mean formed by all information of the samples involved in the input variables regarding KNN); and / or

[0036] • The feature information respectively contained in each of the f feature maps can be separately subtracted from the arithmetic mean formed by the feature information for that feature map; and / or

[0037] • The feature information assigned to each of the n positions by all feature maps can be separately subtracted from the arithmetic mean formed by the feature information belonging to that position for all feature maps.

[0038] As described above, the normalizer can "cycle" anywhere in the KNN because the output variable of the normalizer has the same dimension as its input variable and can thus replace these input variables during further processing in the KNN.

[0039] In a particularly advantageous design, at least one normalizer receives the weighted sum of the input variables of the processing layer as the input variable. The output variable of this normalizer is introduced into a non-linear activation function to form the output variable of the processing layer. If normalizers are connected at this point in many or even all processing layers, the behavior of the non-linear activation functions within the KNN can be largely normalized because these activation functions always act on values of substantially the same order of magnitude.

[0040] In another particularly advantageous design, at least one normalizer receives the output variable of the first processing layer as the input variable, where the output variable is formed by applying a non-linear activation function. The output variable of this normalizer is introduced as an input variable into a further processing layer, which weighted-sums these input variables according to trainable parameters. If many or even all transitions between adjacent processing layers in the KNN are made through normalizers, the order of magnitude of the input variables respectively participating in this weighted sum can be substantially normalized within the KNN. This ensures better convergence during training.

[0041] As explained above, in the described KNN case, the accuracy of classifying, regressing, and / or semantically segmenting real and / or simulated physical measurement data can be significantly improved. In particular, for example, the accuracy can be measured by means of validation input variables that were not used during training and are known as "ground truth" for the validation output variables (i.e., for example, the rated classification to be achieved or the rated regression value to be achieved). In addition, the sensitivity to "adversarial examples" is reduced. Thus, in a particularly advantageous design, the KNN is configured as a classifier and / or a regressor.

[0042] A KNN configured as a classifier can be used, for example, to identify objects and / or object states sought in the corresponding application range among the input variables of the KNN. Thus, for example, autonomous agents such as robots or at least partially autonomous vehicles must identify objects in their environment in order to be able to act appropriately in situations characterized by a specific constellation of objects. For example, a KNN configured as a classifier can also identify features (such as injuries) in the context of medical imaging from which a medical diagnosis can be derived. Similarly, such a KNN can also be used in the context of optical inspection to check whether manufactured products or other work results (such as weld seams) are normal.

[0043] Semantic segmentation of physical measurement data can be formed, for example, by classifying the components of the measurement data according to what type of object they belong to.

[0044] Physical measurement data can in particular be, for example, image data recorded by spatially resolving the sensing of electromagnetic waves in the visible range or, for example, using a thermal imager in the infrared range. The spatially resolved components of the image data can be, for example, pixels, stixels (columnar pixels), or voxels (volume pixels) depending on the specific space in which these images are located, i.e., depending on the dimension of the image data. For example, the physical measurement data can also be obtained by detecting the reflection of an interrogation radiation in the range of radar measurements, lidar measurements, or ultrasonic measurements.

[0045] In the mentioned applications, a KNN configured as a regressor can also be used instead of or in combination with it. In this function, the KNN can provide conclusions about continuous variables sought in the corresponding application range. Examples of such variables are the size and / or speed of an object, as well as continuous evaluation metrics of product quality (such as roughness or the number of defects in a weld seam) or continuous evaluation metrics of features (such as the percentage of tissue considered damaged) that can be used for medical diagnosis.

[0046] Thus, in general, it is particularly advantageous for the KNN to be configured as a classifier and / or a regressor for identifying and / or quantitatively evaluating objects and / or states sought in the corresponding application range among the input variables of the KNN.

[0047] Advantageously, the KNN is configured as a classifier for identifying from physical measurement data obtained by observing the traffic situation in the vehicle's own environment using at least one sensor

[0048] • traffic signs, and / or

[0049] • pedestrians, and / or

[0050] • other vehicles, and / or

[0051] • other objects characterizing the traffic situation.

[0052] This is one of the most important tasks of at least partial automated driving. In the field of robotics or in the case of general autonomous agents, environmental perception also plays an important role.

[0053] The effect that can be achieved by using a normalizer in the KNN described above is in principle independent of whether the normalizer is a unit encapsulated in a certain form. It is only important that the intermediate products generated during processing are normalized at appropriate places in the KNN, and the results of this normalization are used to replace the intermediate products during further processing of the KNN.

[0054] Thus, the present invention generally relates to a method for operating a KNN having a large number of processing layers connected in series, each processing layer being configured to process input variables into output variables according to the trainable parameters of the KNN.

[0055] Within the scope of this method, a set of variables determined during processing is extracted from the KNN as input variables for normalization in at least one processing layer and / or between at least two processing layers. The input variables for normalization are transformed into one or more input vectors using a pre-given transformation, where each input variable exactly enters one input vector.

[0056] One or more input vectors are normalized into one or more output vectors based on a normalization function, where the normalization function has at least two different mechanisms and switches between the mechanisms according to the norm of the input vector at a point and / or in a region, the position of the point and / or the region depending on a pre-given parameter ρ.

[0057] The output vectors are translated into the normalized output variables using the inverse of the pre-given transformation, the normalized output variables having the same dimension as the normalized input variables. Then processing continues in the KNN, where the normalized output variables replace the previously extracted normalized input variables.

[0058] All of the disclosures given above related to the function of the normalizer also clearly apply to this method.

[0059] According to the above description, the present invention also relates to a system which is configured to control other technical systems based on an evaluation of physical measurement data using KNN. The system includes at least one sensor for recording physical measurement data, the aforementioned KNN, and a control unit. The control unit is configured to form a control signal from the output variables of the KNN for a vehicle or other autonomous agent (such as a robot), a classification system, a quality control system for mass-produced products, and / or a medical imaging system. All of the systems mentioned benefit from the fact that the KNN has learned to classify, regress, and / or perform semantic segmentation better, especially compared to a KNN that relies on "batch normalization" or "ELU" activation functions.

[0060] The sensor may include, for example, one or more image sensors for light of any visible or invisible wavelength and / or at least one radar sensor, lidar sensor, or ultrasonic sensor.

[0061] According to the above description, the present invention also relates to a method for training and operating the above KNN. In the context of this method, learning input variables are fed to the KNN. The learning input variables are processed by the KNN into output variables. An evaluation of the output variables is determined according to a cost function, which indicates how well the output variables match the learning output variables belonging to the learning input variables.

[0062] The trainable parameters of the KNN are optimized together with at least one of the previously described parameters ρ, which characterizes the transition between two mechanisms of the normalization function. The goal of this optimization is to obtain the following output variables during the further processing of the learning input variables, and it is expected that the evaluation of these input variables by the cost function will be better. This does not mean that each optimization step must be an improvement in this regard; on the contrary, the optimization can also learn from the "wrong path" that initially leads to deterioration.

[0063] In the case where the number of trainable parameters typically ranges from several thousand to several million, one or more additional parameters ρ are generally not important in the training cost for the KNN. This is in contrast to the optimization of discrete parameters (such as the batch size of "batch normalization"). As previously explained, the optimization of such discrete parameters makes it necessary to re-traverse the complete training of the KNN for each candidate value of the discrete parameter. Therefore, by including the additional parameter ρ as a continuous parameter in the training method as well, the total cost is significantly reduced compared to "batch normalization".

[0064] Furthermore, training the parameters of the KNN together with one or more additional parameters ρ can also utilize the synergistic effect between the two trainings. Thus, for example, during learning, the change in the trainable parameters (which directly control the processing of the input variables from the processing layer to the output variables) can advantageously interact with the change in the additional parameter ρ acting on the normalization function. Through this "combined force", for example, particularly "difficult cases" of classification and / or regression can be overcome.

[0065] The physical measurement data recorded using at least one sensor can be fed as input variables to the trained KNN. These input variables can then be processed by the trained KNN into output variables. Then, manipulation signals can be formed from the output variables for a vehicle or other autonomous agent (such as a robot), a classification system, a quality control system for mass-produced products, and / or a medical imaging system. Finally, the manipulation signals can be used to manipulate the vehicle, the classification system, the quality control system for mass-produced products, and / or the medical imaging system.

[0066] According to the above description, the present invention also relates to an additional method, which includes the complete action chain from providing the KNN to manipulating the technical system.

[0067] This additional method starts with providing the KNN. Then, the trainable parameters of the KNN and optionally at least one parameter ρ for optimizing the transition between the two mechanisms of the normalization function are trained such that the learning input variables are processed by the KNN into output variables that match the learning output variables belonging to the learning input variables according to the cost function.

[0068] The physical measurement data recorded using at least one sensor is fed as input variables to the trained KNN. These input variables are processed by the trained KNN into output variables. Manipulation signals are formed from the output variables for a vehicle or other autonomous agent (such as a robot), a classification system, a quality control system for mass-produced products, and / or a medical imaging system. The manipulation signals are used to manipulate the vehicle, the classification system, the quality control system for mass-produced products, and / or the medical imaging system.

[0069] The improved learning ability of the KNN described above has the following effect in this context: by manipulating the corresponding technical system, an action appropriate for the situation represented by the physical measurement data is triggered with a greater probability.

[0070] These methods can be implemented, in particular, fully or partially by a computer. Accordingly, the invention also relates to a computer program with machine-readable instructions which, when executed on one or more computers, cause the one or more computers to carry out one of the described methods. In this sense, an embedded system of a vehicle control device and a technical device which can also execute machine-readable instructions should also be regarded as a computer.

[0071] The invention also relates to a machine-readable data carrier and / or a download product with the computer program. A download product is a digital product which can be transmitted via a data network, i.e., can be downloaded by a user of the data network, and which can, for example, be offered for immediate download in an online store.

[0072] Furthermore, a computer can be equipped with the computer program, the machine-readable data carrier or the download product. BRIEF DESCRIPTION OF THE DRAWINGS

[0073] Other measures for improving the invention are shown in more detail below together with the description of the preferred embodiments of the invention based on the drawings.

[0074] Figure 1 An embodiment of KNN 1 is shown;

[0075] Figure 2 An embodiment of the normalizer 3 is shown;

[0076] Figure 3 An exemplary tensor 31' of the input variable 31 with the normalizer 3 is shown;

[0077] Figure 4 An embodiment of a system 10 with the KNN 1 is shown;

[0078] Figure 5 An embodiment of a method 100 for training and running the KNN 1 is shown;

[0079] Figure 6 An embodiment of a method 200 with a complete action chain from providing the KNN 1 to controlling a technical system is shown. DETAILED DESCRIPTION

[0080] Figure 1The KNN 1 exemplarily shown herein includes three processing layers 21-23. Each processing layer 21-23 receives input variables 21a-23a and processes them into output variables 21b-23b. The input variable 21a of the first processing layer 21 is also the input variable 11 of the overall KNN 1. The output variable 23b of the third processing layer 23 is simultaneously the output variables 12, 12' of the overall KNN 1. The actual KNN 1, especially for use in classification or other computer vision applications, is significantly deeper and includes dozens of processing layers 21-23.

[0081] Figure 1 Shown herein are the possibilities of two examples regarding how the normalizer 3 can be introduced into the KNN 1.

[0082] One possibility is to feed the output variable 21b of the first processing layer 21 as the input variable 31 to the normalizer 3, and then feed the output variable 35 of the normalizer as the input variable 22a to the second processing layer 22.

[0083] Schematically shown within the block 22 is the processing performed in the second processing layer 22, including a second possibility of binding one or more normalizers 3. First, the input variables 22a are summed into one or more weighted sums according to the trainable parameters 20 of the KNN 1, which is represented by the sum symbol. The result is fed as the input variable 31 to the normalizer 3. The output variable 35 of the normalizer 3 is calculated as the output variable 22b of the second processing layer 22 using a non-linear activation function (represented as the ReLU function herein). Figure 1 Shown herein is the output variable 35 of the normalizer 3 being calculated as the output variable 22b of the second processing layer 22 using a non-linear activation function (represented as the ReLU function herein).

[0084] Multiple different normalizers 3 can be used in the same KNN 1. In particular, each normalizer 3 can then have its own parameter ρ for the transition between the mechanisms of its normalization function 33. Additionally, each normalizer 3 can also be coupled with its own dedicated preprocessing.

[0085] Figure 2 An embodiment of the normalizer 3 is shown. The normalizer 3 uses a translation element 3a that implements a pre-given transformation 3a' to translate the input variable 31 of the normalizer 3 into one or more input vectors 32. These input vectors 32 are fed to the normalization element 3b and normalized there into an output vector 34. The output vector 34 is translated in the back-translation element 3c according to the inverse 3a'' of the pre-given transformation 3a' into the output variable 35 of the normalizer 3, and these output variables have the same dimension as the input variable 31 of the normalizer 3.

[0086] How the normalization of the input vector 32 to the output vector 34 is carried out is shown in detail within the block 3b. The normalization function 33 used has two mechanisms 33a and 33b, in each of which the normalization function exhibits qualitatively different behavior and in particular affects the input vector 32 with different intensities. The norm 32a of the corresponding input vector 32 interacts with at least one pre-given parameter ρ to determine which of the mechanisms 33a and 33b to use. For the sake of illustration, this is shown as a binary decision in Figure 2 . However, in practice it is particularly advantageous that the mechanisms 33a and 33b smoothly transition into each other, in particular in a differentiable manner with respect to the parameter ρ.

[0087] Figure 3 An exemplary tensor 31' of the input variable 31 of the normalizer 3 is shown. In this example, the tensor 31' is organized as a stack of f feature maps 31a. Thus, the index i on the feature map 31a ranges from 1 to f. Each feature map 31a assigns feature information 31c to n positions 31b respectively. Thus, the index j on the position 31b ranges from 1 to n.

[0088] Figure 3 Two possibilities of how the input vector 32 can be formed are shown exemplarily in. According to the first possibility, all the feature information 31c of the feature map 31a (here the feature map 31a for i = 1) are combined separately in one input vector 32. According to the second possibility, all the feature information 31c belonging to the same position 31b (here the position 31b for j = 1) are combined separately in the input vector 32. For the sake of clarity, a third possibility not shown in Figure 3 consists in writing all the feature information 31c from the entire tensor 31' into a single input vector 32.

[0089] Figure 4 An embodiment of the system 10 is shown by means of which additional technical systems 50 - 80 can be controlled. At least one sensor 6 is provided for recording physical measurement data 6a. The measurement data 6a are fed as input variable 11 to the KNN 1, which can in particular be in its trained state 1*. The output variable 12' provided by the KNN 1, 1* is processed in the evaluation unit 7 as a control signal 7a. This control signal 7a is intended for controlling a vehicle or other autonomous agent (such as a robot) 50, a classification system 60, a quality control system 70 for mass-produced products and / or a medical imaging system 80.

[0090] Figure 5It is a flowchart of an embodiment of method 100 for training and running KNN 1. In step 110, learning input variable 11a is fed to KNN 1. In step 120, learning input variable 11a is processed by KNN 1 into output variable 12, where the behavior of KNN 1 is characterized by trainable parameter 20. In step 130, it is evaluated according to cost function 13 how well output variable 12 matches learning output variable 12a belonging to learning input variable 11a. In step 140, trainable parameter 20 is optimized with the aim of obtaining the following output variable 12 during further processing of learning input variable 11a by KNN 1, and a better evaluation 130a is determined for said output variable 12 in step 130.

[0091] Figure 6 It is a flowchart of an embodiment of method 200 having a complete action chain from providing KNN 1 to controlling the systems 50, 60, 70, 80.

[0092] In step 210, KNN 1 is provided. In step 220, the trainable parameter 20 of KNN 1 is trained so as to produce the trained state 1* of KNN 1. In step 230, physical measurement data 6a determined using at least one sensor 6 is fed as input variable 11 to the trained KNN 1*. In step 240, output variable 12' is formed by the trained KNN 1*. In step 250, a control signal 7a is formed from output variable 12'. In step 260, control signal 7a is used to control one or more of the systems 50, 60, 70, 80.

Claims

1. An artificial neural network KNN (1), configured as a classifier for image data, the KNN having a large number of processing layers (21 - 23) connected in series, each processing layer being configured to process an input variable (21a - 23a) into an output variable (21b - 23b) according to the trainable parameters (20) of the KNN (1), wherein at least one normalizer (3) is connected in at least one processing layer (21 - 23) and / or between at least two processing layers (21 - 23), wherein the normalizer (3) · includes a translation element (3a), the translation element being configured to translate an input variable (31) introduced into the normalizer (3) into one or more input vectors (32) using a pre - given transformation (3a'), wherein each input variable (31) exactly enters one input vector (32); · includes a normalization element (3b), the normalization element being configured to normalize the one or more input vectors (32) into one or more output vectors (34) based on a normalization function (33), wherein the normalization function (33) has at least two different mechanisms (33a, 33b) and switches between the mechanisms (33a, 33b) according to the norm (32a) of the input vector (32) at a point and / or in a region, the position of the point and / or the region depending on a pre - given parameter ρ; · includes a back - translation element (3c), the back - translation element being configured to translate the output vector (34) into an output variable (35) using the inverse (3a”) of the pre - given transformation (3a'), the output variable having the same dimension as the input variable (31) fed to the normalizer (3), wherein at least one normalization function (33) is configured to keep input vectors (32) with a norm (32a) less than the parameter ρ unchanged and to normalize input vectors (32) with a norm (32a) greater than the parameter ρ to a uniform norm (32a) while maintaining the direction, wherein the parameter ρ is a continuous parameter.

2. The KNN (1) according to claim 1, wherein the switching of at least one normalization function (33) between different mechanisms (33a, 33b) is controlled by a Softplus function, and when the norm (32a) of the input vector (32) is equal to the parameter ρ, the argument of the Softplus function has a zero - crossing point.

3. The KNN (1) according to any one of claims 1 to 2, wherein at least one pre - given transformation (3a') includes combining all the feature information (31c) from a tensor (31') of the input variable (31) into one or more input vectors (32), where f feature maps (31a) are combined in the tensor, and each feature map assigns feature information (31b) to n different positions (31c).

4. The KNN (1) according to claim 3, wherein for each of the f feature maps (31a), at least one predefined transformation (3a') includes combining the feature information (31c) regarding all positions (31b) included in the feature map (31a) in an input vector (32) assigned to the feature map (31a).

5. The KNN (1) according to any one of claims 3 to 4, wherein for each of the n positions (31b), at least one predefined transformation (3a') includes combining the feature information (31c) assigned to the position (31b) through all the feature maps (31a) in an input vector (32) assigned to the position (31b).

6. The KNN (1) according to any one of claims 3 to 5, wherein at least one predefined transformation (3a') includes combining all the feature information (31c) from the tensor (31) in a unique input vector (32).

7. The KNN (1) according to any one of claims 3 to 6, wherein at least one predefined transformation (3a') includes subtracting the arithmetic mean formed for all the feature information (31c) from all the feature information (31c) respectively.

8. The KNN (1) according to any one of claims 3 to 7, wherein at least one predefined transformation (3a') includes subtracting the arithmetic mean (31c) formed for the feature map (31a) from the feature information (31c) respectively included in each of the f feature maps (31a).

9. The KNN (1) according to any one of claims 3 to 8, wherein at least one predefined transformation (3a') includes subtracting the arithmetic mean formed for all the feature maps (31a) from the feature information (31c) respectively assigned to each of the n positions (31b) through all the feature maps (31a).

10. The KNN (1) according to any one of claims 1 to 9, wherein at least one normalizer (3) obtains a weighted sum of the input variables (21a - 23a) of the processing layers (21 - 23) as the input variable (31), and the output variable (35) of the normalizer (3) is introduced into a non - linear activation function to form the output variables (21b - 23b) of the processing layers (21 - 23).

11. The KNN (1) according to any one of claims 1 to 10, wherein at least one normalizer (3) obtains the output variables (21b - 23b) of the first processing layers (21 - 23) as the input variable (31), the output variables being formed by applying a non - linear activation function, and wherein the output variable (35) of the normalizer (3) is introduced as the input variables (21a - 23a) into a further processing layer (21 - 23), the further processing layer summing up the input variables (21a - 23a) weighted by the trainable parameters (20).

12. The KNN (1) according to claim 1, further configured as a classifier for identifying from physical measurement data obtained by observing traffic conditions in the environment of its own vehicle using at least one sensor · traffic signs, and / or · pedestrians, and / or · other vehicles, and / or · other objects characterizing the traffic conditions.

13. A system (10) comprising at least one sensor (6) for recording physical measurement data (6a), a KNN (1, 1*) according to any one of claims 1 to 12, and a control unit (7), the physical measurement data (6a) being introduced as input variables (11) into the KNN, the control unit being configured to form control signals (7a) for a vehicle or other autonomous agent (50), a classification system (60), a quality control system (70) for mass-produced products, and / or a medical imaging system (80) from the output variables (12') of the KNN (1).

14. A method (100) for training and operating a KNN (1) according to any one of claims 1 to 12, having the steps: · feeding (110) learning input variables (11a) to the KNN (1); · processing (120) the learning input variables (11a) by the KNN (1) into output variables (12); · determining (130) an evaluation (130a) of the output variables (12) according to a cost function (13), the evaluation indicating how well the output variables (12) match learning output variables (12a) belonging to the learning input variables (11a); · optimizing (140) the trainable parameters (20) of the KNN (1) together with at least one parameter ρ, the parameter ρ characterizing the transition between two mechanisms (33a, 33b) of a normalization function (33), the optimization being aimed at obtaining the following output variables (12) during further processing (120) of the learning input variables (11a), the evaluation of the input variables (130a) by the cost function (13) being expected to be better; · feeding (230) physical measurement data (6a) recorded by at least one sensor (6) as input variables (11) to the trained KNN (1*), and processing (240) the physical measurement data by the trained KNN (1*) into output variables (12'); · forming (250) control signals (7a) for a vehicle or other autonomous agent (50), a classification system (60), a quality control system (70) for mass-produced products, and / or a medical imaging system (80) from the output variables (12'); · using the control signals (7a) to control (260) the vehicle (50), the classification system (60), the quality control system (70) for mass-produced products, and / or the medical imaging system (80).

15. A computer program product comprising machine-readable instructions which, when executed on one or more computers, cause the one or more computers to implement the KNN(1) according to any one of claims 1 to 12 and to perform the method (100) according to claim 14.

16. A machine-readable data carrier and / or download product having the computer program product according to claim 15.

17. A computer equipped with the computer program product according to claim 15 and / or equipped with the machine-readable data carrier and / or download product according to claim 16.