METHOD AND DEVICE FOR CLASSIFYING SENSOR DATA AND FOR DETERMINING A CONTROL SIGNAL FOR CONTROLLING AN ACTUATOR
Patent Information
- Application Number
- DE502019013480
- Authority / Receiving Office
- DE · DE
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2019-10-01
- Filing Date
- 2019-11-28
- Publication Date
- 2025-07-10
- Estimated Expiration
- 2039-11-28
AI Technical Summary
Existing neural network training methods face challenges in efficiently classifying input signals due to issues with data augmentation, batch normalization, and the complexity of hyperparameter tuning, which can lead to overfitting and reduced generalizability.
The introduction of a scaling layer in the neural network that maps input signals onto a predeterminable value range, defined by a norm, simplifies the architecture search and training process by projecting inputs onto a ball with a fixed center and radius, thereby controlling the scale of input signals.
This approach enhances the efficiency of neural network training by reducing overfitting, improving generalizability, and simplifying the training process, while also allowing for memory-efficient compression of neural networks.
Description
[0001] The invention relates to a method for classifying input signals, a method for providing a control signal, a computer program, a machine-readable storage medium and an actuator control system. State of the art
[0002] From "Improving neural networks by preventing co-adaptation of feature detectors," arXiv preprint arXiv:1207.0580v1, Geoffrey E. Hinton, Nitish Srivastava, Alex Krizhevsky, IIya Sutskever, Ruslan R. Salakhutdinov (2012), a method for training neural networks is known in which feature detectors are randomly omitted during training. This method is also known as "dropout."
[0003] From "Batch Normalization: Accelerating Deep Network Training by Reducing Internal Covariate Shift", arXiv preprint arXiv:1502.03167v3, Sergey Loffe, Christian Szegedy (2015), a method for training neural networks is known in which input variables in a layer are normalized for a small batch (English: "mini-batch") of training examples. ELAD HOFFER ET AL: "Norm matters: efficient and accurate normalization schemes in deep networks", ARXIV.ORG, CORNELL UNIVERSITY LIBRARY, 201 OLIN LIBRARY CORNELL UNIVERSITY ITHACA, NY 14853, 5.March 2018, discloses a method for classifying input signals which were determined as a function of an output signal of a sensor, by means of a neural network, wherein the neural network comprises a scaling layer, wherein the scaling layer maps an input signal present at the input of the scaling layer onto an output signal present at the output of the scaling layer in such a way that this mapping corresponds to a projection of the input signal onto a predefinable value range, wherein the parameters which characterize the mapping are predefinable. Advantages of the invention
[0004] The method with the features of the independent claims has the advantage that an architecture search of the neural network is simplified.
[0005] Advantageous further training is the subject of dependent claims. Disclosure of the invention
[0006] With a sufficient amount of training data, so-called "deep learning" methods, i.e. (deep) artificial neural networks, can be used to efficiently create a mapping between an input space V 0 and an exit room V k This can, for example, be a classification of sensor data, especially image data, i.e. a mapping of sensor data or image data to classes. The approach is based on k - 1 many hidden rooms V 1 , ... , V k -1. Furthermore, k many illustrations f i< : V i- 1 → V i ( i = 1 ... k ) between these spaces. Each of these images f i< is usually referred to as a layer. Such a layer f i< is typically weighted w i ∈ W i< with a suitably chosen room W i< parameterized. The weights w 1 , ... , w k the k many layers f i< are also referred to as weights w ∈ W := W 1< × ... × W k< and the picture of the entrance room V 0 to the exit room V k as f w : V 0 → V k , which are derived from the individual illustrations f i< (with weights explicitly indicated as subscripts w i ) results in f w x : = f w k k ∘ … ∘ f w 1 1 x .
[0007] With a given probability distribution D , which on V 0 × V k is defined, the task of training the neural network is to determine weights w ∈ W to determine such that an expected value Φ of a cost function L Φ w = E x D y D ∼ D L f w x D , y D is minimized. Here, the cost function L a measure of the distance between the function f w determined mapping of an input variable x D to one size f w ( x D ) in the exit room V k and an actual output value y D in the exit room V k .
[0008] A "deep neural network" can be understood as a neural network with at least two hidden layers.
[0009] To minimize this expected value Φ, gradient-based methods can be used that determine a gradient ∇Φ with respect to the weights w. This gradient ∇Φ is usually calculated using training data ( x j , y j ), i.e. by ∇ w L ( f w ( x j , y j )), where the indices j from a so-called epoch. An epoch is a permutation of the labels {1, ..., N} of the available training data points.
[0010] To expand the training data set, so-called data augmentation (also called augmentation) can be used. For each index j from the era instead of the pair ( x j , y j ) an augmented pair ( x a , y j ), where the input signal x j by an augmented input value x a ∈ α ( x j ) is replaced. α ( x j ) can represent a number of typical variations of the input signal x j (including the input signal x j itself), which allows a classification of the input signal x j , i.e. the output signal of the neural network, unchanged.
[0011] However, this epoch-based sampling is not entirely consistent with the definition in equation (1), since each data point is selected exactly once during an epoch. The definition in equation (1), on the other hand, assumes independently drawn data points. This means that while equation (1) assumes a drawing of the data points "with replacement," epoch-based sampling performs a drawing of the data points "without replacement." This can lead to the assumptions of mathematical convergence proofs not being met (because if one draws N Examples from a set N many data points, the probability of drawing each of these data points exactly once is smaller than e − N 2 (for N > 2), while in epoch-based sampling this probability is always equal to 1.
[0012] If data augmentation is used, this statistical effect can be further enhanced, since in each epoch an element of the set α ( x j ) and, depending on the augmentation function α , it cannot be ruled out that α ( x j ) ≈ α ( x i ) for i ≠ j. A statistically correct representation of the augmentations using the quantity α ( x j ) is difficult here, since the effect is not available for each entry date x j must be equally pronounced. For example, a rotation may have no effect on circular objects, but may have a strong effect on general objects. Therefore, the size of the set α ( x j ) from the date of receipt x j dependent, which can be problematic for adverserial training methods.
[0013] Finally, the number Nof the training data points is usually a complex parameter to set. N If the value is too large, the duration of the training process can be unduly extended. NIf the value chosen is too small, convergence cannot be guaranteed, since mathematical proofs of convergence are generally based on assumptions that are then not met. Furthermore, it is not clear at what point training should reliably end. If one takes a portion of the data points as the evaluation data set and determines the quality of convergence using this evaluation data set, this can lead to over-fitting of the weights w with regard to the data points of the evaluation data set. This not only reduces data efficiency but can also degrade the performance of the network when applied to data other than the training data. This can lead to a reduction in what is known as "generalizability".
[0014] To reduce overfitting, information stored in the hidden layers can be randomly thinned out using the "dropout" method mentioned above.
[0015] To improve the randomization of the training process, statistical parameters can be µ and σ introduced via so-called mini-batches, which are probabilistically updated during the training process. During inference, the values of these parameters µ and σ chosen as fixed values, for example as estimates from training by extrapolating the exponential decay behavior.
[0016] If the layer with index i a batch normalization layer, the corresponding weights w i = ( µ i , σ i ) are not updated during gradient descent, ie these weights w i are therefore treated differently than the weights w k the remaining layers k This increases the complexity of an implementation.
[0017] In addition, the size of the mini-batches is a parameter that generally influences the training result and therefore must be adjusted as well as possible as an additional hyperparameter, for example in the context of a (possibly complex) architecture search.
[0018] In a first aspect of the invention, the neural network comprises a scaling layer, wherein the scaling layer maps an input signal present at the input of the scaling layer to an output signal present at the output of the scaling layer in such a way that this mapping corresponds to a projection of the input signal onto a predeterminable value range, wherein parameters characterizing the mapping are predeterminable. The value range is defined by a norm. In this way, the scaling layer ensures that the scale of the input signal is limited with respect to this norm.
[0019] In this context, "specifiable" means that these parameters are adjusted during training of the neural network depending on a gradient, whereby this gradient is determined in the usual way depending on an output signal of the neural network and a corresponding desired output signal.
[0020] This means that firstly in a training phase the predeterminable parameters are adapted depending on a method for training the neural network, wherein during training an adaptation of these predeterminable parameters takes place depending on the output signal of the neural network when the input signal of the neural network and the associated desired output signal are supplied, wherein the adaptation of the predeterminable parameters takes place depending on a determined gradient which is dependent on the output signal of the neural network and the associated desired output signal.
[0021] According to the invention, the scaling layer maps an input signal present at the input of the scaling layer onto an output signal present at the output of the scaling layer in such a way that this mapping corresponds to a projection onto a ball, with the center c and / or radius ρ this ball can be fixed.
[0022] Here the mapping is given by the equation y = argmin N 1( y-c )≤ ρ N 2 ( x - y) with a first standard ( N 1 ) and a second standard ( N 2 ). The term "norm" is to be understood in the mathematical sense.
[0023] In a particularly efficient further training course, it can be provided that the first standard ( N 1 ) and the second standard ( N 2 ) are chosen equally.
[0024] Alternatively or additionally, the first standard ( N 1 ) a L ∞< -norm. This norm can also be calculated very efficiently, especially when the first norm ( N 1 ) and the second standard ( N 2 ) are not chosen equally.
[0025] Alternatively, it can be provided that where the first standard ( N 1 ) a L1< norm. This choice of the first norm favors the sparsity of the output signal of the scaling layer. This is advantageous, for example, for the compression of neural networks, since weights with the value 0 have no contribution to the output value of their layer.
[0026] Therefore, a neural network with such a layer can be used particularly memory-efficiently, especially in conjunction with a compression method.
[0027] In the variants described for the first standard ( N 1 ) it can advantageously be provided that the second standard ( N 2 ) a L 2< norm. This makes the procedures particularly easy to implement.
[0028] It is particularly advantageous if the solution of the equation y = argmin N 1( y-c )≤ ρ N 2 ( x - y ) using a deterministic Newton method.
[0029] Surprisingly, it was found that this method is particularly efficient when an input signal with many important, i.e. heavily weighted, features is present at the input of the scaling layer.
[0030] Embodiments of the invention are explained in more detail below with reference to the accompanying drawings. In the drawings: Figure 1 schematically shows a structure of an embodiment of a control system; Figure 2 schematically shows an embodiment for controlling an at least partially autonomous robot; Figure 3 schematically shows an embodiment for controlling a manufacturing system; Figure 4 schematically shows an embodiment for controlling a personal assistant; Figure 5 schematically shows an embodiment for controlling an access system; Figure 6 schematically shows an embodiment for controlling a monitoring system; Figure 7 schematically shows an embodiment for controlling a medical imaging system; Figure 8 schematically shows a training system; Figure 9 schematically shows a structure of a neural network; Figure 10 schematically shows information forwarding within the neural network; Figure 11 in a flowchart an embodiment of a training method; Figure 12 in a flowchart an embodiment of a method for estimating a gradient;Figure 13 shows a flowchart of an alternative embodiment of the method for estimating the gradient; Figure 14 shows a flowchart of an embodiment of a method for scaling the estimated gradient; Figure 15 shows flowcharts of embodiments for implementing a scaling layer within the neural network. Figure 16 shows a flowchart of a method for operating the trained neural network. Description of the embodiments
[0031] Figur 1 shows an actuator 10 in its environment 20 interacting with a control system 40. Actuator 10 and environment 20 are collectively referred to as an actuator system. At preferably regular time intervals, a state of the actuator system is detected by a sensor 30, which may also be provided by a plurality of sensors. The sensor signal S—or, in the case of multiple sensors, one sensor signal S each—of the sensor 30 is transmitted to the control system 40. The control system 40 thus receives a sequence of sensor signals S. The control system 40 uses these to determine control signals A, which are then transmitted to the actuator 10.
[0032] The sensor 30 is any sensor that detects a state of the environment 20 and transmits it as a sensor signal S. It can be, for example, an imaging sensor, in particular an optical sensor such as an image sensor or a video sensor, or a radar sensor, or an ultrasonic sensor, or a LiDAR sensor. It can also be an acoustic sensor that receives, for example, structure-borne sound or voice signals. Likewise, the sensor can be a position sensor (such as GPS) or a kinematic sensor (for example, a single- or multi-axis acceleration sensor). A sensor that characterizes an orientation of the actuator 10 in the environment 20 (for example, a compass) is also possible. A sensor that detects a chemical composition of the environment 20, for example, a lambda sensor, is also possible.Alternatively or additionally, the sensor 30 may also comprise an information system that determines information about a state of the actuator system, such as a weather information system that determines a current or future state of the weather in the environment 20.
[0033] The control system 40 receives the sequence of sensor signals S from the sensor 30 in an optional receiving unit 50, which converts the sequence of sensor signals S into a sequence of input signals x (alternatively, the sensor signal S can also be directly adopted as the input signal x). The input signal x can, for example, be an excerpt or a further processing of the sensor signal S. The input signal x can, for example, comprise image data or images, or individual frames of a video recording. In other words, the input signal x is determined as a function of the sensor signal S. The input signal x is fed to a neural network 60.
[0034] The neural network 60 is preferably parameterized by parameters θ, for example comprising weights w, which are stored in a parameter memory P and are provided by this.
[0035] The neural network 60 determines output signals y from the input signals x. Typically, the output signals y encode classification information of the input signal x. The output signals y are fed to an optional conversion unit 80, which determines therefrom control signals A, which are fed to the actuator 10 in order to control the actuator 10 accordingly.
[0036] The actuator 10 receives the control signals A, is controlled accordingly, and performs a corresponding action. The actuator 10 can include a control logic (not necessarily structurally integrated) that determines a second control signal from the control signal A, which is then used to control the actuator 10.
[0037] In further embodiments, the control system 40 comprises the sensor 30. In still further embodiments, the control system 40 alternatively or additionally comprises the actuator 10.
[0038] In further preferred embodiments, the control system 40 comprises one or more processors 45 and at least one machine-readable storage medium 46 on which instructions are stored which, when executed on the processors 45, cause the control system 40 to carry out the method for operating the control system 40.
[0039] In alternative embodiments, a display unit 10a is provided alternatively or in addition to the actuator 10.
[0040] Figur 2 shows an embodiment in which the control system 40 is used to control an at least partially autonomous robot, here an at least partially automated motor vehicle 100.
[0041] The sensor 30 may be one of the sensors associated with Figur 1 The sensors mentioned may preferably be one or more video sensors, preferably arranged in the motor vehicle 100, and / or one or more radar sensors, and / or one or more ultrasonic sensors, and / or one or more LiDAR sensors, and / or one or more position sensors (for example GPS).
[0042] The neural network 60 can, for example, detect objects in the environment of the at least partially autonomous robot from the input data x. The output signal y can be information that characterizes where objects are present in the environment of the at least partially autonomous robot. The output signal A can then be determined depending on this information and / or according to this information.
[0043] The actuator 10, preferably arranged in the motor vehicle 100, can be, for example, a brake, a drive, or a steering system of the motor vehicle 100. The control signal A can then be determined such that the actuator or actuators 10 are controlled such that the motor vehicle 100, for example, prevents a collision with the objects identified by the neural network 60, in particular if the objects belong to certain classes, e.g., pedestrians. In other words, the control signal A can be determined depending on the determined class and / or according to the determined class.
[0044] Alternatively, the at least partially autonomous robot may also be another mobile robot (not shown), for example, one that moves by flying, swimming, diving, or walking. The mobile robot may also be, for example, an at least partially autonomous lawnmower or an at least partially autonomous cleaning robot. In these cases, too, the control signal A can be determined such that the drive and / or steering of the mobile robot are controlled such that the at least partially autonomous robot, for example, prevents a collision with the objects identified by the neural network 60.
[0045] In a further alternative, the at least partially autonomous robot can also be a gardening robot (not shown) that uses an imaging sensor 30 and the neural network 60 to determine a type or condition of plants in the environment 20. The actuator 10 can then be, for example, a chemical applicator. The control signal A can be determined depending on the determined type or condition of the plants in such a way that an amount of chemicals corresponding to the determined type or condition is applied.
[0046] In yet further alternatives, the at least partially autonomous robot can also be a household appliance (not shown), in particular a washing machine, a stove, an oven, a microwave, or a dishwasher. The sensor 30, for example an optical sensor, can detect a state of an object treated by the household appliance, for example, in the case of the washing machine, a state of laundry located in the washing machine. The neural network 60 can then determine a type or state of this object and characterize it by the output signal y. The control signal A can then be determined in such a way that the household appliance is controlled depending on the determined type or state of the object. For example, in the case of the washing machine, it can be controlled depending on the material of the laundry inside.Control signal A can then be selected depending on which material of the laundry was determined.
[0047] Figur 3 shows an embodiment in which the control system 40 is used to control a production machine 11 of a production system 200 by controlling an actuator 10 controlling this production machine 11. The production machine 11 can be, for example, a machine for punching, sawing, drilling, and / or cutting.
[0048] The sensor 30 may be one of the sensors associated with Figur 1 The sensors mentioned may be, preferably, an optical sensor that detects, for example, properties of manufactured products 12. It is possible for the actuator 10 controlling the manufacturing machine 11 to be controlled depending on the determined properties of the manufactured product 12, so that the manufacturing machine 11 accordingly carries out a subsequent processing step for this manufactured product 12. It is also possible for the sensor 30 to determine the properties of the manufactured product 12 processed by the manufacturing machine 11 and, depending thereon, to adapt the control of the manufacturing machine 11 for a subsequent manufactured product.
[0049] Figur 4 shows an example in which the control system 40 is used to control a personal assistant 250. The sensor 30 can be one of the sensors associated with Figur 1 The sensors mentioned above may be acoustic sensors. The sensor 30 is preferably an acoustic sensor that receives voice signals from a user 249. Alternatively or additionally, the sensor 30 may also be configured to receive optical signals, for example, video images of a gesture from the user 249.
[0050] Depending on the signals from sensor 30, control system 40 determines a control signal A from personal assistant 250, for example, by the neural network performing gesture recognition. This determined control signal A is then transmitted to personal assistant 250, which is thus controlled accordingly. This determined control signal A can, in particular, be selected such that it corresponds to a presumed desired control by user 249. This presumed desired control can be determined depending on the gesture recognized by neural network 60. Control system 40 can then, depending on the presumed desired control, select control signal A for transmission to personal assistant 250 and / or select control signal A for transmission to the personal assistant according to the presumed desired control 250.
[0051] This corresponding control can, for example, include the personal assistant 250 retrieving information from a database and presenting it in a manner that the user 249 can understand.
[0052] Instead of the personal assistant 250, a household appliance (not shown), in particular a washing machine, a stove, an oven, a microwave or a dishwasher, can also be provided in order to be controlled accordingly.
[0053] Figur 5 shows an embodiment in which the control system 40 is used to control an access system 300. The access system 300 may include a physical access control, for example, a door 401.
[0054] The sensor 30 may be one of the sensors associated with Figur 1 The sensors mentioned may be, preferably, an optical sensor (for example, for capturing image or video data) configured to detect a face. This captured image can be interpreted by means of the neural network 60. For example, the identity of a person can be determined. The actuator 10 may be a lock that, depending on the control signal A, releases the access control or not, for example, opens the door 401 or not. For this purpose, the control signal A can be selected depending on the interpretation of the neural network 60, for example, depending on the determined identity of the person. Instead of physical access control, logical access control can also be provided.
[0055] Figur 6 shows an embodiment in which the control system 40 is used to control a monitoring system 400. Of the Figur 5 This embodiment differs from the illustrated embodiment in that, instead of the actuator 10, the display unit 10a is provided, which is controlled by the control system 40. For example, the neural network 60 can determine whether an object detected by the optical sensor is suspicious, and the control signal A can then be selected such that this object is displayed in highlighted color by the display unit 10a.
[0056] Figur 7 shows an example in which the control system 40 is used to control a medical imaging system 500, for example, an MRI, X-ray, or ultrasound device. The sensor 30 can, for example, be an imaging sensor, and the display unit 10a is controlled by the control system 40. For example, the neural network 60 can determine whether an area recorded by the imaging sensor is conspicuous, and the control signal A can then be selected such that this area is displayed in a highlighted color by the display unit 10a.
[0057] Figur 8 schematically shows an embodiment of a training system 140 for training the neural network 60 by means of a training method.
[0058] A training data unit 150 determines suitable input signals x, which are fed to the neural network 60. For example, the training data unit 150 accesses a computer-implemented database in which a set of training data is stored and, for example, randomly selects input signals x from the set of training data. Optionally, the training data unit 150 also determines desired, or "actual," output signals y T associated with the input signals x, which are fed to an evaluation unit 180.
[0059] The artificial neural network x is configured to determine corresponding output signals y from the input signals x supplied to it. These output signals y are fed to the evaluation unit 180.
[0060] The evaluation unit 180 can, for example, use a cost function (loss function) dependent on the output signals y and the desired output signals y T a performance of the neural network 60 can be characterized. The parameters θ can be dependent on the cost function be optimized.
[0061] In further preferred embodiments, the training system 140 comprises one or more processors 145 and at least one machine-readable storage medium 146 having stored thereon instructions which, when executed on the processors 145, cause the control system 140 to execute the training method.
[0062] Figur 9 shows an example of a possible structure of the neural network 60, which in the exemplary embodiment is provided as a neural network. The neural network comprises a plurality of layers S 1 , S 2 , S 3 , S 4 , S 5 to derive from the input signal x applied to an input of an input layer S1, to determine the output signal y, which is available at an output of an output layer S 5. Each of the layers S 1 , S 2 , S 3 , S 4 , S 5 is set up to generate a (possibly multidimensional) input signal x, z 1 , z 3 , z 4 , z 6 , which is located at an entrance of the respective layer S 1 , S 2 , S 3 , S 4 , S 5 is applied, a -- a (possibly multi-dimensional) output signal z 1 , z 2 , z 4 , z 5 , y to determine that at an output of the respective layer S 1 , S 2 , S 3 , S 4 , S 5. Such output signals are also called feature maps, especially in image processing. It is not necessary for the layersS 1 , S 2 , S 3 , S 4 , S 5 are arranged such that all output signals that enter further layers as input signals enter each layer from a preceding layer into an immediately subsequent layer. Alternatively, skip connections or recurrent connections are also possible. It is of course also possible for the input signal x to enter several of the layers, or for the output signal y of the neural network 60 to be composed of output signals from a plurality of layers.
[0063] The output layer S5 can be given, for example, by an Argmax layer (i.e. a layer which selects from a plurality of inputs with respective associated input values a designation of the input whose associated input value is the largest among these input values), one or more of the layers S 1 , S 2 , S 3 can be given, for example, by convolutional layers.
[0064] According to the invention, a layer S 4 is designed as a scaling layer, which is designed such that a signal at the input of the scaling layer ( S 4 ) input signal (x) to a signal at the output of the scaling layer ( S4 ) is mapped to the output signal (y) present at the output, such that the output signal (y) present at the output is a rescaling of the input signal (x), wherein parameters characterizing the rescaling can be fixedly specified. Embodiments of methods, the scaling layer S 4 can perform are listed below in connection with Figur 15 described.
[0065] Figur 10 schematically illustrates the information transfer within the neural network 60. Schematically shown here are three multidimensional signals within the neural network 60, namely the input signal x, as well as subsequent feature maps z 1 , z 2 . In the exemplary embodiment, the input signal x has a spatial resolution of n x 1 × n y 1 Pixels, the first feature map z 1 of n x 2 × n y 2 pixels, the second feature map z 2 of n x 3 × n y 3 pixels. In the exemplary embodiment, the resolution of the second feature map isz 2 is lower than the resolution of the input signal x, but this is not necessarily the case.
[0066] Also shown is a feature, e.g. a pixel, ( i , j ) 3 of the second feature map z 2 . If the function that the second feature map z 2 from the first feature map z 1, for example, by a convolutional layer or a fully connected layer, it is also possible that a plurality of features of the first feature map z 1 in determining the value of this characteristic ( i , j ) 3. However, it is also possible that only one feature of the first feature card z 1 in determining the value of this characteristic ( i , j ) 3.
[0067] By "input" it can advantageously be understood that there is a combination of values of the parameters that characterize the function with which the second feature map z 2 from the first feature map z 1 is determined, and from values of the first feature map z 1 such that the value of the feature ( i , j ) 3 depends on the value of the incoming attribute. The totality of these incoming attributes is Figur 10 referred to as area Be.
[0068] In determining each characteristic ( i,j ) 2 of the area Be, in turn, includes one or more features of the input signal x. The set of all features of the input signal x that are used to determine at least one of the features ( i,j ) 2 of the area Be are called the receptive field rF of the feature ( i,j ) 3 . In other words, the receptive field rF of the feature ( i ,j ) 3 all those features of the input signal x which are directly or indirectly (in other words: at least indirectly) involved in the determination of the feature ( i , j ) 3, ie whose values exceed the value of the characteristic ( i , j ) 3 can influence.
[0069] Figur 11 shows in a flowchart the sequence of a method for training the neural network 60 according to an embodiment.
[0070] First (1000) a training data set X including couples ( x i , y i ) from input signals x i and corresponding output signals y i provided. A learning rate η is initialized, for example to η = 1.
[0071] Furthermore, an initial quantity G and a second quantity N initialized, for example, if in step 1100 the Figur 12 illustrated embodiment of this part of the method is used. If in step 1100 the Figur 13 illustrated embodiment of this part of the method can be used to initialize the first set G and the second set N be waived.
[0072] The initialization of the first set G and the second set N can be done as follows: The first quantity G , which are those pairs ( x i , y i ) of the training dataset X that have already been drawn during a current epoch of the training procedure is initialized as an empty set. The second set N, those couples ( x i , y i ) of the training dataset X that have not yet been drawn during the current epoch, is initialized by selecting all pairs ( x i , y i ) of the training data set X be assigned.
[0073] Now (1100) is calculated using pairs ( x i , y i ) from input signals x i and corresponding output signals y i of the training data set X a gradient g the parameter regarding the parameters θ estimated, so g = ∇ θ . Examples of this method are described in connection with Figur 12 and 13 respectively.
[0074] Then (1200) optionally a scaling of the gradient g Examples of this method are described in connection with Figur 14 described.
[0075] Afterwards (1300) an optional adjustment of a learning rate η The learning rate can be η for example, a predefined learning rate reduction factor Dη (e.g. Dη = 1 / 10) (i.e. η ← η · Dη ), provided that the number of epochs passed through is divisible by a predefined number of epochs, for example 5.
[0076] Then (1400) the parameters θ using the determined and possibly scaled gradient g and the learning rate η updated. For example, the parameters θ replaced by θ - η · g.
[0077] It is now checked (1500) by means of a predefined convergence criterion whether the method has converged. For example, depending on an absolute change in the parameters θ (e.g. between the last two epochs) to decide whether the convergence criterion is met or not. For example, the convergence criterion can be met if and only if a L 2< -Norm about the change of all parameters θ between the last two epochs is smaller than a predefined convergence threshold.
[0078] If it is decided that the convergence criterion is met, the parameters θ as learned parameters, and the process ends. If not, the process branches back to step 1100.
[0079] Figur 12 illustrates in a flowchart an exemplary procedure for determining the gradient g in step 1100.
[0080] First (1110) a predefined number bs in pairs ( x i , y i ) of the training dataset X (without replacement) drawn, i.e. selected and added to a stack B (English: "batch"). The specified number bs is also called batch size. The batch B is initialized as an empty set.
[0081] For this purpose, it is checked (1120) whether the stack size bs is greater than the number of pairs ( x i , y i ), which in the second set N are present.
[0082] Is the stack size bs not greater than the number of pairs ( x i , y i ), which in the second set N are available, bs many couples ( x i , y i ) randomly from the second set N drawn (1130), thus selected and added to the stack B added.
[0083] Is the stack size bs greater than the number of pairs ( x i , y i ), which in the second set N are present, all pairs of the second set N, their number with s is drawn (1140), i.e. selected and added to the stack B added, and the remaining, i.e. bs - s many, from the first crowd G drawn, i.e. selected and added to the stack B added.
[0084] For all parameters θ then (1150) at step (1130) or (1140) it is optionally decided whether these parameters θ should be skipped in this training session or not. For example, for each shift ( S 1 , S 2 , ... , S 6 ) separately define a probability with which parameter θ this layer. For example, this probability can be set for the first layer ( S 1 ) 50% and reduced by 10% with each subsequent layer.
[0085] Using these defined probabilities, for each of the parameters θ decide whether he will be ignored or not.
[0086] For each pair ( x i , y i ) of the stack B Now (1155) it is optionally decided whether the respective input signal x i augmented or not. For each corresponding input signal x i to be augmented, an augmentation function is selected, preferably randomly, and applied to the input signal x i The augmented input signal x i then replaces the original input signal x i . If the input signal x i For an image signal, the augmentation function can be provided, for example, by a rotation by a predeterminable angle.
[0087] Then (1160) for each pair ( x i , y i ) of the stack B the corresponding (and possibly augmented) input signal x i selected and fed to the neural network 60. The parameters to be transferred θ of the neural network 60 are deactivated during the determination of the corresponding output signal, e.g., by temporarily setting them to the value zero. The corresponding output signal y ( x i ) of the neural network 60 is assigned to the corresponding pair ( x i , y i ) Depending on the output signals y ( x i ) and the respective output signals y i of the couple ( x i , y i ) as the desired output signal y T a cost function is determined.
[0088] Then (1165) for all pairs ( x i , y i ) of the stack B together the complete cost function = Σ i ∈ B and each of the parameters that cannot be skipped θ the corresponding component of the gradient g e.g., using backpropagation. For each of the parameters to be passed θ the corresponding component of the gradient g set to zero.
[0089] Now it is checked (1170) whether the check in step 1000 determined that the stack size bs is greater than the number of pairs ( x i , y i ), which in the second set N are present.
[0090] It was determined that the stack size bs not greater than the number of pairs ( x i , y i ), which in the second set N are present, (1180) all pairs ( x i , y i ) of the stack B the first set G added and from the second set N removed. It is now checked (1185) whether the second set N is empty. If the second set N empty, a new epoch begins (1186). For this purpose, the first set G initialized again as an empty set, and the second set N is reinitialized by entering all pairs ( x i , y i ) of the training data set X and the process branches to step (1200). If the second set N not empty, branches directly to step (1200).
[0091] It was determined that the stack size bs is greater than the number of pairs ( x i , y i ), which in the second set N are present, the first set G reinitialized (1190) by removing all pairs ( x i , y i ) of the stack B be assigned, the second set N is reinitialized by entering all pairs ( x i , y i ) of the training data set X and then the pairs ( x i , y i ), which are also in the stack B are removed. A new epoch then begins, and the process branches to step 1200. This ends this part of the procedure.
[0092] Figur 13 illustrates in a flowchart another exemplary method for determining the gradient g in step 1100. First, parameters of the method are initialized (1111). In the following, the mathematical space of the parameters θ with WThe parameters include θ so np many individual parameters, the space W a np -dimensional space, for example W = ℝ np . An iteration counter n is set to the value n = 0 initialized, a first size m 1 is then used as m 1 = 0 ∈ W (also as np -dimensional vector), a second size m 2 = 0 ∈ W ⊗ W (also as np × np -dimensional matrix).
[0093] Then (1121) a pair ( x i , y i ) from the training dataset X selected and augmented if necessary. This can be done, for example, in such a way that for each input signal x i of the couples ( x i , y i ) training dataset X a number µ ( α ( x i )) possible augmentations α ( x i ) is determined, and each pair ( x i , y i ) a position size p i = ∑ j < i p j ∑ j p j is assigned. A random number is then φ ∈ [0; 1] uniformly distributed, the position size p i selected that the inequality chain p i ≤ φ < p i + 1 The corresponding index i then denotes the selected pair ( x i , y i ), an augmentation α i the input variable x i can be randomly selected from the set of possible augmentations α ( x i ) and applied to the input variable x i be applied, ie the selected pair ( x i , y i ) is replaced by ( α i ( x i ), y i ) replaced.
[0094] The input signal x i is fed to the neural network 60. Depending on the corresponding output signal y ( x i ) and the output signal y i of the couple ( x i , y i ) as the desired output signal y T the corresponding cost function determined. For the parameters θ a gradient in this regard d e.g. determined by backpropagation, i.e. d = ∇ θ ( y ( x i ), y i ).
[0095] Then (1131) iteration counters n , first size m 1 and second size m 2 updated as follows: n ← n + 1 t = 1 n m 1 = 1 − t ⋅ m 1 + t ⋅ d m 2 = 1 − t ⋅ m 2 + t ⋅ d ⋅ d T
[0096] Subsequently (1141) components C a,b a covariance matrix C provided as C a , b = 1 n m 2 − m 1 ⋅ m 1 T a , b .
[0097] From this, the (vector-valued) first quantity m 1 a scalar product S formed, so S = m 1 , C − 1 m 1 .
[0098] It is understood that for the sufficiently accurate determination of the scalar product S with equation (8) not all entries of the covariance matrix C or the inverse C-1< must be present simultaneously. It is more memory efficient to store the necessary entries during the evaluation of equation (8) C a,b the covariance matrix C to determine.
[0099] Then it is checked (1151) whether this scalar product S satisfies the following inequality: S ≥ λ 2 , where λ is a predeterminable threshold value that corresponds to a confidence level.
[0100] If the inequality is satisfied, the current value of the first quantity m 1 as estimated gradient g and the process branches back to step (1200).
[0101] If the inequality is not satisfied, the program can return to step (1121). Alternatively, it can also be checked (1171) whether the iteration counter n a preset maximum iteration value n max If this is not the case, the process branches back to step (1121), otherwise the estimated gradient g the zero vector 0 ∈ W (1181), and the process branches back to step (1200). This concludes this part of the procedure.
[0102] This procedure ensures that m 1 an arithmetic mean of the determined gradients d above the drawn pairs ( x i , y i ) and m 2 an arithmetic mean of a matrix product d · d T< the determined gradients d over the drawn pairs ( x i , y i ).
[0103] Figur 14 shows an embodiment of the method for scaling the gradient g in step (1200). In the following, each component of the gradient g with a couple ( , l ), where ∈ {1, ..., k} a layer of the corresponding parameter θ referred to, and l ∈ {1, ..., dim( V i )} a numbering of the corresponding parameter θ within the -th layer. Is the neural network as in Figur 10 illustrated for processing multidimensional input data x with corresponding feature maps in the -th layer, the numbering is l advantageously by the position of that feature in the feature map given with which the corresponding parameter θ is associated.
[0104] Now (1220) for each component of the gradient g a scaling factor For example, this scaling factor can by the size of the receptive field rF of the l corresponding feature of the feature map of the -th layer. The scaling factor can alternatively be determined by a ratio of the resolutions, i.e. the number of features, the -th layer in relation to the input layer.
[0105] Then (1220) each component of the gradient g with the scaling factor scaled, so g ι , l ← g ι , l / Ω ι , l .
[0106] Is the scaling factor given by the size of the receptive field rF, overfitting of the parameters θ avoid. If the scaling factor Given the ratio of the resolutions, this is a particularly efficient approximate estimate of the size of the receptive field rF.
[0107] Figur 15 illustrates the process that the scaling layer S 4 is executed.
[0108] The scaling layer S4 is set up to project the input of the scaling layer S 4 applied input signal x to a sphere with radius ρ and center c This is characterized by a first standard N 1 ( y - c ), which defines a distance of the center c from the output of the scaling layer S 4 applied output signal y and a second standard N 2 ( x - y ), which has a distance of the input of the scaling layer S 4 input signal x from the output of the scaling layer S 4 applied output signal y. In other words, this triggers the output of the scaling layer S 4 applied output signal y the equation y = argmin N 1 y − c ≤ ρ N 2 x − y .
[0109] Figur 15 a) illustrates a particularly efficient first embodiment in the case that the first standard N 1 and a second standard N2 are equal. They are denoted by ∥·∥ in the following.
[0110] First (2000) a value at the input of the scaling layer S 4 applied input signal x, a center parameter c and a radius parameter ρ provided.
[0111] Then (2100) a signal is applied to the output of the scaling layer S 4 applied output signal y determined to y = c + ρ ⋅ x − c max ρ x − c .
[0112] This concludes this part of the procedure.
[0113] Figuren 15b) und 15c ) illustrate embodiments for particularly advantageously selected combinations of the first standard N 1 and the second standard N 2 .
[0114] Figur 15 b) illustrates a second embodiment for the case that in the condition (12) to be fulfilled the first norm N 1 (·) is given by the maximum norm ∥·∥ ∞ and the second norm N2 (·) is given by the 2-norm ∥·∥ 2 . This combination of norms is particularly efficient to compute.
[0115] First (3000) analogous to step (2000) the input of the scaling layer S 4 applied input signal x, the center parameter c and the radius parameter ρ provided.
[0116] Then (3100) the components y i of the output of the scaling layer S 4 applied output signal y is determined to y i = c i + ρ falls x i − c i > ρ c i − ρ falls x i − c i < − ρ , x i sonst where i the components are referred to here.
[0117] This method is particularly computationally efficient. This concludes this part of the process.
[0118] Figur 15 c) illustrates a third embodiment for the case that in the condition (12) to be fulfilled the first norm N 1 (·) is given by the 1-norm ∥·∥ 1 and the second norm N2 (·) is given by the 2-norm ∥·∥ 2 . This combination of norms leads to the fact that at the input of the scaling layer S 4 input signal x as many small components as possible are set to the value zero.
[0119] First (4000) analogous to step (2000) the input of the scaling layer S 4 applied input signal x, the center parameter c and the radius parameter ρ provided.
[0120] Then (4100) a sign quantity ε i determined to ϵ i = + 1 falls x i ≥ c i − 1 falls x i < c i and the components x i of the input of the scaling layer S 4 applied input signal x are replaced by x i ← ϵ i ⋅ x i − c i .
[0121] An auxiliary parameter γ is initialized to the value zero.
[0122] Then (4200) a set N determined as N = { i | x i > γ } and a distance measure D= Σ i ∈ N ( x i - γ ).
[0123] Then (4300) it is checked whether the inequality D > ρ is fulfilled.
[0124] If this is the case (4400), the auxiliary parameter γ replaced by γ ← γ + D − ρ N , and it branches back to step (4200).
[0125] If inequality (16) is not satisfied (4500), the components y i of the output of the scaling layer S 4 applied output signal y determined to y i = c i + ϵ i ⋅ x i − γ +
[0126] The notation (·) + means in the usual way ξ + = ξ falls ξ > 0 0 sonst .
[0127] This concludes this part of the procedure. This procedure corresponds to a Newton method and is particularly computationally efficient, especially when many of the components of the input to the scaling layer S 4 applied input signal x are important.
[0128] Figure 16illustrates one embodiment of a method for operating the neural network 60. First (5000), the neural network is trained using one of the described methods. Then (5100), the control system 40 is operated with the thus trained neural network 60 as described. This concludes the method.
[0129] It is understood that the neural network is not limited to feedforward neural networks, but that the invention can be applied equally to any type of neural network, in particular recurrent networks, convolutional neural networks, autoencoders, Boltzmann machines, perceptrons or capsule neural networks.
[0130] The term "computer" encompasses any device capable of executing specified computational instructions. These computational instructions can be in the form of software, hardware, or a combination of software and hardware.
[0131] It is further understood that the methods may not only be implemented entirely in software as described. They may also be implemented in hardware, or in a hybrid form of software and hardware.
Claims
1. Computer-implemented method for classifying input signals (x), which were ascertained dependent on an output signal (S) of a sensor (30), by means of a neural network (60), wherein the sensor (30) is an imaging sensor, in particular an optical sensor such as an image sensor or a video sensor, or a radar sensor or an ultrasonic sensor or a LiDAR sensor or an acoustic sensor or a position sensor or a kinematic sensor, wherein a control signal (A) is ascertained dependent on an output signal (y) of the neural network (60), where an actuator (10) is controlled on the basis of the control signal (A), wherein the neural network (60) comprises a scaling layer (S4), wherein the scaling layer maps an input signal (z4) present at the input of the scaling layer (S4) onto an output signal (z5) present at the output of the scaling layer (S4), in such a way that this map corresponds to a projection of the input signal (z4) onto a predeterminable value range, characterized in that parameters (ρ, c) that characterize the map are predeterminable, and in that the map is represented by the equation z5 = argminN1(z5-c)≤ρN2(z4 - z5) with a first norm (N1) and a second norm (N2), and in that in a training phase the predeterminable parameters (ρ, c) were adapted on the basis of a method for training the neural network (60), and in that during training these predeterminable parameters (ρ, c), were adapted on the basis of an output signal (y) of the neural network (60) when an input signal (x) of the neural network (60) and an associated desired output signal (yT) were supplied, and the predeterminable parameters were adapted on the basis of an ascertained gradient (g) that depended on the output signal (y) of the neural network (60) and the associated desired output signal (yT).
2. Method according to Claim 1, wherein the first norm (N1) and the second norm (N2) are chosen to be the same.
3. Method according to Claim 1, wherein the first norm (N1) is an L∞ norm.
4. Method according to Claim 1, wherein the first norm (N1) is an L1 norm.
5. Method according to Claim 3 or 4, wherein the second norm (N2) is an L2 norm.
6. Method according to Claim 5, wherein the equation z5 = argminN1(z5-c)≤ρN2(z4 - z5) is solved by means of a deterministic Newton's method.
7. Computer program comprising instructions which, when the program is executed by a computer, cause said computer to carry out the method according to any of Claims 1 to 6.
8. Machine-readable storage medium (46, 146) on which the computer program according to Claim 7 is stored.
9. Actuator control system (40) configured to carry out the method according to any of Claims 1 to 6.