Method for training a neural network
By introducing a new regularization term that considers the convolution structure, calculating the norms of the convolution kernel and inputs to determine the convolution regularization value, the problem of insufficient generalization ability caused by the convolution layer structure in the prior art is solved, and better generalization ability of neural networks is achieved.
Patent Information
- Application Number
- CN202010616556.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2019-07-02
- Filing Date
- 2020-07-01
- Publication Date
- 2025-06-27
- Estimated Expiration
- 2040-07-01
AI Technical Summary
The prior art fails to effectively consider the structure of the convolutional layer when training neural networks, resulting in insufficient generalization capabilities.
A new regularization term is introduced, taking into account the convolution structure of the neural network, and the convolution regularization value is determined by calculating the convolution kernel norm and the input norm, so that regularization is performed when training the neural network.
By considering the structure of the convolutional layer, the generalization ability of the neural network during the training process is improved, making the performance of the neural network more stable on different data sets.
Smart Images

Figure CN112183740B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a method for training a neural network having convolutional layers and determining a convolutional regularization value. Background Art
[0002] For controlling at least partially automated systems, such as self-driving vehicles or robots as examples of mobile platforms, deep learning methods have been developed, in which neural networks are often used. Neural networks have shown excellent performance in various practical tasks such as speech recognition, machine translation, image classification, image segmentation, in playing video games, in playing board games, or in predicting protein binding. An important part of such at least partially automated systems is their ability to perceive complex situations related to their environment, so that many of the above examples can be adapted to this task.
[0003] A prerequisite for the safe and efficient operation of such at least partially automated systems is, for example, the interpretation of the environment of the mobile platform for, for example, decision-making processes such as trajectory planning and trajectory control of the mobile platform. Machine learning methods such as neural networks are increasingly used to solve such tasks. The task of machine learning methods is to learn, with the aid of training data, to detect the function of the training data as well as possible. A cost function is used here to evaluate the quality of the learned function. It has proven advantageous, when training such machine learning methods, to regularize the weights. This means that additional regularization terms are added additively to the cost function. Summary of the Invention
[0004] In the case of classification problems, adding additional regularization terms to the cost function results in having to minimize the sum of the loss function and the regularization terms when training the neural network. Examples of regularization terms are the so-called L1 or L2 regularization. A typical feedforward neural network can be understood as a sequence of blocks, each of which consists of a linear operation and a subsequent non-linearity. Such linear blocks can be described by matrices. In the case of convolutional networks, i.e., neural convolutional networks, these linear mappings are parameterized by weights having a certain number of indices.
[0005] L1 regularization sums the absolute values of the elements of the matrix, and L2 regularization squares the absolute values of the elements of the matrix before this summation. Thereby, the regularization terms are used to control the size of the function class. Here, the regularization factor balances the set of possible functions for solving the task of the neural network.
[0006] In the case of the hitherto known regularization terms, the "convolutional" structure of the neural network has not been taken into account. Typically, the weights of the layers are regularized instead of the linear mappings of the neural convolutional network.
[0007] According to the present invention, there are described a method for training a neural network using a large number of training runs, the use of a neural network, a neural network, a device, and a computer program product and a computer-readable storage medium, which features at least partially have the effects mentioned. Advantageous configurations are the subject of the following description.
[0008] The present invention is based on the knowledge that the regularization term should take into account the convolutional structure of the neural network such that the linear mapping of the convolutional layer is regularized. The present invention introduces a new regularization term that takes into account the convolutional structure of the neural network.
[0009] According to one aspect of the present invention, there is described a method for training a neural network using a large number of training runs, wherein the neural network has at least one convolutional layer, and wherein at least one convolution is performed by means of the convolutional layer in at least one training run.
[0010] In a step of the method, a plurality of input feature maps are provided from the output of the previous layer of the neural network.
[0011] In a further step, at least one output feature map of the convolutional layer is formed by means of a plurality of convolutional kernels and the plurality of input feature maps, wherein a convolutional kernel is assigned to each combination consisting of one of the input feature maps and each output feature map of the at least one output feature map, and each of the convolutional kernels has a plurality of convolutional weights.
[0012] In a further step, a plurality of convolutional kernel norms are determined for all combinations consisting of the input feature maps and the output feature maps, wherein the convolutional kernel norms are formed according to the absolute values of the convolutional weights of the convolutional kernels.
[0013] In a further step, a plurality of input norms are determined for all output feature maps, wherein the input norms are formed according to the convolutional kernel norms of all input feature maps, and the definition of the convolutional kernel norms is different from the definition of the input norms. In a further step, a convolutional norm is determined according to the input norms of all output feature maps in order to determine a convolutional regularization value for other training runs.
[0014] Neural networks provide a framework for many different algorithms for machine learning, collaboration, and processing complex data inputs. Such neural networks learn to perform tasks according to examples without typically programming the neural network with task-specific rules.
[0015] Such a neural network is based on a collection of connected units or nodes called artificial neurons. Each connection can transmit a signal from one artificial neuron to another. The artificial neuron receiving the signal can process the signal and then activate other artificial neurons connected to that artificial neuron.
[0016] In the case of a conventional implementation of a neural network, the signals are real numbers at the junctions of the artificial neurons, and the output of an artificial neuron is calculated by a non-linear function of the sum of the inputs to that artificial neuron. The connections of the artificial neurons typically have weights that are adapted as the learning progresses. The weights increase or decrease the strength of the signal at the junction. An artificial neuron can have a threshold such that a signal is only output if the total signal exceeds that threshold.
[0017] Typically, a large number of artificial neurons are grouped into layers. Different layers perform possibly different types of transformations on their inputs. Possibly after passing through these layers multiple times, the signal migrates from the first layer (i.e., the input layer) to the last layer (i.e., the output layer).
[0018] As a supplement to the account of neural networks, the structure of an artificial Convolutional Neural Network consists of one or more convolutional layers, optionally followed by Pooling Layers. A sequence of layers can be used with or without normalization layers (e.g., batch normalization), zero-padding layers, dropout layers, and activation functions (e.g., the rectified linear unit ReLU, sigmoid function, tanh function, or softmax function).
[0019] In principle, these units can be repeated any number of times until repeated enough times and then called a deep convolutional neural network. Such a convolutional neural network can have a sequence of the following layers that scan an input grid down to a lower resolution in order to obtain the desired information and store redundant information.
[0020] The data or signal at the input of such a neural network can be divided into coordinate data and feature data, where the feature data is assigned to the coordinate data. In the case of a convolution operation running in a convolutional layer, the amount of the coordinate data becomes smaller, and the amount of the feature data assigned to the coordinate data typically increases. Here, the feature data is typically aggregated in so-called Feature-Maps within the layers of the neural network.
[0021] The last convolutional layer extracts the most complex features, which are arranged in multiple feature maps and generate an output image when an input image or input signal is applied at the input end of the neural network. In addition, the last convolutional layer preserves the spatial information that may be lost in subsequent fully connected layers when necessary, where the fully connected layers are used for classification.
[0022] An "autoencoder" should be understood as an artificial neural network that enables the learning of specific patterns contained in input data. Autoencoders are used to generate a compressed or noise-free representation of the input data by correspondingly extracting important features (such as specific categories) from a generalized background.
[0023] Autoencoders use three or more layers:
[0024] · An input layer, such as a two-dimensional image.
[0025] · Multiple significantly smaller layers that form an encoding for reducing data.
[0026] · An output layer whose dimension corresponds to that of the input layer, i.e., each output parameter in the output layer has the same meaning as the corresponding parameter in the input layer. Autoencoders can also have convolutional layers, and the method according to the present invention can be used to train autoencoders.
[0027] A feedback neural network (English: Recurrent Neural Network, RNN, recursive neural network) is a neural network that, contrary to a feedforward network, also has connections from the neurons of one layer to the neurons of the same layer or a previous layer. Here, this structure is particularly suitable for discovering time-coded information in data.
[0028] When training a neural network, typically a distinction is made between a training phase and a testing phase, which is also referred to as a propagation phase. In the training phase, which consists of a large number of training runs, the neural network learns based on a training data set. Thus, the weights between the individual neurons are usually modified. Here, the learning rule describes how the neural network makes these changes. In the case of supervised learning (monitored or supervised learning), the correct output is pre-given as a "teaching vector", and based on this teaching vector, the parameters of the neural network or weights such as convolutional kernel weights are optimized. In contrast, no parameters or weights are changed during the testing phase. Instead, here it is checked whether the network has learned correctly based on the modified weights from the training phase. For this purpose, data is presented at the input of the neural network and it is checked which outputs the neural network has calculated. Here, it is checked whether the neural network has detected the training material using the output stimuli that have been shown to the neural network. By presenting new stimuli, it can be determined whether the network has solved the task in a generalized manner.
[0029] Since the structure of the convolutional layer is taken into account when determining the convolutional regularization value in the method, improved generalization can be achieved when training the neural convolutional network with regularization using the convolutional regularization value as an additional additive term to the cost function (English: loss - function).
[0030] The cost function measures how well the currently existing neural network maps a given data set. When training a neural network, the weights are gradually changed so that the cost function is made minimal and thus the training data set is (almost) completely mapped by the neural network. The task of the regularization term is to control the number of solutions determined in this way, which are calculated by minimizing the cost function. The higher the regularization term, the fewer solutions to the minimization task, which allows for improved generalization to be expected. The convolutional regularization term proposed in the present invention takes into account the convolutional structure of the neural network in a special way.
[0031] A neural convolutional network trained in such a way can contribute to improved generalization in a large number of applications. For example, such a neural convolutional network can be trained for tasks where the input signal of the neural network is one - dimensional, for example in the case of audio signals including speech or noise processing, two - dimensional, for example for evaluating, classifying or processing images (including scans from a lidar (LIDAR) system or a radar (RADAR) system), or in the case of three - dimensional input signals, for example for the analysis and classification of magnetic resonance tomography (MRT) scans. For these different applications of the described method, the method for training a neural network using the convolutional regularization value can be adapted accordingly, as will be further described below.
[0032] Throughout the description of the present invention, the sequence of method steps is shown in a manner that facilitates understanding of the method. However, those skilled in the art will recognize that many of these method steps can also be traversed in a different order and result in the same or corresponding outcomes. In this sense, the order of these method steps can be changed accordingly. Numbers are assigned to several features to improve readability or make the assignment more explicit, but this does not imply the existence of specific other features.
[0033] According to one aspect, it is proposed that the input feature map and the output feature map each have a feature component and a coordinate component. Here, the coordinate component describes the spatial dimension, and the feature component carries information related to the features at the spatial position described by the coordinate component of the input signal or input feature map or output feature map of the neural network.
[0034] According to one aspect, it is proposed that the input feature map and the output feature map have different feature component dimensions. As an example, in particular, values of three feature components (such as the intensities of red, green, and blue) can be assigned to the input feature map, and values of a two-dimensional image with X and Y coordinates can be assigned to the input feature map as the coordinate component. Then, the output feature map can have, for example, other feature maps (Feature-Maps) that extract structures from the image of the input feature map.
[0035] According to another aspect, it is proposed that the convolution regularization value is determined by multiplying the convolution norm by the frequency of all convolution kernels used to calculate the output map. By additionally multiplying the convolution norm by the frequency of all convolution kernels used to calculate the output map, the structure of the neural network is considered in a more detailed manner to achieve a further improvement in the generalization of the neural network trained in this way.
[0036] According to another aspect, in order to determine the frequency of the convolution kernels used to calculate at least one output map, the spatial size of the input feature map is compared with the spatial size of the convolution kernel.
[0037] According to another aspect, the total convolution regularization value for the neural network is determined by determining the convolution regularization value for each layer of the neural network and adding the sum of these convolution regularization values for all layers to the cost function used to train the neural network.
[0038] It can thus be achieved that the structure of all convolution layers of the neural network is considered when calculating the convolution regularization value for improving the generalization of the neural network, and the method described determines a regularization term corresponding to the normal regularization term when applied to other layers of the neural network.
[0039] In particular, the method can be performed in such a way that a total convolutional regularization value composed of the convolutional regularization values of the respective selected layers of the neural network is determined. This results in the possibility of adapting the training of the neural network to special requirements.
[0040] According to another aspect, it is proposed that after at least one training step, at least some parameters and / or weights of the neural network are adapted by applying the convolutional regularization value to the determination of the cost function. As already described above, applying the convolutional regularization value to the determination of the cost function is adding the convolutional regularization value to the normal cost function. During the training of the neural network, the parameters or convolutional weights of the neural network are adapted by means of the cost function.
[0041] In the case of a classification problem, this leads to the following minimization task:
[0042]
[0043] where V represents the cost function, R represents the regularization term, with a freely selectable or optimizable parameter λ. Here, the function f describes the neural network. The parameters x i , y i represent the training data. Since the function f(x i ) describes the neural network, f(x i ) should be equal to y i . The loss function V is used to calculate the difference between f(x i ) and y i . The losses of all training data are summed in the sum.
[0044] According to another aspect, it is proposed that the parameters and / or convolutional weights are adapted in each training run until a quality criterion is met. If a certain quality criterion is met, other data can be used to check whether the neural network has been generalized and thus has been sufficiently trained.
[0045] According to another aspect, the neural network is a network for classifying input data. When classifying input data by means of a neural network, convolutional layers are typically used for the classification, so that the generalization improved by means of the convolutional regularization value can improve the classification by means of the neural network.
[0046] According to another aspect, it is proposed to add a convolutional regularization value to the cost function at least every two training steps. To adapt the training to the specific requirements of the neural network, it may be advantageous to add the convolutional regularization value to the cost function only every second training step. In particular, the convolutional regularization value can be freely chosen to be added to the cost function in a large number of steps to train the neural network. For this purpose, additional weights can be added to the convolutional regularization value, and the additional weights can change their magnitudes, for example, according to the classification quality in the current step. With such additional factors, the convolutional regularization value can also be added only every second step or at other regular intervals.
[0047] According to another aspect, it is proposed that the convolutional kernel norm, the input norm, and the convolutional norm are L_p norms, and at least two of these norms have different values for the parameter p. By using different parameters when selecting these norms, the convolutional regularization value can map the structure of the convolutional operation and thus improve the generalization of the neural network.
[0048] According to another aspect, it is proposed that the convolutional kernel norm is the L1 norm, the input norm is the L2 norm, and the convolutional norm is the L1 norm. For this choice of these norms, it can be shown that the structure of the convolutional layer is better considered.
[0049] According to one aspect, it is proposed that the previous layer is formed by the input layer of the neural network.
[0050] According to one aspect, the dimension of the coordinate components of the input feature map of the neural network is 1, 2, or 3. Here, one-dimensional input data can correspond to an audio signal, two-dimensional input data can correspond to an image, and three-dimensional input data can be suitable for magnetic resonance tomography (MRT) scan imaging.
[0051] According to one aspect, the neural network is a feedforward network, a recurrent network, a neural convolutional network, an autoencoder network, a multi-layer network, or an encoder-decoder network. Therefore, the method can be advantageously applied to a large number of types of neural networks.
[0052] According to one aspect, the subsequent network layer is a pooling layer or non-linear.
[0053] According to another aspect, the convolutional regularization value for the two-dimensional input data of the neural network is determined by the relationship described in Equation 1:
[0054]
[0055] The variables α1 and α2 of the convolutional regularization value ||A|| describe the size of the input graph with a first dimension and a second dimension; k and l are kernel indices extending to κ and ι and thus describe the size of the kernel in two dimensions; i and j are indices of the output feature map and the input feature map, and i and j extend to γ and η. In addition, g ijkl is the weight of the convolutional kernel, and p and q represent different norms.
[0056] A feedforward neural network can be understood as a chain of linear operations A and element-wise non-linearities.
[0057] f(x) = φ L (A( L (φ L-1 (A L-1 (...)))))
[0058] Another way of writing this is:
[0059]
[0060] Define the following parameters for each layer. The layer receives η features as input signals and outputs γ features. The spatial part of the layer is described by the variable α1 × α2 ×... × α D For sound signals, for example, D = 1, and for images, for example, D = 2.
[0061] The linear operation Ak of the convolutional network can now be written as a sum of tensor products. For two-dimensional signals, for example, the following applies:
[0062]
[0063] Here, g ijkl is the weight of the layer. To define this matrix, auxiliary matrices must also be introduced:
[0064]
[0065] Here, the matrix consists of a zero block (with i columns), an (α - κ + 1) × (α - κ + 1) identity matrix, and again a zero block (with κ - ι + 1 columns). Thus, these matrices are given by:
[0066] R i := [0,...,1,..,0]
[0067] S J := [0,...,1,...,0] T
[0068] U k+1 := P(α,κ,k)
[0069] Tl+1 : =P(β, ι, l)
[0070] According to the weight g of the neural network ijkl The following function is calculated for each layer:
[0071]
[0072] This term is for each layer of linear operator A k Determine and sum and add to the cost function as a regularization term. The following cost function is used to train the network value:
[0073]
[0074] Typically, p=2 and q=1 are chosen.
[0075] According to another aspect, a convolution regularization value ||A|| is determined for the D-dimensional input data of the neural network by the relationship described in Formula 2:
[0076]
[0077] If a higher stride, dilation or padding is used in the neural network, the frequency with which the respective convolution weights are used changes and the pre-factor (α1-κ1+1)...(α D -κ D +1).
[0078] To adapt the prefactor, the frequency with which the weight is "used" in the multiplication during the convolution is selected. That is, in a convolution with a stride of 2 the prefactor is halved
[0079]
[0080] A neural network which has been trained according to one of the above-described methods is described. This is in particular a neural convolutional network, since the described convolution regularization values take into account the structure of the convolutional layers contained in the neural convolutional network.
[0081] The use of a neural network that has been trained according to one of the above-described methods for classifying or detecting signals is described. In particular, a neural network trained in this way can be used for characterizing the vehicle environment, for tomographic imaging or for object detection.
[0082] There are a large number of different applications for this method of training a neural network using convolution regularization values, to which the method must be adapted as necessary. In this case, the method must be adapted to the different dimensions of the application of the neural network in the manner described below.
[0083] The application of the method described herein and the neural network trained using the method can be used individually or in combination, for example, when representing an environment of at least a partially automated mobile platform.
[0084] Here, a mobile platform can be understood as a mobile at least partially automated system and / or a driver assistance system of a vehicle. An example can be an at least partially automated vehicle or a vehicle with a driver assistance system. That is, in this context, in terms of at least partially automated functions, the at least partially automated system includes a mobile platform, but the mobile platform also includes vehicles and other mobile machines including driver assistance systems. Other examples of mobile platforms can be driver assistance systems with multiple sensors, mobile multi-sensor robots (such as robotic vacuum cleaners or lawn mowers), multi-sensor monitoring systems, manufacturing machines, personal assistants, or access control systems. Each of these systems can be a fully or partially automated system.
[0085] If the application of the neural network involves one-dimensional signals, such as audio signals including speech signals or motor noise, the feature map of the input signal can be represented as a vector. Thus, the convolutional kernel has components and the convolutional regularization value is determined by the following formula
[0086]
[0087] where the definitions of the parameters are as shown above. In this way, for example, what was said or who said it can be identified.
[0088] For two-dimensional signals, such as color images (RGB images), X-ray images, scans of lidar (LIDAR) systems, or radar maps, the feature map of the input signal is a matrix. Thus, the kernel has components and the convolutional regularization value is determined by the following formula
[0089]
[0090] where the definitions of the parameters are as shown above. As an example of the application of two-dimensional signals, image signals can be detected, such as traffic sign recognition or object recognition (such as pedestrian recognition). Then, the task of the neural network is to determine whether a traffic sign or a pedestrian can be detected in the image of the environment and, if necessary, classify the traffic sign or the pedestrian.
[0091] For three-dimensional signals such as magnetic resonance tomography (MRT) scans, the feature map of the input signal can be represented in matrix form. Thereby, the convolutional kernel has components and the convolutional regularization value can be calculated in the following form:
[0092] ||A|| γ,η:p,q :=(α1-κ1+1)(α2-κ2+1)(α3-κ3+1)
[0093]
[0094] The definition of the parameters is as shown above. As an application of a three-dimensional input signal, the tasks of the neural network may exemplarily include identifying tumors from MRI images and classifying tumor types.
[0095] According to another aspect, the use of the neural network for classification involves an input signal, wherein a control signal for controlling the at least partially automated vehicle and / or a warning signal for warning a vehicle occupant is sent according to the result of the classification. Examples of such input signals may be sound sources, or may be image signals. Thus, general input signals of different dimensions used with the method may be as described above.
[0096] A device is described which has a neural network which has been trained according to one of the methods described above.
[0097] A computer program is proposed, which comprises instructions which, when the program is executed by a computer, cause the computer to carry out one of the above-mentioned methods.
[0098] A machine-readable storage medium is proposed, on which the described computer program is stored. BRIEF DESCRIPTION OF THE DRAWINGS
[0099] The embodiments of the present invention are Figures 1 to 3 and explained in more detail below.
[0100] Figure 1 It shows forming an output feature map with the aid of a plurality of input feature maps;
[0101] Figure 2 At least one output feature map of a convolution layer is formed according to a plurality of input feature maps by means of a plurality of convolution kernels; and
[0102] Figure 3 The method steps for determining a convolution norm with the aid of a convolution kernel are shown. DETAILED DESCRIPTION
[0103] exist Figure 1Exemplarily schematically shows how at least one output feature map 150 of the convolutional layer is formed from a plurality of input feature maps 110, 120, 130 provided according to the output of the previous layer of the neural network. Here, for different positions constituted by the values of the i-th convolutional kernels 115, 125, 135 and the corresponding positions of the pixels 155, each pixel 155 of the i-th output feature map 150 is calculated according to the input feature maps 110, 120, 130.
[0104] In Figure 2 Schematically shows how in method step S2, at least one output feature map 150 of the convolutional layer is formed from a plurality of input feature maps 110, 120, 130 provided according to the output of the previous layer of the neural network in method step S1, each by means of a plurality of convolutional kernels 115, 125, 135; 116, 126, 136; or 117, 127, 137 and the plurality of input feature maps 110, 120, 130, wherein a convolutional kernel 115, 125, 135; 116, 126, 136; or 117, 127, 137 is assigned to each combination constituted by one of the input feature maps 110, 120, 130 and each output feature map in the at least one output feature map 150, and each of the convolutional kernels 115, 125, 135; 116, 126, 136; or 117, 127, 137 has a plurality of convolutional weights.
[0105] Furthermore, in Figure 2 Indicates how in an exemplary case the frequency of all convolutional kernels for calculating the output map can be determined. The starting point is the size of the input feature maps 110, 120, 130 indicated by two spatial coordinates, that is, by the coordinate components of the input feature maps - in this two-dimensional example by the width 170 of the input feature map 110 and the height 160 of the input feature map 110. In this example, the extent of the convolutional kernel 115 is represented by the width 145 and the height 140. As schematically shown by three positions 115, 116, 117 of the same convolutional kernel in the first input map 110, when comparing the spatial dimensions of the convolutional kernel 115 and the spatial dimensions of the input map 110, the defined frequency of applying the convolutional kernel 115 to the input feature map is obtained so as to apply the convolutional kernel 115 to the entire input feature map 110 without overlap. Thus, for example, this frequency is calculated as (α1 - k + 1)(α2 - ι + 1).
[0106] Here, α1 or α2 is the extent 160, 170 of the described input feature map 110, while κ or ι is the extent 140, 145 of the described convolutional kernel 115.
[0107] In Figure 3Schematically shows how the convolution kernel norms 320a to 320g are determined from the convolution kernels 310a to 310g of the method in method step S3, where the convolution kernel norms 320a to 320g are formed based on the absolute values of the convolution weights of the convolution kernels 310a to 310g. This thus corresponds to a spatial aggregation by forming the convolution kernel norms 320a to 320g.
[0108] Subsequently, the convolution kernel norms 320a to 320g can be multiplied by the frequencies of all the convolution kernels used to calculate the output map, which is indicated by the set of convolution kernel norms 330a to 330g in Figure 3 This multiplication method step is optional for determining the convolution norm or the convolution regularization value and can also be performed at other positions in the method.
[0109] In method step S4, the input norms 340a to 340d are determined based on the set of convolution kernel norms 320a to 320g, 330a to 330g, where the input norms are determined based on the convolution kernel norms 320a to 320g, 330a to 330g of all the input feature maps 110, 120, 130, and the definition of the convolution kernel norms is different from the definition of the input norms. In other words, this corresponds to an aggregation along the input dimension of the feature maps.
[0110] In method step S5, the convolution norm 350 for determining the convolution regularization value is determined based on the input norms 340 of all the output feature maps (e.g., input norms 340a to 340d). In other words, this corresponds to an aggregation along the output dimension of the feature maps.
Claims
1. A method for training a neural network using a large number of training runs, wherein the neural network has at least one convolutional layer, and at least one convolution is performed by means of the convolutional layer in at least one training run, wherein, The input signals of the neural network include one-dimensional input signals, two-dimensional input signals, and / or three-dimensional input signals. Among them, the one-dimensional input signal is an audio signal, the two-dimensional input signal is an image, and the three-dimensional input signal is a magnetic resonance tomography. Wherein, the neural network is trained for the following tasks. When the input signal of the neural network is used one-dimensionally in the case of an audio signal, it includes speech or noise processing. When used two-dimensionally, it is used for evaluating, classifying, or processing images. Or in the case of a three-dimensional input signal, it is used for the analysis and classification of magnetic resonance tomography. The method has the following steps: Providing a plurality of input feature maps from the output of the previous layer of the neural network (S1); Forming at least one output feature map of the convolutional layer by means of a plurality of convolutional kernels and the plurality of input feature maps (S2), wherein a convolutional kernel is assigned to each combination formed by one of the input feature maps and each output feature map in the at least one output feature map, and each of the convolutional kernels has a plurality of convolutional weights; Determining a plurality of convolutional kernel norms for all combinations formed by the input feature maps and the output feature maps (S3), wherein the convolutional kernel norm is formed according to the absolute values of the convolutional weights of the convolutional kernel; Determining a plurality of input norms for all output feature maps (S4), wherein the input norm is formed according to the convolutional kernel norms of all input feature maps, and the definition of the convolutional kernel norm is different from the definition of the input norm; Determining the convolutional norm (S5) according to the input norms of all output feature maps in order to determine the convolutional regularization value for other training runs.
2. The method according to claim 1, wherein Determining the convolutional regularization value by multiplying the convolutional norm by the frequency of all convolutional kernels used to calculate the output feature map.
3. The method according to claim 2, wherein In order to determine the frequency of the convolutional kernel used to calculate the at least one output feature map, the spatial size of the input feature map is compared with the spatial size of the convolutional kernel.
4. The method according to claim 2 or 3, wherein Determining the total convolutional regularization value for the neural network by determining the convolutional regularization value for each layer of the neural network and adding the sum of the convolutional regularization values of all layers to the cost function used to train the neural network.
5. The method according to any one of claims 1 to 3, wherein, After at least one training step, adapting at least some of the parameters and / or weights of the neural network by applying the convolutional regularization value to the determination of the cost function.
6. The method according to any one of claims 1 to 3, wherein The convolutional kernel norm, the input norm, and the convolutional norm are L_p norms, and at least two of these norms have different values for the parameter p.
7. The method according to any one of claims 1 to 3, wherein The previous layer is formed by the input layer of the neural network.
8. The method according to any one of claims 1 to 3, wherein The neural network is a feedforward network, a recurrent network, a neural convolutional network, an autoencoder network, a multi-layer network, or an encoder-decoder network.
9. The method according to any one of claims 1 to 3, wherein Determining the convolutional regularization value ||A|| for the two-dimensional input data of the neural network through the relationship described in Equation 1: Among them, the variables α1 and α2 of the convolutional regularization value ||A|| illustrate the sizes of the input graphs with the first and second dimensions; k and l are kernel indices extending to κ and ι and thus illustrate the sizes of the kernel in two dimensions; i and j are indices of the output feature map and the input feature map, i and j extend to γ and η, and in addition, g ijkl is the weight of the convolutional kernel, and p and q represent different norms.
10. A neural network that has been trained according to the method of any one of the preceding claims 1 to 9.
11. Use of a neural network that has been trained according to the method of any one of claims 1 to 9.
12. Use of the neural network according to claim 11 for classifying an input signal, wherein, Send a control signal for controlling at least a partially automated vehicle and / or a warning signal for warning vehicle occupants according to the result of the classification.
13. A device having a neural network that has been trained corresponding to the method according to any one of claims 1 to 9.
14. A computer program product having a computer program that includes instructions that, when the computer executes the computer program, cause the computer to perform the method according to any one of claims 1 to 9.
15. A machine-readable storage medium storing a computer program that includes instructions that, when the computer executes the computer program, cause the computer to perform the method according to any one of claims 1 to 9.
Citation Information
Patent Citations
Neural network training method, device, computer system and mobile device
CN108496188A
A neural network pruning method based on rhombus convolution
CN109376859A