Apparatus and computer-implemented method for steering with artificial neural networks
By using a piecewise continuously differentiable activation function in a convolutional neural network, the problem of gradient instability in deep artificial neural networks is solved, improving the learning speed and the stability of manipulation, and it is suitable for feedback calculation in recurrent neural networks.
Patent Information
- Application Number
- CN202010331758.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2020-04-08
- Filing Date
- 2020-04-24
- Publication Date
- 2026-01-23
- Estimated Expiration
- 2040-04-24
AI Technical Summary
In existing technologies, deep artificial neural networks are prone to burst or vanishing activations during the training and inference phases, leading to instability in gradient methods and affecting the digital instability of model learning and manipulation, which is particularly pronounced in recurrent neural networks.
Employing activation functions that are at least piecewise continuously differentiable and monotonically increasing, with a finite or odd number of fixed points and alternating increasing and decreasing derivative regions, this method is used to determine the output parameters of convolutional neural networks, avoiding gradient bursts or vanishing, and is suitable for feedback calculations in recurrent neural networks.
It improves learning speed and training stability, ensuring numerical stability during training and manipulation phases, and is particularly suitable for feedback computation in recurrent neural networks.
Smart Images

Figure CN111860762B_ABST
Abstract
Description
Technical Field
[0001] This invention is based on a device and computer-implemented method for manipulation using artificial neural networks. Information is stored in the artificial neural network, which can map input parameters, such as those from sensors, to output parameters for manipulation. Background Technology
[0002] Activation functions are a crucial component of deep artificial neural networks. In certain architectures, and particularly in recurrent architectures of artificial neural networks, these activation functions lead to burst or vanishing activation. In this sense, burst or vanishing activation means that the states of the individual units of the neural network, i.e., neurons, take values that contribute to the digital instability of the neural network, both during training and during inference (i.e., the application phase). During training, burst or vanishing activation leads to digitally unstable gradients in the gradient method used to determine the parameters of the artificial neural network. These burst or vanishing gradients during training may cause the model, i.e., the parameters stored in the artificial neural network, to no longer be learned. Consequently, the information stored in the artificial neural network becomes unusable. The result can be undesirable manipulation. Furthermore, burst or vanishing activation itself can lead to undesirable characteristics of the system during manipulation, i.e., during the inference phase. This is particularly pronounced in recurrent neural networks, as these networks reinforce instability through their repetitive feedback computations.
[0003] It is hoped that this will provide a correspondingly improved control by utilizing artificial neural networks. Summary of the Invention
[0004] This is achieved through a computer-based method for running an image classifier using an artificial neural network, a device for manipulating an artificial neural network, a computer-readable storage medium, and a computer program product.
[0005] A computer-implemented method for running an image classifier using an artificial neural network, particularly a convolutional neural network (CNN), characterized in that an output parameter (608) of the artificial neural network is determined, wherein the output parameter is determined based on input data for the artificial neural network and an activation function of the artificial neural network and characterizes the classification of the input data, wherein the activation function is at least piecewise continuously differentiable and monotonically increasing, wherein the activation function has at least three fixed points, wherein the number of fixed points of the activation function is either finite and odd or countable and discrete, and wherein the derivative of the activation function is alternately monotonically increasing and monotonically decreasing in the region between adjacent fixed points.
[0006] In this case, the input data can be provided as image data, which may include, for example, frames of output signals from a video sensor, radar sensor, LiDAR sensor, or ultrasonic sensor.
[0007] In this context, classification can also be understood as region-by-region, especially pixel-by-pixel, classification as semantic segmentation, and / or as classification as detecting whether an object has been identified in an image fragment.
[0008] The computer-implemented method for controlling based on measured output parameters can also specify the use of control parameters to control robots, vehicles, household appliances, tools, production machines, human assistance systems, access control systems, or systems for transmitting information, wherein the control parameters are determined based on the output parameters.
[0009] In the case of recurrent neural networks (RNNs), where the RNN has feedback from its individual units to itself within a layer, an activation function is iteratively applied to the internal state of that layer. This activation function helps avoid vanishing or bursty gradients, especially when the individual units of the layer are fully interconnected and inputs from other layers of the RNN are also provided. A fixed point is understood as a point for which the activation function maps values x from the defined set of the activation function to the same values x from the range of values of the activation function. For example, the activation function has a first fixed point F1 for a first value x1 in the defined set of the activation function, a second fixed point F2 for a second value x2 in the defined set of the activation function, and a third fixed point F3 for a third value x3 in the defined set of the activation function. In this example, the first value x1 is less than the second value x2, and the second value is less than the third value x3. In this example, the derivative of the activation function is monotonically increasing in a first region between the first value x1 and the second value x2, and monotonically decreasing in a second region between the second value x2 and the third value x3. In this example, the first and second regions are adjacent. This means that the activation function has no other fixed points between the first value x1 and the second value x2, and no other fixed points between the second value x2 and the third value x3. Compared to regular activation functions, this type of activation function has the following advantages: the gradient of the activation function is almost entirely different from zero. This type of activation function is particularly advantageous for the application phase of recurrent neural networks, i.e., for the interferential phase, because these activation functions have stable convergence properties, especially in the case of repeated applications, i.e., repeated computation in a feedback manner. Furthermore, this type of activation function improves learning speed, especially during training. Fixed points can be easily pre-given. This makes iterative selection of the activation function particularly easy during training.
[0010] Preferably, the derivative of the activation function is less than 1 for any value less than the minimum value of the defined set, where the minimum value defines a fixed point for the activation function; or the derivative of the activation function is less than 1 for any value greater than the maximum value of the defined set, where the maximum value defines a fixed point for the activation function. This type of activation function is particularly suitable for replacing the conventional sigmoid function.
[0011] Preferably, the derivative is specified to be greater than zero, especially over the range of values for the defined set of fixed points where the activation function is not defined. Compared to conventional activation functions, this type of activation function has the advantage that its gradient is different from zero, at least outside the fixed points. This further improves the learning speed.
[0012] Preferably, the set of values for the activation function is defined by a range of numbers between a lower and upper limit, specifically by a range from zero (inclusive) to one (inclusive). This allows the activation function to be used directly in conventional artificial neural networks.
[0013] Preferably, the activation function is specified to be continuously differentiable. This allows for particularly efficient computation using gradient methods.
[0014] Preferably, the activation function is a smooth function or a piecewise linear function, wherein the fixed point having the minimum value of all fixed points of the activation function has a value of zero, and the fixed point having the maximum value of all fixed points has a value of 1. This function is particularly well-suited to replace the sigmoid function.
[0015] Preferably, the activation function is a piecewise linear function, wherein the activation function has at least two segments between two directly adjacent fixed points, the segments being linear functions with different slopes. This function achieves properties similar to rundung to a fixed point, but without influencing gradients through gradients with zero values. To this end, the activation function has a first slope in the first segment of the segment and a second slope in the second segment. The first slope is much smaller than the second slope. The first segment is much larger than the second segment. Thus, values from the defined range of the activation function in the first segment are mapped to values from the range of values of the activation function, which are very close to each other and very close to a fixed point.
[0016] Preferably, the value of the fixed point is a power of 2. This is especially suitable for assigning a value from the range of values of the activation function, particularly for two segments of a linear function with different slopes.
[0017] Specifications for a device for manipulation using an artificial neural network, the device comprising an output for manipulating robots, vehicles, household appliances, tools, production machines, human assistance systems, access control systems, or systems for transmitting information using manipulation parameters, wherein the device includes a computing device, a memory for the artificial neural network, and an input for input data, wherein the device is configured to determine the manipulation parameters according to the described method. The device is suitable for training artificial neural networks and for manipulation after successful training.
[0018] The present invention relates to a computer-readable storage medium, characterized in that the computer-readable storage medium includes computer-readable instructions, which, when executed by a computer, perform the method according to any one of the embodiments. Attached Figure Description
[0019] Other advantageous embodiments are derived from the following description and accompanying drawings. In the drawings...
[0020] Figure 1 A schematic diagram of a device with an artificial neural network is shown.
[0021] Figure 2 A schematic diagram of the first activation function is shown.
[0022] Figure 3 A schematic diagram of the second activation function is shown.
[0023] Figure 4 A schematic diagram of the third activation function is shown.
[0024] Figure 5 A schematic diagram of the fourth activation function is shown.
[0025] Figure 6 The steps in the method used for manipulation are shown. Detailed Implementation
[0026] exist Figure 1 The diagram schematically illustrates a device 100 with an artificial neural network.
[0027] The device 100 includes an output 102 for controlling a robot, vehicle, household appliance, tool, production machine, human assistance system, access control system or system for transmitting information using control parameters 104.
[0028] Device 100 includes a computing device 106, a memory 108 for an artificial neural network, and an input terminal 110 for input data 112. The input data is, for example, sensor data, which is submitted to the input layer of the artificial neural network. For this purpose, the input data is normalized, for example. In this example, output parameters are submitted to the output layer of the artificial neural network for the output terminal 102 to determine control parameters 104. Determining control parameters 104 is used, for example, to control an actuator.
[0029] Device 100 is configured to determine control parameters 104 according to the method described below. Device 100 is configured, on the one hand, to train an artificial neural network, and on the other hand, after successful training, to control robots, vehicles, household appliances, tools, production machines, assistive systems for humans, access control systems, or systems for transmitting information. Device 100 can also be configured in a self-learning manner during operation.
[0030] Artificial neural networks include, for example, deep artificial neural networks, especially those constructed as recurrent neural networks (RNNs). RNNs preferably have feedback in some layers. RNNs can be constructed as fully connected RNNs. They can also incorporate a large number of Long Short-Term Memory (LSTM) modules or gated recurrent unit (GRU) modules.
[0031] The artificial neural network uses an activation function in unit 114. This activation function can be the same for layer 116 or for all layers with activation functions, or different activation functions can be used in different units or layers. Feedback 118 can be set from a unit in one layer to a unit in a previous layer or from a unit in one layer to itself. Figure 1 In this diagram, the unit containing the transfer function is denoted by ∑. A unit 120 containing parameters is also provided. The parameters are the weights used in the artificial neural network. These parameters can be learned during training or, in the case of a self-learning artificial neural network, during runtime using gradient methods.
[0032] Figures 2 to 5 An exemplary activation function is shown. The activation function is plotted in a coordinate system in which the values y from the range of values of the activation function are plotted relative to the values x from the defined range of the activation function.
[0033] An activation function has at least three fixed points. The activation function can be piecewise continuously differentiable and monotonically increasing. The number of fixed points of the activation function can be either finite and odd (ungerade) or countable and discrete.
[0034] The derivative of the activation function alternately increases and decreases monotonically in the region between adjacent fixed points.
[0035] A fixed point is understood to be a point for which the activation function maps a value x from the definition set of the activation function to the same value x from the range of values of the activation function.
[0036] For example, Figure 2 The first activation function 200 shown has a first fixed point F1 when the first value x1 is defined in the first set of definitions of the first activation function 200. The first activation function 200 has a second fixed point F2 when the second value x2 is defined in the first set of definitions of the first activation function 200. The first activation function 200 has a third fixed point F3 when the third value x3 is defined in the first set of definitions of the first activation function 200.
[0037] In this example, the first value x1 at the origin of the coordinate system is less than the second value x2, and the second value is less than the third value x3.
[0038] In this example, the derivative of the first activation function 200 is monotonically increasing in the first region 202 between the first value x1 and the second value x2, and monotonically decreasing in the second region 204 between the second value x2 and the third value x3.
[0039] In this example, the first region 202 and the second region 204 are adjacent. This means that the first activation function 200 has no other fixed points between the first value x1 and the second value x2, and the first activation function 200 has no other fixed points between the second value x2 and the third value x3.
[0040] The first activation function 200 is continuously differentiable.
[0041] exist Figure 3 The second activation function 300 is shown. In this example, the second activation function 300 has the same fixed point as the first activation function 200. Unlike the first activation function 200, the second activation function 300 is a piecewise linear function.
[0042] The second activation function 300 has at least two segments between two directly adjacent fixed points, each segment being a linear function with a different slope. Figure 3 The diagram shows a first segment 302 between the first fixed point F1 and the second fixed point F2, and a second segment 304 between the second fixed point F2 and the third fixed point F3.
[0043] The first segment 302 is preferably adjacent to another segment with the same slope at the first fixed point F1. Preferably, the other segment with the same slope is adjacent to the second segment 304 at the third fixed point F3.
[0044] exist Figure 4 The third activation function 400 shown is piecewise linear and has the following fixed points, which in this example have integer values x. In this example, n fixed points F1, ... F2 are shown. n The slope of the piecewise linear third activation function 400 is described as for the second activation function 300. This creates alternating regions of monotonically increasing and monotonically decreasing values between fixed points.
[0045] exist Figure 5 The fourth activation function 500 shown is piecewise linear and has the following fixed points, which in this example have a value x, a power of 2 (Potenz von Zwei). In this example, n fixed points F1, ... F2 are shown.n .
[0046] The set of values for the described activation function is preferably defined by a range of numbers between a lower limit and an upper limit. The set of values is preferably defined by a range from zero (inclusive) to one (inclusive).
[0047] Preferably, for values in the definition set of the activation function that are less than the minimum value of the definition set, the derivative of the described activation function is less than 1, wherein a fixed point of the activation function is defined for the minimum value. Preferably, for values in the definition set of the activation function that are greater than the maximum value of the definition set, the derivative of the described activation function is less than 1, wherein a fixed point of the activation function is defined for the maximum value.
[0048] In particular, within the range of values of the definition set of fixed points for which the activation function is not defined, the derivative of the described activation function is greater than zero.
[0049] The following text is based on Figure 6 Describe a method using one of the described activation functions.
[0050] The computer-implemented method is used to train artificial neural networks or to control robots, vehicles, household appliances, tools, production machines, human assistance systems, access control systems, or systems for transmitting information using manipulation parameters 104.
[0051] In step 602, input data 112 is detected. Input data 112 is, for example, sensor data.
[0052] In step 604, the input data 112 is submitted to the input layer of the artificial neural network. The input data 112 is, for example, standardized beforehand.
[0053] In step 606, the output parameters of the artificial neural network are determined. The output parameters are determined based on the input data used for the artificial neural network and the activation function of the artificial neural network. The activation function is at least piecewise continuously differentiable and monotonically increasing. The activation function has at least three fixed points. The number of fixed points of the activation function is either finite and odd or countable and discrete. The derivative of the activation function is alternately monotonically increasing and monotonically decreasing in the region between adjacent fixed points. One of the described activation functions is used in this example.
[0054] Optionally, a gradient method can be used in this step to determine the parameter updates for the artificial neural network based on the derivative of the activation function.
[0055] Of particular advantage is that, in recurrent neural networks, the state is determined by utilizing gradient methods and repeatedly calculating the state of at least one neuron of the recurrent neural network in a feedback manner based on the activation function. This avoids bursty or vanishing gradients during the training phase and during the intervention phase, i.e., when applying the artificial neural network.
[0056] In step 608, the manipulation parameters are determined based on the output parameters of the artificial neural network.
[0057] In step 610, the robot, vehicle, household appliance, tool, production machine, human assistance system, access control system, or information transmission system are controlled using the control parameter 104.
[0058] Next, proceed to step 602.
[0059] For example, the method can be started to adjust or control the actuator based on sensor data, and the method can be terminated when the adjustment or control is turned off.
[0060] Sensor data includes, for example, LiDAR, radar, wheel speed sensors, video or audio data, data from acceleration or yaw rate sensors, or data containing information about the transmission status, steering angle, or accelerator pedal position. For instance, artificial neural networks are trained to control partially autonomous motor vehicles based on sensor data.
[0061] Actuators can be, for example, active steering devices, engine control devices, or valves used therein. Actuators can also be inlet control devices, such as personnel separation devices or door closing and opening mechanisms that release or de-release the inlet based on personnel identified in an image or audio signal.
[0062] The described device and method enable training to be developed particularly efficiently. This is because these activation functions can potentially be used very quickly with self-learning artificial neural networks.
Claims
1. A computer-implemented method for running an image classifier using an artificial neural network, characterized in that, The output parameters of the artificial neural network are determined based on input data for the artificial neural network and an activation function of the artificial neural network, and characterize the classification of the input data, wherein the input data can be given as image data, wherein the activation function is at least piecewise continuously differentiable and monotonically increasing, wherein the activation function has at least three fixed points, wherein the number of fixed points of the activation function is either finite and odd or countable and discrete, and wherein the derivative of the activation function is alternately monotonically increasing and monotonically decreasing in the region between adjacent fixed points.
2. The method according to claim 1, characterized in that, The derivative of the activation function is less than 1 for any value less than the minimum value of the defined set, where the minimum value of the defined set defines a fixed point of the activation function. Alternatively, the derivative of the activation function is less than 1 for any value greater than the maximum value of the defined set, where the maximum value of the defined set defines a fixed point of the activation function.
3. The method according to claim 1 or 2, characterized in that, The derivative is greater than zero.
4. The method according to claim 3, characterized in that, The derivative is greater than zero in the range of values for a fixed set of points for which no activation function is defined.
5. The method according to claim 1 or 2, characterized in that, The value set of the activation function is defined by a numerical range between the lower and upper limits.
6. The method according to claim 5, characterized in that, The value set of the activation function is defined by a range of numbers from zero (inclusive) to one (inclusive).
7. The method according to claim 1 or 2, characterized in that, The activation function is continuously differentiable.
8. The method according to claim 1 or 2, characterized in that, The activation function is a smooth function or a piecewise linear function, wherein the fixed point having the minimum value of all fixed points of the activation function has a value of zero, and the fixed point having the maximum value of all fixed points has a value of 1.
9. The method according to claim 1 or 2, characterized in that, The activation function is a piecewise linear function, wherein the activation function has at least two segments between two directly adjacent fixed points, and the segments are linear functions with different slopes.
10. The method according to claim 1 or 2, characterized in that, The value of the fixed point is a power of 2.
11. The method according to claim 1 or 2, characterized in that, The state of a neuron in a recurrent neural network is determined by using a gradient method and repeatedly calculating the state in a feedback manner according to the activation function.
12. A device (100) for manipulation using an artificial neural network, characterized in that, The device (100) includes an output (102) for controlling a robot, vehicle, household appliance, tool, production machine, human assistance system, access control system or system for transmitting information using control parameters (104), wherein the device (100) includes a computing device (106), a memory (108) for the artificial neural network and an input (110) for input data (112), wherein the device (100) is configured to determine the control parameters (104) according to any one of claims 1 to 11.
13. A computer-readable storage medium, characterized in that, The computer-readable storage medium includes computer-readable instructions that, when executed by a computer, perform the method according to any one of claims 1 to 11.
14. A computer program product, characterized in that... The storage medium according to claim 13 has a computer program stored on it.
Citation Information
Patent Citations
Partial discharge signal processing method and apparatus employing neural network
CN105210088A
Neural-network image classification method based on SPeLUs function
CN108596235A