Neural Networks Online Training

By separating spatial and temporal gradient components for efficient computation and update during online training, the method addresses the inefficiencies of existing BPTT methods, enabling effective and efficient online training of neural networks.

JP7682255B2Active Publication Date: 2025-05-23INTERNATIONAL BUSINESS MACHINE CORPORATION
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2023502937
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2020-07-21
Filing Date
2021-07-06
Publication Date
2025-05-23
Estimated Expiration
2041-07-06

AI Technical Summary

Technical Problem

Existing methods for training recurrent neural networks, such as gradient-based training using backpropagation through time (BPTT), face challenges with system locking and inefficiencies in online learning scenarios due to the need to record and propagate error signals through deep time-unfolded networks.

Method used

The proposed method separates spatial and temporal gradient components, allowing for independent computation and update at each time instance, which facilitates efficient online training and implementation on hardware accelerators like memristor arrays.

Benefits of technology

This approach enables effective online training of neural networks, reducing computational complexity and maintaining performance comparable to BPTT, while allowing for real-time updates and reduced memory requirements.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007682255000058
    Figure 0007682255000058
  • Figure 0007682255000059
    Figure 0007682255000059
  • Figure 0007682255000060
    Figure 0007682255000060
Patent Text Reader

Abstract

The present invention is particularly directed to a computer-implemented method for training parameters of a recurrent neural network. The network comprises one or more layers of neuron units. Each neuron unit has an internal state, sometimes referred to as a unit state. The method includes providing training data to the recurrent neural network, the training data including an input signal and an expected output signal. The method further includes calculating a spatial gradient component for each neuron unit and a temporal gradient component for each neuron unit. The method further includes updating the temporal gradient component and the spatial gradient component for each neuron unit at each time instance of the input signal. The calculation of the spatial gradient component and the temporal gradient component may be performed independently of each other. The present invention further relates to neural networks and related computer program products.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical field]

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS This application is a non-provisional adaptation of U.S. Provisional Application No. 63 / 054,247, entitled "ONLINE TRAINING OF RECURRENT NEURAL NETWORKS," filed on July 21, 2020, and incorporated herein by reference in its entirety for all purposes.

[0002] The present invention is directed, inter alia, to a computer-implemented method for training neural networks, in particular recurrent neural networks.

[0003] The invention further relates to an associated neural network and an associated computer program product. [Background technology]

[0004] During the last few years, the number of applications making use of artificial neural networks (ANNs) has grown rapidly. In particular, in tasks such as speech recognition, language translation or building neural computers, recurrently connected ANNs, so-called RNNs, have demonstrated phenomenal performance levels.

[0005] Recurrent neural networks (RNNs) have played a key role in recent advances in artificial intelligence. One known technique for training RNNs is gradient-based training, which utilizes backpropagation of error through time (BPTT).

[0006] However, BPTT has limitations because it needs to record all past activity by unfolding the network in time, which can become very deep with increasing input sequence length. For example, a 2-second long spoken input sequence with a time step of 1 ms results in an unfolded network 2000 layers deep.

[0007] Therefore, propagating errors backwards in time can result in a system locking problem, making BPTT unusable for online learning scenarios. Variants that allow online training have recently regained interest in the research community. One known approach focuses on approximating BPTT via online algorithms. Another approach takes inspiration from ecology and investigates spiking neural networks (SNNs).

[0008] Thus, there remains a need for advantageous methods for training neural networks, especially for online training. Summary of the Invention

[0009] According to one aspect, the invention is embodied as a computer-implemented method for training a neural network. The network comprises one or more layers of neuron units. Each neuron unit has an internal state, which may also be referred to as a unit state. The method includes providing training data to the neural network, the training data including an input signal and an expected output signal. The method further includes calculating a spatial gradient component for each neuron unit and calculating a temporal gradient component for each neuron unit. The method further includes updating the temporal gradient component and the spatial gradient component for each neuron unit at each time instance of the input signal.

[0010] Thus, the method according to the embodiment of the present invention is based on the separation of spatial and temporal gradient components. This can facilitate a deeper understanding of the feedback mechanism. Furthermore, it can facilitate an efficient implementation on hardware accelerators such as memristor arrays. The method according to the embodiment of the present invention can be used in particular for online training. The method according to the embodiment of the present invention can be used in particular for training the training parameters of a neural network.

[0011] The method according to an embodiment of the present invention processes time data as an input signal. Time data may be defined as data representing states or values ​​in time, or in other words data relating to time instances. The input signal may in particular be a continuous input data stream. The input signal is processed by the neural network at time instances, or in other words at time steps.

[0012] According to one embodiment, the computation of the spatial and temporal gradient components is performed independently of each other, which has the advantage that these gradient components can be computed in parallel to reduce computation times.

[0013] According to an embodiment, the spatial gradient component establishes a training signal and the temporal gradient component establishes a qualification trace.

[0014] Methods according to embodiments of the present invention may be used in particular for low complexity devices such as Internet of Things (IoT) devices as well as edge artificial intelligence (AI) devices.

[0015] According to an embodiment, the method comprises updating training parameters of the neural network at specific or predefined time instances, in particular at each time instance, the updating may in particular be performed as a function of spatial and temporal gradient components.

[0016] The training parameters that may be trained according to the embodiment include, in particular, the input weights and / or recurrent weights of the neuron unit. By updating the training parameters at each time instance, the neuron unit learns at each time instance, or in other words, at each time step.

[0017] According to an embodiment, the spatial gradient component is based on connectivity parameters of the neural network, for example the connectivity of the individual neuronal units. According to an embodiment, the connectivity parameters in particular describe parameters of the architecture of the neural network. According to an embodiment, the connectivity parameters may be defined as the number or set of transmission lines that allow information exchange between the individual neuronal units. According to an embodiment, the spatial gradient component is a component that takes into account the spatial aspects of the neural network, in particular the interdependence between the individual neuronal units at each time instance.

[0018] According to an embodiment, the time gradient component is based on the temporal dynamics of the neuronal unit. According to an embodiment, the time gradient component is a component that takes into account the temporal dynamics of the neuronal unit, in particular the temporal evolution of the internal state / unit state.

[0019] According to an embodiment, the method includes calculating, at each time instance, a spatial gradient component for each of one or more layers, and calculating, at each time instance, a temporal gradient component for each of one or more layers. Thus, at each time instance / time step, the method calculates a temporal gradient component and a spatial gradient component for each layer. The spatial gradient component / training signal may be specific for each layer and propagates from the last layer to the input layer without going back in time, i.e., it represents the spatial gradient through the network architecture.

[0020] According to embodiments, each layer can compute its own temporal gradient components / eligibility traces that depend only on the contributions of the respective layers, i.e., they represent the temporal gradients through time for the same layer. According to embodiments, spatial gradient components may be shared for two or more layers.

[0021] According to an embodiment, the method may be used for single layer networks as well as multi-layer networks.

[0022] According to an embodiment, the method may be applied to recurrent neural networks, spiking neural networks, and hybrid networks comprising or consisting of units with unit states and units without unit states.

[0023] According to embodiments, the methods and parts of the methods may be implemented in neuromorphic hardware, in particular in an array of memristor devices.

[0024] For shallow networks, the method according to the embodiment of the present invention can maintain equal gradients as a backpropagation through time (BPTT) technique.

[0025] According to an embodiment of another aspect of the present invention, a neural network, in particular a recurrent neural network, is provided. The neural network comprises one or more layers of neuron units. Each neuron unit has an internal state, which may also be denoted as a unit state. The neural network is configured to perform a method comprising providing training data to the neural network, the training data comprising an input signal and an expected output signal. The method further comprises calculating a spatial gradient component for each neuron unit and calculating a temporal gradient component for each neuron unit. The method further comprises updating the temporal gradient component and the spatial gradient component for each neuron unit at each time instance of the input signal. The calculation of the spatial gradient component and the temporal gradient component may be performed independently of each other.

[0026] Depending on the embodiment, the neural network may be a recurrent neural network, a spiking neural network, or a hybrid neural network.

[0027] According to an embodiment of another aspect of the present invention, a computer program product for training a neural network is provided. The computer program product comprises a computer-readable storage medium having program instructions embodied therewith, the program instructions being executable by the neural network to cause the neural network to perform a method comprising receiving training data comprising an input signal and an expected output signal. The method comprises a further step of calculating a spatial gradient component for each neuronal unit and calculating a temporal gradient component for each neuronal unit. The further step comprises updating the temporal gradient component and the spatial gradient component for each neuronal unit at each time instance of the input signal. According to an embodiment, the calculation of the spatial gradient component and the temporal gradient component may be performed independently of each other.

[0028] Embodiments of the invention are described in more detail below, by way of illustrative and non-limiting examples, with reference to the accompanying drawings, in which: [Brief description of the drawings]

[0029]

Figure 1

Figure 2

Figure 3

Figure 4a

Figure 4b

Figure 5

Figure 6

Figure 7

Figure 8

Figure 9

Figure 10

Figure 11

[0030] An embodiment of the present invention provides a method for training, in particular online training, of neural networks, in particular recurrent neural networks (RNNs). The method may be as follows, also denoted OSTL: The method according to an embodiment of the present invention provides an advantageous algorithm that can be used for online learning applications by separating spatial and temporal gradients.

[0031] Figure 1 illustrates the gradient flow of a computer-implemented method for training a neural network 100 according to one embodiment of the present invention. In the case of Figure 1, the neural network 100 is assumed to be a recurrent neural network (RNN) with a single layer 110 with neuron units 111. The neural network is evolved for three time steps t.

[0032] Each neuron unit 111 has an internal state S, 120. The method relies on the input signal x t , 131 and expected output signal 132. The method then comprises providing the neural network with training data including the spatial gradient component L t , 141 and the time gradient component e t , 142. Furthermore, at each time instance t of the input signal 131, the temporal gradient component 142 and the spatial gradient component 141 are updated for each neuron unit 111.

[0033] The goal of learning / training is to learn the parameters θ of the neural network so that it can predict the current output signal y at time t. t and the input signal x t The error E between t The aim of the study is to train the method to minimize

[0034] In an RNN, the network error E at time t is t is often the output of a neuron unit in the output layer, y t is a function of E t =f(y t In addition, many neuronal units in an RNN have internal states s on which their outputs depend. t y t =f(s t ) This internal state of a neuron unit is in turn expressed as its input signal x via trainable input weights W and trainable recurrent weights H, respectively. t which in turn may be a recursive function itself that depends recursively on its output signal.

[0035] According to the embodiment, the equation governing the internal state is s t =f(x t ,s t-1 ,y t-1 ,W,H), for example, s t =Wx t +Hy t-1 It can be formulated as follows: For simplicity of notation, all trainable parameters of the RNN 100 may be collectively described by the variable θ as follows: t =f(x t ,s t-1 ,y t-1 ,θ).

[0036] Moreover, the output y t The notation for y may be extended according to an embodiment to allow for a direct dependence on trainable parameters, i.e., y t =f(s t ,θ), for example, y t =σ(s t +b).

[0037] Using this notation, the change in parameter θ required to minimize E is given by, based on the principles of gradient descent,

number

[0038] From this, an embodiment of the present invention uses a back propagation through time (BPTT) technique as a starting point for differentiation, and calculates dE / dθ as

number

number

[0039] Equation 3 can be rewritten in recursive form as follows:

number

number

number

[0040] Thus, according to an embodiment, the computation of the spatial and temporal gradient components may be performed independently of each other.

[0041] In the standard RNN example, the explicit form of these equations is:

number

[0042] According to the embodiment, the notation is inspired by standard nomenclature in biological systems, where synaptic weight changes are often decomposed into a learning signal and an eligibility trace. In the simplest case, the eligibility trace is a low-pass filtered version of neural activity, while the learning signal represents a spatially propagated reward signal. Thus, according to the embodiment, in Eq. 6, e t,θ The time gradient, denoted as L, is associated with the eligibility trace and is expressed as t The spatial gradient, denoted as , can be associated with the training signal.

[0043] Similar to ecosystems, the parameter change dE / dθ from Equation 5 is calculated as the sum over time of the product of the fitness trace and the training signal. This allows the parameter updates to be calculated online, as shown in Figure 1.

[0044] Furthermore, note that the derivatives in Equation 6 are exact.

[0045] As can be seen from FIG. 1, at each time step the time gradient may be combined with the spatial gradient of this time step, without having to go back to the beginning of the input sequence / input signal as required according to known backpropagation through time techniques.

[0046] Figure 2 illustrates a gradient flow of a computer-implemented method for training a neural network 200 according to one embodiment of the present invention. In the case of Figure 2, the neural network 200 is assumed to be a recurrent neural network (RNN) with multiple layers.

[0047] More specifically, Figure 2 shows the gradient flow for a two-layer RNN with a first layer 210 with neuron unit 211 and a second layer 220 with neuron unit 221. The layers 210 and 220 are unrolled for three time steps to separate the spatial and temporal gradients.

[0048] Each neuron unit 211 has an internal state S 1 , 230. Each neuron unit 221 has an internal state S 2 , 231. The method includes: t , 141 and expected output signals 142 to the neural network 200. The method then calculates for each neuron unit 211 a spatial gradient component L 1 t , 151, and calculate the spatial gradient component L for each neuronal unit 221. 2 t , 152. Furthermore, the method calculates for each neuronal unit 211 the time gradient component e 1 t , 161, and calculate the time gradient component e for each neuron unit 221. 2 t , calculate 162.

[0049] Furthermore, at each time instance t of the input signal 141, the temporal gradient components 161, 162 and the spatial gradient components 151, 152 are updated for each neuron unit 211, 221, respectively.

[0050] Many prior art applications rely on more complex multi-layer architectures. To extend the method according to the embodiment of the present invention to deep architectures, state s t and output y t The definition of the error E in the deep architecture may be revised as follows: t is just a function of the last output layer k, i.e., E t =f(y k t ), where each layer l has its own trainable parameters θ l The input of layer l is the output of the previous layer y l-1 t For the first layer, the external input is used, y 0 t =x t It is.

[0051] Therefore, the definition is:

number

[0052] For single-layer neural networks, separation of spatial and temporal components occurs when following the differentiation outlined by Equations 3-5.

[0053] However, in the case of a multi-layer architecture, the term ds in Eq. t / dθ includes different layers l and m, e.g., ds l t / dθ m , which results in cross-layer dependencies (see sidebar).

[0054] To maintain the benefits discussed above, a clear separation of spatial and temporal gradients is also introduced for multi-layer architectures according to embodiments of the present invention. Thus, similar steps as described above for single-layer RNNs are performed using generalized state and output equations 8 and 9. Following detailed differentiation in the supplementary notes, the following eligibility traces and training signals are obtained for layer l:

number

number

number

[0055] As can be seen by comparing Equations 5 to 13, the learning signal L l t Eligibility Tracing l t,θ The approach according to embodiments of the present invention for multiplying with remains the same for deep networks.

[0056] Learning signal L l t is layer-specific and propagates from the last layer to the input layer without going back in time, i.e., it represents the spatial gradient through the network architecture. Furthermore, each layer has its own eligibility trace e l t,θ We compute,x,=,y,=1,y,=1,y,=2 ...

[0057] However, additional terms are also included in Equation 13, which involve a mixture of spatial and temporal gradients and generally require going back in time. These terms are collected in the remainder term R.

[0058] In order to maintain the separation between the spatial and temporal gradients, Equation 13 is simplified according to the embodiment by omitting the term R. Thus, the following formulation for multi-layer networks is obtained according to the embodiment:

number

[0059] Therefore, according to an embodiment of the present invention, the remainder term R is intentionally omitted and the mixed spatial and temporal gradient components are not taken into account during learning / training. However, the inventors' research has provided insight that this is an advantageous approach. In particular, it is known what is omitted by such an approach. Furthermore, the inventors' simulations provide empirical evidence that performance comparable to BPTT can be achieved even without these terms, as will be further explained below. Moreover, according to an embodiment, the remainder term R may also be approximated, thus allowing a better approximation of the gradient from Equation 13.

[0060] FIG. 3 shows a spiking neuron unit SNU 310 of a spiking neural network 300. With reference to FIG. 3, it is shown that a method according to an embodiment can be applied to a spiking neural network (SNN). The dashed lines in FIG. 3 indicate connections with time lags, and the bold lines indicate parameterized connections. The SNU 310 includes a block input 320, a block output 321, a reset gate 322, and a membrane potential 323.

[0061] Historically, SNNs were often trained with some form of spike-timing dependent plasticity, and recently gradient-based training for SNNs has been proposed, for example in Wozniak, S., Pantazi, A., Bohnstingl, T., and Eleftheriou, E. Deep learning incorporating biologically-inspired neural dynamics.arXiv, December 2018, URL: https: / / arxiv.org / abs / 1812.07040.

[0062] Such a method aims to bridge the ANN world with the SNN world by recreating SNN dynamics with ANN-based building blocks to form a spiking neuron unit SNU 310. The SNU 310 of the spiking neural network 300 receives multiple input signals.

[0063] This approach enables SNU to enable gradient-based learning, which makes it possible to harness the power of known optimization techniques for ANNs while reproducing the dynamics of leaky integrate-and-fire (LIF) neuron models well known in neuroscience.

[0064] As shown above, the method according to the embodiment of the present invention may be used for general-purpose RNNs, but it can also be applied according to the embodiment to train deep SNNs formulated as RNNs. This is shown below. We start with the expressions for the states and outputs of the SNU layer l and compare them with (Wozniak et al., 2018).

[0065]

number

[0066] By using equations 15 and 16,

number

number

number

number

[0067] Mean squared error loss function, e.g.,

number

number

[0068] For a deep neural network with k layers consisting of RNNs or recurrent SNUs, the method according to the embodiment of the present invention requires O(kn 4 ) time complexity. This time complexity is determined by the network structure itself, and is mainly determined by the recurrent matrix H l When a feed-forward architecture is used according to the embodiment, H l The terms containing vanish, and the SNU formula becomes:

number

[0069] These formulas are then transformed into the following eligibility trace:

number

number

number

[0070] This reduces the time to O(kn 4 ) to O(kn 2 ) the time complexity is significantly reduced. Using a feed-forward SNU network architecture does not necessarily impede solving time tasks. Such networks have long been used in SNNs, and it is because the network is a layer-type recurrent matrix H l This implies that the recursion should depend on the internal state of the unit, which is implemented using self-recursion, rather than on the internal state of the unit itself.

[0071] It should be noted that according to an embodiment, the training signal may be calculated without the matrix W, for example based on some randomization or approximation of W. More specifically, the training signal may be calculated based on a different matrix that is not used in the forward path. In other words, the forward path may use the matrix W, and the training signal is calculated for a different matrix B. The matrix B may or may not be trainable.

[0072] According to the embodiment, the method presented above may also be used for hybrid networks. In this respect, a very common scenario in deep RNNs or SNNs is that they are often combined with a layer of stateless neurons at the output, for example a sigmoid layer or a softmax layer. The method according to the embodiment of the invention can also be applied to train these hybrid networks, including one or more layers of stateless neurons, without any modification. In detail, the equations for the states and outputs of these layers are:

number

number

number

number

[0073] Note that stateless layers do not introduce any remainder term R. This has the effect that when we add such layers to a network, even between RNN layers, the gradients relative to the next layer remain unchanged.

[0074] Fig. 4a shows the test results of the method according to the embodiment of the present invention compared to the backpropagation through time (BPTT) technique. More specifically, Fig. 4a relates to music prediction based on the JSB dataset, as presented in the literature: Boulanger-Lewandowski, N., Bengio, Y., and Vincent, P., Modeling temporal dependencies in high-dimensional sequences: Application to polyphonic music generation and transcription, In Proceedings of the 29th International Conference on International Conference on Machine Learning, ICML'12, pp. 1881-1888, Madison, WI, USA, 2012. Omnipress. ISBN 9781450312851.

[0075] For this, a standard training / test data split was used. For the test, the hybrid architecture comprises a feed-forward SNU layer with 150 units and a stateless layer sigmoid layer with 88 units on top. To obtain a baseline, the same network with all its hyperparameters was trained with the method according to the embodiment of the invention and BPTT for 1000 epochs. The Y-axis represents the negative log-likelihood averaged over 10 random initial states. Bar 411 shows the results of training the BPTT method, bar 412 shows the results of training the method according to the embodiment of the invention. Furthermore, bar 413 shows the results of the trial run of the BPTT method, and bar 414 shows the results of the trial run of the method according to the embodiment of the invention.

[0076] As shown in Figure 4a, the results obtained using the method according to the embodiment of the present invention are in fact comparable to those obtained using BPTT. Note that the task proves the equivalence of the gradients of BPTT and the method according to the embodiment of the present invention for a hybrid architecture with a single RNN layer and a stateless layer on top.

[0077] As shown in Fig. 4b, this task may be used to prove the reduced computational complexity of the method according to the embodiment of the invention for feed-forward SNNs. For this purpose, the number of required floating-point operations MFLOPs (y-axis) was measured using the built-in TensorFlow profiler for one parameter updated over different input sequence lengths (x-axis) of the JSB input sequence (see Fig. 4b). As can be seen from line 422, BPTT needs to perform a time expansion and therefore has a linear dependence on the sequence length T, while the method according to the embodiment of the invention illustrated by line 421 does not and therefore remains constant. However, in a practical implementation, updates from the method according to the embodiment of the invention may need to be accumulated over time, which results in the same complexity as BPTT. It should be noted that the initial high cost of the method according to the embodiment of the invention is due to the implementation overhead, since the method according to the embodiment of the invention is not included in the standard toolbox of TensorFlow. Nevertheless, the obtained plot is consistent with the theoretical complexity analysis.

[0078] Figure 5 shows the test results for another task of handwritten digit classification based on the MNIST dataset, introduced in the literature: Lecun, Y., Bottou, L., Bengio, Y., and Haffner, P., Gradient based learning applied to document recognition. Proc., IEEE86(11):2278-2324, November 1998, ISSN1558-2256, doi:10.1109 / 5.726791.

[0079] Again, a standard training / test data split was used. According to the test, a five-layer feed-forward architecture of the SNU with 256 units was adopted and trained for 50 epochs averaged over 10 random initial states. Similar to the tasks shown with reference to FIGS. 4a and 4b, the accuracy of the method according to an embodiment of the present invention coincides with that of BPTT. The y-axis represents the accuracy (percentage), the x-axis represents the number of epochs, line 510 represents the result of BPTT, and line 520 represents the result of the method according to an embodiment of the present invention.

[0080] FIG. 6 shows how the method according to an embodiment of the present invention can be implemented in neuromorphic hardware. The neuromorphic hardware may specifically include a crossbar array including a plurality of row lines 610, a plurality of column lines 620, and a plurality of junctions 630 disposed between the plurality of row lines 610 and the plurality of column lines 620. Each junction 630 includes a resistive change memory element 640, particularly a series arrangement of access elements including a resistive change memory element and access terminals for accessing the resistive change memory element. The resistive change element may be, for example, a phase change memory element, a conductive bridge random access memory element (CBRAM), a metal oxide resistive change random access memory element (RRAM), a magnetic resistive change random access memory element (MRAM), a ferroelectric random access memory element (FeRAM), or an optical memory element.

[0081] According to an embodiment, the input weights and the recurrent weights may be arranged in the neuromorphic device, particularly as the resistance states of the resistive change elements.

[0082] According to such an embodiment, the trainable input weight W l and the trainable recurrent weight H l are mapped to the resistive change memory element 640.

[0083] FIG. 7 shows a simplified schematic diagram of a neural network 700 according to an embodiment of the present invention. The neural network 700 comprises an input layer 710 comprising a plurality of neuron units 10, one or more hidden layers 720 comprising a plurality of neuron units 10, and an output layer 730 comprising a plurality of neuron units 10. The neural network 700 comprises a plurality of electrical connections 20 between the neuron units 10. The electrical connections 20 connect the outputs of neurons from one layer, e.g., the input layer 710, to the inputs of neuron units from the next layer, e.g., one of the hidden layers 720. The neural network 700 may in particular be embodied as a recurrent neural network.

[0084] Thus, the network 700 comprises recurrent connections from one layer to neuronal units from the same or previous layer, as indicated diagrammatically by arrows 30 .

[0085] Figure 8 shows a flowchart of method steps of a computer-implemented method for training parameters of a recurrent neural network.

[0086] The method begins at step 810 .

[0087] Training data is received by, or in other words provided to, the neural network in step 820. The training data includes input signals and expected output signals.

[0088] In step 830, the neural network calculates the spatial gradient components for each neuronal unit.

[0089] In step 840, the neural network calculates the time gradient components for each neuronal unit.

[0090] In step 850, the neural network updates the temporal and spatial gradient components for each neuronal unit at each time instance of the input signal.

[0091] According to one embodiment, updates to the parameters of the neural network can be accumulated and withheld until a subsequent time step T. The computation of the spatial and temporal gradient components is performed independently of each other.

[0092] Steps 820-850 are repeated in a loop 860. More specifically, steps 820-850 may be repeated at specific or predefined time instances, in particular at each time instance.

[0093] Referring to FIG. 9, an exemplary embodiment of a computing system 900 for performing a method according to an embodiment of the present invention is shown. The computing system 900 can form a neural network according to an embodiment. The computing system 900 can be operable in numerous other general-purpose or special-purpose computing system environments or configurations. Examples of well-known computing systems, environments, or configurations or combinations thereof that may be suitable for use with the computing system 900 include, but are not limited to, personal computer systems, server computer systems, thin clients, thick clients, handheld or laptop devices, multiprocessor systems, microprocessor-based systems, set-top boxes, programmable consumer electronics, network PCs, minicomputer systems, mainframe computer systems, and distributed cloud computing environments that include any of the above systems or devices.

[0094] The computing system 900 may be described in the general context of computer system executable instructions, such as program modules, executed by a computer system. Generally, program modules may include routines, programs, objects, components, logic, data structures, etc. that perform particular tasks or implement particular abstract data types. The computing system 900 may be depicted in the form of a general-purpose computing device. Components of the server computing system 900 may include, but are not limited to, one or more processors or processing units 916, a system memory 928, and a bus 918 that couples various system components, including the processor 916, to the system memory 928.

[0095] Bus 918 represents one or more of any of several types of bus structures, including a memory bus or memory controller, a peripheral bus, an accelerated graphics port, and a processor or local bus using any of a variety of bus architectures. By way of example, and not limitation, such architectures include Industry Standard Architecture (ISA) bus, Micro Channel Architecture (MCA) bus, Enhanced ISA (EISA) bus, Video Electronics Standards Association (VESA) local bus, and Peripheral Component Interconnect (PCI) bus.

[0096] Computing system 900 typically includes a variety of computer system readable media. Such media can be any available media that can be accessed by computing system 900 and includes both volatile and nonvolatile media, removable and non-removable media.

[0097] The system memory 928 may include computer system readable media in the form of volatile memory, such as random access memory (RAM) 930 and / or cache memory 932. The computing system 900 may further include other removable / non-removable, volatile / non-volatile computer system storage media. By way of example only, a storage system 934 may be provided for reading from and writing to a non-removable, non-volatile magnetic medium (not shown, typically referred to as a "hard drive"). Although not shown, a magnetic disk drive may be provided for reading from and writing to a removable, non-volatile magnetic disk (e.g., a "floppy disk"), and an optical disk drive may be provided for reading from and writing to a removable, non-volatile optical disk, such as a CD-ROM, DVD-ROM, or other optical media. In such an instance, each may be connected to the bus 918 by one or more data media interfaces. As further depicted and described below, memory 928 may include at least one program product having a set (e.g., at least one) program module configured to perform functions of an embodiment of the present invention.

[0098] A set (at least one) of program modules 942, as well as programs / utilities 940 having an operating system, one or more application programs, other program modules, and program data, may be stored in memory 928, by way of example and not limitation. Each of the operating system, one or more application programs, other program modules, and program data, or any combination thereof, may include an implementation of a networked environment. The program modules 942 generally perform the functions and / or methods of the embodiments of the invention described herein. The program modules 942 may, in particular, perform one or more steps of a computer-implemented method for training a recurrent neural network, such as one or more steps of the method described with reference to FIG. 1, FIG. 2, and FIG. 8.

[0099] The computing system 900 may also communicate with one or more external devices 915, such as a keyboard, a pointing device, a display 924, one or more devices that allow a user to interact with the computing system 900, or any device (e.g., a network card, a modem, etc.) that allows the computing system 900 to communicate with one or more other computing devices, or a combination thereof. Such communication may occur via an input / output (I / O) interface 922. Still further, the computing system 900 may communicate with one or more networks, such as a local area network (LAN), a general wide area network (WAN), or a public network (e.g., the Internet), or a combination thereof, via a network adapter 920. As depicted, the network adapter 920 communicates with other components of the computing system 900 via a bus 918. Although not shown, it should be understood that other hardware and / or software components may be used in conjunction with the computing system 900. Examples include, but are not limited to, microcode, device drivers, redundant processing units, external disk drive arrays, RAID systems, tape drives, and data archive storage systems.

[0100] The present invention may be a system, method, and / or computer program product at any possible level of integration of technical detail, and may include one or more computer-readable storage media having computer-readable program instructions for causing a processor to carry out aspects of the present invention.

[0101] A computer readable storage medium may be a tangible device capable of holding and storing instructions for use by an instruction execution device. A computer readable storage medium may be, for example, but not limited to, an electronic storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination of the foregoing. A non-exhaustive list of more specific examples of computer readable storage media includes the following: portable computer diskettes, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), static random access memory (SRAM), portable compact disk read-only memory (CD-ROM), digital versatile disk (DVD), memory sticks, floppy disks, mechanically encoded devices such as punch cards or ridge structures in grooves having instructions recorded thereon, and any suitable combination of the foregoing. Computer-readable storage media as used herein should not be interpreted as signals that are inherently ephemeral, such as radio waves or other freely propagating electromagnetic waves, electromagnetic waves propagating through a waveguide or other transmission medium (e.g., light pulses through a fiber optic cable), or electrical signals transmitted through wires.

[0102] The computer readable program instructions described herein may be downloaded from the computer readable storage medium to the respective computing / processing device or to an external computer or storage device via a network, such as the Internet, a local area network, a wide area network, or a wireless network, or a combination thereof. The network may comprise copper transmission cables, optical transmission cables, wireless transmissions, routers, firewalls, switches, gateway computers, and / or edge servers. A network adapter card or network interface in each computing / processing device receives the computer readable program instructions from the network and forwards the computer readable program instructions for storage in the computer readable storage medium in the respective computing / processing device.

[0103] The computer readable program instructions for carrying out the operations of the present invention may be assembler instructions, instruction set architecture (ISA) instructions, machine instructions, machine dependent instructions, microcode, firmware instructions, state setting data, configuration data for an integrated circuit, or source or object code written in any combination of one or more programming languages, including object oriented programming languages ​​such as Smalltalk®, C++, and procedural programming languages ​​such as the “C” programming language or similar programming languages. The computer readable program instructions may be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the latter scenario, the remote computer may be connected to the user's computer via any type of network, including a local area network (LAN) or a wide area network (WAN), or the connection may be to an external computer (e.g., via the Internet using an Internet Service Provider). In some embodiments, electronic circuitry including, for example, a programmable logic circuit, a field programmable gate array (FPGA), or a programmable logic array (PLA) may execute computer readable program instructions by utilizing state information of the computer readable program instructions to personalize the electronic circuitry to perform aspects of the invention.

[0104] Aspects of the present invention are described herein with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer readable program instructions.

[0105] These computer readable program instructions may be provided to a general purpose computer, special purpose computer, or other programmable data processing apparatus generating machine such that the instructions, executed via a processor of the computer or other programmable data processing apparatus, create means for implementing the functions / operations specified in one or more blocks of the flowcharts and / or block diagrams. These computer readable program instructions may also be stored in a computer readable storage medium capable of directing a computer, programmable data processing apparatus, or other device, or combination thereof, to function in a particular manner such that the computer readable storage medium having the instructions stored thereon comprises an article of manufacture containing instructions implementing aspects of the functions / operations specified in one or more blocks of the flowcharts and / or block diagrams.

[0106] The computer readable program instructions may also be loaded into a computer, other programmable data processing apparatus, or other device to generate a computer-implemented process such that a series of operational steps are performed on the computer, other programmable apparatus, or other device, such that the instructions, which execute on the computer, other programmable apparatus, or other device, implement the functions / operations specified in one or more blocks of the flowcharts and / or block diagrams.

[0107] The flowcharts and block diagrams in the figures illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present invention. In this regard, each block in the flowcharts or block diagrams may represent a module, segment, or portion of instructions that includes one or more executable instructions for implementing a specified logical function. In some alternative implementations, the functions noted in the blocks may occur out of the order noted in the figures. For example, two blocks shown in succession may in fact be executed substantially in parallel, or the blocks may sometimes be executed in reverse order, depending on the functionality involved. It should be noted that each block of the block diagrams and / or flowchart illustrations, and combinations of blocks in the block diagrams and / or flowchart illustrations, may be realized by a dedicated hardware-based system that executes the specified functions or operations or executes a combination of dedicated hardware and computer instructions.

[0108] The description of various embodiments of the present invention is presented for illustrative purposes, but is not intended to be exhaustive or limited to the disclosed embodiments. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the described embodiments. The terms used in this specification have been selected to best explain the principles of the embodiments, practical applications, or technical improvements to the technology found in the market, or to enable other skilled in the art to understand the embodiments disclosed herein.

[0109] In general, modifications described for one embodiment may be applied to another embodiment as appropriate.

[0110] In the following, a detailed differentiation of the method according to the embodiment of the invention for deep neural networks, in particular for recurrent networks including multi-layer architectures, is provided as a supplementary explanation. Many prior art applications rely on multi-layer networks, in which the error Et is just a function of the last output layer k, i.e.

number

[0111]

number

number

number

number

number

number

number

number

number

number

number

number

number

number

number

number

number

number

number

number

Claims

1. 1. A computer-implemented method for training a neural network, the neural network comprising a layer of neuronal units, each neuronal unit having an internal state (unit state), the method comprising: providing training data to the neural network, the training data including input signals and expected output signals; calculating a spatial gradient component for each said neuronal unit; calculating a time gradient component for each said neuronal unit; updating the temporal gradient component and the spatial gradient component for each neuronal unit at each time instance of the input signal; updating a predefined set of training parameters of the neural network as a function of the spatial gradient components and the temporal gradient components; Including, Calculating the spatial gradient components comprises: [0010] Calculating Calculating the time gradient component [0025] Calculating Where: t denotes each said time instance; Let y t denote the current output signal at time instance t, Let L t denote the spatial gradient component at the time instance t, E t denotes the error of the neural network, in particular the error between the expected output signal at the time instance t and the current output signal; Let s t denote the unit state at the time instance t, Let θ denote the training parameters of the neural network, e t,θ denotes the temporal gradient component at the time instance t; Computer-implemented method.

2. 1. A computer-implemented method for training a neural network, the neural network comprising multiple layers of neuronal units, each neuronal unit having an internal state (unit state), the method comprising: providing training data to the neural network, the training data including input signals and expected output signals; calculating a spatial gradient component for each said neuronal unit; calculating a time gradient component for each said neuronal unit; updating the temporal gradient component and the spatial gradient component for each neuronal unit at each time instance of the input signal; updating a predefined set of training parameters of the neural network as a function of the spatial gradient components and the temporal gradient components; Including, Calculating the spatial gradient components comprises: [0030] Calculating Calculating the time gradient component [0045] Calculating where t denotes each said time instance; l denotes each layer of the plurality of layers, Let L l t denote the spatial gradient component of layer l at time instance t, Let y l t denote the current output signal of layer l at time instance t, Let E t denote the error of the neural network, in particular the error between the expected output signal at the time instance t and the current output signal; Let s l t denote the unit state of layer l at the time instance t, k denotes the last or output layer of the neural network; Let m′ denote the hidden layers of said neural network, ranging from 1 to (k−l+1), Let θ denote the training parameters of the neural network, e l t,θ denotes the temporal gradient component of layer l at the time instance t; Computer-implemented method.

3. The method further comprising: calculating the spatial gradient components for each of the plurality of layers at each of the time instances; calculating the temporal gradient components for each of the plurality of layers at each time instance; The computer-implemented method of claim 2 , comprising:

4. and updating a predetermined set of training parameters of the neural network as a function of the spatial gradient components and the temporal gradient components, wherein updating the training parameters comprises: [0050] where R is the remainder term. The computer-implemented method of claim 2 .

5. The computer-implemented method of claim 4 , wherein the remainder term R is approximated using a combination of a qualification trace and a training signal.

6. The computer-implemented method of claim 1 or 2, wherein the calculation of the spatial gradient component and the temporal gradient component is performed independently of each other.

7. A computer-implemented method as described in claim 1 or 2, wherein updating a predefined set of training parameters of the neural network comprises updating the predefined set of training parameters of the neural network at a particular or predefined time instance as a function of the spatial gradient component and the temporal gradient component.

8. A computer-implemented method as described in claim 1 or 2, wherein updating a predefined set of training parameters of the neural network comprises updating the predefined set of training parameters of the neural network at each time instance as a function of the spatial gradient components and the temporal gradient components.

9. the spatial gradient components are based on connectivity parameters of the neural network; the time gradient component is based on parameters relating to the temporal dynamics of the neuronal unit; 3. A computer-implemented method according to claim 1 or 2.

10. and updating a predetermined set of training parameters of the neural network as a function of the spatial gradient components and the temporal gradient components, the updating of the training parameters comprising: [006] where α is the learning rate.

3. A computer-implemented method according to claim 1 or 2.

11. 3. The computer-implemented method of claim 1 or 2, wherein the neural network is selected from the group consisting of a recurrent neural network, a hybrid network, a spiking neural network, and a generalized recurrent network, the generalized recurrent network in particular comprising or consisting of long-short-term memory units and gated recurrent units.

12. A computer-implemented method for training a neural network, the neural network comprising one or more layers of neuronal units, each neuronal unit having an internal state (unit state), the method comprising: providing training data to the neural network, the training data including input signals and expected output signals; calculating a spatial gradient component for each said neuronal unit; calculating a time gradient component for each said neuronal unit; updating the temporal gradient component and the spatial gradient component for each neuronal unit at each time instance of the input signal; updating a predefined set of training parameters of the neural network as a function of the spatial gradient components and the temporal gradient components; the neural network comprises a plurality of layers of the neuronal units; Calculating the spatial gradient components comprises: [0070] Calculating t denotes each said time instance; l denotes each layer of the plurality of layers, Let L l t denote the spatial gradient component of layer l at time instance t, Let y k t denote the current output signal of layer k, Let E t denote the error of the neural network, in particular the error between the expected output signal at the time instance t and the current output signal; Let s k t denote the unit state of layer k, k denotes the last or output layer of the neural network; Let m′ denote the hidden layers of said neural network, ranging from 1 to (k−l+1), Calculating the time gradient component [0080] Calculating t denotes each said time instance; l denotes each layer of the plurality of layers, y t denote the current output signal at said time instance t; s t Let t denote the current state of the unit at time instance t, Let θ denote the training parameters of the neural network, [0097] A computer-implemented method.

13. A computer program for training a recurrent neural network, the computer program causing a computer to carry out the method according to any one of claims 1 to 12.

14. 1. A computing system for training parameters of a neural network, comprising: the neural network comprises one or more layers of neuron units, each having an internal state; The computing system includes one or more computer processors and a system memory; The system memory stores a computer program according to claim 8, The steps of the method are performed by the one or more computer processors.

23. A computing system configured to:

Citation Information

Patent Citations

  • Neuromorphic circuit, neuromorphic array learning method and program

    WO2020129204A1