Adjustment Based on Weights in a Neural Network
The method addresses inefficiencies in training deep neural networks by dynamically updating weight values based on importance values and adjustment matrices, reducing training time and memory requirements while improving accuracy.
Patent Information
- Application Number
- JP2023530205
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2020-11-30
- Filing Date
- 2021-11-04
- Publication Date
- 2025-06-20
- Estimated Expiration
- 2041-11-04
AI Technical Summary
Existing methods for training deep neural networks require manual adjustment of hyperparameters and additional memory for storing intermediate values, leading to inefficiencies in convergence rate and training time.
A method for training neural networks that determines importance values for nodes based on current weight values, generates an adjustment matrix, and updates weight values using a combination of gradient values and adjustment matrix elements, eliminating the need for manual hyperparameter adjustment and additional memory.
This approach reduces training time, memory requirements, and improves accuracy by dynamically adjusting weight updates based on internal network dynamics, without the need for manual hyperparameter tuning or additional memory for intermediate values.
Smart Images

Figure 0007696429000071 
Figure 0007696429000072 
Figure 0007696429000073
Abstract
Description
Technical Field
[0001] The present invention generally relates to a computer-implemented method for training a neural network, and more particularly to a neural network including nodes and weighted connections between selected nodes.
Background Art
[0002] In the research community and the IT (information technology) organizations of companies, AI (artificial intelligence) and machine learning are currently the most advanced technologies that are mainstream. Multiple methods have been applied as effective tools for machine learning. For certain types of problems, artificial neural networks (ANNs) or deep neural networks (DNNs) may be well-suited for the technical architectures to support the applications of artificial intelligence.
[0003] A neural network requires training, which may be supervised, semi-supervised, or unsupervised, before it can be used for inference tasks such as classification or prediction. Currently, supervised learning techniques, which usually require multiple annotated training data, are often used. During training, based on the input data, the neural network generates one or more output signals that can be compared with the desired results (i.e., annotations). A function between these two may be used to adjust the connection or edge weight coefficient values between the nodes of different layers of the neural network.
[0004] Backpropagation is the most currently used algorithm for training deep neural networks in a wide variety of tasks. To address the problems of weight transport and symmetry in backpropagation (BP), various techniques have been developed. Some of them are feedback alignment (FA), direct feedback alignment (DFA), and indirect feedback alignment (IFA).
[0005] To minimize the loss function, BP, FA, DFA, and IFA rely on the method of stochastic gradient descent (SGD) and other optimizers. Basically, SGD can direct a combined set of learned weight values towards a global minimum or a local minimum within the space of variations. Thereby, the convergence of SGD strongly depends on the learning rate (η). Learning rate scheduling techniques can improve this convergence, but they may also require time-consuming manual adjustment and adaptation of hyperparameters, maintaining an equal learning rate for all parameters at each step.
[0006] Some optimizers, such as optimizers based on momentum, can perform heterogeneous updates, but they may also need to store the momentum estimates in memory (i.e., the main memory of the underlying computer system). This can sometimes be the reason for the higher computational cost of such techniques.
[0007] Therefore, the existing modifications of classical SGD introduced to improve the convergence rate disadvantageously require either manual adjustment of hyperparameters or additional memory.
[0008] Parameter training for image search or image classification is known. It is known to perform iterative calculations on an objective function using model parameters, and this objective function is a cost function used for image training.
[0009] The increasing complexity of deep learning architectures can result in increasingly long training times, requiring weeks or even months, due to the "vanishing gradient". It is known to train deep neural networks using a learning rate specific to each layer and network, which can adapt to the curvature of the function and increase the learning rate at low-curvature points. SUMMARY OF THE INVENTION
[0010] In one aspect of the present invention, a method, computer program product, and system for training a neural network include: (i) determining a set of importance values for a set of nodes of the neural network based on corresponding weight values of weighted connections between selected nodes of the set of nodes; (ii) determining an adjustment matrix including connection values that depend on the determined importance values of the set of nodes; (iii) determining a first updated value of a first weight value of a first weighted connection by a combination of a gradient value derived from a feedback signal regarding the first weighted connection and corresponding elements of the adjustment matrix, wherein the feedback signal represents a function of a desired activity and a current activity of the first weighted connection during a first training cycle; and (iv) applying the update to the weighted connections including the first weighted connection according to the adjustment matrix including the first updated value.
[0011] According to one aspect of the present invention, a computer-implemented method for training a neural network that may include nodes and weighted connections between selected ones of the nodes may be provided. Thereby, a function of desired and current activations during training results in a feedback signal that may be used to adjust the weight values of the connections for each weight value update cycle.
[0012] This method may include, for each update cycle, determining an importance value for each node based on the current weight value of the connection, and determining an adjustment matrix that includes values depending on the determined importance values. Further, this method may include, for each update cycle, determining a locally updated value specific to each weight value of the connection by a combination of a gradient value derived from a feedback signal regarding the connection and a determined corresponding element of the adjustment matrix, and applying the updates to the connections during all update cycles.
[0013] According to another aspect of the present invention, a neural network training system for training a neural network that may include nodes and weighted connections between selected ones of the nodes may be provided. Thereby, a function of desired and current activations during training results in a feedback signal that may be used to adjust the weight values of the connections.
[0014] This system may include a memory and a processor, and the memory stores program code portions to enable the processor to, for each update cycle, determine an importance value for each node based on the current weight value of the connection, determine an adjustment matrix that includes values depending on the determined importance values, determine a locally updated value specific to each weight value of the connection by a combination of a gradient value derived from a feedback signal regarding the connection and a determined corresponding element of the adjustment matrix, and apply the updates to the connections during all update cycles.
[0015] It should be noted that embodiments of the present invention are described with reference to various subjects. In particular, some embodiments are described with reference to method-type claims, while other embodiments are described with reference to apparatus-type claims. However, those skilled in the art will appreciate from the foregoing description and the following description that, unless otherwise noted, any combination of features belonging to one type of subject, in addition to any combination between features related to different subjects, particularly between features of method-type claims and features of apparatus-type claims, is also considered to be disclosed within this document.
[0016] The aspects defined above and further aspects of the present invention will become apparent from the examples of the embodiments described hereinafter and will be described with reference to the examples of the embodiments, but the present invention is not limited thereto.
[0017] Some embodiments of the present invention will be described below with reference to the following drawings by way of example only.
Brief Description of the Drawings
[0018]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
DETAILED DESCRIPTION OF THE INVENTION
[0019] In this specification, the training of a neural network including nodes and weighted connections between selected nodes among the nodes is described. A function of desired and current activations during training results in a feedback signal used to adjust the weight values of the connections. During a weight value update cycle, the process determines importance values of various nodes based on the current weight values of the connections, and determines an adjustment of the feedback signal specific to each weight value of the connection by a combination of the determined corresponding elements of the gradient value and the adjustment matrix derived from the feedback signal regarding the connection. During the update cycle, the updates are applied to the connections.
[0020] Some embodiments of the present invention recognize that in the prior art, the disadvantage of requiring the storage of accumulated intermediate values in memory remains unresolved. Therefore, there is a need to overcome the currently known limitations in the training of deep neural networks, thereby reducing not only the training time but also the amount of memory required and improving the accuracy of inference.
[0021] In the context of this description, the following rules, terms, or expressions, or combinations thereof may be used.
[0022] The term "neural network" (NN) may refer to a network of nodes and connections between nodes that is inspired by the brain and can be trained for inference, as opposed to procedural programming. The nodes may be organized within layers, and the connections may have weight values that represent the selective strength of the relationships between selected ones of the nodes. The weight values define the parameters of the neural network. The neural network may be trained using sample data, for example, for the classification of data received at the input layer of the neural network, and the result of the classification may be made available at the output layer of the neural network, together with a confidence value. A neural network that includes a plurality of hidden layers (in addition to the input layer and the output layer) is typically referred to as a deep neural network (DNN).
[0023] The term "node" may refer to a computational node that represents a computational function (e.g., ReLU) having an output value that may depend on a plurality of received signals via a plurality of connections from layers upstream of the node within the neural network.
[0024] The term "weighted connection" may refer to a link or edge between nodes within the neural network. The strength of the link between nodes in different adjacent layers may be represented by the weight value of the connection.
[0025] The term "feedback signal" may refer to a signal derived from the current output and the expected output of the neural network, and may simply be the difference between two values. The feedback signal may be a more complex function and may be a function of the actual output and the target output. The feedback signal may be used, for example, by backpropagation, to further adapt the weight values of the neural network.
[0026] The term "Hadamard product" refers to a special mathematical product of two matrices of equal dimension, which can generate another matrix of the same dimension as the operands, and each element i, j is the product of the elements i, j of the original two matrices.
[0027] The term "weight value" may represent the strength value of the connection between two nodes within a neural network. The weight value of the connection may be signed and may be defined such that it is excitatory when positive and inhibitory when negative. The connection may be strong when the absolute value of the weight is relatively large and weak otherwise.
[0028] The term "update cycle" may indicate the period (which may also be shown as a parameter) before the weight values of a neural network can be updated. This update may be performed (in a simple case) after a single training sample or multiple training samples. This period may be adaptable or configurable or both, and may be self-optimized according to the learning results of the neural network. The update of the weight values may be performed, for example, after 32 or 64 training samples, and may depend on the total number of training samples and the algorithm for changing the period of the update cycle during the training process, or may depend on the exceeding of a threshold value. When the update of the weights can be performed after a group of training samples called a "mini-batch", the adjustment may be performed for each sample, the gradients adjusted for each sample are accumulated, and after the last sample, the weights are updated.
[0029] The term "importance value" may indicate a numerical value assigned to a selected node within a selected layer of a neural network. In one embodiment, the importance value may be derivable as the sum of all weight values of the received connections to the selected node (or, for example, the absolute value (in the mathematical sense) of the weight values). In an alternative embodiment, the importance value may be the sum of the weight values of all the outgoing connections of the node. Generally, the greater the sum of the absolute weight values, the greater the importance. The importance value may be regarded as the responsibility of a particular node for influencing the signal moving from the input layer to the output layer through the NN.
[0030] The term "adjustment matrix" may indicate a matrix of scalar values in the range [1, 2] derived from the importance values. The adjustment matrix may be regarded as the importance values normalized to a particular range.
[0031] The term "gradient value derived from the feedback signal" may indicate a direction in a multi-dimensional space that ambiguously indicates the direction of the minimum value within the parameter space of the learning neural network. Generally, when the learning process can be completed (i.e., when the loss function can be optimized), the minimum value can be reached.
[0032] The term "connections received by the node" may indicate all connections that end at the node. Since the node is positioned behind the synapses represented by the connections, it may be indicated as postsynaptic in this case.
[0033] In contrast to the received connections, the term "outgoing connections of the node" may indicate all connections that "leave" from the selected node. Such a node may be indicated as presynaptic since it is positioned before the synapses (i.e., connections) within the NN.
[0034] The term "upstream layer" may be indicated (visible from the selected layer) such that those layers are positioned closer to the input layer of the NN.
[0035] The term "backpropagation" (BP) may denote an algorithm widely used in training feedforward neural networks for supervised learning. In fitting (i.e., training) a neural network, the backpropagation method determines the gradient of the loss function with respect to the weights of the NN for single (or multiple) input-output training data. Here, as an example, gradient descent such as stochastic gradient descent (SGD), a variant thereof, should be mentioned. Generally, the backpropagation algorithm may function by determining the gradient of the loss function with respect to each weight by means of the chain rule that determines the gradient of one layer at a time, and in the chain rule, iteration is performed in the reverse direction from the last NN layer in order to avoid redundant calculations of intermediate terms.
[0036] The term "direct feedback alignment" (DFA) may also denote a form of error propagation for training neural networks. It has been discovered that the weight values used for the propagation of the error of the output in the reverse direction (when compared with the desired output) do not have to be symmetric with the weight values used for the propagation of the activation in the forward direction. In fact, for the network to learn how to make use of the feedback, randomly generated feedback weights function uniformly well enough. The principle of feedback alignment can be used for training in layers from zero initial conditions, more independently of other parts of the network. Thereby, the error propagated by the fixed random feedback may link directly from the output layer to each hidden layer. This simple method may reach a training error of almost zero even in convolutional networks and very deep networks without any backpropagation of error.
[0037] The term "feedback alignment" (FA) may be regarded as a simpler form of DFA, in which access to the layers for weight value update is performed stepwise in the upstream direction of the NN.
[0038] The term "Indirect Feedback Alignment" (IFA) may denote another form of updating the weight values of the connections between the nodes of the NN. Here, the error may be propagated from the output layer to the first hidden layer, where the weight values of the connection layer may be updated, and then may affect the next downstream layer.
[0039] The term "Stochastic Gradient Descent" (SGD) may denote a known iterative method for optimizing an objective function having appropriate smoothness properties. Thus, Stochastic Gradient Descent may be regarded as a probabilistic approximation of gradient descent optimization, since it can replace the actual gradient (computed from the entire data set) with an estimate (computed from a randomly selected subset of the data) made from the actual gradient.
[0040] The term "Adam Optimizer Algorithm" may denote a known extension to Stochastic Gradient Descent and is often used in computer vision and natural language processing. This algorithm may be used to replace the classical SGD method and iteratively update the weight values of the network. This algorithm may achieve a high convergence rate. While the SGD method maintains a single learning rate for all weight updates during complete training, the Adam method may determine individual adaptive learning rates for different parameters (weight values) from estimates of the first and second moments of the gradient.
[0041] The term "Nesterov Accelerated Gradient" (NAG) may denote a known modified SGD method, named after its inventor. The NAG method includes a gradient descent step for a momentum-like term that is not the same as the term used in classical momentum. This algorithm is known to converge faster and be more efficient than the classical SGD method.
[0042] The term "RMSprop" may also refer to an optimization technique based on gradients used in the training of neural networks. The gradients of very complex functions such as neural networks tend to vanish or grow explosively as data propagates through the function. RMSprop was developed as a stochastic technique for mini-batch learning. RMSprop may use a moving average of the squared gradients to normalize the gradients. This normalization balances the step size (momentum) by reducing the step in the case of large gradients to prevent explosive increase and increasing the step in the case of small gradients to prevent vanishing.
[0043] Details of each figure will be described below. All indications in the figures are schematic diagrams. First, a block diagram of an embodiment of a computer-implemented method of the present invention for training a neural network will be described. Thereafter, in addition to embodiments of a neural network training system, further embodiments will be described.
[0044] FIG. 1 shows a block diagram of a preferred embodiment of a method 100 for training a neural network (in particular, for training connection / edge parameters / weight values). A neural network includes nodes and weighted connections between selected ones of the nodes. Generally, the concepts proposed herein relate to deep neural networks that include multiple hidden layers. It is not necessary for each node of a layer to be connected to an adjacent (i.e., upstream or downstream) layer. This can be easily implemented by setting the weight value associated with the connection to zero (in other words, "removing the connection") with a probability selected for each input presentation.
[0045] Method 100, based on the current weight values of the connections, for each node, an importance value (e.g., I1, I2,..., I nincluding determining (102). In particular, it means connecting a specific node. The position of the node may be pre-synaptic (i.e., the transmitting connection) or post-synaptic (i.e., the receiving connection to the node).
[0046] Method 100 also includes determining (104) an adjustment (i.e., regulation) matrix that includes values that depend on the determined importance values. This will be explained in more detail below.
[0047] Furthermore, method 100 determines (106) a locally updated value specific to each weight value of the connection by a combination (in particular, the Hadamard product) of the gradient value derived from the feedback signal regarding the connection and the determined corresponding element of the adjustment matrix, and includes applying (108) the updates to the connections during all update cycles.
[0048] Thereby, the term update cycle does not necessarily indicate a cycle after each training sample. Instead, one or more training samples can be considered as a group of training samples, for example, every 32 or 64 training samples, before an update cycle can be executed. Thereby, the updates are accumulated, but the weight values are effectively changed only (i) after a certain time or (ii) after the adjustment exceeds a specific threshold. A simple solution is that the proposed method can determine the adjustment, the update is a function of the adjustment, and a certain degree of freedom is given when the update is applied.
[0049] Figure 2 shows a block diagram of an embodiment of a neural network 200 including a plurality of layers 202, 204, 206, 208, 210, 212. However, an actual deep neural network, or a deep neural network actually used, may include a much larger number of nodes and layers. In this example, layer 204 represents the upstream layer l-1 with respect to node layer l206. The downstream layers with respect to layer l206 are layer l+1 208 and layer l+2 210. Layer 212 represents the output layer of neural network 200, and layer 202, which is a layer of nodes all symbolized by circles, represents the input layer of the neural network.
[0050] Furthermore, different types of connections between nodes are shown. Strong connections (i.e., strong synapses) having relatively large weight values are represented by thick solid lines. Weak synapses may be represented by dashed lines. Strengthening a connection may be represented by another type of dashed line. Finally, weakening the connection between two nodes in different layers may be represented by a dashed line including short partial lines and long partial lines. Generally, the upstream direction within neural network 200 is from right to left. Thus, the downstream direction within neural network 200 is from the left side (input side) towards the output layer 212 (i.e., the right side).
[0051] Figure 3 shows an embodiment of a matrix 300 useful for deriving the importance value of a particular node. Thereby, importance or importance value can be understood as the average responsibility of a node with respect to the propagation of input signals to the neural network and thus with respect to the final error based on the current distribution of weight values. Here, the weight value is w j,i (where j corresponds to "post" and i corresponds to "pre"). Vector 302 represents the importance values of the nodes of a particular layer within the neural network. In the example shown, in particular, the importance (importance value) of the post-synaptic node is constructed in the absolute term by adding the received weight value to a particular node i. Here, the importance value of node i is Ii That is. Therefore, the first component of the vector 302 is the importance value of the first (topmost) node of each layer of nodes in the neural network.
[0052] FIG. 4 shows a determination 400 of update values, also shown as a matrix useful for the next step of determining weight updates. The determination of the weight update factor is shown as a vector 402. The vector 402 shows an operation of an importance vector to obtain local update factor values (also shown as local modulation factor (i.e., "M")) that include amounts restricted to the range [1, 2]. To obtain the local update matrix 404, the local update factor values are repeated the same number of times as the size of the layer. As an example, the top left value M1 represents the first local update value for all connections located to receive at the first node (i.e., of the node after the first synapse).
[0053] FIG. 5 shows the steps of a determination 500 of weight updates based on the importance of a particular node in the "local" embodiment. Based on the preparatory procedures described in FIGS. 3 and 4, in the case of the example shown here of the four nodes in the layer, the updated weight values can be determined here by multiplying the local update matrix by the gradient Δw determined by the Hadamard product, i.e., by the selected learning rule, to adjust or affect the update of the weight values, resulting in a matrix 502.
[0054] In contrast to the local embodiments, due to the higher complexity of weight updates in the non-local embodiments, FIG. 6 shows a diagram 600 of the equations used for non-local embodiments of weight value updates. Thereby, the determination of the importance coefficient and the determination of the local update value (or adjustment coefficient value) are the same when compared to the local version of the previous figure. In the non-local version of the concept proposed herein, in the third step, the feedback signal is adjusted by the local update / adjustment coefficient. The diagram of FIG. 6 shows how the adjustment in the propagating version is applied in a neural network with two hidden layers trained using BP. Thereby, the following symbols are used. W1, W2, W3 are feed-forward weights. x, h1, h2, y are layer activations. M1, M2, M3 are adjustment matrices for each layer. f is the activation function. f’ is the derivative of the activation function. _ T denotes the transpose of a vector or matrix. t is the time step. η is the learning rate.
Number
[0055] Note also that "x" represents the input layer, while "target" represents the target (i.e., the desired output of the neural network). Nodes within different layers are symbolized by circles, and the number of circles / nodes (actually four per layer) is merely a symbolic representation. The layers of nodes may have a uniform size or may contain a different number of nodes per layer.
[0056] Simulation results based on the MNIST database (the well-known Modified National Institute of Standards and Technology containing handwritten data samples) for the proposed concept embodiments indicate that significantly fewer training epochs are required to achieve the same accuracy compared to the classical SGD process. It was also found that the same holds true when the MNIST database is replaced by the Fashion MNIST database.
[0057] In the simulation setup, as already mentioned, the MNIST database and the Fashion MNIST database were used. As the training algorithm, backpropagation was used, and the determination of the update adjustment value (i.e., the adjustment value) was grouped by the post-synaptic nodes and propagated upstream. The training / validation split ratio was 90:10. As the activation function for the nodes, ReLU was used within an NN containing 10 layers and 256 neurons / nodes per layer. The dropout rate was 0.1. The results shown are the average values over 3 simulation runs.
[0058] As can be easily seen, the proposed concept of adjustment / update values shows a faster convergence rate during training and a higher classification accuracy during testing. Therefore, at the start of training, a baseline algorithm with a constant learning rate equal to the average learning rate of the adjusted SGD (in the form of post-synaptic importance coefficient determination) was used.
[0059] When comparing the test accuracies of the MNIST database and the Fashion MNIST database, the results shown in Table 1 below are obtained.
[0060]
Table 1
[0061] Regarding alternative embodiments of the proposed concept, significant differences can be shown between classical SGD and regulated SGD. The regulation of SGD can be applied to various training algorithms that rely on SGD, such as (i) backpropagation, (ii) feedback alignment (FA), and its variants direct feedback alignment (DFA) and indirect feedback alignment (IFA).
[0062] In the preliminary investigation results setting, the extended MNIST database (for DFA) and MNIST (for FA) were used as the datasets. The training algorithms were DFA and FA. The update of the adjustment value was determined by grouping pre (DFA) or post (FA) and propagating upstream (both). Here too, the training / validation split ratio was 90:10. As the activation functions, ReLU (for DFA) and tanh (for FA) were used. The neural network was composed of three hidden layers each containing 256 nodes. There was no dropout, and the results are the average values over 5 simulation runs. The accuracy can be significantly improved compared to the training of classical SGD and the validation of SGD.
[0063] Table 2 shows the training data using DFA for the extended MNIST database, as well as the training and validation using FA for the MNIST database. The following results in Table 2 were measured.
[0064]
Table 2
[0065] The proposed computer implementation method for training a neural network can provide multiple advantages, technical effects, contributions, or improvements, or combinations thereof, including the following: (i) In contrast to existing modifications of classical SGD introduced to improve the convergence rate of connection weight values within a neural network during training, the concept proposed herein requires no manual adjustment of hyperparameters and no additional memory for storing intermediate values used in the calculation of weight updates. (ii) The proposed concept can adjust, affect, or update the weight value updates based on the internal dynamics of the neural network in a manner that can be achieved by directly adjusting the learning rate of specific parameters in a heterogeneous manner, and thus the proposed concept can effectively enable training with fewer training cycles and may also require less memory compared to momentum-based methods. Furthermore, the proposed concept can enable higher classification accuracy during testing. (iii) This method may be dynamically adapted to various data sets and architectures of neural networks regardless of the depth of the NN, the size of the hidden layer, and the activation function used. (iv) Different from other modifications of SGD such as momentum-based optimizers, the proposed adjustment of SGD does not require storing the accumulated gradients from previous training steps with additional memory cost for use in state-of-the-art concepts such as algorithms and AI chips inspired by neuromorphic biology. (v) The accuracy of the inference results can be improved, or the number of training epochs required to reach the target accuracy can be reduced, or both.
[0066] The proposed computer implementation method for training a neural network can provide a plurality of advantages, technical effects, contributions, or improvements, or combinations thereof, including the following: (i) it can be stated that the adjustment helps to improve the performance of the trained model when the complexity of the model increases and classical SGD cannot utilize the model; (ii) determining the importance value for each node may include constructing the sum of the weight values of the connections received by the node, and the weight values used to construct the sum may be absolute values (without sign); (iii) this post-synaptic determination of importance can lead to better results in training the neural network than the pre-synaptic determination of importance (this term reflects that connections, links, or edges within the neural network represent synapses, considering the fundamental idea of the mechanism based on the mammalian brain); (iv) determining the importance value for each node may include determining the sum of the weight values of the connections of the node or emanating from the node, and this connection may be represented as pre-synaptic as the connection leaves each node, i.e., exits from each node, and absolute weight values may be used; or (v) the adjustment value specific to each weight value of the connections within one layer of the connections affects at least one of the upstream layers of the connections within the neural network (this is specifically the "adjustment value" from the adjustment matrix, not the "updated value" that also affects the upstream layer and can be used to calculate the updated value), and such types of embodiments may be indicated as "non-local", or a combination thereof.
[0067] The proposed computer implementation method for training a neural network can provide multiple advantages, technical effects, contributions, or improvements, or a combination thereof, including that the updated value specific to each weight value of the connections within one layer of connections is neutral (i.e., has no influence) with respect to all upstream layers of the connections within the neural network. That is, the proposed computer implementation method does not have to affect the feedback signal updated by the adjustment matrix. In contrast to the "non-local" version, this embodiment may be shown as "local". Thus, in other words, in the embodiment of the local version of the proposed concept, the update of the weight of the connection from the pre-synaptic node pre in layer l-1 to the post-synaptic node a in layer l obtained by applying the adjustment may be expressed as follows.
[0068] [Number] where η is the learning rate, [Number] is the adjustment coefficient of all connections to the post-synaptic node a in layer l.
[0069] The proposed computer implementation method for training a neural network can provide a plurality of advantages, technical effects, contributions, or improvements, or combinations thereof, including the following. (i) The feedback signal may be based on one selected from a group including backpropagation (BP), feedback alignment (FA), particularly variants thereof, such as direct feedback alignment (DFA), and indirect feedback alignment (IFA). Thus, the proposed concept can function well enough using the standard feedback signal of the neural network that is often used recently. Or, (ii) The update of the connection weight value specific to each node within each layer may include multiplying the gradient value derived from the feedback signal by an adjustment coefficient (specifically, an adjustment coefficient value). This can function correctly for both the local version and the non-local version using the adjustment coefficient. Or both.
[0070] The proposed computer implementation method for training a neural network has a postsynaptic decision, and the importance value
Number
Number
[0071]
Number
Number
[0072] The proposed computer implementation method for training a neural network is the importance value for each node a in layer l of the neural network at time step t.
Number
Number
Number
Number
[0073]
Number
[0074] The proposed computer-implemented method for training a neural network is (i) with non-local updates - this method involves adjusting coefficient values [Number] by directly applying [Number] and may include determining [Number] is the weight value connecting node pre(pre) in layer l-1 to node a (post=a) in layer l at time step t, or (ii) for training, one method that can be selected from the group including stochastic gradient descent, Adam optimizer method, Nesterov accelerated gradient, and RMSprop may be used, or both, which may provide multiple advantages, technical effects, contributions, or improvements, or combinations thereof. This makes the concepts proposed herein basically independent of the optimizer used. Other optimizers can also be used normally.
[0075] For the sake of completeness, FIG. 7 shows an embodiment of the neural network training system 900 of the present invention for training a neural network. The neural network includes nodes 906 and weighted connections 908 between selected nodes among the nodes 906, and the function of the desired and current activities during training results in a feedback signal used to adjust the weight values of the connections. The system 900 also includes a memory 902 and a processor 904. Thereby, for each update cycle, the memory 902 enables the processor to determine the importance value for each node based on the current weight value of the connection (e.g., by the first decision unit 910), and determine an adjustment matrix (e.g., by the second decision unit 912) that includes values depending on the determined importance value.
[0076] Furthermore, the stored program code portion enables the processor 904 to determine a locally updated value specific to each weight value of the connection (e.g., by a third decision unit) based on the combination of the gradient value derived from the feedback signal regarding the connection and the determined corresponding element of the adjustment matrix, and apply the update to the connection (e.g., by the application unit 912) during all update cycles.
[0077] Accordingly, all elements of the neural network training system 900 (in particular, all modules and units in addition to the memory 902 and the processor 904) should be considered to have the option of being implemented in hardware. In that case, the units and modules are electrically connected for the exchange of data and signals. In particular, specific memories for the nodes 906 and the connections 908 may exist.
[0078] According to the above-described implementation options, in particular, the first decision unit 910, the second decision unit 912, the third decision unit 914, and the application unit 916 may be implemented entirely in hardware or in a combination of software elements and hardware elements. All active and passive units, modules, and components of the system 900 may exchange signals and data directly with each other or may utilize the internal bus system 918 of the system 900.
[0079] Embodiments of the present invention may in fact be implemented with any kind of computer, regardless of the platform suitable for storing or executing or both the program code. FIG. 8 shows, by way of example, a computing system 1000 suitable for executing program code related to the proposed method.
[0080] Computing system 1000 is merely an example of a suitable computer system, and is not intended to suggest any limitation as to the scope of use or functionality of the embodiments of the invention described herein, whether or not computing system 1000 is capable of implementing, executing, or both, any of the functions shown above. There are components in computing system 1000 that can operate with a number of other general-purpose or special-purpose computing system environments or configurations. Examples of well-known computing systems, environments, or configurations, or combinations thereof, suitable for use with computing system / server 1000 include, but are not limited to, personal computer systems, server computer systems, thin clients, thick clients, handheld or laptop devices, microprocessor systems, microprocessor-based systems, set-top boxes, programmable consumer electronics, network PCs, microcomputing systems, mainframe computer systems, and distributed cloud computing environments including any of these systems or devices. Computing system / server 1000 may be described in the general context of computer-executable instructions, such as program modules, being executed by a computer system. Generally, program modules may include routines, programs, objects, components, logic, data structures, etc. that perform particular tasks or implement particular abstract data types. Computing system / server 1000 may be implemented in a distributed cloud computing environment where tasks are performed by remote processing devices linked through a communications network. In a distributed cloud computing environment, program modules may be located in both local and remote computer system storage media including memory storage devices.
[0081] As shown in the figure, the computer system / server 1000 is shown in the form of a general-purpose computing device. The components of the computer system / server 1000 may include, but are not limited to, one or more processors or processing units 1002, a system memory 1004, and a bus 1006 that couples various system components including the system memory 1004 to the processor 1002. The bus 1006 represents one or more of any of a plurality of types of bus structures including a memory bus or memory controller, a peripheral bus, an accelerated graphics port, and a processor or local bus that uses any of a variety of bus architectures. By way of example, such architectures include, but are not limited to, an ISA (Industry Standard Architecture) bus, an MCA (MicroChannel Architecture) bus, an EISA (Enhanced ISA) bus, a VESA (Video Electronics Standards Association) local bus, and a PCI (Peripheral Component Interconnects) bus. The computer system / server 1000 typically includes various computer system readable media. Such media may be any available media that is accessible by the computer system / server 1000 and includes both volatile and nonvolatile media, removable and non-removable media.
[0082] System memory 1004 may include a computer system readable medium in the form of volatile memory, such as random access memory (RAM) 1008, cache memory 1010, or both. The computer system / server 1000 may further include other removable / non-removable, volatile / non-volatile computer system storage media. By way of example only, a storage system 1012 may be provided for reading from and writing to a non-removable, non-volatile magnetic medium (not shown, typically called a “hard drive”). Although not shown, a magnetic disk drive for reading from and writing to a removable, non-volatile magnetic disk (e.g., a “floppy (R) disk”), and an optical disk drive for reading from and writing to a removable, non-volatile optical disk such as a CD-ROM, DVD-ROM, or other optical media may be provided. In such instances, each may be connected to the bus 1006 by one or more data media interfaces. As further shown and described below, the memory 1004 may include at least one program product including a series of (e.g., at least one) program modules configured to execute the functions of embodiments of the present invention.
[0083] For example, a program / utility having a series of (at least one) program modules 1016 may be stored in the memory 1004, but is not limited thereto, and an operating system, one or more application programs, other program modules, and program data may also be stored. Each of the operating system, one or more application programs, other program modules, and program data, or combinations thereof, may include an implementation of a network environment. The program modules 1016 typically execute the functions or methods of embodiments of the present invention described herein, or both.
[0084] The computer system / server 1000 may communicate with one or more external devices 1018 such as a keyboard, a pointing device, a display 1020, one or more devices that enable a user to interact with the computer system / server 1000, or any device (e.g., a network card, a modem, etc.) that enables the computer system / server 1000 to communicate with one or more other computing devices, or a combination thereof. Such communication may occur via an input / output (I / O) interface 1014. Further, the computer system / server 1000 may communicate with one or more networks such as a local area network (LAN), a general wide area network (WAN), or a public network (e.g., the Internet), or a combination thereof, via a network adapter 1022. As shown, the network adapter 1022 may communicate with other components of the computer system / server 1000 via a bus 1006. Although not shown, it should be understood that other hardware components or software components or both may be used in conjunction with the computer system / server 1000. Examples include, but are not limited to, microcode, device drivers, redundant processing units, external disk drive arrays, RAID systems, tape drives, and data archive storage systems.
[0085] Further, a neural network training system 900 may be connected to the bus system 1006.
[0086] The description of various embodiments of the present invention is presented for illustrative purposes and is not intended to be exhaustive or limited to the disclosed embodiments. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the described embodiments. The terms used herein are selected to best explain the principles of the embodiments, the practical application, or the technical improvements over technologies found in the marketplace, or to enable other skilled artisans to understand the embodiments disclosed herein.
[0087] The present invention may be embodied as a system, a method, or a computer program product, or a combination thereof. The computer program product may include a computer-readable storage medium including computer-readable program instructions for causing a processor to execute aspects of the present invention.
[0088] This medium may be an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system for a propagation medium. Examples of computer-readable media include semiconductor memory or solid-state memory, magnetic tape, removable floppy (R) disk, random access memory (RAM), read-only memory (ROM), rigid magnetic disk, and optical disk. Current examples of optical disks include compact disk read-only memory (CD-ROM), compact disk read / write (CD-R / W), and Blu-ray disk.
[0089] A computer-readable storage medium can be a tangible device that can hold and store instructions for use by an instruction execution device. The computer-readable storage medium can be, for example, an electronic storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination thereof, but is not limited thereto. A non-exhaustive list of more specific examples of computer-readable storage media includes portable floppy (R) disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), static random access memory (SRAM), portable compact disk read-only memory (CD-ROM), digital versatile disk (DVD), memory stick, floppy (R) disk, mechanically encoded devices such as punch cards or raised structures in grooves in which instructions are recorded, and any suitable combination thereof. As used herein, a computer-readable storage medium should not be construed to be a signal per se that is transient, such as a radio wave or other freely propagating electromagnetic wave, an electromagnetic wave propagating through a waveguide or other transmission medium (e.g., an optical pulse passing through an optical fiber cable), or an electrical signal transmitted via a wire.
[0090] The computer-readable program instructions described herein can be downloaded from a computer-readable storage medium to each computing / processing device, or to an external computer or external storage device via a network (e.g., the Internet, a local area network, a wide area network, or a wireless network, or a combination thereof). This network may comprise copper transmission cables, optical transmission fibers, wireless transmission, routers, firewalls, switches, gateway computers, or edge servers, or a combination thereof. A network adapter card or network interface within each computing / processing device receives the computer-readable program instructions from the network and transfers them for storage on a computer-readable storage medium within each computing / processing device.
[0091] Computer-readable program instructions for carrying out the operations of the present invention may be source code or object code written in any combination of one or more programming languages, including assembly instructions, instruction-set-architecture (ISA) instructions, machine instructions, machine-dependent instructions, microcode, firmware instructions, state-setting data, or object-oriented programming languages such as Smalltalk(R), C++, and conventional procedural programming languages such as the "C" programming language or similar programming languages. The computer-readable program instructions may be executed entirely on the user's computer, partially on the user's computer as a stand-alone software package, partially on the user's computer and a remote computer, or entirely on the remote computer or server. In the latter scenario, the remote computer may be connected to the user's computer via any type of network, including a local area network (LAN) or a wide area network (WAN), or the connection may be made to an external computer (e.g., via the Internet using an Internet service provider). In some embodiments, to carry out aspects of the present invention, an electronic circuit, including, for example, a programmable logic circuit, a field-programmable gate array (FPGA), or a programmable logic array (PLA), may execute computer-readable program instructions for customizing the electronic circuit by utilizing the state information of the computer-readable program instructions.
[0092] Aspects of the present invention will be described herein with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer readable program instructions.
[0093] These computer readable program instructions may be provided to a processor of a general purpose computer, special purpose computer, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, create means for implementing the functions / acts specified in one or more blocks of the flowchart and / or block diagram. These computer readable program instructions may be stored in a computer readable storage medium that includes instructions for causing a computer, programmable data processing apparatus, or other device to function in a particular manner so that the product comprises an article of manufacture including instructions for implementing the aspects of the functions / acts specified in one or more blocks of the flowchart and / or block diagram.
[0094] The computer readable program instructions may also be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable apparatus, or other device such that the instructions which execute on the computer, other programmable apparatus, or other device implement the functions / acts specified in one or more blocks of the flowchart and / or block diagram.
[0095] The flowcharts and block diagrams in the figures, or both, illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present invention. In this regard, each block in the flowchart or block diagram may represent a module, segment, or portion of instructions that comprises one or more executable instructions for implementing the specified logical function. In some alternative implementations, the functions noted in the blocks may occur out of the order noted in the figures. For example, two blocks shown in succession may, in fact, be executed substantially concurrently, depending on the functionality involved, or may sometimes be executed in the reverse order. It should also be noted that each block of the block diagrams or flowcharts, or both, and combinations of blocks in the block diagrams or flowcharts, or both, can be implemented by a dedicated hardware-based system that performs the specified functions or operations, or a combination of dedicated hardware and computer instructions.
[0096] The terms used herein are for the purpose of describing particular embodiments only and are not intended to be limiting of the invention. As used herein, the singular forms "a", "an", and "the" are intended to include the plural forms as well, unless the context clearly indicates otherwise. It will be further understood that the terms "comprises" and / or "comprising", when used in this specification, specify the presence of stated features, integers, steps, operations, elements, or components, or combinations thereof, but do not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components, or groups thereof, or combinations thereof.
[0097] All means or steps and corresponding structures, materials, acts, and equivalents within the scope of the following claims are intended to include any structure, material, or act for performing functions in combination with other claimed elements specifically claimed. The description of the invention has been presented for purposes of illustration and description, but is not intended to be exhaustive or limited to the invention in the disclosed form. Many modifications and variations will be apparent to those of ordinary skill in the art without departing from the scope and spirit of the invention. Embodiments are chosen and described in order to best explain the principles of the invention and its practical application, and to enable others of ordinary skill in the art to understand the invention for various embodiments with various modifications as are suited to the particular use contemplated.
[0098] In summary, the concept of the present invention can be summarized by the following paragraphs.
[0099] A method for training a neural network, the neural network including nodes and weighted connections between selected ones of the nodes, a function of desired and current activations during training yielding a feedback signal used to adjust the weight values of the connections, the method comprising, for each weight value update cycle, determining an importance value for each node based on the current weight values of the connections to the node, determining an adjustment matrix including values depending on the determined importance values, determining a locally updated value specific to each weight value of the connections by a combination of a gradient value derived from the feedback signal for the connections and the determined corresponding elements of the adjustment matrix, and applying the updates to the connections during all update cycles.
[0100] The method, wherein determining the importance value for each node includes constructing a sum of the weight values of the connections received by the node.
[0101] The method, wherein determining the importance value for each node includes determining a sum of the weight values of the connections transmitted by the node.
[0102] The importance value for each node a in layer l of the neural network at time step t [Number] is the sum of the absolute strengths of all the weight values of layer l including the postsynaptic neuron with a being a [Number] is the sum of the absolute strengths of all the weight values of layer l including the postsynaptic neuron with a being a
[0103] The importance value for each node a in layer l of the neural network at time step t [Number] is the sum of the absolute strengths of all the weight values of layer l+1 including the presynaptic neuron with a being a [Number] is the sum of the absolute strengths of all the weight values of layer l+1 including the presynaptic neuron with a being a
[0104] At time step t, for node a in layer l, the local update value of the adjustment matrix [Number] is the importance value of node a [Number] is determined as a value limited by 1 and multiplied by 2 to the ratio of the importance value of node a to the maximum importance among all the neurons in layer l
[0105] The local update value specific to each weight value of the connections within one layer of the connections [Number] A method that also affects at least one of the layers upstream of the connection within the neural network.
[0106] An update value specific to each weight value of the connections within one layer of the connection
Number
[0107] The update of the weight value of the connection specific to each node within each layer is the gradient value derived from the feedback signal multiplied by an adjustment coefficient value
Number
[0108] Local update value
Number
Number
Number
[0109] The method according to any of the preceding clauses, wherein the feedback signal is based on one selected from the group including backpropagation, feedback alignment, direct feedback alignment, and indirect feedback alignment.
[0110] For training, the method according to any of the preceding claims, wherein one method selected from the group including stochastic gradient descent, Adam optimizer method, Nesterov accelerated gradient, and RMSprop is used.
[0111] A neural network training system for training a neural network, wherein the neural network includes nodes and weighted connections between selected nodes among the nodes, and a function of desired activity and current activity during training results in a feedback signal used to adjust the weight values of the connections, the system comprising a memory and a processor, the memory storing a program code portion for enabling the processor to determine an importance value for each node based on the current weight value of the connection for each update cycle, determine an adjustment matrix including values depending on the determined importance value, determine an updated value specific to each weight value of the connection by a combination of the gradient value derived from the feedback signal regarding the connection and the determined corresponding element of the adjustment matrix, and apply the update to the connection during all update cycles.
[0112] A system that enables the program code portion to enable the processor to also construct the sum of the weight values of the connections received by the node during the determination of the importance value for each node.
[0113] A system that enables the program code portion to enable the processor to also determine the sum of the weight values of the connections transmitted by the node during the determination of the importance value for each node.
[0114] The importance value for each node a in layer l of the neural network at time step t
Equation
Number
[0115] The importance value for each node a in layer l of the neural network at time step t
Number
Number
[0116] At time step t, for node a in layer l, the local update coefficient value of the adjustment matrix
Number
Number
[0117] The local update coefficient value specific to each weight value of the connections in one layer of the connections
Number
[0118] The local update coefficient value specific to each weight value of the connections in one layer of the connections
Number
[0119] During the update of the connection weight values specific to each node within each layer, a program code portion enables the processor to multiply the gradient value derived from the feedback signal by an adjustment coefficient value
Number
[0120] Adjustment coefficient value
Number
Number
Number
[0121] A system where the feedback signal is based on one selected from a group including backpropagation, feedback alignment, direct feedback alignment, and indirect feedback alignment.
[0122] For training, a system where one method selected from a group including stochastic gradient descent, Adam optimizer method, Nesterov accelerated gradient, and RMSprop is used.
[0123] A computer program product for training a neural network, wherein the neural network includes nodes and weighted connections between selected ones of the nodes, and a function of a desired activity and a current activity during training yields a feedback signal that is used to adjust the weight values of the connections, the computer program product comprising a computer-readable storage medium having program instructions embodied thereon, the program instructions being executable by one or more computing systems or controllers, and causing the one or more computing systems to: determine an importance value for each node based on the current weight values of the connections; determine an adjustment matrix that includes values that depend on the determined importance values; determine an updated value specific to each weight value of the connections by a combination of a gradient value derived from the feedback signal for the connections and a corresponding element of the determined adjustment matrix; and apply the updates to the connections during all update cycles.
Claims
1. A method for training a neural network, comprising: determining a set of importance values for a set of nodes of the neural network based on corresponding weight values of weighted connections between selected nodes of the set of nodes; determining a local update matrix that includes a value depending on the determined importance values of the set of nodes; determining a first updated value of a first weight value of the first weighted connection by a combination of a gradient value derived from a feedback signal regarding the first weighted connection and a corresponding element of the local update matrix, wherein the feedback signal represents a function of a desired activity and a current activity of the first weighted connection during a first training cycle; applying an update to the weighted connections including the first weighted connection according to a matrix including the first updated value.
2. training cycles alternate with update cycles, and the applying of the update occurs after the first training cycle during a first update cycle, according to Claim 1.
3. The determining of the set of importance values includes constructing a sum of weight values of identified weighted connections received by the set of nodes, according to Claim 1 or Claim 2.
4. The importance value for each node a in layer l of the neural network at time step t 【Equation 1】 is the sum of the absolute strengths of all weight values 【Equation 2】 of layer l including the postsynaptic neurons, according to Claim 3.
5. determining the importance value comprises determining the sum of the weighted connection weight values originating from the set of nodes, the method according to any one of claims 1 to 4.
6. the importance value for each node a in layer l of the neural network at time step t 【Equation 3】 is the sum of the absolute strengths of all the weight values of layer l+1 containing the presynaptic neurons 【Equation 4】 the method according to claim 5.
7. At time step t, for node a in layer l, the first update value of the local update matrix 【Equation 5】 is the importance value of node a 【Equation 6】 multiplied by 2 and the ratio with the maximum importance value among all the neurons in layer l to form a product, the product is determined to be limited by 1 at the lower limit, the method according to any one of claims 1 to 6.
8. the first update value specific to each weight value of the weighted connections in one layer of the weighted connections 【Equation 7】 affects at least one upstream layer of the weighted connections in the neural network, the method according to claim 7.
9. the first update value 【Equation 8】 directly applying and 【Equation 9】 further includes determining 【Equation 10】 The method according to claim 8, wherein the weight value connects the node pre(pre) in layer l-1 to the node post in layer l at time step t.
10. The update value specific to each weight value of the weighted connections in one layer of the weighted connections 【Equation 11】 The method according to claim 7, wherein the update value is neutral with respect to all upstream layers of the weighted connections in the neural network.
11. Adjusting the weight value of the connection specific to each node in each layer, multiplying the gradient value derived from the feedback signal by an adjustment coefficient value 【Equation 12】 The method according to claim 10, comprising multiplying.
12. The feedback signal is (a) backpropagation, (b) feedback alignment, (c) direct feedback alignment, and (d) an element of the group consisting of indirect feedback alignment, the method according to claim 1.
13. The training is (a) stochastic gradient descent, (b) Adam optimizer method, (c) Nesterov accelerated gradient method, and (d) the method according to any one of claims 1 to 12, which is executed by a method selected from the group consisting of the RMSprop method.
14. A neural network training system for training a neural network, wherein the neural network includes nodes and weighted connections between selected ones of the nodes, and a function of a desired activity and a current activity during training results in a feedback signal used to adjust the weight values of the connections, the system comprising: a memory; and a processor, wherein the memory stores program code portions that enable the processor, for each update cycle, to: determine an importance value for each node based on the current weight values of the connections; determine an adjustment matrix that includes values that depend on the determined importance values; determine an updated value specific to each weight value of the connections by a combination of a gradient value derived from the feedback signal for the connections and the determined corresponding element of the adjustment matrix; and apply the updates to the connections during all update cycles. A neural network training system storing the program code portions. **Claim 15** To determine the importance value for each node, the program code portion enables the processor to: further construct the sum of the weight values of the connections received by the node. The neural network training system according to claim 14. **Claim 16** The importance value for each node a in layer l of the neural network at time step t **Equation 13** is the sum of the absolute strengths of all the weight values of layer l that includes the postsynaptic neurons **Equation 14** The neural network training system according to claim 15. **Claim 17** Determining the importance value for each node includes The neural network training system according to claim 14, comprising determining the sum of the weight values of the connections transmitted by the node. **Claim 18** The importance value for each node in layer l of the neural network at time step t **Equation 15** is the sum of the absolute strengths of all the weight values of layer l+1 including the presynaptic neurons **Equation 16** The neural network training system according to claim 17. **Claim 19** At time step t, for node a in layer l, the local update coefficient value of the adjustment matrix **Equation 17** is the importance value of node a **Equation 18** multiplied by 2 and formed into a product with the ratio of the maximum importance among all the neurons in layer l, and the product is determined to be limited by 1 at the lower limit. The neural network training system according to claim 14. **Claim 20** The local update coefficient value specific to each weight value of the connection in one layer of the connection **Equation 19** **Equation 19** affects at least one upstream layer of the connection in the neural network. The neural network training system according to claim 19. **Claim 21** The adjustment coefficient value **Equation 20** is directly applied to **Equation 21** further comprising determining 【Number 22】 wherein, for a time step t, is the weight value connecting node pre(pre) in layer l - 1 to the said node a(post = a) in layer l, the neural network training system according to claim 20.
22. the local update coefficient value specific to each weight value of the connection within one layer of the connection 【Number 23】 wherein is neutral with respect to at least one upstream layer of the connection within the neural network, the neural network training system according to claim 19.
23. the update value specific to each weight value of the connection within one layer of the connection 【Number 24】 wherein is neutral with respect to all upstream layers of the connection within the neural network, the neural network training system according to claim 19.
24. updating the weight value of the connection specific to each node within each layer, multiplying the gradient value derived from the feedback signal by an adjustment coefficient value 【Number 25】 comprising, the neural network training system according to claim 23.
25. A computer program for training a neural network, wherein the neural network includes nodes and weighted connections between selected ones of the nodes, and a function of a desired activity and a current activity during training yields a feedback signal used to adjust the weight values of the connections, causing a computer to, determine an importance value for each node based on the current weight value of the connection, determining an adjustment matrix that includes values depending on the determined importance values; determining updated values specific to each weight value of the connection by a combination of the gradient value derived from the feedback signal regarding the connection and the determined corresponding element of the adjustment matrix; a computer program that causes the updates to be applied to the connection during all update cycles.
Citation Information
Patent Citations
Neural network and neural network training method
JP2017511948A
Filter Specificity as a Training Criterion for Neural Networks
JP2018520404A