Neural network training device, method, and program

The neural network learning device and method address the inefficiency in learning by using DA-STDP to selectively enhance synapses involved in reward generation, thereby improving learning efficiency in large-scale networks.

WO2025109753A1PCT designated stage expired Publication Date: 2025-05-30NT T INC
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
PCT/JP2023/042191
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2023-11-24
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

Existing methods for learning in spiking neural networks, such as Izhikevich's method, fail to distinguish between different rewards generated by multiple sub-purposes in large-scale networks, leading to inefficient learning.

Method used

A neural network learning device and method that utilize dopamine-modulated STDP (DA-STDP) to enhance the connection strength of synapses only when they contribute to the generation of a specific reward, by calculating the connection distance from the synapse to the reward generation neuron and changing the connection strength based on this distance.

Benefits of technology

This approach allows for intensive strengthening of synapses that contribute to reward generation, improving learning efficiency in large-scale neural networks by distinguishing between multiple rewards and enhancing relevant synapses accordingly.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure JP2023042191_30052025_PF_FP_ABST
    Figure JP2023042191_30052025_PF_FP_ABST
Patent Text Reader

Abstract

According to one aspect of the present invention, when training a neural network provided with a plurality of general neurons that have a first function for determining, according to a neuron potential representing the activity of a neuron, whether or not the neuron is firing, and outputting a signal corresponding to the result of the determination, a plurality of reward generation neurons that have the first function and a second function for controlling the increase or decrease of a reward according to the presence or absence of firing as determined by the first function, and outputting the controlled reward, and a plurality of synapses that respectively connect the general neurons and the reward generation neurons, each synapse performs: a process in which each time a reward generation neuron outputs a reward, the synapse calculates the pathway connection distance from the synapse to the reward generation neuron; and a process in which the connection strength of the synapse is changed in accordance with a function including said connection distance and according to the time from when one general neuron or reward generation neuron connected to the synapse fires to when the other general neuron or reward generation neuron connected to the synapse fires.
Need to check novelty before this filing date? Find Prior Art

Description

Neural network learning device, method and program

[0001] One aspect of the present invention relates to the learning of a spiking neural network that realizes the function of the brain of a living organism such as a human, and in particular to a learning device, a learning method, and a program for a neural network that performs learning using dopamine-modulated STDP.

[0002] Unlike deep neural networks, which digitally represent the activity of neurons by whether they fire or not, spiking neural networks are a technology that analogically represents the activity of neurons within the network by the magnitude of the neuronal potential. The brains of living things, including humans, operate as spiking neural networks. Because spiking neural networks are closer to the workings of living brains, they may be able to achieve advanced intelligence that cannot be achieved with deep neural networks.

[0003] Unlike deep neural networks, spiking neural networks do not use backpropagation learning with training data because, just as living brains do not have this function, spiking neural networks do not have the backpropagation learning function.

[0004] Spiking neural networks use spike-timing-dependent plasticity (STDP), similar to that found in living brains, for learning. Similarly, dopamine-modulated STDP (DA-STDP) can also be used to enhance learning.

[0005] STDP is a mechanism in which, when two neurons are connected by a synapse, if the pre-neuron fires and then the post-neuron fires, the strength of the synaptic connection increases, and if they fire in the opposite order, the strength of the synaptic connection decreases. The brains of living things that do not have a backpropagation function can learn because they use STDP.

[0006] DA-STDP is a mechanism that amplifies the degree of change in synaptic connection strength caused by STDP when dopamine flows into the synapse. Because dopamine is a substance produced when living things receive a reward, DA-STDP reinforces learning in the direction of obtaining more rewards.

[0007] The method of determining which past action reflects the reward received by an organism is called the delayed reward problem, and a complete solution has yet to be found. The reward may reflect an action taken one second ago, or an action taken 30 seconds ago. Or each action may have contributed a certain percentage to the reward. Calculating the degree to which each action contributed to the reward is equivalent to solving the delayed reward problem. Because multiple rewards occur consecutively over the course of an organism's life, solving the delayed reward problem becomes increasingly difficult in situations where a reward occurs multiple times, not just once.

[0008] A method for solving delayed reward problems in spiking neural networks is proposed in Non-Patent Document 1. This method introduces the "reward contribution probability." Reward contribution probability is a parameter that represents the likelihood of contributing to a later reward, and will be referred to as c hereafter, following Non-Patent Document 1. One c exists for each synapse. c gradually decreases exponentially over time. When either the neuron before or after the synapse fires, the degree of STDP change at the synapse is added to c , increasing the absolute value of c . Otherwise, c continues to decrease. When a reward is generated and dopamine is produced, dopamine amplifies the change in synaptic strength due to STDP. The key to this method is to make the amplification rate proportional to c . Synapses of neurons that fire just before reward generation have a large c , so they are more influenced by dopamine, while synapses of neurons that fire well before reward generation have a small c , so they are only slightly influenced by dopamine. In summary, this method determines whether a behavior contributed to a reward based on the length of time that elapsed between the behavior and receiving the reward. This method is called Izhikevich's method.

[0009] Eugene M. Izhikevich, “Solving the Distal Reward Problem through Linkage of STDP and Dopamine Signaling”, Cerebral Cortex, Volume 17, Issue 10, October 2007, Pages 2443-2452.

[0010] However, in Izhikevich's method, when dopamine is produced, its effects are distributed equally across all synapses in the network, which is fine for small, single-function networks, but in large, multi-function networks like those in mammalian brains, network-wide dopamine effects are inappropriate.

[0011] For example, when an animal hunts for food, its primary goal is to approach prey without being detected, and to achieve this goal, it has multiple secondary goals, such as being quiet, staying downwind, and being in a position that is difficult for the prey to see. In this case, dopamine is produced when the animal behaves quietly, and also when it is downwind, but the synapses that contribute to the quiet behavior and the downwind position are different. For proper learning to occur, these synapses need to be strengthened separately for each secondary goal. However, Izhikevich's method cannot distinguish between these synapses for each secondary goal.

[0012] This invention was made with the above-mentioned circumstances in mind, and aims to provide a technology that enables learning in which the strength of synapses that contributed to the generation of a reward is selectively strengthened each time a reward is generated.

[0013] In order to solve the above problem, one aspect of the neural network learning device or learning method of the present invention is to train a neural network comprising: a plurality of general neurons having a first function of determining whether or not to fire in accordance with a neuron potential representing neuron activity and outputting a signal corresponding to the result; a plurality of reward-generating neurons having the first function and a second function of controlling an increase or decrease in reward in accordance with whether or not to fire due to the first function and outputting the controlled reward; and a plurality of synapses connecting the general neurons and the reward-generating neurons, respectively; and the device or method executes the following processes at the synapses: each time a reward-generating neuron outputs the reward, the device or method calculates the connection distance of the path from the own synapse to the reward-generating neuron; and, depending on the time elapsed between the firing of one of the general neurons or the reward-generating neuron connected to the own synapse and the firing of the other of the general neurons or the reward-generating neuron connected to the own synapse, changes the connection strength of the own synapse in accordance with a function including the connection distance.

[0014] According to one aspect of the present invention, the connection strength of a synapse is strengthened only when it is involved in the reward it contributed to, allowing the synapse to distinguish between multiple rewards. Since only the synapses involved in generating the reward are strengthened selectively, it is possible to improve the learning efficiency in a large-scale neural network in which rewards are generated from multiple reward-generating neurons.

[0015] That is, according to one aspect of the present invention, a technology can be provided that enables learning in which, each time a reward is generated, synapses that contributed to the generation of the reward are selectively strengthened.

[0016] FIG. 1 is a block diagram showing an example of the functional configuration of a neural network training device according to a first embodiment of the present invention. FIG. 2 is a flowchart illustrating an example of the operation of a general neuron in the neural network training device shown in FIG. 1. FIG. 3 is a flowchart illustrating an example of the operation of a reward-generating neuron in the neural network training device shown in FIG. 1. FIG. 4 is a flowchart illustrating an example of the operation of a synapse in the neural network training device shown in FIG. 1. FIG. 5 is a diagram used to explain the operation of the neural network training device shown in FIG. 1. FIG. 6 is a diagram illustrating a first example of an STDP function. FIG. 7 is a block diagram showing an example of the functional configuration of a neural network training device according to a second embodiment of the present invention. FIG. 8 is a flowchart illustrating an example of the operation of a synapse in the neural network training device shown in FIG. 7. FIG. 9 is a diagram illustrating a second example of an STDP function. FIG. 10 is a diagram illustrating a third example of an STDP function. FIG. 11 is a block diagram showing an example of a hardware configuration for realizing the neural network training devices according to the first and second embodiments.

[0017] Hereinafter, an embodiment of the present invention will be described with reference to the drawings.

[0018] First Embodiment (Configuration Example) FIG. 1 is a block diagram showing an example of the functional configuration of a neural network learning device according to a first embodiment of the present invention.

[0019] The neural network learning device NL according to the first embodiment includes, as its functional units, a plurality of general neurons 11, 12, ..., a plurality of reward-generating neurons 21, 22, ..., synapses 31, 32, ... connecting these neurons, an input unit 41, and an output unit 42.

[0020] The general neurons 11, 12, . . . have the function of updating the neuron potential in response to the input of signals from adjacent preceding and succeeding synapses 31, 32, . . . , and firing to output a signal when the neuron potential exceeds a predetermined value.

[0021] The reward-generating neurons 21, 22, ... have the function of generating and outputting rewards in addition to the firing control function based on the neuron potential that the general neurons 11, 12, ... have. The reward-generating function normally decreases its value according to a monotonically decreasing function, and updates the reward value so that the value increases by a fixed number each time the neuron potential exceeds a predetermined value and fires.

[0022] The synapses 31, 32, ... have the function of changing the ease of transmission of signals between neurons through learning, and are equipped with a synapse connection distance calculation processing unit 311 and a synapse connection strength change processing unit 312 as characteristic functional units of this invention.

[0023] Each time a reward-generating neuron 21, 22, ... generates dopamine in response to the generation of a reward, the synaptic connection distance calculation processing unit 311 calculates the synaptic connection distance by counting the number of synapses present on the path from its own synapse to the reward-generating neuron 21, 22, .... An example of a method for calculating this synaptic connection distance will be described in the operation example.

[0024] The synaptic connection strength change processor 312 defines the ease with which a synapse signal passes as connection strength. The synaptic connection strength change processor 312 then changes the connection strength of the synapse using a predetermined function including the synaptic connection distance, depending on the time interval between when one of the two neurons adjacent to the synapse fires and when the other neuron fires. An example of calculating this synaptic connection strength will also be described in the operation example.

[0025] Although each neuron is connected to all other neurons by synapses, it is not necessary for it to be connected to all other neurons, and there may be neurons that are not connected. Also, in Figure 1, for simplicity, the rewards generated by the reward-generating neurons 21, 22, ... are shown as being transmitted to only some synapses, but in reality, the rewards are transmitted to all synapses 31, 32, ....

[0026] FIG. 11 is a block diagram showing an example of the hardware configuration of an information processing device that constitutes the neural network learning device NL according to the first embodiment.

[0027] The information processing device constituting the neural network learning device NL includes a control unit 1A that uses a hardware processor such as a central processing unit (CPU). A storage unit including a program storage unit 2 and a data storage unit 3, and an input / output interface unit (hereinafter referred to as input / output I / F unit) 4 are connected to the control unit 1A via a bus 5.

[0028] The control unit 1A performs processing related to the general neurons 11, 12, ..., the reward-generating neurons 21, 22, ..., and the synapses 31, 32, ..., among the functional units shown in Fig. 1. All of these processing operations are realized by causing a hardware processor in the control unit 1 to execute application programs stored in the program storage unit 2. Note that some or all of the above processing operations may be realized using hardware such as an LSI (Large Scale Integration) or an ASIC (Application Specific Integrated Circuit).

[0029] The program storage unit 2 is configured by combining, for example, a non-volatile memory such as a HDD (Hard Disk Drive) or SSD (Solid State Drive) as a storage medium that can be written to and read at any time, and a non-volatile memory such as a ROM (Read Only Memory), and stores programs necessary for the control unit 1A to execute each of the above-mentioned functional units, in addition to middleware such as an OS (Operating System).

[0030] The data storage unit 3 is configured, for example, by combining a non-volatile memory such as an HDD or SSD as a storage medium that can be written to and read from at any time with a volatile memory such as a RAM (Random Access Memory), and is used to temporarily store parameter values ​​generated during the processing of the control unit 1A.

[0031] An input device DIN such as a keyboard, mouse, storage medium, etc., and an output device DOUT such as a display capable of displaying display data are connected to the input / output I / F unit 4. The input / output I / F unit 4 performs the functions of the input unit 41 and the output unit 42, and accepts control commands and parameters input from the input device DIN, and outputs output data generated by the control unit 1A to the output device DOUT.

[0032] (Example of Operation) Next, an example of operation of the neural network learning device NL configured as above will be described.

[0033] Here, as an example, as shown in Figure 5, we will explain the operation of a general neuron 15 and a presynapse 31A connected to this general neuron 15 when a reward-generating neuron 21 generates dopamine in response to the generation of a reward.

[0034] (1) Operation of General Neuron FIG. 2 is a flowchart showing an example of the operation of the general neuron 15.

[0035] When the general neuron 15 receives a signal output from the connected presynapse 31A in step S10, it updates the neuron potential in step S11. The neuron potential is expressed by an arbitrary function whose parameters are the neuron potential at the immediately preceding timing and the signal input from the synapse 31A. The signal input from the synapse 31A is either "1" or "0."

[0036] For example, the Izhikevich model consisting of the following three equations is used as the arbitrary function: Note that the arbitrary function may be a function other than the function defined by the Izhikevich model.

[0037] Here, v is the neuron potential, t is time, u is a variable introduced for use in calculations, I is the input, and a, b, c, and d are constants that can be set arbitrarily. The input I is obtained by multiplying the sum of the signals output from the connected synapses by an arbitrary magnification, which is usually set to "1."

[0038] After updating the neuron potential v according to the above function formula, the general neuron 15 subsequently determines in step S12 whether the neuron has fired as a result of updating the neuron potential v. The determination of whether the neuron has fired is made according to the condition defined by the conditional expression shown in the third line of the above function formula. That is, if the neuron potential v is, for example, 30 or more, it is determined that the neuron has fired, and if the neuron potential v is less than 30, it is determined that the neuron has not fired. Note that the conditional expression indicates that v and u are changed specifically only when the neuron has fired.

[0039] Based on the result of the determination of whether or not firing has occurred in step S12, the general neuron 15 generates an output value of "1" if firing has occurred in step S13, and generates an output value of "0" if not firing in step S14. Then, in step S15, the general neuron 15 outputs the output value of "1" or "0" to the other synapse connected thereto.

[0040] Thereafter, the general neuron 15 similarly repeats the series of operations from step S10 to step S15 described above until the neural network learning device NL completes the learning operation.

[0041] (2) Operation of the Reward-Generating Neuron FIG. 3 is a flowchart showing an example of the operation of the reward-generating neuron 21.

[0042] In step S20, the reward-generating neuron 21 decreases the reward value each time the processing loop is repeated. The decrease characteristic may be any monotonically decreasing function, for example, the inverse of an exponential function. In step S21, the reward-generating neuron 21 outputs the decreased reward value to all general neurons.

[0043] Following the reduction and output process of the reward value, the reward-generating neuron 21 performs the process of updating the neuron potential, the process of firing based on the updated neuron potential, and the process of generating and outputting a signal based on the result, just like the general neuron described above.

[0044] That is, when the reward-generating neuron 21 receives a signal from one of its connected synapses in step S22, it updates the neuron potential in step S23 and determines whether or not it has fired based on the updated neuron potential in step S24. If it has fired, it generates an output value of "1" in step S25, and if it has not fired, it generates an output value of "0" in step S27. Then, the reward-generating neuron 21 outputs the output value of "1" or "0" to the other connected synapse in step S28.

[0045] Moreover, only when it is determined in step S24 that the reward-generating neuron 21 has fired, the reward value is increased by an arbitrarily determined value in step S26.

[0046] Thereafter, the reward-generating neuron 21 similarly repeats the series of operations from step S20 to step S28 described above until the neural network learning device NL completes the learning operation.

[0047] (3) Operation of Synapse FIG. 4 is a flowchart showing an example of the operation of the target synapse 31A.

[0048] First, in step S30, the synapse 31A calculates the distance from the synapse 31A to each reward-generating neuron that generated dopamine using the synapse connection distance calculation processing unit 311.

[0049] For example, as shown in Figure 5, if reward-generating neuron 21 generates dopamine, the number of synapses on the shortest path from this reward-generating neuron 21 to its own synapse 31A (the path indicated by the thick arrow in Figure 5) is counted, and the counted value is taken as the connection distance between synapse 31A and reward-generating neuron 21. In this example, the connection distance is "3".

[0050] Next, in step S31, the synapse 31A determines whether the connected pre-neuron 12 or post-neuron 15 has fired. If the result of this determination is that neither neuron has fired, the synapse 31A returns to the synapse connection distance update process in step S30.

[0051] In response to this, assume that either the connected pre-neuron 12 or the connected post-neuron 15 fires. In this case, in step S32, the synapse 31A calculates a spike timing dependent plasticity (STDP) value according to a prepared STDP function.

[0052] Figure 6 shows an example of an STDP function, with the horizontal axis representing the value obtained by subtracting the most recent firing time of the pre-neuron from the most recent firing time of the post-neuron, and the vertical axis representing the STDP value. In this STDP function, the STDP value is positive if the pre-neuron fires before the post-neuron fires, and negative if they fire in the reverse order. The shorter the firing interval, the larger the absolute value of the STDP value. Note that if the firing interval is "0", the STDP value will be "0" as a special case.

[0053] Next, in step S33, the synapse 31A executes processing to update the synapse connection strength by the synapse connection strength change processing unit 312 as follows.

[0054] That is, the synapse 31A first calculates the distance influence for each reward-generating neuron 21, 22, ... using a function that decreases according to the synaptic connection distance previously calculated in step S30. Here, the inverse of the synaptic connection distance is used as an example of the function, but other functions may also be used. The synapse 31A then multiplies the reward value for each reward-generating neuron 21, 22, ... by the calculated distance influence. The reward value used here is the value that each reward-generating neuron 21, 22, ... most recently transmitted to all synapses.

[0055] The synapse 31A then calculates the sum of the values ​​obtained by multiplying the distance influence degrees calculated for all reward-generating neurons 21, 22, ... by the reward values. The synapse 31A then multiplies the calculated sum by the most recent STDP value, and determines the result as the connection strength increment. Finally, the synapse 31A adds the increment in connection strength to the current synaptic connection strength. Note that the increment in connection strength can be a negative value, so the synaptic connection strength may decrease in the synaptic connection strength change process.

[0056] When the synapse connection strength change process is completed, the synapse 31A returns to step S30 and executes the series of processes from step S30 to step S33 again. Thereafter, the synapse 31A similarly repeats the processing operations from step S30 to step S33 until the neural network learning device NL finishes the learning operation.

[0057] (Effects) As described above, in the first embodiment, the functions of the synapses 31, 32, ... are newly provided with a synaptic connection distance calculation processing unit 311 and a synaptic connection strength change processing unit 312. Then, when the reward-generating neurons 21, 22, ... generate a reward, the synaptic connection distance calculation processing unit 311 calculates a synaptic connection distance representing the distance from each synapse 31, 32, ... to the reward-generating neurons 21, 22, ..., and further, the synaptic connection strength change processing unit 312 changes the connection strength of the synapses 31, 32, ... when a reward is generated by the reward-generating neurons 21, 22, ..., by, for example, decreasing the rate at which the change in strength is amplified in accordance with the synaptic connection distance.

[0058] Therefore, the connection strength of the synapses 31, 32, ... is strengthened only when they are involved in, for example, the rewards to which they contributed. As a result, in the existing Izhikevich method, when rewards are generated by multiple reward-generating neurons, the synapses cannot distinguish between the multiple rewards. However, according to the first embodiment, the synapses 31, 32, ... can distinguish between the multiple rewards. Furthermore, since only the synapses involved in generating the rewards are strengthened in a focused manner, it is possible to improve the learning efficiency in a large-scale neural network in which rewards are generated from multiple reward-generating neurons 21, 22, ....

[0059] Second Embodiment (Configuration Example) FIG. 7 is a block diagram showing an example of the functional configuration of a neural network learning device NL according to a second embodiment of the present invention.

[0060] 7, the same parts as those in Fig. 1 are denoted by the same reference numerals, and detailed explanations thereof will be omitted. The neural network learning device NL according to the second embodiment is also configured by the information processing device shown in Fig. 11, similar to the first embodiment.

[0061] In the second embodiment, each synapse 31, 32, ... includes a reward contribution possibility calculation processing unit 313 in addition to the synapse connection distance calculation processing unit 311 and synapse connection strength change processing unit 312 described in the first embodiment.

[0062] The reward contribution possibility calculation processor 313 has a function of calculating a value representing the magnitude of the possibility that the synapse 31, 32, ... will later contribute to the generation of a reward in the reward-generating neuron 21, 22, .... Specifically, the reward contribution possibility calculation processor 313 performs a process of decreasing the value indicating the reward contribution possibility according to a monotonically decreasing function when none of the neurons before and after connected to the own synapse fires, and on the other hand, adding the STDP value to the value indicating the reward contribution possibility when either of the neurons before and after fires.

[0063] (Operation Example) Next, the operation of the neural network learning device NL configured as above will be described. Note that the operation of the general neurons 11, 12, ... and the reward-generating neurons 21, 22, ... is the same as the operation described in the first embodiment, so only the operation of the synapses 31, 32, ... will be described here.

[0064] 8 is a flowchart showing the operation of the target synapse 31B. In FIG. 8, the same parts as those in FIG. 4 are denoted by the same reference numerals.

[0065] First, in step S40, the synapse 31B causes the value indicating the reward contribution possibility to decrease according to a decreasing function each time the processing loop is repeated by the reward contribution possibility calculation processing unit 313. The decreasing characteristic may be any monotonically decreasing function, and for example, the reciprocal of an exponential function is used.

[0066] Next, in step S30, the synapse 31B calculates the connection distance from its own synapse 31B to each reward-generating neuron that generated dopamine in response to the generation of reward, using the synapse connection distance calculation processor 311. The method for calculating this connection distance is the same as the method described in the first embodiment.

[0067] Next, in step S31, the synapse 31B determines whether or not either of the neurons 12 and 15 connected thereto has fired. If the result of this determination is that neither of the neurons 12 and 15 has fired, the synapse 31B returns to step S40 and decreases the value indicating the possibility of reward contribution.

[0068] In response to this, assume that either the preceding or following neuron 12 or 15 fires. In this case, in step S32, the synapse 31B calculates an STDP value according to a pre-prepared STDP function. This STDP value calculation process is also performed using the same processing method as described in the first embodiment.

[0069] When the new STDP value is calculated, the synapse 31B proceeds to step S41, where the reward contribution possibility calculation processing unit 313 executes the process of updating the value indicating the reward contribution possibility as follows.

[0070] That is, the synapse 31B adds the STDP value calculated in step S32 to the value indicating the possibility of reward contribution. Note that the STDP value may be either positive or negative, so the value indicating the possibility of reward contribution may either increase or decrease.

[0071] Next, in step S33, the synapse 31B causes the synapse connection strength change processing unit 312 to execute processing for updating the synapse connection strength as follows.

[0072] That is, the synapse 31B first calculates the distance influence for each reward-generating neuron 21, 22, ... using a function that decreases according to the synaptic connection distance previously calculated in step S30. Here, the inverse of the synaptic connection distance is used as an example of the function, but other functions may also be used. The synapse 31B then multiplies the reward value for each reward-generating neuron 21, 22, ... by the calculated distance influence. In this case, the reward value used is the value most recently transmitted by each reward-generating neuron 21, 22, ... to all synapses.

[0073] The synapse 31B then calculates the sum of the values ​​obtained by multiplying the reward values ​​by the distance influence degrees calculated for all reward-generating neurons 21, 22, etc. Then, the synapse 31B multiplies the calculated sum by the most recent STDP value and the value indicating the reward contribution possibility calculated in step S41, and sets the result as the connection strength increase.

[0074] Finally, the synapse 31B adds the increment of the connection strength to the current synapse connection strength. Note that the increment of the connection strength may be a negative value, so the synapse connection strength may decrease in the process of changing the synapse connection strength.

[0075] When the synapse connection strength change process is completed, the synapse 31B returns to step S40 and executes the series of processes from step S40 to step S33 again. Thereafter, the synapse 31B similarly repeats the processing operations from step S40 to step S33 until the neural network learning device NL finishes the learning operation.

[0076] (Effect) As described above, in the second embodiment, the synapse 31B updates the value indicating the reward contribution possibility when either the neuron before or after it fires, and when calculating the synapse connection strength, the synapse 31B calculates an increase in connection strength by multiplying the sum of the values ​​calculated for each reward-generating neuron 21, 22, ... by the distance influence degree and the reward value by the most recent STDP value and the value indicating the reward contribution possibility, and then adds the calculated increase to the synapse connection strength.

[0077] Therefore, the closer the firing times of the neurons before and after connected to the synapses 31, 32, ... are to the reward generation time, and the larger the STDP values ​​of the synapses 31, 32, ..., the stronger the synaptic connection strength of the synapses 31, 32, .... In other words, the greater the possibility that the synapses 31, 32, ... contributed to the generation of the reward, the stronger the synaptic connection strength of the synapses 31, 32, .... This makes it possible to more effectively train the neural network for rewards.

[0078] [Examples] Several examples of the neuron potential update process performed in the general neurons 11, 12, . . . and the reward-generating neurons 21, 22, .

[0079] Example 1 Example 1 is an example in which the Izhikevich model is used as a function used to calculate a neuron potential. The following five types of parameters are set for the Izhikevich model.

[0080] a=0.02, b=0.2, c=-65, d=2~8 a=0.1, b=0.2, c=-65, d=2 a=0.02, b=0.2, c=-50, d=2 a=0.02, b=0.25, c=-65, d=2 a=0.02, b=0.2, c=-55, d=4.

[0081] Example 2 Example 2 is a case where the Hodgkin-Huxley model is used as a function used to calculate neuron potentials. The function formula of the Hodgkin-Huxley model is expressed as follows:

[0082]

[0083] Each parameter α in the above function formula m , β m , α h , β h , α n , β n is expressed as follows:

[0084]

[0085] Here, C m =1.0, g Na = 120, g K = 36, g L =0.3, E Na =50.0, E K =-77, E L =-54.387. stim is the input, and is obtained by multiplying the sum of the signals of the input synapses by an arbitrary magnification, which can usually be set to "1". Symbols other than this are variables used only in the calculation.

[0086] (Example 3) Example 3 is a case where a leaky integrate and fire model is used as a function used to calculate a neuron potential. The function formula of this leaky integrate and fire model is shown as follows: τ m (dV(t) / dt)=(-V(t)+E rest ) + I(t).

[0087] where V(t) is the neuron potential to be determined, tm=10, E rest = 60. I(t) is the input, which is obtained by multiplying the sum of the input synaptic signals by an arbitrary magnification factor. The magnification factor is usually set to "1".

[0088] Other Embodiments (1) In the first and second embodiments, the basic STDP function shown in Fig. 6 is used as the STDP function. However, the present invention is not limited to this. For example, a function in which the positive and negative signs of the basic STDP function are reversed as shown in Fig. 9, or a function that is symmetrical as shown in Fig. 10 may be used as the STDP function.

[0089] (2) There are two types of general neurons: excitatory neurons, which act as basic general neurons by causing the next neuron connected to fire when they fire, and inhibitory neurons, which, in contrast to excitatory neurons, suppress the firing of the next neuron connected to them when they fire.

[0090] In the first and second embodiments, an excitatory neuron is used as a general neuron, but an inhibitory neuron may be used instead of the excitatory neuron, or both excitatory and inhibitory neurons may be used together. The operation of an inhibitory neuron is the same as that of an excitatory neuron, except that the output is "-1" instead of "1."

[0091] (3) In the first and second embodiments, the synaptic connection distance is calculated based on a single path from the synapse to the reward-generating neuron. However, there may be multiple paths. In this case, the following methods can be used to calculate the connection distance:

[0092] (3-1) If there is one path from the synapse to the reward-generating neuron, the distance is calculated as a simple distance.

[0093] (3-2) If there are multiple paths from the synapse to the reward-generating neuron, calculate the distance of the shortest path among them.

[0094] (3-3) If there are multiple paths from a synapse to a reward-generating neuron, information should be transmitted more easily when there are multiple paths. Therefore, the distance calculated based on the shortest path is corrected so that it becomes shorter the more paths there are.

[0095] (3-4) If there are multiple paths from a synapse to a reward-generating neuron, the reciprocal of the distance of each path is summed up, and the reciprocal of the sum is used as the distance.

[0096] (4) In addition, the functional configuration and operation of the neural network learning device can be modified in various ways without departing from the spirit of the present invention.

[0097] Although the embodiments of the present invention have been described in detail above, the above description is merely an example of the present invention in every respect. It goes without saying that various improvements and modifications can be made without departing from the scope of the present invention. In other words, when implementing the present invention, specific configurations according to the embodiments may be appropriately adopted.

[0098] In short, this invention is not limited to the above-described embodiments, and in the implementation stage, the components can be modified and embodied without departing from the spirit of the invention. Furthermore, various inventions can be formed by appropriately combining multiple components disclosed in the above-described embodiments. For example, some components may be omitted from all the components shown in the embodiments. Furthermore, components from different embodiments may be appropriately combined.

[0099] DESCRIPTION OF SYMBOLS 1A, 1B... Control unit 2... Program storage unit 3... Data storage unit 4... Input / output I / F unit 5... Bus 11, 12... General neurons 21, 22... Reward-generating neurons 31, 32... Synapses 41... Input unit 42... Output unit 311... Synaptic connection distance calculation processing unit 312... Synaptic connection strength change processing unit 313... Reward contribution possibility calculation processing unit

Claims

1. A plurality of general neurons having a first function of determining the presence or absence of firing according to a neuron potential representing the activity of a neuron and outputting a signal corresponding to the result, a plurality of reward generation neurons having the first function and a second function of controlling an increase or decrease in reward according to the presence or absence of firing by the first function and outputting the controlled reward, and a plurality of synapses respectively connecting the general neurons and the reward generation neurons, wherein the synapse includes: a first processing unit that calculates a connection distance of a path from the self-synapse to the reward generation neuron each time the reward generation neuron outputs the reward; and a second processing unit that changes the connection strength of the self-synapse according to a function including the connection distance according to the time from when one of the general neurons or the reward generation neurons connected to the self-synapse fires until the other of the general neurons or the reward generation neurons connected to the self-synapse fires. A neural network learning device.

2. The synapse further includes a third processing unit that controls a value indicating the possibility that the self-synapse contributes to the generation of the reward of the reward generation neuron according to the presence or absence of firing of the general neuron or the reward generation neuron connected to the self-synapse, and the second processing unit reflects, in the connection strength, the value indicating the possibility of contributing to the generation of the reward controlled by the third processing unit when changing the connection strength of the self-synapse. The neural network learning device according to claim 1.

3. A learning method for a neural network comprising: a plurality of general neurons having a first function of determining the presence or absence of firing according to a neuron potential representing the activity of a neuron and outputting a signal corresponding to the result; a plurality of reward generation neurons having a second function of controlling an increase or decrease of a reward according to the first function and the presence or absence of firing by the first function and outputting the controlled reward; and a plurality of synapses connecting between the general neurons and the reward generation neurons, wherein in the synapse, every time the reward generation neuron outputs the reward, a process of calculating a connection distance of a path from the self-synapse to the reward generation neuron is performed, and according to a function including the connection distance according to a time from when one of the general neurons or the reward generation neurons connected to the self-synapse fires until the other of the general neurons or the reward generation neurons connected to the self-synapse fires, a process of changing a connection strength of the self-synapse is performed.

4. A program for causing a processor included in the neural network learning apparatus according to any one of claims 1 or 2 to execute at least one of each process performed by the first processing unit, the second processing unit, or the third processing unit included in the neural network learning apparatus.

Citation Information

Patent Citations

  • Method and apparatus for neural learning of natural multi-spike sequences in a spiking neural network

    JP2014532907A

  • plasticity synapse management

    JP2017515207A

  • Network traversal using neuromorphic instantiations of spike-timing dependent plasticity

    JP2018136919A

  • Synaptic circuit and neural network device

    JP2022129049A