Learning device, subtraction circuit and activation function circuit

The learning device employs stochastic computing and feedback mechanisms in neural network circuits to minimize size and power consumption while maintaining accuracy, addressing the challenges of existing technologies.

JP7727315B2Active Publication Date: 2025-08-21HOKKAIDO UNIVERSITY
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
JP2021118326
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2021-07-16
Publication Date
2025-08-21
Estimated Expiration
2041-07-16

AI Technical Summary

Technical Problem

Existing learning devices with neural networks face challenges in reducing circuit size and power consumption while maintaining accuracy, as previous technologies either include unnecessary circuits or result in complex configurations.

Method used

The learning device incorporates a neural network with circuits for each neuron, including summation, subtraction, and activation functions, utilizing stochastic computing and feedback mechanisms to minimize circuit size and power consumption.

Benefits of technology

This configuration significantly reduces circuit size and power consumption while preserving accuracy, utilizing probabilistic computing to optimize circuit paths and learning functions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007727315000011
    Figure 0007727315000011
  • Figure 0007727315000012
    Figure 0007727315000012
  • Figure 0007727315000013
    Figure 0007727315000013
Patent Text Reader

Abstract

To provide a learning apparatus configured to significantly reduce the circuit scale and power consumption thereof while suppressing reduction in accuracy as a learning apparatus.SOLUTION: In a learning apparatus including a neural network for inference and learning using stochastic computing, each of neuron circuits NR constituting the neural network includes at least a summation arithmetic unit AD, a shunt suppression subtraction unit SB and an activation function unit AF each adapted to the stochastic computing. The learning apparatus includes a back propagation unit which applies weight corresponding to output data output from an output layer, to feed back at least the output data to an input layer or a hidden layer constituting the neural network. The summation arithmetic unit AD includes a plurality of OR gates OR1, etc., adapted to the stochastic computing, and generates positive data and negative data as a result of the summation, to be output to the shunt suppression subtraction unit SB.SELECTED DRAWING: Figure 4
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to the technical field of a learning device and a subtraction circuit and an activation function circuit included therein, more particularly to a learning device that includes a neural network and performs inference and learning, and a subtraction circuit and an activation function circuit included in the learning device. [Background technology]

[0002] In recent years, research and development into learning devices that include neural networks has been actively conducted. While the majority of such research and development uses so-called non-stochastic operations, there are also a few that use stochastic operations. Examples of documents showing prior art related to such learning devices that include neural networks that use stochastic operations include Non-Patent Document 1 and Non-Patent Document 2 shown below.

[0003] Non-Patent Document 1 discloses a probabilistic neural computer characterized by an activation function using a so-called state machine, and Non-Patent Document 2 discloses a neurochip using probabilistic logic characterized by an activation function using a digital comparator. [Prior art documents] [Non-patent literature]

[0004] [Non-Patent Document 1] Bradley D. Brown and Howard C. Card, “Stochastic Neural Computation I: Computational Elements” IEEE Transactions on Computers, VOL. 50, NO. 9, September 2001 [Non-patent document 2] Shigeo Sato, Ken Nemoto, Shunsuke Akimoto, Mitsunaga Kinjo, and Koji Nakajima, “Implementation of a New Neurochip Using Stochastic Logic”, IEEE Transactions on Neural Networks, VOL. 14, NO. 5, September 2003 Summary of the Invention [Problem to be solved by the invention]

[0005] Generally, in a learning device, the circuitry that handles the processing in the learning stage has the largest circuit size and consumes the most power, and therefore, reducing the circuit size and power consumption is an urgent issue.

[0006] However, when applying either the technology disclosed in the above-mentioned Non-Patent Document 1 or the technology disclosed in the above-mentioned Non-Patent Document 2 to address these issues, sufficient effects cannot be obtained because unnecessary circuits are included (when Non-Patent Document 1 is applied) or the circuit configuration is inevitably complex (when Non-Patent Document 2 is applied).

[0007] The present invention has been made in consideration of the above-mentioned demands and problems, and one example of the object of the present invention is to provide a learning device that can significantly reduce the circuit size and power consumption while suppressing a decrease in accuracy as a learning device, as well as a subtraction circuit and an activation function circuit included in the learning device. [Means for solving the problem]

[0008] In order to solve the above-mentioned problems, the invention described in claim 1 is a learning device for performing inference and learning including a neural network, wherein a circuit corresponding to each neuron constituting the neural network includes at least a summation operation circuit such as a summation operation unit, a subtraction circuit such as a shunt suppression subtraction unit, and an activation function circuit such as an activation function unit, and the device is provided with feedback means such as a first backpropagation unit that feeds back output data from an output layer constituting the neural network to an input layer or hidden layer constituting the neural network while weighting the output data accordingly, and circuit is composed of a plurality of OR gates connected in a manner corresponding to stochastic computing, and generates positive data and negative data as the summation results, and performs the subtraction. circuit Output to Between the OR gates, a number of registers corresponding to the number of the OR gates are connected. It is configured as follows.

[0009] According to the invention of claim 1, a circuit corresponding to each neuron constituting a neural network includes at least a summation circuit, a subtraction circuit, and an activation function circuit, and the summation circuit is made up of a plurality of OR gates connected in a manner corresponding to stochastic computing, and generates positive and negative data as the summation result and outputs them to the subtraction circuit. Therefore, by using stochastic computing, it is possible to significantly reduce the circuit size and power consumption while suppressing a decrease in accuracy as a learning device.

[0011] Also Since a number of registers corresponding to the number of OR gates are connected between the OR gates, the circuit path (critical path) can be reduced even when the number of neurons is large, making it possible to further miniaturize the circuit scale.

[0012] In order to solve the above problem, claims 2The invention described in the above is a learning device that performs inference and learning including a neural network, wherein the circuits corresponding to the neurons that make up the neural network include at least a summation calculation circuit such as a summation calculation unit, a subtraction circuit such as a shunt suppression subtraction unit, and an activation function circuit such as an activation function unit, and the output data from the output layer that makes up the neural network is fed back to the input layer or hidden layer that makes up the neural network while weighting the output data accordingly. Return means of return a feedback means such as a first backpropagation unit having a learning function the subtraction circuit teeth , supporting probabilistic computing Perform shunt suppression subtraction, which is a simulated subtraction Inverter and AND gate connected in this manner Made up of , the inverter circuit and the AND gate generates a product of the positive data output from the summation circuit and the generated inverted data, and outputs the product to the activation function circuit as the subtraction result in the subtraction circuit. The learning function is a learning function based on a learning rule corresponding to the backpropagation method. It is configured as follows.

[0013] Claim 2 According to the invention described in the above, a circuit corresponding to each neuron constituting a neural network includes at least a summation circuit, a subtraction circuit, and an activation function circuit, A feedback means having a learning function is provided, Subtraction circuits support probabilistic computing Perform shunt suppression subtraction, which is a simulated subtraction Inverter and AND gate connected in this manner Made up of The negative data from the summation circuit is inverted to generate a product of the inverted data and the positive data from the summation circuit, and the result is output to the activation function circuit as the subtraction result. In this case, the learning function of the feedback means is a learning function based on a learning rule corresponding to the error back propagation method. Therefore, probabilistic computing Shunt suppression subtraction, which is a simulated subtraction corresponding to Using The learning function in the feedback means is a learning rule corresponding to the error backpropagation method. This makes it possible to significantly reduce the circuit size and power consumption of the learning device while suppressing a decrease in accuracy.

[0014] In order to solve the above problem, claims 3The invention described in is a learning device that performs inference and learning including a neural network, wherein the circuits corresponding to each neuron that constitutes the neural network include at least a summation operation circuit such as a summation operation unit, a subtraction circuit such as a shunt suppression subtraction unit, and an activation function circuit such as an activation function unit, and the device is equipped with feedback means such as a first backpropagation unit that feeds back output data from an output layer that constitutes the neural network to an input layer or a hidden layer that constitutes the neural network while weighting the output data accordingly, and the activation function circuit, which includes two OR gates and one AND gate connected in a manner corresponding to stochastic computing, is configured to generate and output output data as the neuron based on the subtraction result output from the subtraction circuit and a clock signal.

[0015] Claim 3 According to the invention described in (2), a circuit corresponding to each neuron constituting a neural network includes at least a summation circuit, a subtraction circuit, and an activation function circuit, and the activation function circuit includes two OR gates and one AND gate connected in a manner corresponding to stochastic computing, and generates and outputs output data as a neuron based on the subtraction result output from the subtraction circuit and a clock signal. Therefore, by using stochastic computing, it is possible to significantly reduce the circuit size and power consumption while suppressing a decrease in accuracy as a learning device.

[0016] In order to solve the above problem, claims 4 The invention described in claim 2 The subtraction circuit included in the learning device described in the above item 1 includes the inverter and the AND gate connected in a manner corresponding to the stochastic computing, the inverter generates the inverted data, and the AND gate is configured to generate a product of the positive data output from the summation circuit and the generated inverted data, and output the product to the activation function circuit as the subtraction result.

[0017] Claim 4 According to the invention described in claim2 The inverters and AND gates constituting the subtraction circuit included in the learning device described in are connected in accordance with stochastic computing, the inverter generates inverted data, and the AND gate generates a product of the positive data from the summation circuit and the inverted data, which is output to the activation function circuit as the subtraction result of the subtraction circuit. Therefore, by using stochastic computing in the configuration of the subtraction circuit, it is possible to significantly reduce the circuit size and power consumption while suppressing a decrease in accuracy as a learning device.

[0018] In order to solve the above problem, claims 5 The invention described in claim 3 is the activation function circuit included in the learning device described in claim 3, which includes two of the OR gates and one of the AND gates connected in a manner corresponding to the stochastic computing, and is configured to generate and output the output data based on the output subtraction result and the clock signal.

[0019] Claim 5 According to the invention described in claim 3 The activation function circuit included in the learning device described in the above has two OR gates and one AND gate connected in accordance with stochastic computing, and generates and outputs output data as a neuron based on the subtraction result output from the subtraction circuit and a clock signal. Therefore, by using stochastic computing in the configuration of the activation function circuit, it is possible to significantly reduce the circuit size and power consumption while suppressing a decrease in the accuracy of the learning device. [Effects of the Invention]

[0020] Claim 1 toAccording to the first aspect of the present invention, a circuit corresponding to each neuron constituting a neural network includes at least a summation circuit, a subtraction circuit, and an activation function circuit, and the summation circuit is composed of a plurality of OR gates connected in a manner corresponding to stochastic computing, and generates positive and negative data as the summation result and outputs them to the subtraction circuit. Therefore, by using stochastic computing, it is possible to significantly reduce the circuit size and power consumption while suppressing a decrease in accuracy as a learning device. Furthermore, since registers corresponding in number to the number of OR gates are connected between the OR gates, the circuit path (critical path) can be reduced even when there are a large number of neurons, making it possible to further miniaturize the circuit scale.

[0021] Also, claims 2 According to the second aspect of the present invention, a circuit corresponding to each neuron constituting a neural network includes at least a summation circuit, a subtraction circuit, and an activation function circuit, A feedback means having a learning function is provided, Subtraction circuits support probabilistic computing Perform shunt suppression subtraction, which is a simulated subtraction Inverter and AND gate connected in this manner Made up of The negative data from the summation circuit is inverted to generate a product of the inverted data and the positive data from the summation circuit, and the result is output to the activation function circuit as the subtraction result. In this case, the learning function of the feedback means is a learning function based on a learning rule corresponding to the backpropagation method. Therefore, by using shunt suppression subtraction, which is a simulated subtraction corresponding to stochastic computing, and by using a learning rule corresponding to the backpropagation method as the learning function of the feedback means, it is possible to significantly reduce the circuit size and power consumption while suppressing a decrease in accuracy as a learning device.

[0022] Furthermore, claims 3 According to the third aspect of the present invention, a circuit corresponding to each neuron constituting a neural network includes at least a summation circuit, a subtraction circuit, and an activation function circuit, and the activation function circuit includes two OR gates and one AND gate connected in a manner corresponding to stochastic computing, and generates and outputs output data as a neuron based on the subtraction result output from the subtraction circuit and a clock signal. Therefore, by using probabilistic computing, it is possible to significantly reduce the circuit size and power consumption of the learning device while suppressing a decrease in accuracy.

[0023] As stated above, In any of the above aspects, by using probabilistic computing, it is possible to significantly reduce the circuit size and power consumption of the learning device while suppressing a decrease in accuracy. [Brief explanation of the drawings]

[0024] [Figure 1] 1A and 1B are diagrams conceptually illustrating a probabilistic computing system according to the principles of the present invention, in which (a) is a circuit diagram conceptually illustrating the probabilistic computing system, and (b) is a table showing each operation, etc., in the probabilistic computing system. [Figure 2] 1 is a diagram conceptually showing the configuration of a multilayer perceptron according to the principle of the present invention; [Figure 3] 1 is a block diagram showing a configuration of a learning device according to an embodiment; [Figure 4] 1A and 1B are block diagrams showing the configuration of a neuron circuit constituting a learning device of an embodiment, in which FIG. 1A is a block diagram showing the configuration of the neuron circuit, and FIG. 1B is a circuit diagram of a summation calculation unit constituting the neuron circuit. [Figure 5] 1A and 1B are circuit diagrams of a shunt suppression subtraction unit and other components that constitute a neuron circuit constituting a learning device of an embodiment, where (a) is a circuit diagram of the shunt suppression subtraction unit, (b) is a circuit diagram of an activation function unit that constitutes the neuron circuit, and (c) is a graph showing the difference between the output of the activation function unit and its theoretical value. [Figure 6] 1A and 1B are circuit diagrams showing the configuration of a first backpropagation unit constituting a learning device of an embodiment, in which (a) is a circuit diagram showing the configuration of the first backpropagation unit, (b) is a circuit diagram of a differentiation circuit constituting the first backpropagation unit, and (c) is a circuit diagram of a sign determination circuit constituting the first backpropagation unit. [Figure 7] 1A and 1B are block diagrams showing the general configuration of a decoder and other components that make up the learning device of the embodiment, where FIG. 1A is the block diagram and FIG. 1B is a table showing input data to the decoder. [Figure 8]1A and 1B are circuit diagrams showing the configuration of a second back-propagation unit constituting a learning device of an embodiment, in which (a) is a block diagram showing the configuration of the second back-propagation unit, (b) is a circuit diagram of an absolute value determination circuit constituting the second back-propagation unit, (c) is a circuit diagram of a sign determination circuit constituting the second back-propagation unit, and (d) is a circuit diagram of a sign judgment circuit constituting the sign determination circuit. [Figure 9] Figure (I) shows the learning results of the learning device of the embodiment, where (a) shows the learning results of AND logic by the learning device, and (b) shows the learning results of OR logic by the learning device. [Figure 10] Figure (II) shows the learning results by the learning device of the embodiment, where (a) is a diagram showing the learning results of NAND logic by the learning device, and (b) is a diagram showing the learning results of NOR logic by the learning device. [Figure 11] 11A and 11B are diagrams showing the learning results of the learning device of the embodiment (III), where (a) is a diagram showing the learning results of exclusive OR (XOR) logic by the learning device, and (b) is a diagram showing the learning results of exclusive NOR (XNOR) logic by the learning device. [Figure 12] 4A and 4B are diagrams showing the learning results of the learning device of the embodiment, where (a) is a diagram showing the learning results of a linear regression problem by the learning device, and (b) is a diagram showing the learning results of a nonlinear regression problem by the learning device. DETAILED DESCRIPTION OF THE INVENTION

[0025] (1) The principle of the present invention First, before describing the embodiments of the present invention, the principles of the present invention will be described with reference to Figures 1 and 2. Figure 1 is a diagram conceptually showing a probabilistic computing system according to the principles of the present invention, and Figure 2 is a diagram conceptually showing the configuration of a multilayer perceptron according to the principles of the present invention.

[0026] The inventors of the present invention have made the present invention with the aim of configuring a learning device that uses the theory of probabilistic computing, which has been studied for some time, to achieve the above-mentioned problems of significantly reducing the circuit size and power consumption in a learning device that performs inference and learning including a neural network.

[0027] Here, the probabilistic computing system SC based on the theory of probabilistic computing applied to the present invention is as shown in the circuit diagram of Figure 1(a) showing its outline, in which n-bit integer data input from outside is "E", data obtained by converting this "E" into a binary time series is "X" (in the following explanation, this "X" will be simply referred to as "input data"), and output data is "O", and is configured by connecting arithmetic circuits C1 and C2 and an n-bit counter C3 in series. Note that the probabilistic computing system SC shown in Figure 1(a) has a length of 2 n Each value is expressed by the proportion of the value "1" in the random bit string. Various operations such as multiplication by the probabilistic computing system SC are performed by the circuit and logical operations shown in Figure 1(b).

[0028] It is known that the above-described probabilistic computing system SC can perform various operations with only a small number of logic gates. Therefore, by applying such a probabilistic computing system SC to the learning device of the present invention, it is possible to significantly reduce the circuit size and power consumption.

[0029] Next, the configuration of a multilayer perceptron that constitutes the learning device of the present invention, to which the theory of probabilistic computing is applied, will be explained using Figure 2. Figure 2 is a conceptual diagram showing the configuration of a multilayer perceptron that is the principle of the present invention.

[0030] As shown in FIG. 2, a multilayer perceptron MLP, for example, with three layers, that constitutes the learning device of the present invention is composed of an input layer IN, a hidden layer HD, and an output layer OUT. In this case, the input layer IN consists of, for example, k nodes to which input data X is input. The hidden layer HD consists of, for example, j nodes. The output layer OUT consists of, for example, i nodes that each output output data Y (output data Y as a neural network). In the following description, the nodes that constitute the input layer IN of the multilayer perceptron MLP of the present invention will be referred to as "input nodes," the nodes that constitute the hidden layer HD of the multilayer perceptron MLP will be referred to as "hidden layer nodes," and the nodes that constitute the output layer OUT of the multilayer perceptron MLP will be referred to as "output nodes." Weights W are used between each input node that constitutes the input layer IN and each hidden layer node that constitutes the hidden layer HD. j,k The hidden layer nodes of the hidden layer HD are connected by corresponding connection lines. The hidden layer nodes of the output layer OUT are connected by weighting lines W i,j are connected by corresponding connection lines.

[0031] Furthermore, the forward calculation performed on input data X in the multilayer perceptron MLP of the present invention is expressed by the following equations (1) and (2), where "g" is the so-called activation function.

number

number

[0032] (2) Regarding the embodiment of the present invention Next, an embodiment of the present invention corresponding to the principle of the present invention explained using Figures 1 and 2 will be described using Figures 3 to 8. Note that Figure 3 is a block diagram showing the configuration of a learning device of the embodiment, Figure 4 is a block diagram showing the configuration of a neuron circuit constituting the learning device, and Figure 5 is a circuit diagram of a shunt suppression subtraction unit and the like constituting the neuron circuit. Also, Figure 6 is a circuit diagram showing the configuration of a first backpropagation unit constituting the learning device, Figure 7 is a block diagram showing the configuration of a decoder and the like constituting the learning device, and Figure 8 is a circuit diagram showing the configuration of a second backpropagation unit constituting the learning device.

[0033] As shown in Fig. 3, the learning device S of the embodiment is a learning device that outputs output data Y that reflects the results of inference and learning on input data X, and is configured by connecting a hidden layer circuit HDC, an output layer circuit OUTC, a first backpropagation unit 10, a first decoder 11, a first weight storage unit 12, a first encoder 13, a second backpropagation unit 20, a second decoder 21, a second weight storage unit 22, and a second encoder 23 in the manner shown in Fig. 3. The meanings of the symbols in Fig. 3 are the same as those used in Figs. 1 and 2 and the above formulas (1) to (5) (the same applies to Figs. 4 to 8 below). In this configuration, the hidden layer circuit HDC corresponds to the hidden layer HD in the multilayer perceptron MLP illustrated in Fig. 2, and the output layer circuit OUTC corresponds to the output layer OUT in the multilayer perceptron MLP. The input layer IN in the multilayer perceptron MLP is the hidden layer HD in Fig. 3, which stores input data X. k is expressed as a single input line.

[0034] On the other hand, the learning parameters for the learning device S are as follows: the number of input nodes in the corresponding multilayer perceptron MLP (see Figure 2) is "3 (including bias)," the number of hidden layer nodes in the multilayer perceptron MLP is "4," the number of output nodes in the multilayer perceptron MLP is "1," the learning rate is 0.3, and the bit string length of each data is 255 bits. The logics to be learned by the learning device S are six: "AND logic," "NAND logic," "OR logic," "NOR logic," "exclusive OR logic," and "exclusive NOR logic."

[0035] In the learning device S, the first backpropagation unit 10, the first decoder 11, the first weight storage unit 12, and the first encoder 13 implement a learning function from the output layer circuit OUTC to the hidden layer circuit HDC by the backpropagation method (see equations (3) to (5)). Similarly, the second backpropagation unit 20, the second decoder 21, the second weight storage unit 22, and the second encoder 23 implement a learning function from the hidden layer circuit HDC to the input layer IN by the backpropagation method (see equations (3) to (5)). Furthermore, the first backpropagation unit 10, the first decoder 11, the first weight storage unit 12, and the first encoder 13, as well as the second backpropagation unit 20, the second decoder 21, the second weight storage unit 22, and the second encoder 23, each correspond to an example of the "feedback means" of the present invention.

[0036] Next, the configuration of a neuron circuit included in the hidden layer circuit HDC or the output layer circuit OUTC and corresponding to each neuron in the multilayer perceptron MLP will be described with reference to Figures 4 and 5. Figure 4 is a block diagram showing the configuration of a neuron circuit constituting the learning device of the embodiment, and Figure 5 is a circuit diagram of a shunt suppression subtraction unit and the like constituting the neuron circuit.

[0037] As shown in Fig. 4(a), a neuron circuit NR corresponding to one neuron included in the hidden layer circuit HDC or the output layer circuit OUTC is composed of a multiplier MX, a summation calculation unit AD, a shunt suppression subtraction unit SB, and an activation function unit AF. In this case, the summation calculation unit AD corresponds to an example of the "summation calculation circuit" of the present invention, the shunt suppression subtraction unit SB corresponds to an example of the "subtraction circuit" of the present invention, and the activation function unit AF corresponds to an example of the "activation function circuit" of the present invention.

[0038] In the above configuration, input data X and weighting W are input to the multiplier MX. Next, the summation calculation unit AD is a summation calculation unit that applies the theory of incomplete addition in the stochastic computing system SC shown in Figure 1. Furthermore, the shunt suppression subtraction unit SB is a subtraction unit that performs simulated subtraction that applies the theory of shunt suppression in the stochastic computing system SC. Furthermore, the activation function unit AF is a circuit that applies an activation function corresponding to the stochastic computing system SC. Finally, output data Y corresponding to the input data X and weighting W is output from the activation function unit AF.

[0039] Next, the detailed configurations of the summation calculation unit AD, the shunt suppression subtraction unit SB, and the activation function unit AF that constitute the neuron circuit NR will be described with reference to FIG. 4(b) and FIG.

[0040] First, the detailed configuration of the sum calculation unit AD will be explained using Fig. 4(b). As shown in Fig. 4(b), the sum calculation unit AD to which the output data from the multiplier MX is input is configured by connecting OR gates OR1 to OR gates OR12 and registers R1 to R11 in the manner shown in Fig. 4(b).

[0041] In this case, the so-called weighted summation used in conventional neuron circuits has high accuracy in the summation, but the output is normalized. Therefore, if this is generally used in a multilayer perceptron, a decoder must be inserted into the calculation process, which leads to an increase in the circuit size. In contrast, in the imperfect summation by the summation calculation unit AD of the embodiment (the theoretical formula is the following formula (6)), there is little deviation between the result of the imperfect summation and the theoretical summation result within a range smaller than a certain threshold.

number

[0042] In this case, it is possible to configure the summation calculation unit AD corresponding to the probabilistic computing system SC of the embodiment using only the OR gates OR1 to OR12. However, by connecting the OR gates OR1 and the like via the registers R1 to R11 as shown in Figure 4(b), it is possible to reduce the increase in the so-called critical path (i.e., shorten the critical path) due to the increase in the number of neuron circuits NR constituting the hidden layer circuit HDC or the output layer circuit OUTC.

[0043] Generally, the so-called latency (also referred to as response time or delay time (in communication)) caused by the inclusion of a register R1 or the like does not pose a problem in a steady-state probabilistic computing system SC. On the other hand, incorporating a register into a logic operation circuit is a technique commonly used as the basic structure of general digital circuits, digital integrated circuits that embody such circuits, or FPGAs (Field-Programmable Gate Arrays) that constitute the learning device S of the embodiment. Therefore, the summation operation unit AD of the embodiment is configured to combine this FPGA with the probabilistic computing system SC so as to cancel out the weaknesses of both. In the following explanation, for the sake of simplicity, the input data X m The summation processing by the summation calculation unit AD is simply

number

[0044] Next, the detailed configuration of the shunt suppression subtraction unit SB will be described with reference to FIG. 5(a). As shown in FIG. 5(a), the shunt suppression subtraction unit SB, to which the output data from the summation calculation unit AD is input, is configured by connecting one inverter NT and an AND gate AND1 in the manner shown in FIG. 5(a). Note that if the shunt suppression subtraction unit SB performs simulated subtraction corresponding to the stochastic computing system SC instead of the subtraction unit of the prior art, the circuit size (circuit area) can be reduced, but the accuracy of the subtraction inevitably decreases. For this reason, in the learning device S of the embodiment, the learning rules in the first backpropagation unit 10 and the second backpropagation unit 20 are devised as described below to suppress the decrease in accuracy.

[0045] Next, the detailed configuration of the activation function unit AF will be explained using Fig. 5(b). As shown in Fig. 5(b), the activation function unit AF, to which the output data from the shunt suppression subtraction unit SB is input, is configured by connecting registers R20 to R22, OR gates OR20 and OR21, and AND gate AND2 in the manner shown in Fig. 5(b).

[0046] Here, the characteristics of an OR gate such as OR gate OR20 can generally be defined as the following equation (7).

number

number

number

[0047] Next, the detailed configuration of the first back propagation unit 10 of the embodiment will be described with reference to Fig. 6. Fig. 6 is a circuit diagram showing the configuration of the first back propagation unit 10.

[0048] The first backpropagation unit 10 and the second backpropagation unit 20 of the embodiment each realize a learning function using the error backpropagation method described above, which includes the error function E of the above equation (3), as shown in the following equation (10) corresponding to the above equation (3).

number

[0049] For this reason, the first backpropagation unit 10 is configured to realize the following equation (11), as shown in FIG. 6(a), by connecting a multiplier 10A, a differentiation circuit 10B, a sum calculation unit 10C having a configuration similar to that of the sum calculation unit AD, a sign determination circuit 10D, an inverter 10E and an inverter 10F, and an exclusive OR circuit 10G in the manner shown in FIG. 6(a).

number

[0050] In the above configuration, the coefficients enclosed by the dashed line in the above formula (11) are changed from the conventional definition of the update difference ΔW. As a result, in the learning function using the backpropagation method of the embodiment, the change in the learning rule for suppressing a decrease in subtraction accuracy by using simulated subtraction in the shunt suppression subtraction unit SB is realized.

[0051] Next, the configurations of the first decoder 11, first weight storage unit 12, and first encoder 13 of the embodiment connected to the first backpropagation unit 10 will be described using FIG. 7 as an example for handling 8-bit data. FIG. 7 is a block diagram showing a schematic configuration of the decoder and other components constituting the learning device of the embodiment. The configurations and functions of the first decoder 11, first weight storage unit 12, and first encoder 13 of the embodiment are essentially the same as those of the second decoder 21, second weight storage unit 22, and second encoder 23 of the embodiment, except that the former processes input data from the first backpropagation unit 10, while the latter processes input data from the second backpropagation unit 20. Therefore, both configurations are shown together in FIG. 7.

[0052] As shown in FIG. 7(a), the first decoder 11 receives two pieces of output data U / D and output data EN from the first backpropagation unit 10. At this time, the states of these two pieces of output data U / D and output data EN can be expressed by three values ​​(1, 0, and −1). Therefore, the data required for the decoder 11, which is configured with an 8-bit counter, to decode the output data U / D and output data EN is only the sign and absolute value of the state of the pulse train, as shown in FIG. 7(b). The output data from the decoder 11 is stored in a first weight storage unit 12, which is an 8-bit register, for each learning process (1 EPOCH). The stored value in the first weight storage unit 12 is then compared with an 8-bit random number sequence generated by a linear feedback shift register SR (not shown in FIG. 3) in a first encoder 13, which is configured with an 8-bit comparator. The comparison result in the first encoder 13 is then fed back to the output layer circuit OUTC, as shown in FIG. 3.

[0053] Next, the detailed configuration of the second back propagation unit 20 of the embodiment will be described with reference to Fig. 8. Fig. 8 is a circuit diagram showing the configuration of the second back propagation unit 20.

[0054] As mentioned above, the equation (3) the above In order to realize the learning function by the error backpropagation method shown in equation (10) together with the first backpropagation unit 10, the second backpropagation unit 20 of the embodiment is configured by connecting multiplexers 20A and 20B, sign decision circuits 20C1 to 20CNY-1 for each bit in a data string, and absolute value decision circuits 20D1 to 20DNY-1 for each bit in the manner shown in FIG. 8(a) in order to realize the following equation (12).

number

[0055] In the above configuration, the coefficients enclosed by the dashed line in the above formula (12) are changed from the conventional definition of the update difference ΔW. As a result, similar to the first backpropagation unit 10, in the learning function using the error backpropagation method of the embodiment, the change in the learning rule for suppressing a decrease in subtraction accuracy by using simulated subtraction in the shunt suppression subtraction unit SB is realized.

[0056] The configurations and functions of the second decoder 21, second weight storage unit 22, and second encoder 23 of the embodiment connected to the second backpropagation unit 20 are similar to the configurations and functions of the first decoder 11, first weight storage unit 12, and first encoder 13 of the embodiment described with reference to Fig. 7, as described above, and therefore will not be described again. The result of comparing the value stored in the second weight storage unit 22 in the second encoder 23 with the 8-bit random number sequence generated by the linear feedback shift register SR is fed back to the hidden layer circuit HDC as shown in Fig. 3. [Example]

[0057] Next, as an example corresponding to the embodiment, experimental results (simulation results) by the inventors of the present invention regarding the learning results by the learning device S of the embodiment will be described with reference to Fig. 9 to Fig. 12. Fig. 9 to Fig. 12 are diagrams each showing the learning results by the learning device S.

[0058] That is, as an example, the inventors of the present invention have obtained the learning results of AND logic (see Figure 9(a)), OR logic (see Figure 9(b)), NAND logic (see Figure 10(a)), NOR logic (see Figure 10(b)), exclusive OR (XOR) logic (see Figure 11(a)), and exclusive NOR (XNOR) logic (see Figure 11(b)) using the learning device S of the embodiment, respectively, in terms of the relationship between the number of learning sessions (EPOCH) and the accuracy rate.

[0059] Similarly, the learning results of the learning device S of the embodiment for the linear regression problem (see FIG. 12(a)) and the learning results of the nonlinear regression problem (see FIG. 12(b)) are obtained.

[0060] In this case, the learning parameters when the learning result shown in Figure 12(a) was obtained were: the number of input nodes in the corresponding multilayer perceptron MLP (see Figure 2) was "9 (including bias)," the number of hidden layer nodes in the multilayer perceptron MLP was "32," the number of output nodes in the multilayer perceptron MLP was "8," the learning rate was 0.3, and the bit string length of each data was 255 bits. Furthermore, the training data when the learning result was obtained had an input of x = n (n is an integer equal to or greater than 0) and an output of y = x (see Figure 12(a) left) and y = (1 / 2)x (see Figure 12(a) right). In Figure 12(b), the "+" marks represent the learning results (inference results), and the lines represent the correct answer data (training data).

[0061] 12(b) was obtained, the learning parameters were: the number of input nodes in the corresponding multilayer perceptron MLP was "9 (including bias)," the number of hidden layer nodes in the multilayer perceptron MLP was "128," the number of output nodes in the multilayer perceptron MLP was "8," the learning rate was 0.3, and the bit string length of each data was 255 bits. Furthermore, the training data used to obtain the learning results was: input x = 10n (n is an integer equal to or greater than 0), and output y = cos(x) (see left side of FIG. 12(b)) and y = sin(x) (see right side of FIG. 12(b)).

[0062] As is clear from any of Figures 9 to 12, even in the learning device S of the embodiment, i.e., the learning device S in which the configuration of the probabilistic computing system SC has been applied, thereby achieving a significant reduction in circuit size and power consumption, the accuracy of the learning results is not reduced.

[0063] As described above, according to the configuration of the learning device S of the embodiment, each neuron circuit NR (see FIG. 4, etc.) constituting the learning device S includes at least a summation calculation unit AD, a shunt suppression subtraction unit SB, and an activation function unit AF, each corresponding to a stochastic computing system SC, and the summation calculation unit AD is composed of a plurality of OR gates OR1, etc., corresponding to stochastic computing, and generates positive and negative data as the summation calculation results and outputs them to the shunt suppression subtraction unit SB. Therefore, by using stochastic computing, it is possible to significantly reduce the circuit size and power consumption of the learning device S while suppressing a decrease in accuracy.

[0064] Furthermore, in the sum calculation unit AD, between the OR gates OR1, etc., registers R1, etc., are connected in a number corresponding to the number of the OR gates OR1, etc., so that even when the number of neurons is large, the circuit path (critical path) can be reduced, making it possible to further miniaturize the circuit scale.

[0065] Furthermore, the shunt suppression subtraction unit SB generates the product of the inverted data obtained by inverting the negative data from the summation calculation unit AD and the positive data from the summation calculation unit AD, and outputs the subtraction result to the activation function unit AF.Therefore, by using stochastic computing, it is possible to significantly reduce the circuit size and power consumption of the learning device S while suppressing a decrease in accuracy.

[0066] Furthermore, an activation function unit AF including an OR gate OR20, an OR gate OR21, and an AND gate AND2 generates output data Y as a neuron based on the subtraction result output from the shunt suppression subtraction unit SB and the clock signal. m By using probabilistic computing, it is possible to significantly reduce the circuit size and power consumption of the learning device S while suppressing a decrease in accuracy.

[0067] Furthermore, the configurations of the shunt suppression subtraction unit SB and the activation function unit AF of the embodiment can be utilized to configure separate neural networks using them, and even in these cases, it is possible to significantly reduce the circuit size and power consumption while suppressing a decrease in accuracy as a learning device including the separate neural networks. [Industrial Applicability]

[0068] As described above, the present invention can be used in the field of learning devices, and particularly when applied to the field of learning devices that perform inference and learning including multilayer perceptrons (MLPs), particularly significant effects can be obtained. [Explanation of symbols]

[0069] SC Probabilistic Computing System MLP Multilayer Perceptron IN Input layer HD hidden layer OUT output layer HDC hidden layer circuit OUTC Output layer circuit 10 First Backpropagation Section 10A, 20Da multiplier 10B Differential circuit 10D, 20Cb, 20Cc sign determination circuit 10E, 10F, NT, NT1, NT2, NT3, NT4, NT5, NT6, NT7 Inverter 10G, 20Db, 20Dc, 20Ca exclusive OR circuit 11 First decoder 12 First weight storage section 13 First Encoder 20 Second Backpropagation Section 20A, 20B Multiplexer 20C1, 20CN Y -1 Sign determination circuit 20D1, 20DN Y -1 Absolute value determination circuit 21 Second decoder 22 Second weight storage section 23 Second Encoder NR neuron circuit MX multiplier AD, 10C summation calculation section SB Shunt suppression subtractor AF activation function part OR1, OR2, OR3, OR4, OR5, OR6, OR7, OR8, OR9, OR10, OR11, OR12, OR20, OR21, OR30 OR gate R1, R2, R3, R4, R5, R6, R7, R8, R9, R10, R11, R20, R21, R22, R30, R31 Registers AND1, AND2, AND3, AND4, AND5, AND6, AND7, AND8, AND9, AND10, AND11 AND gates

Claims

1. A learning device that performs inference and learning including a neural network, a circuit corresponding to each neuron constituting the neural network includes at least a summation circuit, a subtraction circuit, and an activation function circuit; a feedback means for feeding back output data from an output layer constituting the neural network to an input layer or a hidden layer constituting the neural network while weighting the output data accordingly; the summation circuit is composed of a plurality of OR gates connected in a manner corresponding to stochastic computing, and generates positive data and negative data as the summation results and outputs them to the subtraction circuit; A learning device characterized in that a number of registers corresponding to the number of said OR gates are connected between said OR gates.

2. A learning device that performs inference and learning including a neural network, a circuit corresponding to each neuron constituting the neural network includes at least a summation circuit, a subtraction circuit, and an activation function circuit; a feedback means having a learning function for feeding back output data from an output layer constituting the neural network to an input layer or a hidden layer constituting the neural network while weighting the output data accordingly; the subtraction circuit comprises inverters and AND gates connected in a manner to perform shunt suppression subtraction, which is a simulated subtraction compatible with stochastic computing; the inverter inverts the negative data output from the summation circuit to generate inverted data; the AND gate generates a product of the positive data output from the summation circuit and the generated inverted data, and outputs the product to the activation function circuit as the subtraction result in the subtraction circuit; The learning device is characterized in that the learning function is a learning function based on a learning rule corresponding to the error back propagation method.

3. A learning device that performs inference and learning including a neural network, a circuit corresponding to each neuron constituting the neural network includes at least a summation circuit, a subtraction circuit, and an activation function circuit; a feedback means for feeding back output data from an output layer constituting the neural network to an input layer or a hidden layer constituting the neural network while weighting the output data accordingly; A learning device characterized in that the activation function circuit, which includes two OR gates and one AND gate connected in a manner corresponding to probabilistic computing, generates and outputs output data as the neuron based on the subtraction result output from the subtraction circuit and a clock signal.

4. The subtraction circuit included in the learning device according to claim 2, the inverter and the AND gate connected in a manner corresponding to the probabilistic computing; the inverter generates the inverted data; the AND gate generates a product of the positive data output from the summation circuit and the generated inverted data, and outputs the product to the activation function circuit as the subtraction result.

5. The activation function circuit included in the learning device according to claim 3, two of said OR gates and one of said AND gates connected in a manner corresponding to said probabilistic computing; an activation function circuit for generating and outputting the output data based on the output subtraction result and the clock signal;

Citation Information

Patent Citations

  • Neuro imitating circuit

    JP1993067068A

  • Method for learning neural network element by probabilistic calculation method

    JP2009129302A