Information processing apparatus, and program
Patent Information
- Application Number
- JP2021138412
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2021-08-26
- Publication Date
- 2025-06-02
- Estimated Expiration
- 2041-08-26
AI Technical Summary
Existing neural network circuit designs face inefficiencies in machine learning and circuit optimization, necessitating improved design efficiency.
An information processing device that converts machine learning parameters of a first type neural network into those of a second type, generating manufacturing and estimation information for the second type neural network, utilizing conversion processing to optimize circuit design.
Enhances the efficiency of neural network circuit design by converting and optimizing machine learning parameters, leading to improved performance and reduced resource consumption.
Smart Images

Figure 00000017_0000 
Figure 00000018_0000 
Figure 00000019_0000
Abstract
Description
Technical Field
[0001] The present invention relates to an information processing apparatus and a program.
Background Art
[0002] Among neural network circuits that have been studied in recent years, there are those that use a circuit that multiplies weights corresponding to each of a plurality of input signals, accumulates the results of multiplying the weights, and performs non-linear conversion by an activation function to output. This circuit is a circuit that mimics neurons, and in a neural network circuit using this circuit, a plurality of such circuits are prepared and connected to each other to perform machine learning.
[0003] In this example, machine learning is performed on the weights and the connectivity between circuits that mimic neurons. However, because of the high costs of storing and reading the weights and performing the product-sum operation on input signals, various methods for performing efficient machine learning have been studied (Non-Patent Document 1).
Prior Art Documents
Non-Patent Documents
[0004]
Non-Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0005] Thus, in the past, when manufacturing neural network circuits, there were many factors to consider, such as improving the efficiency of machine learning and optimizing the circuit design. Currently, efforts are being made to improve performance by repeatedly trying out designs and machine learning processes, and improving design efficiency has been a challenge.
[0006] This invention has been made in view of the above circumstances, and one of its objectives is to provide an information processing device and program that can improve the efficiency of designing neural network circuits. [Means for solving the problem]
[0007] One aspect of the present invention, which solves the problems of the above-mentioned conventional example, is an information processing device that includes: an acceptance means for accepting input of machine learning parameters of a first type of neural network that has been trained to produce an output for a predetermined input; a conversion processing means for converting the accepted machine learning parameters of the first type of neural network into machine learning parameters of a second type of neural network that is of a different type from the first type of neural network; a generation means for generating manufacturing information for manufacturing the second type of neural network based on the converted machine learning parameters; an estimation means for generating estimation information regarding at least one of the scale or performance of the second type of neural network manufactured according to the manufacturing information; and a means for outputting the manufacturing information and the estimation information. [Effects of the Invention]
[0008] According to the present invention, by converting other types of neural networks that have already undergone machine learning to obtain machine learning parameters for a new neural network, and by generating and outputting estimated information such as its size and performance, it can contribute to improving the efficiency of neural network circuit design. [Brief explanation of the drawing]
[0009] [Figure 1]This is a block diagram showing an example configuration of an information processing device according to an embodiment of the present invention. [Figure 2] This is an explanatory diagram illustrating an example of a first type of neural network that is the target of processing by an information processing device according to an embodiment of the present invention. [Figure 3] This is an explanatory diagram illustrating an example of a second type of neural network generated by an information processing device according to an embodiment of the present invention. [Figure 4] This is an explanatory diagram illustrating another example of a second type of neural network generated by an information processing device according to an embodiment of the present invention. [Figure 5] This is a functional block diagram showing an example of an information processing device according to an embodiment of the present invention. [Figure 6] This is a flowchart illustrating an example of the operation of an information processing device according to an embodiment of the present invention. [Figure 7] Another flowchart illustrating an example of operation of an information processing device according to an embodiment of the present invention. [Figure 8] This is an explanatory diagram showing an example of the contents of a performance database used by an information processing device according to an embodiment of the present invention. [Figure 9] This is an explanatory diagram showing an example of a first type of neural network that is the target of processing in an example of the operation of an information processing device according to an embodiment of the present invention. [Figure 10] This is an explanatory diagram showing an example of a neuron cell circuit of a second type of neural network generated in an example of operation of an information processing device according to an embodiment of the present invention. [Modes for carrying out the invention]
[0010] Embodiments of the present invention will be described with reference to the drawings. As illustrated in Figure 1, the information processing device 1 according to an embodiment of the present invention is configured to include a control unit 11, a storage unit 12, an operation unit 13, a display unit 14, and an input / output unit 15.
[0011] The control unit 11 is a program control device such as a CPU, and operates according to a program stored in the memory unit 12. In this embodiment, the control unit 11 accepts input of machine learning parameters for a first type of neural network that has been trained to produce output information for predetermined input information. Here, the first type of neural network is, for example, a deep learning neural network, and is configured to include an input layer 20a, at least one intermediate layer 20b, c... and an output layer 20z, as illustrated in Figure 2.
[0012] Furthermore, each layer of the neural network here includes input nodes 21-1, 21-2…, 21-n and output nodes 22-1, 22-2…, 22-m, as illustrated in Figure 2. Each output node 22 (when it is not necessary to distinguish between output nodes 22-1, 22-2…, the suffixes -1, -2… are omitted and it will be referred to as output node 22) is composed of a multiply-accumulate unit 221 and a nonlinear function operation unit 222. Here, the multiply-accumulate unit 221 of the i-th (i=1,2,…,m) output node 22-i uses weights wij (j=1,2,…,n) set by machine learning and bias b to multiply the input values xj (j=1,2,…,n) from each of the input nodes 21-1, 21-2… by the corresponding weight wij and accumulate, and then adds bias b to the accumulated result.
[0013] The nonlinear function calculation unit 222 takes a predetermined nonlinear function (e.g., a sigmoid function) as input, receives the result of the calculation performed by the sum-of-accumulate unit 221, obtains the output value of the nonlinear function, and outputs the obtained output value. This output is either input to the input node 21 of the next layer of the neural network, or (in the case of the output of the final layer, which has no subsequent layers) is output externally.
[0014] Also, the control unit 11 converts the received machine learning parameters of the first type of neural network (weights wij and biases b, etc. of each neural network) into machine learning parameters of a second type of neural network that is different in type from the first type of neural network. Here, an example of the specific configuration of the second type of neural network will be described later.
[0015] Based on the machine learning parameters obtained by the conversion, the control unit 11 generates manufacturing information for manufacturing the second type of neural network. An example of this manufacturing information is described in a hardware description language such as VHDL (Verilog Hardware Description Language). In this example, the manufacturing information represents what the hardware of the second type of neural network is like. Since the manufacturing information in this example, the expression of the hardware by it, and the method of manufacturing the hardware based on the manufacturing information are widely known, the description here is omitted.
[0016] Also, the control unit 11 generates estimation information regarding at least one of the scale or performance of the second type of neural network manufactured according to the manufacturing information, and outputs the manufacturing information and the estimation information. The detailed operation content of this control unit 11 will be described later.
[0017] The storage unit 12 is a memory device, a disk device, etc., and holds the program executed by the control unit 11. Also, this storage unit 12 operates as a work memory of the control unit 11.
[0018] The operation unit 13 includes a mouse, a keyboard, etc., receives the user's operation, and outputs information representing the content of the operation to the control unit 11. Also, this input unit 13 receives data from an external storage device and outputs it to the control unit 11.
[0019] The display unit 14 includes, for example, a display, etc., and outputs information according to an instruction input from the control unit 11.
[0020] The input / output unit 15 is an interface such as USB (Universal Serial Bus) and accepts information from the outside, such as machine learning parameters for a first type of neural network, and outputs it to the control unit 11. The input / output unit 15 also outputs information to external devices according to instructions input from the control unit 11.
[0021] Here, an example of a second type of neural network will be described. In this embodiment, an example of this second type of neural network includes, for example, an input circuit section 30, at least one machine learning circuit 40, and an output circuit section 50, as illustrated in Figure 3(a). Here, the machine learning circuit 40 further includes at least one neuron cell circuit (NC) 41, the neuron cell circuit 41 includes at least an input section 410, an adder section 412 for accumulating the input data, and a nonlinear function calculation section 413, and further includes an inverter 411 for inverting the sign of the input data and outputting it in predetermined cases, which will be described later.
[0022] Specifically, the input unit 410, as illustrated in Figure 3(b), has at least one input port 4101, and accepts input data through each input port 4101. The input unit 410 either outputs the data received from each input port 4101 directly to the adder unit 412, or outputs it to the adder unit 412 via an inverter 411. The adder unit 412 then accumulates the data from each input port 4101 of the input unit 410, either as is or with its sign inverted via the inverter 411.
[0023] The nonlinear function calculation unit 413 takes the cumulative value of the data output by the adder unit 412 as input and outputs the result of calculating a predetermined nonlinear function for that input. In this example, the nonlinear function calculation unit 413 also outputs the result of calculating a predetermined nonlinear function for a value obtained by multiplying the cumulative value output by the adder unit 412 by a predetermined weight value. This nonlinear function calculation unit 413 may be formed by a predetermined arithmetic unit, but it may also be formed using a nonvolatile memory element such as ROM as follows.
[0024] That is, an example of this non-linear function operation unit 413 is a non-volatile memory element, and at its memory address a, the value of A·f(a·Δq) obtained by multiplying the output value of a predetermined non-linear function f corresponding to the input value q (= a·Δq) by a predetermined weight information A (A is a predetermined positive real value) is stored.
[0025] Here, Δq is, for example, using the maximum value Vmax that the adder unit 412 can output, the minimum value Vmin, and the domain of definition xmin, xmax of the function f (where xmin < xmax), Δq = (xmax - xmin) / (Vmax - Vmin) and is obtained as such. However, the calculation of Δq is not limited to this. As long as the value of the non-linear function f corresponding to the input value is output when the input value from Vmin to Vmax within the above range is input, Δq may be determined by other calculation methods. Alternatively, the domain of definition xmin, xmax of the function f may be set so that Δq = 1.
[0026] In the second type of neural network in this example, the machine learning parameters are the accumulation method of the data output from each input port 4101 of the input unit 410 (such as which input port 4101's output sign is inverted for accumulation), and the non-linear function calculated by the non-linear function operation unit 413 (in the case of the above example, corresponding to the value stored in each address of the memory element).
[0027] Also, the neuron cell circuit 41 in the second type of neural network is not limited to the above. For example, as illustrated in FIG. 4, it may include an input unit 410, a pair of adder units 412'a, b, a multiplication and addition unit 414 that multiplies weights, and a non-linear function operation unit 413', and may further include an inverter 411 that inverts the sign of the data received by the input unit 410 in a predetermined case to be described later. Note that the same reference numerals are given to those having the same configuration as in FIG. 3.
[0028] Even in the example of FIG. 4, the input unit 410 outputs the data input from each of the input ports 4101 as it is to the adder unit 412'a, or outputs it to the adder unit 412'b via the inverter 411. The adder unit 412'a accumulates the data output as it is by the input unit 410 and outputs an accumulated value of data to be multiplied by a positive weight (conveniently referred to as P data). Also, the adder unit 412'b accumulates the data with the sign inverted, which is output by the input unit 410 via the inverter 411, and outputs an accumulated value of data to be multiplied by a negative weight (conveniently referred to as N data).
[0029] The multiply-add unit 414 multiplies the P data output by the adder unit 412'a by a predetermined weight value Wp (where Wp is a positive real value). Also, the multiply-add unit 414 multiplies the N data output by the adder unit 412'b by a predetermined weight value Wn (where Wn is a positive real value, which may be equal to the value of Wp or different from the value of Wp). Then, the multiply-add unit 414 adds the value obtained by multiplying the P data by the weight value Wp and the value obtained by multiplying the N data by the weight value Wn and outputs the result.
[0030] The non-linear function operation unit 413' takes the value output by the multiply-add unit 414 as an input and outputs the operation result of a predetermined non-linear function for the input. This non-linear function operation unit 413' may also be formed by a predetermined arithmetic unit, or may be formed as follows using a non-volatile memory element such as a ROM.
[0031] In this example, the non-linear function operation unit 413' is a non-volatile memory element, and the value of f(a·Δq), which is the output value of a predetermined non-linear function f corresponding to the input value q (=a·Δq), is stored at the memory address a.
[0032] Here, Δq is calculated using, for example, the maximum value Vmax and the minimum value Vmin that the multiply-add unit 414 can output, and the domain of definition xmin, xmax of the function f (where xmin < xmax), Δq=(xmax - xmin) / (Vmax - Vmin) This is how it was calculated. However, the calculation of Δq is not limited to this; if the value of the nonlinear function f corresponding to the input value is output when the input value is Vmin to Vmax within the above range is input, then Δq may be determined by other calculation methods. Alternatively, the domain xmin,xmax of the function f may be set such that Δq=1.
[0033] Next, the operation of the control unit 11 according to one example of this embodiment will be described. In one example of this embodiment, the control unit 11 functionally implements a configuration including a receiving unit 51, a conversion processing unit 52, a generation unit 53, an estimation unit 54, and an output unit 55, as illustrated in Figure 5, according to a program stored in the storage unit 12.
[0034] The receiving unit 51 accepts input of machine learning parameters for a first type of neural network to be processed. For illustrative purposes, specifically, it accepts input of learning parameters for a deep learning neural network that has been trained to produce an output for a given input.
[0035] The conversion processing unit 52 converts the machine learning parameters of the first type of neural network received by the receiving unit 51 into machine learning parameters of the second type of neural network. In one example of this embodiment, the conversion processing unit 52 obtains the machine learning parameters of the second type of neural network using the machine learning parameters of the first type of neural network received by the receiving unit 51, along with separately input training data.
[0036] This training data can consist of pairs of input data and corresponding output data (training data) that a second type of neural network should produce.
[0037] Furthermore, assuming that the second type of neural network is as illustrated in Figure 3, the conversion processing unit 52 generates manufacturing information for configuring each neuron cell circuit 41 while referring to the learning parameters of the deep learning neural network received by the receiving unit 51.
[0038] Specifically, as illustrated in Figure 6, the conversion processing unit 52 sequentially selects each layer of the first type of neural network received by the receiving unit 51 as the processing layer, in order of proximity to the input layer 20a (S11: Selection of processing layer). The conversion processing unit 52 virtually generates and initializes at least one neuron cell circuit 41 corresponding to the output node 22-j (j=1,2…,m) of the selected processing layer (S12). The conversion processing unit 52 sequentially selects the virtually generated neuron cell circuits 41 and obtains weight values for the input nodes 21-1,21-2…,21-n connected to the output node 22-j corresponding to the selected neuron cell circuit 41 (S13).
[0039] The conversion processing unit 52 performs pruning by referring to the weight values of the input nodes 21-1, 21-2…, 21-n connected to each output node 22-j in the first type of neural network. For example, the conversion processing unit 52 excludes input nodes 21 whose weight values exceed a predetermined threshold, and virtually generates input ports 4101 corresponding to the input nodes 21 whose weight values exceed the threshold (S14: pruning process). As already explained, the threshold here may be set by the user, for example. The threshold for this weight value may also be determined by excluding a predetermined percentage of the lower number of input nodes 21 from the distribution of the weight values of each referenced input node 21. For example, if the lower 20% (this percentage may also be set by the user) of input nodes 21 are excluded, the smallest weight value among the weight values of the input nodes 21 that are not excluded may be used as the threshold. Furthermore, the threshold value may be set differently for each processing layer, depending on the position of the processing layer in the first type of neural network (its positional relationship to a predetermined layer, such as an input layer or output layer).
[0040] As another example, the threshold value and the like of this pruning process may be set so as to be minimized, for example, in the layer immediately before the layer immediately before the output layer. In this case, for example, when converting a deep learning neural network in which a fully connected layer is arranged immediately before the output layer to obtain a second type of neural network, the closer the layer (excluding the fully connected layer) is to the fully connected layer, the more input ports 4101 of the corresponding neuron cell circuit 41 will be reduced.
[0041] Also generally, when the first type of neural network is a deep learning neural network having the first layer as an input layer and the Nth layer as an output layer (N is an integer of 3 or more), the conversion processing unit 52 may determine the threshold value and the like of the above pruning process as follows. That is, when setting the neuron cell circuit 41 corresponding to the output node of the ith layer (where i is an integer such that 1 ≤ i ≤ N) of the first type of neural network layer, the conversion processing unit 52 controls such that when the value i representing the depth of the layer is within a range of at least a predetermined integer value J (where 1 < J ≤ N), the larger the value i, the smaller the number of input ports set by the pruning process (for example, the larger the threshold value) may be.
[0042] The conversion processing unit 52 sets a virtual adder unit 412, and when the weight value of the input node 21 corresponding to the input port 4101 virtually generated in step S14 is positive, connects the output of the input port 4101 directly to the adder unit 412. Also, when the weight value of the input node 21 corresponding to the input port 4101 virtually generated in step S14 is negative, the conversion processing unit 52 sets an inverter 411 and connects the output of the input port 4101 to the adder unit 412 via the inverter 411 (S15: Setting of the accumulation method).
[0043] Note that when the neuron cell circuit 41 in the second type of neural network is the one illustrated in FIG. 4, in this step S15, instead of the above-described processing, the following processing will be performed.
[0044] The conversion processing unit 52 that constitutes the neuron cell circuit 41 in the example of Figure 4 sets up a virtual pair of adder units 412'a and 412'b and a multiplication / addition unit 414. When the weight value of the input node 21 corresponding to the input port 4101 virtually generated in step S14 is positive, the output of the input port 4101 is connected directly to the adder unit 412'a. Furthermore, when the weight value of the input node 21 corresponding to the input port 4101 virtually generated in step S14 is negative, the conversion processing unit 52 sets up an inverter 411 and connects the output of the input port 4101 to the adder unit 412'b via the inverter 411.
[0045] Here, the adder unit 412'a accumulates the data output directly by the input unit 410 and outputs P data, which is the cumulative value of the data to be multiplied by a positive weight. The adder unit 412'b accumulates the data with the sign inverted, output by the input unit 410 via the inverter 411, and outputs N data, which is the cumulative value of the data to be multiplied by a negative weight.
[0046] The multiplication and addition unit 414 is set to multiply the P data output by the adder unit 412'a by a predetermined weight value Wp (where Wp is a positive real value), and also multiply the N data output by the adder unit 412'b by a predetermined weight value Wn (where Wn is a positive real value and may be equal to or different from Wp), and then add the value obtained by multiplying the P data by the weight value Wp and the value obtained by multiplying the N data by the weight value Wn and output the result. In this way, the conversion processing unit 52 that constitutes the neuron cell circuit 41 in the example of Figure 4 performs the setting of the accumulation method in step S15.
[0047] The conversion processing unit 52 sets up a virtual nonlinear function calculation unit 413, takes the output of the adder unit 412 or the multiplication / addition unit 414 set up in step S15 as the input to the nonlinear function calculation unit 413, and takes the output of the nonlinear function calculation unit 413 as the output of the neuron cell circuit 41 selected in step S12 (S16). The nonlinear function calculation unit 413 set up here is a nonvolatile memory element, and its memory address a stores the value obtained by multiplying f(a·Δq), which is the output value of a predetermined nonlinear function f corresponding to the input value q (=a·Δq), by a predetermined positive real value A (Δq has already been explained, so a repeated explanation is omitted). The specific details of the nonlinear function f will be described later.
[0048] The conversion processing unit 52 repeatedly performs the processes from steps S13 to S16 for each neuron cell circuit 41 related to the processing target layer virtually generated in step S12 (i.e., for each output node 22 of the processing target layer). This sets the weight values of the group of neuron cell circuits 41 of the second type of neural network corresponding to the input layer 20a of the first type of neural network up to the processing target layer selected in step S11.
[0049] Then, if the corresponding output layer is not output layer 20z (if there is another layer), the conversion processing unit 52 returns to step S12, selects the next layer as the corresponding output layer, and repeats the processing from S13 to S16. Furthermore, if the corresponding output layer is output layer 20z, the conversion processing unit 52 uses machine learning to set the nonlinear function calculated by the nonlinear function calculation unit 413 of each neuron cell circuit 41 that has been set up to this point (S17). Specifically, here, a linear combination of multiple distinct nonlinear functions fk(x) (k=1,2…) (a function obtained by linearly combining the above distinct nonlinear functions fk(x)):
number
number
[0050] In this example, the transformation processing unit 52 performs machine learning of nonlinear functions as follows: The transformation processing unit 52 sequentially selects input data included in the input training data and inputs it into a second type of neural network. The transformation processing unit 52 then obtains the output data of the second type of neural network at that time, compares it with the output data included in the training data corresponding to the input data, and corrects the coefficient ak related to the nonlinear function calculated by the nonlinear function calculation unit 413 of each neuron cell circuit 41 corresponding to the processing layer, as well as the parameters for each individual nonlinear function fk(x), in a method similar to backpropagation. The transformation processing unit 52 performs this process for each input data included in the training data and sets the nonlinear function A·fk(x) (in the example of Figure 3) or fk(x) (in the example of Figure 4) calculated by the nonlinear function calculation unit 413 of each neuron cell circuit 41 (where A is a predetermined positive real value).
[0051] After setting this nonlinear function, the conversion processing unit 52 may output information representing the machine learning parameters of the second type of neural network generated up to this point. For example, for each layer of the first type of neural network received by the receiving unit 51, the conversion processing unit 52 generates information sequentially recording information that identifies at least one neuron cell circuit 41 corresponding to that layer, associated with information that identifies that layer (here, the input layer 20a is set to "1", and thereafter a layer number is used that is incremented by "1" as it approaches the output layer 20z; hereafter, this number will be called the layer number).
[0052] Here, the information used to identify each neuron cell circuit 41 includes information identifying at least one input port 4101, excluding the input ports 4101 reduced by the pruning process (step S14), information on the presence or absence of an inverting unit 411 corresponding to the output of each input port 4101, information on the adder unit 412' and the multiplication / adding unit 414, or the adder unit 412, and information on the nonlinear function calculation unit 413 set in step S17, all of which are recorded in association with each other. The information on the nonlinear function calculation unit also includes information identifying a calculation unit that takes the value output by the adder unit 412 as input, calculates the nonlinear function determined in step S17, and outputs an output value (calculation result of the nonlinear function) of a predetermined number of output bits.
[0053] Furthermore, the conversion processing unit 52 may, for example, instruct the user, perform the following correction processing on at least some of the neuron cell circuits 41 corresponding to each node in each layer of the first type of neural network.
[0054] In other words, the conversion processing unit 52 may correct the number of bits used to represent the output value of the nonlinear function calculated by the nonlinear function calculation unit 413 of at least some of the neuron cell circuits 41 (hereinafter referred to as the function output value) (S18) and output the machine learning parameters of the second type of neural network. In other words, in one example of this embodiment, the conversion processing unit 52 reduces the number of bits used to represent the function output value for at least some of the neuron cell circuits 41 to a number of bits equal to or greater than the minimum number of bits determined by a predetermined method.
[0055] Here, the minimum number of bits may vary depending on how close the layer of the first type of neural network to which the output node corresponding to the neuron cell circuit 41 that is subject to bit reduction belongs (the corresponding layer) is to its output layer 20z. As an example, the conversion processing unit 52 sets the minimum number of bits of the function output value of the neuron cell circuit 41 related to the input layer 20a of the first type of neural network to 1 bit, and thereafter, for the intermediate layers 20b, c, etc., it sets the minimum number of bits of the function output value of the neuron cell circuit 41 related to each layer to 1 bit, 4 bits, 4 bits, etc., increasing as it approaches the output layer 20z. Furthermore, for the output layer 20z, it is also preferable that the conversion processing unit 52 ensures that the minimum number of bits of the function output value of the neuron cell circuit 41 related to the output layer 20z is equal to the number of bits before reduction (i.e., not reduced).
[0056] This is based on the consideration that later stages (layers closer to the output layer 20z) have inherited the characteristics of the earlier stages and have consolidated those characteristics, thus requiring greater expressiveness. Note that this setting of the number of bits for the function output value is just one example, and the number of bits for the function output value may be determined by other methods.
[0057] This step S18 process changes the bit count setting for the output value of the nonlinear function operation unit that was reduced, which is included in the machine learning parameters of the second type of neural network.
[0058] Specifically, as illustrated in Figure 7, the conversion processing unit 52 sequentially selects each layer of the first type of neural network as the processing target layer, starting from the input layer 20a (S21). The conversion processing unit 52 then checks whether the number of bits in the function output value of the neuron cell circuit 41 of the second type of neural network corresponding to the node belonging to the processing target layer is greater than the minimum number of bits defined for the neuron cell circuit 41 (S22). If it is greater (S22: Yes), it reduces the number of bits by "1" (S23).
[0059] The conversion processing unit 52 then sequentially inputs the input data contained in the input training data into the second type of neural network being processed and obtains the respective outputs (S24). The conversion processing unit 52 compares each of the training data corresponding to the input data with the corresponding output of the second type of neural network and determines whether the error, which can be expressed as the sum of the absolute values of the differences between them, falls below a predetermined threshold (S25). If the error falls below the predetermined threshold (S25: Yes), the conversion processing unit 52 returns to step S22 and continues processing.
[0060] On the other hand, if the error in step S25 does not fall below a predetermined threshold (S25: No), the conversion processing unit 52 increases the number of bits in the function output value of the neuron cell circuit 41 of the second type of neural network corresponding to the node belonging to the processing layer by "1" (S26), and returns to step S21. If there is a next layer, it selects the next layer as the processing layer and continues processing. If there is no next layer, the conversion processing unit 52 terminates the processing related to this bit correction.
[0061] Furthermore, in step S22, if the number of bits in the function output value is not greater than the minimum number of bits defined for the neuron cell circuit 41 (S22: No), the conversion processing unit 52 returns to step S21, selects the next layer as the processing target layer if there is a next layer, and continues processing, or terminates the bit correction process if there is no next layer.
[0062] Furthermore, although the conversion processing unit 52 performs the processing corresponding to the next layer after the processing in step S16, this embodiment is not limited to this, and the conversion processing unit 52 may repeat the processing from step S13 to step S17 multiple times (the processing shown by the dashed line in Figure 6). That is, the pruning process, the setting of the accumulation method in the neuron cell circuit 41 corresponding to each node of the processing target layer, and the setting and learning of the nonlinear function may be repeated a predetermined number of times. In this case, the threshold used in the pruning process is changed with each repeated execution. For example, the threshold used in the pruning process is increased by a predetermined percentage (e.g., 10%) with each repeated execution. This increases the number of input ports 4101 that are deleted by the pruning process with each repeated execution. Also, when performing repeated executions, the initial value of the nonlinear function in each neuron cell circuit 41 during the learning of the nonlinear function uses the result of the previous learning of the nonlinear function.
[0063] The generation unit 53 generates manufacturing information for manufacturing a second type of neural network based on the machine learning parameters of the second type of neural network obtained by the conversion processing unit 52.
[0064] Specifically, the generation unit 53 refers to information that identifies the neuron cell circuits 41, which are recorded in association with the layer numbers of each layer of the first type of neural network that was converted, as machine learning parameters for the second type of neural network obtained by the conversion processing unit 52. The generation unit 53 then generates hardware design information (which can be described in a predetermined hardware description language) that describes hardware elements corresponding to each input port 4101, hardware elements that multiply the data input to each input port 4101 by the corresponding weight values, and hardware elements corresponding to an adder unit 412 that accumulates the values multiplied by the weight values, according to the information of the input port 4101 which is associated with the information that identifies the input port 4101 associated with the referenced information.
[0065] Furthermore, the generation unit 53 generates information for hardware design to implement a nonlinear function calculation unit 413 that takes the output of the adder unit 412 as input and outputs the output value of a nonlinear function corresponding to the input value, based on the information of the nonlinear function calculation unit 413 associated with the referenced information.
[0066] The generation unit 53 generates hardware design information for the neuron cell circuits 41, which are recorded in association with each layer number, and connects these sequentially, for example, in the order of the layer numbers, to output information (manufacturing information) for actually manufacturing a second type of neural network using an FPGA or the like.
[0067] The estimation unit 54 refers to the manufacturing information generated by the generation unit 53 and generates estimation information relating to at least one of the size or performance of a second type of neural network to be manufactured according to the manufacturing information.
[0068] Specifically, the estimation unit 54 refers to the manufacturing information generated by the generation unit 53 and extracts, for each layer number, the number of neuron cell circuits 41 (total number of neurons) recorded in association with that layer number, and the average number of input ports 4101 in each of those neuron cell circuits 41 (average number of synapses).
[0069] The estimation unit 54 then generates information such as the total number of neuron cell circuits 41 (total NCs), the amount of power consumed, and the time required for one inference (latency) based on the information extracted here.
[0070] In one example of this embodiment, the estimation unit 54 refers to a performance database (Figure 8) that holds multiple records relating information such as the total number of neurons and average number of synapses for each layer number extracted from manufacturing information of hardware actually manufactured in the past, the sum of the number of neuron cell circuits 41 corresponding to the layer number (total NC count) actually measured for the hardware that was actually manufactured, the amount of power consumed, and the time required for one inference (latency).
[0071] The estimation unit 54 then compares the information on the total number of neurons and the average number of synapses for each layer number contained in each record stored in this performance database with the information on the total number of neurons and the average number of synapses for each layer number of the manufacturing information of the hardware to be estimated, which is extracted by referring to the manufacturing information generated by the generation unit 53, and finds the record that is most similar.
[0072] The estimation unit 54 generates estimation information by accumulating the total number of NCs, power consumption, and latency information contained in the records found for each layer number.
[0073] Furthermore, the similarity between the total number of neurons and the average number of synapses for each layer number can be measured as follows: For example, if the total number of neurons for one layer number to be compared is NNa and the average number of synapses is ASa, and the total number of neurons for the other layer number is NNb and the average number of synapses is ASb, then the similarity S is... S = α|NNa-NNb| + β|ASa-ASb| (where α and β are predetermined positive constants, and |X| represents the absolute value of X), and the smaller the similarity score, the more similar the two are considered to be.
[0074] The output unit 55 outputs at least one of the manufacturing information generated by the generation unit 53 and the estimation information generated by the estimation unit 54.
[0075] [Operation] Next, an example of the operation of the information processing device 1 of this embodiment will be described. In the following, as an example, the first type of neural network to be transformed will include, as illustrated in Figure 9(a), a convolutional layer which is an input layer 20a that accepts image data input, pooling layers which are intermediate layers 20b, c, and d, another convolutional layer, another pooling layer, a fully connected layer which is an intermediate layer 20e preceding the output layer 20z, and the output layer 20z. Furthermore, it will be assumed that this first type of neural network has been machine-trained to perform predetermined estimations using training data.
[0076] The information processing device 1 accepts a model of this first type of neural network (information that identifies each layer) and its machine learning parameters. The machine learning parameters of this first type of neural network include weight information between each of the input nodes 21-1, 21-2… included in each layer and each of the output nodes 22-1, 22-2,… as illustrated in Figure 9(b). Note that Figure 9(b) shows the case where there are only five input nodes for the sake of simplicity, but in reality, the number of input nodes may be much larger. Also, Figure 9(b) shows an example of weight information related to one output node 22.
[0077] In Figure 9(b), the weights from input nodes 21-1, 21-2, 21-3, 21-4, and 21-5 to output node 22 are assumed to be W1=0.08, W2=-0.24, W3=-0.18, W4=0.14, and W5=0.001, respectively.
[0078] The information processing device 1 converts the machine learning parameters of the first type of neural network that it has received into machine learning parameters of the second type of neural network. In the following example, the second type of neural network is assumed to be the one illustrated in Figure 3.
[0079] The information processing device 1 sequentially selects each layer of the first type of neural network that it has accepted as the target of transformation as a processing layer, and virtually generates and initializes at least one neuron cell circuit 41 corresponding to the output node 22-j (j=1,2…,m) of the selected processing layer.
[0080] As a specific example, the information processing device 1 virtually sets up input ports X1, X2, X5 corresponding to input nodes 21-1, 21-2, ..., 21-5 as a neuron cell circuit 41 corresponding to Figure 9(b).
[0081] The information processing device 1 then obtains the weight values for each neuron cell circuit 41 related to the input nodes 21-1, 21-2, ..., 21-5 connected to the corresponding output node 22, and removes the input nodes 21 whose absolute value falls below a predetermined threshold (for example, 0.01 in this case) from the obtained weight values (pruning process). In this example, the input port X5 corresponding to input node 21-5, which corresponds to weight W5, is removed.
[0082] The information processing device 1 refers to the weights from the corresponding input nodes 21-1, 21-2, 21-3, and 21-4 to the output node 22 for the remaining input ports X1, X2, X3, and X4, and sets whether or not to invert the input data based on the sign of each weight. That is, if the sign of the weight of the corresponding input node 21 is negative, the information processing device 1 sets the inverter 411 to invert the data at the input port Xi corresponding to that input node. In this example, the inverter 411 is set to invert the data input to input ports X2 and X3 (Figure 10).
[0083] Furthermore, the information processing device 1 sets the nonlinear functions to be calculated by the nonlinear function calculation unit 413 of each neuron cell circuit 41 using machine learning. This nonlinear function calculation unit 413 takes as input the output of the adder unit 412 in the corresponding neuron cell circuit 41, that is, in the example of Figure 10, the data input to input ports X1 and X4 and the data input to input ports X2 and X3, whose sign has been inverted by the corresponding inverter 411, and the result of accumulating them.
[0084] In one example of this embodiment, for machine learning of this nonlinear function, a function is obtained by linearly combining multiple different nonlinear functions fk(x) (k=1,2…):
number
[0085] Furthermore, the information processing device 1 may, as a pre-machine learning state (initial state), set the coefficient ak for each neuron cell circuit 41 related to the same nonlinear function fk(x) as the nonlinear function calculated by the output node 22 of the corresponding first type of neural network to a value larger than the coefficient ak related to other nonlinear functions fk(x). Specifically, the coefficient ak related to the same nonlinear function fk(x) as the nonlinear function calculated by the output node 22 of the corresponding first type of neural network may be set to a value η that is less than "1" but sufficiently close to "1", and the coefficient ak related to other nonlinear functions fk(x) may be set to a value obtained by dividing η by the number n of other nonlinear functions fk (in this example, the sum of the coefficients ak will be "1").
[0086] The information processing device 1 uses separately prepared training data (data that associates input values with the corresponding output data (training data) of the neural network that should be the correct answer) and sequentially inputs the input data contained in this training data into a second type of neural network that has been virtually set up.
[0087] The information processing device 1 then compares the output value output by a provisionally set second type of neural network with the training data included in the training data corresponding to the input data for each input of input data. Based on this comparison, the information processing device 1 corrects the coefficients ak, etc., of the nonlinear function calculated in the nonlinear function calculation unit of each neuron cell circuit 41 using the backpropagation method.
[0088] In this way, the information processing device 1 obtains machine learning parameters for a second type of neural network by performing two processes: one that fixes a nonlinear function and sets weight values corresponding to the input port 4101, and another that fixes weight values and sets a nonlinear function.
[0089] The information processing device 1 generates manufacturing information for manufacturing each of the neuron cell circuits 41 configured in this way as hardware.
[0090] Furthermore, the information processing device 1 refers to a performance database based on manufacturing information of hardware actually manufactured in the past, as illustrated in Figure 8. It compares the information on the total number of neurons and average number of synapses per layer number contained in each record stored in this performance database with the information on the total number of neurons and average number of synapses per layer number of the manufacturing information of the hardware to be estimated, which is extracted by referring to the generated manufacturing information, and finds the record that is most similar.
[0091] The information processing device 1 then aggregates the total number of NCs, power consumption, and latency information found for each layer number, generates estimated information, and presents it to the user.
[0092] A user of this information processing device 1 may refer to the provided estimate information and instruct the information processing device 1 to further reduce the number of bits used to represent the function output value, which is the output value of the nonlinear function calculation unit of each neuron cell circuit 41, for example, with the aim of reducing power consumption and latency.
[0093] Upon receiving this instruction, the information processing device 1 corrects the number of bits used to represent the function output value, which is the output value of the nonlinear function calculation unit of at least some of the neuron cell circuits 41 included in the second type of neural network obtained by the conversion. Specifically, the information processing device 1 sets the minimum number of bits for the function output value to 1 bit if it is in the input layer 20a, depending on the layer of the first type of neural network to which the output node corresponding to the neuron cell circuit 41 that was subject to bit reduction belongs. Subsequently, the information processing device 1 sets the minimum number of bits for the function output value of the neuron cell circuits 41 corresponding to the hidden layers 20b, c, etc. to 1 bit, 4 bits, 4 bits, etc., increasing as it approaches the output layer 20z.
[0094] The information processing device 1 then uses the training data to correct the number of bits in the function output value, which is the output value of the nonlinear function calculation unit of each neuron cell circuit 41, so that the difference between the output value when input data included in the training data is input and the corresponding training data falls below a predetermined difference, is greater than or equal to the minimum number of bits, and is the smallest possible number of bits.
[0095] The information processing device 1 then generates manufacturing information for manufacturing each corrected neuron cell circuit 41 as hardware, and, by referring to the performance database, it cumulatively calculates the total number of NCs, power consumption, and latency for each corrected layer number to generate estimation information, which is then presented to the user.
[0096] [Iterative cycle of pruning and nonlinear function learning] Furthermore, the information processing device 1 may repeat the pruning process again at the user's instruction or other request. That is, the information processing device 1 may, again, sequentially select each layer of the first type of neural network that has been transformed as a processing target layer, and perform the following processing for each neuron cell circuit 41 corresponding to the output node 22-j (j=1,2…,m) of the selected processing target layer.
[0097] In other words, for each neuron cell circuit 41, the information processing device 1 obtains the weight values for the input nodes 21-1, 21-2, ..., 21-5 connected to the corresponding output node 22, and removes the input nodes 21 whose absolute value falls below a predetermined threshold. Here, the threshold is set to be greater than the threshold used in the previous pruning process. For example, if the threshold in the previous pruning process was 0.01, the threshold in the current pruning process may be set to 0.15.
[0098] In this case, as already mentioned, in the example where the weights from input nodes 21-1, 21-2, 21-3, 21-4, and 21-5 to output node 22 are W1=0.08, W2=-0.24, W3=-0.18, W4=0.14, and W5=0.001 respectively, input port X5 corresponding to input node 21-5, which corresponds to weight W5, and input port X4 corresponding to input node 21-4, which corresponds to weight W4, are excluded.
[0099] The information processing device 1 then refers to the weights from the corresponding input nodes 21-1, 21-2, and 21-3 to the output node 22 for the remaining input ports X1, X2, and X3, and sets the weight values to be multiplied by the data input from each as Wp, Wm, and Wm.
[0100] Next, the information processing device 1 sets the nonlinear functions to be calculated in the nonlinear function calculation unit of each neuron cell circuit 41 using machine learning. At this time, the initial values of each nonlinear function are the same as the learning results of the previous nonlinear function.
[0101] This process reduces the number of input ports 4101, i.e., the average number of synapses. Furthermore, for neuron cell circuits 41 where all input ports 4101 have been removed by the pruning process, the neuron cell circuit 41 itself is removed and replaced with a ground terminal or the like. In this case, the total number of neurons is also reduced.
[0102] As one example, in the operation of transforming the deep learning neural network illustrated in Figure 9(a), the information processing device 1 reduces the number of input ports 4101 of the neuron cell circuit 41 corresponding to the convolutional layer, which is the input layer 20a that accepts image data input, to 30%. Furthermore, thresholds are set in each pruning process so that the reduction amount increases as the layer approaches the fully connected layer, for example, The number of input ports 4101 of the neuron cell circuit 41 corresponding to the intermediate layer 20b (pooling layer) can be increased to 25%. The number of input ports 4101 of the neuron cell circuit 41 corresponding to the intermediate layer 20c (convolutional layer) can be increased to 10%. The number of input ports 4101 of the neuron cell circuit 41 corresponding to the intermediate layer 20d (pooling layer) can be reduced to 10%. Reduce each of them.
[0103] Furthermore, the information processing device 1 reduces the number of input ports 4101 of the neuron cell circuit 41, which corresponds to the fully connected layer that is the intermediate layer 20e preceding the output layer 20z, by 40%. This reduces the overall circuit size.
[0104] According to the example of this embodiment, the user can adjust the pruning threshold, the number of iterations, the number of output bits of the nonlinear function of each neuron cell circuit 41, and obtain various estimation information, thereby repeating the design until the desired conditions are satisfied. The information processing device 1 may also calculate the accuracy of the output value (estimation accuracy) for each of the second type of machine learning parameters of the neural network obtained by repeating the pruning process, etc., using predetermined training data, and include this in the estimation information. [Explanation of Symbols]
[0105] 1 Information processing unit, 11 Control unit, 12 Storage unit, 13 Operation unit, 14 Display unit, 15 Input / Output unit, 21 Input node, 22 Output node, 30 Input circuit unit, 40 Machine learning circuit, 41 Neuron cell circuit, 50 Output circuit unit, 51 Receiving unit, 52 Conversion processing unit, 53 Generation unit, 54 Estimation unit, 55 Output unit, 221 Multiply-accumulate unit, 222 Nonlinear function calculation unit, 410 Input unit, 411 Inverter, 412, 412' Adder unit, 413 Nonlinear function calculation unit, 414 Power-adder, 4101 Input port.
Claims
1. An accepting means for accepting input of machine learning parameters of a first type of neural network that has machine-learned an output for a predetermined input; a conversion processing means for converting the machine learning parameters of the received first type of neural network into machine learning parameters of a second type of neural network that is different in type from the first type of neural network; a generating means for generating manufacturing information for manufacturing the second type of neural network based on the converted machine learning parameters; an estimation means for generating estimation information relating to at least one of the size and performance of a second type of neural network to be manufactured in accordance with the manufacturing information; a means for outputting the manufacturing information and the estimate information; An information processing device comprising:
2. 2. The information processing device according to claim 1, the first type of neural network is a deep learning neural network; the conversion processing means is a conversion processing means for virtually setting a neuron cell circuit for each output node of each layer of the first type neural network, At least one of the neuron cell circuits to be set comprises an input port, an inverter, an adder, and a nonlinear operation unit, and the inverter is an inverter that inverts the sign of input data, When the conversion processing means sets each of the neuron cell circuits, it arranges an input port corresponding to the input node based on the weight information of the input node connected to the output node of the corresponding first type neural network, and for each of the arranged input ports, it sets whether the data input to the input port is output as is, or after setting the inverter, or after inverting, setting the adder to accumulate the directly output data or the data output via the inverter; an information processing device that sets the nonlinear calculation unit to obtain an accumulated value output by the adder and output a function value of a predetermined nonlinear function corresponding to the obtained accumulated value;
3. 3. The information processing device according to claim 2, The information processing device performs pruning processing, in which, when setting each of the neuron cell circuits, the conversion processing means sets only input ports corresponding to input nodes whose weight information relating to the input nodes exceeds a predetermined threshold value, among the input nodes connected to the corresponding output nodes in the first type of neural network.
4. 4. The information processing device according to claim 3, The first type of neural network is a deep learning neural network having a first layer as an input layer and an Nth layer (N is an integer equal to or greater than 3) as an output layer, The conversion processing means controls the number of input ports set by the pruning process to be smaller as the value i representing the depth of the layer is larger, when the value i representing the depth of the layer is at least within a range of a predetermined integer value J (where 1<J≦N), when setting a neuron cell circuit corresponding to an output node of the i-th layer (i is an integer satisfying 1≦i≦N) of the first type neural network layer.
5. 5. The information processing device according to claim 2, The conversion processing means is an information processing device that, when setting each of the neuron cell circuits, sets weight values corresponding to the input ports set for each of the neuron cell circuits, and then sets the nonlinear functions to be calculated by the nonlinear calculation unit set for each of the neuron cell circuits by machine learning processing using predetermined teacher information.
6. 6. The information processing device according to claim 5, The conversion processing means, when performing machine learning on the nonlinear function calculated by the nonlinear calculation unit, sets the nonlinear function as a linear sum of a predetermined number of types of nonlinear functions, and sets the coefficient values to be multiplied by each type of nonlinear function through machine learning, thereby setting the nonlinear function calculated by the nonlinear calculation unit.
7. 7. The information processing device according to claim 2, The conversion processing means is an information processing device that, when setting each of the neuron cell circuits, sets the number of output bits of the nonlinear operation unit of the neuron cell circuit to be set in accordance with the position of the layer of the first type neural network that includes the output node corresponding to the neuron cell circuit to be set.
8. 8. The information processing device according to claim 2, The manufacturing information generated by the generating means is manufacturing information for manufacturing each of the neuron cell circuits set by the conversion processing means as hardware, and is information expressed in a predetermined hardware description language.
9. 9. The information processing device according to claim 2, The neural network system further comprises a means for accessing a performance database that stores information on at least one of the size and performance of previously manufactured second-type neural networks, in association with the number of neuron cell circuits included in the second-type neural network and the number of input ports included in the neuron cell circuits; The estimation means obtains information on at least one of the size and performance of a second type of neural network that is determined to be similar among the second type of neural networks that have been manufactured in the past and are stored in the performance database, using the number of neuron cell circuits set by the conversion processing means and the number of input ports included in the neuron cell circuit, and outputs the obtained information as estimation information.
10. Computer, An accepting means for accepting input of machine learning parameters of a first type of neural network that has machine-learned an output for a predetermined input; a conversion processing means for converting the machine learning parameters of the received first type of neural network into machine learning parameters of a second type of neural network that is different in type from the first type of neural network; a generating means for generating manufacturing information for manufacturing the second type of neural network based on the converted machine learning parameters; an estimation means for generating estimation information relating to at least one of the size and performance of a second type of neural network to be manufactured in accordance with the manufacturing information; a means for outputting the manufacturing information and the estimate information; A program that functions as a