Model Reduction Program, Apparatus, and Method

By identifying and compensating for neurons without input or output connections in neural networks, the model size is reduced while maintaining accuracy, addressing the inefficiencies of existing lightweighting techniques.

JP7700650B2Active Publication Date: 2025-07-01FUJITSU LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
JP2021191164
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2021-11-25
Publication Date
2025-07-01
Estimated Expiration
2041-11-25

AI Technical Summary

Technical Problem

Existing model lightweighting techniques for machine learning models result in reduced computational efficiency and accuracy due to the presence of redundant parameters and loss of useful information when deleting parameters with low influence.

Method used

Identify and delete neurons without input or output connections in a neural network, and compensate for the bias of these neurons by adding it to the bias of connected neurons on the output side, thereby reducing the model size while maintaining accuracy.

Benefits of technology

Improves the efficiency of model lightweighting by reducing the size of machine learning models without compromising their accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007700650000001
    Figure 0007700650000001
  • Figure 0007700650000002
    Figure 0007700650000002
  • Figure 0007700650000003
    Figure 0007700650000003
Patent Text Reader

Abstract

To increase the effect of weight reduction of a machine learning model while suppressing degradation of accuracy of the machine learning model.SOLUTION: A model reduction device performs a forward search of a neural network to specify a first neuron whose input weights are all 0, as a deletion object and correct all of its output weights to 0, and performs a backward search of the neural network to specify a second neuron whose output weights are all 0, as a deletion object and correct all of its input weights to 0. Further, the model reduction device compensates a bias of the first neuron specified as the deletion object by summing up the bias of the first neuron specified as the deletion object in the forward search and a bias of a third neuron connected to the first neuron on an output side. Then, the model reduction device deletes neurons specified as deletion objects from the neural network.SELECTED DRAWING: Figure 6
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The disclosed technology relates to a model reduction program, a model reduction device, and a model reduction method.

Background Art

[0002] Due to the evolution of deep learning technology and the like, machine learning models (hereinafter, also simply referred to as "models") tend to become huge. As the size of the model increases, the computing resources such as memory and processors required for machine learning also increase significantly. On the other hand, environments that require deep learning technology, such as mobile devices, tend to diversify. Also, although a huge model is required at the start of machine learning, as a result of machine learning, the number of parameters finally required for inference may not be large. Therefore, as a countermeasure against the above tendencies, a model lightweighting technology has emerged, in which machine learning of a model is executed in an environment with abundant computing resources such as a server, unnecessary parameters are deleted, and the lightweighted model is used for inference.

[0003] For example, when creating a fuzzy inference model, a method for optimizing the configuration of the fuzzy inference model has been proposed, which deletes meaningless input and output parameters and shortens the calculation time by the fuzzy inference model. This method gives arbitrary input data to the fuzzy inference model, calculates the corresponding output data, creates a plurality of sets of pseudo data, and also constructs a neural network having the same input and output parameters as this fuzzy inference model. Further, this method determines the characteristic values of the neural network by using the pseudo data as teacher data, and calculates the influence degree of each input parameter on each output parameter by using this neural network. Then, this method extracts input parameters with a small influence degree on any output parameter and output parameters with a small influence degree from any input parameter. Then, this method deletes the extracted input and output parameters from the input and output parameters of the fuzzy inference model and corrects the fuzzy inference model.

Prior Art Documents

Patent Document

[0004]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0005] However, like the prior art methods for model lightweighting, simply deleting parameters with low influence may leave redundant parameters in the network configuration. In such cases, the computational efficiency of inference by the generated model decreases. Also, simply deleting parameters with low influence may result in the loss of information useful for inference, and the accuracy of the model after parameter deletion may deteriorate.

[0006] As one aspect, the disclosed technology aims to improve the effect of lightweighting a machine learning model while suppressing a decrease in the accuracy of the machine learning model.

Means for Solving the Problems

[0007] As one embodiment, the disclosed technology identifies, as deletion targets, a first neuron that has no connection from the input layer and a second neuron that has no connection to the output layer in a neural network. Also, the disclosed technology adds the bias of the first neuron to the bias of a third neuron that is connected to the first neuron on the output side. Then, the disclosed technology deletes the identified neurons to be deleted from the neural network.

Effects of the Invention

[0008] As one aspect, it has the effect of being able to improve the effect of lightweighting a machine learning model while suppressing a decrease in the accuracy of the machine learning model.

Brief Description of the Drawings

[0009]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

Figure 9

Figure 10

Figure 11

Figure 12

Figure 13

Figure 14

Figure 15

Figure 16

Figure 17

Figure 18

Figure 19

Embodiments for Carrying Out the Invention

[0010] Hereinafter, with reference to the drawings, an example of an embodiment according to the disclosed technology will be described.

[0011] As shown in FIG. 1, a parameter table representing a neural network, which is a machine learning model, is input to the model reduction device 10 according to this embodiment. In this embodiment, the parameter table input to the model reduction device 10 is a parameter table in which some parameters have been deleted by an existing model lightweighting technology.

[0012] With reference to FIG. 2, an example of an existing model lightweighting technology will be described. In FIG. 2, circles represent neurons of the neural network, and arrows represent connections between neurons. The same applies to the following figures. Also, a weight, which is one of the parameters of the model, is set for the connection between neurons. The existing model lightweighting technology, for example, as shown in FIG. 2, applies a threshold to the weights between neurons, which are the parameters of the model for which machine learning has been executed, and corrects the weights below the threshold to 0. The middle figure in FIG. 2 shows that the weights between neurons represented by the dashed arrows have been corrected to 0. Then, the existing model lightweighting technology outputs a model with reduced parameters as shown in the lower figure in FIG. 2 by removing the parts with weights of 0 as unnecessary parameters.

[0013] In the case of a model lightweighted by existing model lightweighting techniques, as shown in FIG. 3, there may be neurons without input (neuron I indicated by a thick-lined circle in FIG. 3) and neurons not used for output (neuron L indicated by a double-lined circle in FIG. 3) remaining in the model. In this case, the weights between each neuron from the neuron without input to the output layer (the weights in the part of the broken-line arrow in FIG. 3) are unnecessary parameters not used for calculating the output of the model. Similarly, the weights between each neuron from the input layer to the neuron not used for output (the weights in the part of the dotted-line arrow in FIG. 3) are also unnecessary parameters not used for calculating the output of the model.

[0014] Also, each neuron has a bias as a parameter. For example, when the value y output from a neuron is calculated by a simple linear function (y = ax + b), b is the bias term. Here, x is the value output from the previous neuron, and a is the weight between the previous neuron and the target neuron. The bias is a constant value independent of the input obtained as a result of machine learning. When the weight between a neuron without input as described above (for example, I) and the neuron connected to the output side of that neuron is simply deleted, the means for transmitting the bias information of that neuron to the neuron on the output side is lost. As a result, useful information for inference may be lost, and the accuracy of the model after size reduction may deteriorate.

[0015] Therefore, in this embodiment, the parameters are deleted so as to suppress the decrease in the accuracy of the model and improve the effect of model lightweighting, thereby reducing the size of the model. Hereinafter, the functional configuration of the model reduction device 10 according to this embodiment will be described in detail. In the following, as shown in FIG. 4, when neuron i in the n - 1 layer and neuron j in the n layer are in a connection relationship, the weight between neuron i and neuron j is denoted as "w" i,j (n) ". Also, the weight w i,j (n) is referred to as the output weight of neuron i or the input weight of neuron j. Further, the bias of neuron i is denoted as "b" iIt is denoted as "」. The output weight is an example of the "output-side weight" of the disclosed technology, and the input weight is an example of the "input-side weight" of the disclosed technology.

[0016] Functionally, as shown in FIG. 1, the model reduction device 10 includes a correction unit 12, a compensation unit 14, and a deletion unit 16. The correction unit 12 is an example of the "specific unit" of the disclosed technology.

[0017] The correction unit 12 acquires the parameter table input to the model reduction device 10. FIG. 5 shows an example of the parameter table. The example in FIG. 5 is a parameter table of a neural network represented in a graph expression as shown in the upper diagram of FIG. 5. As shown in FIG. 5, the parameter table is provided for each layer. In the parameter table of each layer, as shown by "input" in FIG. 5, the neurons of the corresponding layer correspond to each row. Also, as shown by "output" in FIG. 5, the neurons of the upper layer of the corresponding layer, that is, the neurons whose output values are input to the neurons of the corresponding layer, correspond to each column. Each element of the matrix stores the weight between the neurons corresponding to that row and column. More specifically, each row of the parameter table stores the input weight of the neuron corresponding to that row, and each column of the parameter table stores the output weight of the neuron corresponding to that row. Further, in the parameter table of each layer, the bias of the neuron of the corresponding layer is also stored in the last column of each row.

[0018] In the neural network, the correction unit 12 identifies, as deletion targets, a first neuron having no connection from the input layer and a second neuron having no connection to the output layer. Then, in the parameter table, the correction unit 12 corrects the output weight of the first neuron to 0 and corrects the input weight of the second neuron to 0.

[0019] Specifically, as shown in FIG. 6, the correction unit 12 sequentially identifies the first neurons to be deleted by forward search from the input layer to the output layer in the neural network. More specifically, the correction unit 12 sequentially searches for neurons with all input weights being 0 starting from the input layer. In the example of FIG. 6, based on the fact that all the elements in the row of neuron I in the parameter table of layer n = 2 are 0, the correction unit 12 determines that all the input weights of neuron I are 0 and identifies neuron I as a deletion target. Then, the correction unit 12 corrects all the output weights of neuron I, that is, the weights in the column of neuron I in the parameter table of layer n = 3, to 0. By sequentially searching for neurons with all input weights being 0 through forward search, the correction unit 12 also identifies that all the input weights of neuron M are 0 and corrects all the output weights of neuron M to 0 (the dashed arrows in FIG. 6).

[0020] Similarly, as shown in FIG. 6, the correction unit 12 sequentially identifies the second neurons to be deleted by backward search from the output layer to the input layer in the neural network. More specifically, the correction unit 12 sequentially searches for neurons with all output weights being 0 starting from the output layer. In the example of FIG. 6, based on the fact that all the elements in the column of neuron L in the parameter table of layer n = 4 are 0, the correction unit 12 determines that all the output weights of neuron L are 0 and identifies neuron L as a deletion target. Then, the correction unit 12 corrects all the input weights of neuron L, that is, the weights in the row of neuron L in the parameter table of layer n = 3, to 0. By sequentially searching for neurons with all output weights being 0 through backward search, the correction unit 12 also identifies that all the output weights of neuron G are 0 and corrects all the input weights of neuron G to 0 (the dotted arrows in FIG. 6).

[0021] In addition, the correction unit 12 notifies the compensation unit 14 to execute a process of compensating the bias of the first neuron identified as a deletion target in the forward search.

[0022] Based on the notification from the correction unit 12, the compensation unit 14 compensates the bias of the first neuron to be deleted by adding the bias of the first neuron to the bias of the third neuron connected to the first neuron on the output side. For example, the compensation unit 14 adds the bias of the first neuron to the bias of the third neuron by adding the value obtained by multiplying the bias of the first neuron by the weight between the first neuron and the third neuron.

[0023] More specifically, when the neuron I with bias b I is specified as the first neuron to be deleted, it will be described. As shown by the dashed-dotted line in the upper diagram of FIG. 6 and the lower diagram of FIG. 6, the neuron I of n = 2 layers is connected to each of the neurons L and M of n = 3 layers. Also, the bias of neuron L is b L , the bias of neuron M is b M , the weight between neuron I and neuron L is w I,L (3) , and the weight between neuron I and neuron M is w I,M (3) . In this case, the compensation unit 14 calculates b L and b M as shown below, and updates the values in the bias columns of the rows corresponding to each of neurons L and M in the parameter table of the n = 3 layer. b L ←b L +w I,L (3) b I , b M ←b M +w I,M (3) b I

[0024] The deletion unit 16 deletes the specified neurons to be deleted from the neural network. The neurons to be deleted have all input weights and output weights set to 0 in the parameter table. Specifically, the deletion unit 16 deletes the rows and columns corresponding to the weights of the neurons to be deleted in the parameter table. More specifically, when neuron i in layer n - 1 is the neuron to be deleted, the deletion unit 16 deletes the row of neuron i with all weights set to 0 in the parameter table of layer n - 1 and the column of neuron i with all weights set to 0 in the parameter table of layer n.

[0025] For example, as shown in the left diagram of Fig. 7, assume that neuron D in layer n = 2 is specified as the neuron to be deleted. In this case, the weights of the row "D" in the parameter table of layer n = 2 and the column "D" in the parameter table of layer n = 3 are 0. As shown in the right diagram of Fig. 7, the deletion unit 16 deletes the row "D" in the parameter table of layer n = 2 and the column "D" in the parameter table of layer n = 3. Thereby, the size of the parameter table, that is, the size of the model, is reduced. The deletion unit 16 outputs the parameter table with the reduced size.

[0026] The model reduction device 10 may be implemented by, for example, the computer 40 shown in Fig. 8. The computer 40 includes a CPU (Central Processing Unit) 41, a memory 42 as a temporary storage area, and a non-volatile storage unit 43. The computer 40 also includes input / output devices 44 such as an input unit and a display unit, and an R / W (Read / Write) unit 45 that controls reading and writing of data to and from a non-temporary storage medium 49. The computer 40 also includes a communication I / F (Interface) 46 connected to a network such as the Internet. The CPU 41, the memory 42, the storage unit 43, the input / output devices 44, the R / W unit 45, and the communication I / F 46 are connected to each other via a bus 47.

[0027] The storage unit 43 may be implemented by a HDD (Hard Disk Drive), SSD (Solid State Drive), flash memory, or the like. A model reduction program 50 for causing the computer 40 to function as the model reduction device 10 is stored in the storage unit 43 serving as a storage medium. The model reduction program 50 includes a correction process 52, a compensation process 54, and a deletion process 56.

[0028] The CPU 41 reads the model reduction program 50 from the storage unit 43 and expands it in the memory 42, and sequentially executes the processes included in the model reduction program 50. By executing the correction process 52, the CPU 41 operates as the correction unit 12 shown in FIG. 1. Also, by executing the compensation process 54, the CPU 41 operates as the compensation unit 14 shown in FIG. 1. Further, by executing the deletion process 56, the CPU 41 operates as the deletion unit 16 shown in FIG. 1. As a result, the computer 40 that has executed the model reduction program 50 functions as the model reduction device 10. Note that the CPU 41 that executes the program is hardware.

[0029] Note that the functions realized by the model reduction program 50 can also be realized by, for example, a semiconductor integrated circuit, and more specifically, an ASIC (Application Specific Integrated Circuit) or the like.

[0030] Next, the operation of the model reduction device 10 according to the present embodiment will be described. When a parameter table representing a neural network and having some parameters deleted by an existing model lightweighting technique is input to the model reduction device 10, the model reduction process shown in FIG. 9 is executed in the model reduction device 10. Note that the model reduction process is an example of the model reduction method of the disclosed technique.

[0031] In step S10, the modification unit 12 acquires the parameter table input to the model reduction device 10. Next, in step S20, the modification unit 12 executes forward weight modification processing, identifies a first neuron without a connection from the input layer as a deletion target, and modifies the output weight of the first neuron to 0 in the parameter table. At this time, the compensation unit 14 also executes processing to compensate for the bias of the first neuron. Next, in step S40, the modification unit 12 executes reverse weight modification processing, identifies a second neuron without a connection to the output layer as a deletion target, and modifies the input weight of the second neuron to 0 in the parameter table. Next, in step S60, the deletion unit 16 executes deletion processing and deletes the neuron to be deleted from the neural network. Hereinafter, each of the forward weight modification processing, the reverse weight modification processing, and the deletion processing will be described in detail.

[0032] First, with reference to FIG. 10, the forward weight modification processing will be described.

[0033] In step S21, the modification unit 12 sets a variable n for identifying the layer to be processed in the neural network to 2. Next, in step S22, the modification unit 12 determines whether n exceeds the number of layers N of the neural network. If n does not exceed N, the process proceeds to step S23.

[0034] In step S23, the modification unit 12 acquires a list {c i} of neurons whose input weights in the (n - 1)-th layer are all 0. i is the number of the neuron in the (n - 1)-th layer, and i = 1, 2, ···, I n-1 (I n-1 is the number of neurons in the (n - 1)-th layer). c i is the number of the neuron among the neurons in the (n - 1)-th layer whose input weights are all 0. Specifically, the modification unit 12 adds the numbers of the neurons corresponding to the rows with all 0 weights in the parameter table of the (n - 1)-th layer to the list and acquires {c i}.

[0035] Next, in step S24, the correction unit 12 sets i to 1. Next, in step S25, the correction unit 12 checks whether i is in the list {c i} the maximum number of neurons in n-1 Determine whether i exceeds C n-1 If j does not exceed 1, the process proceeds to step S26. In step S26, the correction unit 12 sets j to 1. j is the number of the neuron in the nth layer, and j=1, 2, . . . , J n (J n (where j is the number of neurons in the nth layer). Next, in step S27, the correction unit 12 corrects j to J n Determine whether j exceeds J n If it does not exceed this, the process proceeds to step S28.

[0036] In step S28, the compensation unit 14 compensates for the bias of the i-th neuron in the n-1 layer by adding the bias of the i-th neuron in the n-1 layer to the bias of the j-th neuron in the n layer. For example, the compensation unit 14 compensates for the bias of the j-th neuron in the n layer by adding the bias of the i-th neuron in the n-1 layer to the bias of the j-th neuron in the n layer. j ←b j +w c_i,j (n) b i and updates the value of the bias column of the row corresponding to the j-th neuron in the parameter table of the n-th layer. Next, in step S29, the correction unit 12 deletes the output weight from the i-th neuron in the n-1-th layer to the j-th neuron in the n-th layer. Specifically, the correction unit 12 deletes the weight w c_i,j is corrected to 0. As a result, the i-th neuron in the n-1th layer has both input and output weights set to 0. Note that "c_i" is written as c i However, c_i=c i The same applies to c_j described below.

[0037] Next, in step S30, the correction unit 12 increments j by 1 and returns to step S27. nIf it exceeds, the process proceeds to step S31. In step S31, the correction unit 12 increments i by 1 and returns to step S25. In step S25, if i exceeds C n-1 it proceeds to step S32. In step S32, the correction unit 12 increments n by 1 and returns to step S22. In step S22, if n exceeds N, the forward weight correction process ends and returns to the model reduction process (Fig. 9).

[0038] If there is no connection relationship between the i-th neuron in the (n - 1)-th layer and the j-th neuron in the n-th layer, the processes in steps S28 and S29 are skipped. Also, if i is not included in the list {c i}, that is, if any of the input weights of the i-th neuron in the (n - 1)-th layer is not 0, the processes in steps S27 to S30 are skipped. Then, in step S31, i can be incremented by 1 and return to step S25.

[0039] Next, referring to Fig. 11, the backward weight correction process will be described.

[0040] In step S41, the correction unit 12 sets the variable n, which specifies the layer to be processed in the neural network, to N - 1. Next, in step S42, the correction unit 12 determines whether n is less than 2. If n is 2 or more, the process proceeds to step S43.

[0041] In step S43, the correction unit 12 obtains the list {c j} of neurons in the n-th layer whose output weights are all 0. c j is the number of the neuron in the n-th layer whose output weights are all 0. Specifically, the correction unit 12 adds the numbers of the neurons corresponding to the columns with all 0 weights in the parameter table of the (n + 1)-th layer to the list to obtain {c j}.

[0042] Next, in step S44, the correction unit 12 sets j to 1. Next, in step S45, the correction unit 12 determines whether j exceeds the maximum value C of the neuron numbers included in the list {c j}. If j does not exceed C n , the process proceeds to step S46. In step S46, the correction unit 12 sets i to 1. Next, in step S47, the correction unit 12 determines whether i exceeds I n . If i does not exceed I n-1 , the process proceeds to step S49. n-1

[0043] In step S49, the correction unit 12 deletes the input weight from the i-th neuron in the (n - 1)-th layer to the j-th neuron in the n-th layer. Specifically, the correction unit 12 corrects the weight w i,c_j (n) stored in the parameter table of the n-th layer to 0. As a result, the input weight and the output weight of the j-th neuron in the n-th layer both become 0.

[0044] Next, in step S50, the correction unit 12 increments i by 1 and returns to step S47. In step S47, if i exceeds I n-1 , the process proceeds to step S51. In step S51, the correction unit 12 increments j by 1 and returns to step S45. In step S45, if j exceeds C n , the process proceeds to step S52. In step S52, the correction unit 12 decrements n by 1 and returns to step S42. In step S42, if n becomes less than 2, the reverse-direction weight correction process ends and returns to the model reduction process (Figure 9).

[0045] Note that if there is no connection relationship between the i-th neuron in the (n - 1)-th layer and the j-th neuron in the n-th layer, the process of step S49 above is skipped. Also, if j is in the list {c j ​If it is not included in}, that is, if any of the output weights of the j-th neuron in the n-th layer is not 0, the processing in steps S47 to S50 is skipped. Then, in step S51, j is incremented by 1 and the process returns to step S45.

[0046] Next, referring to FIG. 12, the deletion process will be described.

[0047] In step S61, the deletion unit 16 sets a variable n, which specifies the layer to be processed in the neural network, to 2. Next, in step S62, the deletion unit 16 determines whether n exceeds the number of layers N of the neural network. If n does not exceed N, the process proceeds to step S63.

[0048] In step S63, the deletion unit 16 obtains a list {c i} of neurons in the (n - 1)-th layer whose input weights are all 0. Specifically, the deletion unit 16 adds the numbers of neurons corresponding to the rows with all 0 weights in the parameter table of the (n - 1)-th layer to the list to obtain {c i}. Next, in step S64, the deletion unit 16 obtains a list {d i} of neurons in the (n - 1)-th layer whose output weights are all 0. Specifically, the deletion unit 16 adds the numbers of neurons corresponding to the columns with all 0 weights in the parameter table of the n-th layer to the list to obtain {d i}.

[0049] Next, in step S65, the deletion unit 16 obtains a list {e i} having as elements the common elements between the list {c i} and the list {d i}. That is, the list {e i} stores the numbers of neurons in the (n - 1)-th layer whose input weights and output weights are both all 0. Next, in step S66, the deletion unit 16 obtains the difference set {f i} between the list {i} containing all the numbers of neurons in the (n - 1)-th layer and the list {e iObtain {}. That is, in the list {f i}, the numbers of the neurons in the (n - 1)-th layer that are not the ones to be deleted are stored.

[0050] Next, in step S67, the deletion unit 16 updates the weight w in the parameter table of the (n - 1)-th layer so that it becomes w h,f_i (n-1) and updates the weight w in the parameter table of the n-th layer so that it becomes w h,i’ (n-1) Here, h is the number of the neurons in the (n - 2)-th layer (h = 1, 2, ···), and i' is the number newly re-numbered as 1, 2, ··· for the numbers included in {f f_i,j (n) and updates the weight w in the parameter table of the n-th layer so that it becomes w i’,j (n) For example, when {f i} = {1, 3}, the third row of the parameter table of the (n - 1)-th layer becomes the second row of the parameter table after deletion, and the third column of the parameter table of the n-th layer becomes the second column of the parameter table after deletion. That is, the rows of the parameter table of the (n - 1)-th layer corresponding to the neurons with the numbers included in the list {e i} and the columns in the parameter table of the n-th layer are deleted. i} and the columns in the parameter table of the n-th layer are deleted.

[0051] Next, in step S68, the deletion unit 16 increments n by 1 and returns to step S62. In step S62, when n exceeds N, the deletion process ends and returns to the model reduction process (Figure 9).

[0052] As described above, the model reduction device according to the present embodiment identifies, as deletion targets, a first neuron that has no connection from the input layer and a second neuron that has no connection to the output layer in a neural network. Further, the model reduction device compensates for the bias of the first neuron by adding it to the bias of a third neuron that is connected to the first neuron on the output side. Then, the model reduction device deletes the identified neurons to be deleted from the neural network. Thereby, it is possible to improve the effect of reducing the weight of the machine learning model while suppressing a decrease in the accuracy of the machine learning model.

[0053] Note that, in the above embodiment, as a process of compensating for the bias of the first neuron, the case where a value obtained by multiplying the bias of the first neuron and the weight between the first neuron and the third neuron is added to the bias of the third neuron has been described, but the present invention is not limited to this. For example, a value obtained by multiplying a value obtained by applying the activation function of the first neuron to the bias of the first neuron and the weight between the first neuron and the third neuron may be added to the bias of the third neuron. In this case, the model reduction device acquires, together with the parameter table, a layer information table and a function table as shown in, for example, FIG. 13. In the example of FIG. 13, in the layer information table, the activation function name used in the layer is defined in association with the layer number. In the function table, the activation function name and the function object used for the calculation of the activation function are defined in association with each other. When adding the bias of neuron i in layer n-1 to neuron j in layer n, for example, the compensation unit of the model reduction device acquires the activation function corresponding to layer n-1 from the layer information table and acquires the function object corresponding to the activation function from the function table. Then, the compensation unit updates the bias b of neuron j by applying the obtained function object to f below. j to update. b j ←b j +w i,j f(b i )

[0054] Also, the above embodiment is applicable to a neural network configured to include a convolutional layer. In this case, the model reduction device acquires, for example, a layer information table and a parameter table as shown in FIG. 14. In the example of FIG. 14, the layer information table defines the attributes of each layer in association with the layer numbers. Note that in FIG. 14, "conv" of the attributes represents a convolutional layer, and "fc" represents a fully connected layer. Also, the parameter table of each layer has a format corresponding to the attributes of that layer. The parameter table for the fc layer is the same as the parameter table described in the above embodiment. In the parameter table of the convolutional layer, weights for the filter size applied to that layer are stored as elements of a matrix corresponding to each neuron. In the example of FIG. 14, an example with a filter size of 3×3 is shown. In this case, the weight corresponding to the k-th element from the left and the l-th element from the top of the filter between the i-th neuron in the n-1 layer and the j-th neuron in the n layer is w i,j,k,l (n) is represented by. For example, w 2,1,2,2 (2) corresponds to the element indicated by the dashed line in the parameter table of FIG. 14.

[0055] In the case of the parameter table of the convolutional layer, the model reduction device identifies neurons corresponding to rows or columns in which all weights, including the weights of each element of the filter, are 0, as neurons with 0 input weights or output weights. For example, in the case of the left diagram of FIG. 15, since the input weights of the third neuron in the n = 2 layer are all 0, the correction unit of the model reduction device identifies the third neuron in the n = 2 layer as a target to be deleted. Then, as shown in the right diagram of FIG. 15, the correction unit corrects the weights in the third column, which are the output weights of the third neuron in the n = 2 layer, to 0 in the parameter table of the n = 3 layer. Then, the deletion unit of the model reduction device deletes the third row including the 3×3 elements of the filter in the parameter table of the n = 2 layer and the third column in the parameter table of the n = 3 layer, which are indicated by the shaded part in the right diagram of FIG. 15. In this way, the disclosed technology can reduce the size of the model even for a neural network having a configuration including a convolutional layer. Note that in FIG. 15, the notation of the column in which the bias value is stored in the parameter table is omitted.

[0056] Also, in the above embodiment, the case where a parameter table in which some parameters have been deleted by an existing model lightweighting technique is input to the model reduction device has been described. However, a parameter table before model lightweighting may be input. In this case, the model reduction device may be provided with an existing model lightweighting function.

[0057] Here, an example of the relationship between the reduction rate of the model size and the accuracy when the disclosed technology is applied will be described. Here, as the neural network, VGG-19-BN of VGGNet with the layer configuration shown in FIG. 16 is used, and as the dataset, CIFAR-10 is used. FIG. 17 shows the accuracy of each model for the cases where the reduction rate of the size is 90% and 98%, respectively. The total data size is calculated by the number of input channels × the number of output channels × the filter size × 4 × 2. In this calculation formula, "4" represents the amount of information of one floating-point variable in bytes, and "2" is multiplied by 2 because one weight parameter contains two types of information, weight information and gradient information. As shown in FIG. 17, it can be seen that in any case of the reduction rate, there is no change in the accuracy of the model before and after the deletion of the parameters, and the influence on the accuracy due to the size reduction is suppressed.

[0058] Regarding the reduced data size of each layer of the neural network in the above example, the case where the reduction rate is 90% is shown in FIG. 18, and the case where the reduction rate is 98% is shown in FIG. 19. In FIGS. 18 and 19, "test_acc" is the accuracy of the prediction of the neural network for the test data, which is the same as "accuracy" in FIG. 17. Also, "train_acc" is the accuracy of the prediction of the neural network for the training data. Note that "accuracy" is the ratio of the value predicted by the neural network to the correct answer.

[0059] Also, in the above embodiment, the mode in which the model reduction program is pre-stored (installed) in the storage unit has been described, but it is not limited to this. The program according to the disclosed technology can also be provided in a form stored in a storage medium such as a CD-ROM, a DVD-ROM, or a USB memory.

[0060] Regarding the above embodiments, the following additional remarks are further disclosed.

[0061] (Additional Remark 1) In a neural network, identify a first neuron without a connection from the input layer and a second neuron without a connection to the output layer as deletion targets, add the bias of the first neuron to the bias of a third neuron connected to the first neuron on the output side, delete the identified neurons to be deleted from the neural network A model reduction program for causing a computer to execute a process including this.

[0062] (Appendix 2) The process of identifying the first neuron as a deletion target includes modifying the weights on the output side of the first neuron to 0, The process of identifying the second neuron as a deletion target includes modifying the weights on the input side of the second neuron to 0, The neurons to be deleted from the neural network are neurons with all weights on the input side and the output side being 0 The model reduction program according to Appendix 1.

[0063] (Appendix 3) The model reduction program according to Appendix 2, which executes a process of identifying the first neuron as a deletion target by forward search from the input layer to the output layer in the neural network, and executes a process of identifying the second neuron as a deletion target by backward search from the output layer to the input layer.

[0064] (Appendix 4) The process of identifying the first neuron and the second neuron as deletion targets includes modifying the corresponding elements in a parameter table that stores the weights between the connected neurons to 0 in the elements of a matrix where one of the connected neurons is assigned to rows and the other is assigned to columns, in the model reduction program according to Appendix 2 or Appendix 3.

[0065] (Appendix 5) The process of deleting neurons with all weights being 0 from the neural network includes deleting the rows and columns corresponding to the weights of the neurons to be deleted in the parameter table, which is the model reduction program described in Supplementary Note 4.

[0066] (Supplementary Note 6) The process of adding the biases includes adding the value obtained by multiplying the bias of the first neuron by the weight between the first neuron and the third neuron to the bias of the third neuron, which is the model reduction program described in any one of Supplementary Notes 1 to 5.

[0067] (Supplementary Note 7) The process of adding the biases includes adding the value obtained by multiplying the value obtained by applying the activation function of the first neuron to the bias of the first neuron by the weight between the first neuron and the third neuron to the bias of the third neuron, which is the model reduction program described in any one of Supplementary Notes 1 to 5.

[0068] (Supplementary Note 8) In a neural network, a specifying unit that specifies, as deletion targets, a first neuron having no connection from the input layer and a second neuron having no connection to the output layer, a compensating unit that adds the bias of the first neuron to the bias of a third neuron connected to the first neuron on the output side, a deleting unit that deletes the specified neurons to be deleted from the neural network, and a model reduction device including the above.

[0069] (Supplementary Note 9) The specifying unit corrects the weights on the output side of the first neuron to 0 and corrects the weights on the input side of the second neuron to 0. The deleting unit deletes neurons from the neural network for which all the weights on the input side and the output side are 0. The model reduction device according to Supplementary Note 8.

[0070] (Appendix 10) The specific unit is a model reduction device according to Appendix 9, which executes a process of specifying the first neuron as a deletion target by forward search from the input layer to the output layer in the neural network, and executes a process of specifying the second neuron as a deletion target by backward search from the output layer to the input layer.

[0071] (Appendix 11) The specific unit is a model reduction device according to Appendix 9 or Appendix 10, and executes a process including modifying a corresponding element in a parameter table storing weights between the connected neurons to 0, where one of the connected neurons is assigned to a row and the other is assigned to a column in a matrix formed by the process of specifying the first neuron and the second neuron as deletion targets.

[0072] (Appendix 12) The deletion unit is a model reduction device according to Appendix 11, and deletes rows and columns corresponding to the weights of neurons to be deleted in the parameter table.

[0073] (Appendix 13) The compensation unit is a model reduction device according to any one of Appendices 8 to 12, and adds a value obtained by multiplying the bias of the first neuron by the weight between the first neuron and the third neuron to the bias of the third neuron.

[0074] (Appendix 14) The compensation unit is a model reduction device according to any one of Appendices 8 to 12, and adds a value obtained by multiplying the value obtained by applying the activation function of the first neuron to the bias of the first neuron by the weight between the first neuron and the third neuron to the bias of the third neuron.

[0075] (Appendix 15) In a neural network, a first neuron without a connection from the input layer and a second neuron without a connection to the output layer are specified as deletion targets. Add the bias of the first neuron to the bias of the third neuron connected to the first neuron on the output side. Delete the identified neuron to be deleted from the neural network. A model reduction method in which a computer executes a process including this.

[0076] (Appendix 16) The process of identifying the first neuron as a deletion target includes modifying the weight on the output side of the first neuron to 0. The process of identifying the second neuron as a deletion target includes modifying the weight on the input side of the second neuron to 0. The neuron to be deleted from the neural network is a neuron in which all of the input-side weight and the output-side weight are 0. The model reduction method described in Appendix 15.

[0077] (Appendix 17) In the neural network, execute a process of identifying the first neuron as a deletion target by a forward search from the input layer to the output layer, and execute a process of identifying the second neuron as a deletion target by a backward search from the output layer to the input layer. The model reduction method described in Appendix 16.

[0078] (Appendix 18) The process of identifying the first neuron and the second neuron as deletion targets includes modifying the corresponding element in the parameter table storing the weights between the connected neurons to 0 in the elements of the matrix obtained by assigning one of the connected neurons to the row and the other to the column. The model reduction method described in Appendix 16 or Appendix 17.

[0079] (Appendix 19) The process of deleting the neuron with all weights being 0 from the neural network includes deleting the row and column corresponding to the weight of the neuron to be deleted in the parameter table. The model reduction method described in Appendix 18.

[0080] (Appendix 20) The process of adding up the biases includes adding the value obtained by multiplying the bias of the first neuron by the weight between the first neuron and the third neuron to the bias of the third neuron, which is the model reduction method according to any one of Appendices 15 to 19.

Explanation of Signs

[0081] 10 Model reduction device 12 Correction unit 14 Compensation unit 16 Deletion unit 40 Computer 41 CPU 42 Memory 43 Storage unit 44 Input / output device 45 R / W unit 46 Communication I / F 47 Bus 49 Storage medium 50 Model reduction program

Claims

1. In a neural network, identify a first neuron having no connection from the input layer and a second neuron having no connection to the output layer as deletion targets, add the bias of the first neuron to the bias of a third neuron connected to the first neuron on the output side, and delete the identified neurons to be deleted from the neural network A model reduction program for causing a computer to execute a process including this.

2. The process of identifying the first neuron as a deletion target includes correcting the weight on the output side of the first neuron to 0, The process of identifying the second neuron as a deletion target includes correcting the weight on the input side of the second neuron to 0, The neurons deleted from the neural network are neurons in which all of the input-side weights and the output-side weights are 0 The model reduction program according to claim 1.

3. In the neural network, execute a process of identifying the first neuron as a deletion target by forward search from the input layer to the output layer, and execute a process of identifying the second neuron as a deletion target by backward search from the output layer to the input layer. The model reduction program according to claim 2.

4. The process of identifying the first neuron and the second neuron as deletion targets includes correcting the corresponding element in a parameter table storing the weights between the connected neurons to 0 in an element of a matrix in which one of the connected neurons is assigned as a row and the other is assigned as a column. The model reduction program according to claim 2 or claim 3.

5. The process of deleting neurons with all weights of 0 from the neural network includes deleting the rows and columns corresponding to the weights of the neurons to be deleted in the parameter table. The model reduction program according to claim 4.

6. The process of adding the biases includes adding a value obtained by multiplying the bias of the first neuron by the weight between the first neuron and the third neuron to the bias of the third neuron. The model reduction program according to any one of claims 1 to 5.

7. The process of adding up the biases includes adding a value obtained by multiplying a value obtained by applying the activation function of the first neuron to the bias of the first neuron and the weight between the first neuron and the third neuron to the bias of the third neuron. The model reduction program according to any one of claims 1 to 5.

8. In a neural network, a specifying unit that specifies, as deletion targets, a first neuron that has no connection from the input layer and a second neuron that has no connection to the output layer, A compensation unit that adds up the bias of the first neuron to the bias of a third neuron that is connected to the first neuron on the output side, A deletion unit that deletes the specified neurons to be deleted from the neural network, A model reduction device including the above.

9. In a neural network, a first neuron that has no connection from the input layer and a second neuron that has no connection to the output layer are specified as deletion targets, The bias of the first neuron is added to the bias of a third neuron that is connected to the first neuron on the output side, Deleting the specified neurons to be deleted from the neural network A model reduction method in which a computer executes a process including the above.

Citation Information

Patent Citations

  • Method for rationalizing constitution of fuzzy inference model

    JP2000322263A

  • Artificial neural network class-based pruning

    JP2018129033A

  • Apparatus and method of compressing neural network

    US20200184333A1