Information processing device, information processing method, and computer-readable recording medium

By specifying the reduction layer according to the calculation quantity distribution of each layer in the neural network, the problem of ineffective calculation quantity reduction caused by the calculation quantity distribution in the prior art is solved, and effective calculation quantity reduction and recognition accuracy protection is achieved.

CN113383347BActive Publication Date: 2025-05-16MITSUBISHI ELECTRIC CORP
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN201980091148.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2019-02-15
Publication Date
2025-05-16
Estimated Expiration
2039-02-15

AI Technical Summary

Technical Problem

When reducing the later-stage computing volume of neural networks, the prior art does not consider the distribution of computing volume, resulting in the inability to effectively reduce the computing volume, affecting the recognition accuracy.

Method used

By specifying the layer that reduces the computational amount according to the computational amount distribution of each layer in the neural network, the effective reduction of the computational amount of the neural network is achieved.

Benefits of technology

The effective calculation amount reduction corresponding to the calculation amount distribution in the neural network is realized, which avoids the reduction of recognition accuracy and meets the real-time processing requirements of resource-limited devices.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113383347B_ABST
    Figure CN113383347B_ABST
Patent Text Reader

Abstract

A processing performance calculation unit (101) calculates the processing performance of an embedded device when a neural network having a plurality of layers is installed. A requirement achievement determination unit (102) determines whether the processing performance of the embedded device when the neural network is installed meets the required processing performance. When the requirement achievement determination unit (102) determines that the processing performance of the embedded device when the neural network is installed does not meet the required processing performance, a reduction layer designation unit (103) designates a layer whose amount of computation is to be reduced from the plurality of layers, i.e., a reduction layer, based on the amount of computation of each layer of the neural network.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to neural networks. Background Art

[0002] In a neural network (hereinafter referred to as a network), large-scale calculations are required. Therefore, when the neural network is directly installed in a device with limited resources such as an embedded device, the neural network cannot be operated in real time. In order to operate the neural network in real time in a device with limited resources, the neural network needs to be lightweight.

[0003] Patent Document 1 discloses a structure for increasing the inference processing speed of a neural network.

[0004] Patent document 1 discloses a structure for reducing the number of product-sum operations in inference processing by reducing the dimension of the weight matrix. More specifically, Patent document 1 discloses the following structure: in order to minimize the reduction in recognition accuracy caused by reducing the amount of calculation, the reduction amount is smaller in the front stage of the neural network, and the reduction amount is larger in the back stage.

[0005] Prior art literature

[0006] Patent Literature

[0007] Patent Document 1: Japanese Patent Application Publication No. 2018-109947 Summary of the invention

[0008] Problems to be solved by the invention

[0009] The technique of Patent Document 1 significantly reduces the amount of computation in the later stages of the neural network. Therefore, in a neural network where the amount of computation in the later stages is smaller than that in the previous stages, the amount of computation in the later stages may be reduced more than necessary.

[0010] The reduction in the amount of calculation affects the recognition accuracy. Therefore, if the amount of calculation in the subsequent stage is reduced more than necessary, the recognition rate may deteriorate and the required recognition accuracy may not be achieved.

[0011] As described above, the technology of Patent Document 1 does not take the distribution of the amount of computation in the neural network into consideration, and therefore has a problem that effective reduction of the amount of computation corresponding to the distribution of the amount of computation cannot be performed.

[0012] One of the main objects of the present invention is to solve the above-mentioned problems. More specifically, the main object of the present invention is to effectively reduce the amount of computation of a neural network based on the distribution of the amount of computation in the neural network.

[0013] Means for solving problems

[0014] The information processing device of the present invention comprises: a processing performance calculation unit, which calculates the processing performance of a device when a neural network having multiple layers is installed; a requirement achievement determination unit, which determines whether the processing performance of the device when the neural network is installed satisfies the required processing performance; and a reduction layer designation unit, which designates a layer for reducing the amount of calculation, i.e., a reduction layer, from among the multiple layers according to the amount of calculation of each layer of the neural network, when the requirement achievement determination unit determines that the processing performance of the device when the neural network is installed does not satisfy the required processing performance.

[0015] Effects of the Invention

[0016] According to the present invention, the reduction layer is specified according to the amount of calculation of each layer, so that it is possible to perform effective reduction of the amount of calculation corresponding to the distribution of the amount of calculation in the neural network. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] Figure 1 This is a diagram showing an example of a neural network and an embedded device according to the first embodiment.

[0018] Figure 2 This is a diagram showing an example of the amount of calculation and processing time of each layer in the first embodiment.

[0019] Figure 3 This is a diagram showing an example of reducing the amount of calculation in the prior art.

[0020] Figure 4 This is a diagram showing the bottleneck of the first embodiment.

[0021] Figure 5 This is a diagram showing an example of reduction in the amount of calculation in the first embodiment.

[0022] Figure 6 This is a flowchart showing an overview of the operation of the first embodiment.

[0023] Figure 7 This is a diagram showing an example of a functional configuration of an information processing device according to Embodiment 1.

[0024] Figure 8 This is a diagram showing a hardware configuration example of the information processing device according to the first embodiment.

[0025] Fig. 9 This is a flowchart showing an example of the operation of the information processing device according to the first embodiment.

[0026] Fig.10 This is a flowchart showing an example of the operation of the information processing device according to the first embodiment.

[0027] Fig.11 This is a diagram showing an example of reduction in the amount of calculation after alleviation in the first embodiment.

[0028] Fig.12 This is a diagram showing an example of additional reduction in the amount of calculation according to the first embodiment.

[0029] Fig.13 This is a diagram showing an example of reduction when there are multiple layers with the same amount of calculation in the first embodiment.

[0030] Fig.14 This is a diagram showing an example of reduction when the difference in the amount of calculation between the layer with the largest amount of calculation and the layer with the second largest amount of calculation in Embodiment 1 is smaller than a threshold value. DETAILED DESCRIPTION

[0031] Hereinafter, embodiments of the present invention will be described using the drawings. In the following description of the embodiments and the drawings, parts with the same reference numerals represent the same parts or corresponding parts.

[0032] Implementation Method 1

[0033] ***summary***

[0034] In this embodiment, the weight reduction of a neural network when the neural network is implemented in a device with limited resources such as an embedded device is described.

[0035] More specifically, in this embodiment, the layer with the largest amount of computation is extracted from the multiple layers of the neural network. Then, the amount of computation of the extracted layer is reduced to meet the required processing performance. In addition, re-learning is performed after the amount of computation is reduced, thereby suppressing the reduction of the recognition rate.

[0036] By repeatedly executing the above steps, according to the present embodiment, a neural network with a small amount of computation that can be installed on a device with limited resources can be obtained.

[0037] ***step***

[0038] Next, the procedure for reducing the weight of the neural network according to the present embodiment will be described with reference to the drawings.

[0039] In the following description and drawings, parts denoted by the same reference numerals represent the same parts or corresponding parts.

[0040] In this embodiment, an example of installing a neural network in an embedded device such as a CPU (Central Processing Unit) is described. In addition, the embedded device executes the processing of the neural network layer by layer. Furthermore, the time taken for the processing of the neural network can be calculated using the following formula.

[0041] Σ(1 layer processing time)

[0042] In addition, the processing time for one layer can be calculated using the following formula.

[0043] Total product and number of operations per layer (OP) / processing capacity of the device (OP / second)

[0044] In addition, the “total product sum number of operations per layer (OP)” can be calculated according to the specifications (parameters) of the network.

[0045] The "processing capacity of the device (OP / sec)" is uniquely determined for each embedded device.

[0046] As described above, it is possible to calculate the processing performance when a neural network is implemented in an embedded device.

[0047] In the following, processing performance refers to “Σ(processing time for 1 layer)”, that is, the time required for the embedded device to process all layers of the neural network (total processing time).

[0048] In the case of "Σ(1-layer processing time) < required processing performance", the required processing performance can be achieved even if the existing neural network is installed in an embedded device.

[0049] On the other hand, in the case of "Σ(1-layer processing time)>required processing performance", when the existing neural network is installed in an embedded device, the required processing performance cannot be achieved.

[0050] When "Σ(1-layer processing time)>required processing performance", it is necessary to change the neural network to reduce the total number of product sum operations.

[0051] Here, assuming Figure 1 A neural network 10 and an embedded device 20 are shown.

[0052] The neural network 10 has an L0 layer, an L1 layer, and an L2 layer. Moreover, the embedded device 20 processes each layer in the order of the L0 layer, the L1 layer, and the L2 layer. In addition, the embedded device 20 has a processing capacity of 10 GOP (Giga Operations) / second.

[0053] Furthermore, it is assumed that the required processing performance of the embedded device 20 is 1 second.

[0054] like Figure 2 As shown, the amount of calculation (the total number of product-sum calculations) of the L0 layer is 100 GOPs, the amount of calculation (the total number of product-sum calculations) of the L1 layer is 0.1 GOP, and the amount of calculation (the total number of product-sum calculations) of the L2 layer is 0.01 GOP.

[0055] If the neural network 10 is directly installed on the embedded device 20, then Figure 2As shown, the processing of L0 layer takes 10 seconds. The processing of L1 layer takes 0.01 seconds. The processing of L2 layer takes 0.001 seconds.

[0056] The total processing time of the L0 layer, L1 layer, and L2 layer is 10.011 seconds, which does not meet the required performance. Therefore, the amount of calculation (the number of product sum calculations) of the neural network 10 needs to be reduced.

[0057] In the technology of Patent Document 1, the amount of calculation is reduced in such a way that “the earlier the neural network stage, the smaller the reduction, and the later the stage, the larger the reduction.” For example, if the total product and the number of calculations are reduced as follows, the required processing performance can be met.

[0058] Reduction in the number of sum-of-product operations at L0 layer: 91%

[0059] Reduction in the number of sum-of-product operations in the L1 layer: 92%

[0060] Reduction in the number of sum-of-product operations in the L2 layer: 93%

[0061] If the above reductions are achieved, Figure 3 As shown, the total number of product-sum calculations for the L0 layer is 9 GOPs, the total number of product-sum calculations for the L1 layer is 0.008 GOPs, and the total number of product-sum calculations for the L2 layer is 0.0007 GOPs. As a result, the total processing time is 0.90087 seconds, which can meet the required processing performance.

[0062] However, since the L2 layer, which originally has a small number of product-sum calculations, is largely reduced, the recognition rate may be reduced.

[0063] like Figure 4 As shown, in this example, the L0 layer becomes a bottleneck and cannot meet the required processing performance.

[0064] Therefore, in this embodiment, if Figure 5 As shown, the amount of computation in the L0 layer, which has the largest number of sum-of-products operations, is reduced.

[0065] Hereinafter, a layer that is a target for reducing the amount of computation will be referred to as a reduction layer.

[0066] In this embodiment, the total product of the reduction layer and the value of the number of operations are calculated so as to satisfy the required processing performance (1 second in this example).

[0067] exist Figure 5 In the example of , the processing time of the L0 layer needs to be reduced to 0.989 seconds. Therefore, the total number of product sum calculations of the L0 layer needs to be reduced to 9.89 GOPs.

[0068] Determine the reduction level and reduction amount as described above (in Figure 5In the example, 90.11 GOP Figure 6 The neural network 10 is changed as shown in step S1 to reduce the total product and number of operations of the reduction layer by the reduction amount.

[0069] The total number of product sum calculations can be reduced by any method, for example, by pruning.

[0070] In addition, the reduction of the amount of calculation also affects the recognition accuracy. Therefore, in this embodiment, Figure 6 As shown in step S2, after the neural network 10 is changed (the amount of calculation is reduced), re-learning is performed.

[0071] If it is determined as a result of re-learning that the desired recognition rate can be achieved, even the modified neural network 10 can satisfy the required processing performance and required recognition accuracy on the embedded device 20 .

[0072] ***Description of the structure***

[0073] Next, the configuration of the information processing device 100 of this embodiment will be described. Note that the operation performed by the information processing device 100 corresponds to an information processing method and an information processing program.

[0074] Figure 7 FIG. 1 shows an example of a functional configuration of the information processing device 100. Figure 8 A hardware configuration example of the information processing device 100 is shown.

[0075] First, refer to Figure 8 A hardware configuration example of the information processing device 100 will be described.

[0076] ***Description of the structure***

[0077] The information processing device 100 of this embodiment is a computer.

[0078] The information processing device 100 includes, as hardware, a CPU 901 , a storage device 902 , a GPU (Graphics Processing Unit) 903 , a communication device 904 , and a bus 905 .

[0079] The CPU 901 , the storage device 902 , the GPU 903 , and the communication device 904 are connected to a bus 905 .

[0080] The CPU 901 and the GPU 903 are ICs (Integrated Circuits) that perform processing.

[0081] The CPU 901 executes a program for realizing the functions of a processing performance calculation unit 101 , an achievement requirement determination unit 102 , a reduction layer designation unit 103 , a network conversion unit 104 , and a recognition rate determination unit 106 , which will be described later.

[0082] The GPU 903 executes a program that realizes the function of the learning unit 105 described later.

[0083] The storage device 902 is a HDD (Hard Disk Drive), a RAM (Random Access Memory), a ROM (Read Only Memory), or the like.

[0084] The storage device 902 stores programs that realize the functions of the processing performance calculation unit 101, the requirement achievement determination unit 102, the reduction layer specification unit 103, the network conversion unit 104, the learning unit 105, and the recognition rate determination unit 106. As described above, the programs that realize the functions of the processing performance calculation unit 101, the requirement achievement determination unit 102, the reduction layer specification unit 103, the network conversion unit 104, and the recognition rate determination unit 106 are read into the CPU 901 and executed by the CPU 901. The program that realizes the function of the learning unit 105 is read into the GPU 903 and executed by the GPU 903.

[0085] exist Figure 8 , a state in which the CPU 901 executes a program that realizes the functions of the processing performance calculation unit 101, the requirement achievement determination unit 102, the reduction layer designation unit 103, the network conversion unit 104, and the recognition rate determination unit 106 is schematically shown. Figure 8 , a state in which the GPU 903 executes a program for realizing the functions of the learning unit 105 is schematically shown.

[0086] The communication device 904 is an electronic circuit that performs communication processing of data.

[0087] The communication device 904 is, for example, a communication chip or a NIC (Network Interface Card).

[0088] Next, refer to Figure 7 A functional configuration example of the information processing device 100 will be described.

[0089] The processing performance calculation unit 101 calculates the processing performance of the embedded device 20 when the neural network 10 is implemented in the embedded device 20 , using the network configuration information 111 and the processing capability information 112 .

[0090] The network configuration information 111 shows Figure 2The total product and the number of operations of each layer of the neural network 10 are shown in FIG. In the network structure information 111, instead of the total product and the number of operations of each layer, the specification of the neural network 10 that can calculate the total product and the number of operations of each layer may be described.

[0091] The processing capability information 112 shows Figure 2 The processing capability (10 GOP / sec) of the embedded device 20 is exemplified in FIG. 1 . In the processing capability information 112 , instead of the processing capability of the embedded device 20 , the specification of the embedded device 20 that can calculate the processing capability of the embedded device 20 may be described.

[0092] In addition, the processing performed by the processing performance calculation unit 101 corresponds to the processing performance calculation processing.

[0093] The requirement achievement determination unit 102 determines whether the processing performance of the embedded device 20 calculated by the processing performance calculation unit 101 satisfies the required processing performance described in the required processing performance information 113 .

[0094] The process performed by the achievement request determination unit 102 corresponds to the achievement request determination process.

[0095] The reduction layer designation unit 103 designates a reduction layer and a reduction amount of the calculation amount of the reduction layer.

[0096] That is, when the requirement-reaching determination unit 102 determines that the processing performance of the embedded device 20 when the neural network 10 is installed does not meet the required processing performance, the reduction layer designation unit 103 designates a layer whose amount of computation is to be reduced, that is, a reduction layer, from among a plurality of layers, based on the amount of computation of each layer of the neural network 10. More specifically, the reduction layer designation unit 103 designates the layer with the largest amount of computation as the reduction layer. In addition, the reduction layer designation unit 103 determines the amount of computation reduction of the reduction layer in such a manner that the processing performance of the embedded device 20 when the neural network 10 with the reduced amount of computation is installed meets the required processing performance.

[0097] The process performed by the reduction layer specifying unit 103 corresponds to the reduction layer specifying process.

[0098] The network conversion unit 104 converts the neural network 10 so as to reduce the amount of calculation of the reduction layer specified by the reduction layer specifying unit 103 by the reduction amount determined by the reduction layer specifying unit 103 .

[0099] The learning unit 105 uses the learning data set 114 to learn the neural network 10 converted by the network conversion unit 104 .

[0100] The recognition rate determination unit 106 analyzes the learning result of the learning unit 105 , and determines whether the recognition rate of the converted neural network 10 satisfies the required recognition rate described in the required recognition rate information 115 .

[0101] When the recognition rate of the converted neural network 10 satisfies the required recognition rate and the processing performance of the embedded device 20 when the converted neural network 10 is installed satisfies the required processing performance, the requirement-reaching determination unit 102 outputs the lightweight network structure information 116 .

[0102] The lightweight network structure information 116 shows the total product and the number of operations of each layer of the neural network 10 after conversion.

[0103] ***Description of the action***

[0104] Next, refer to Fig. 9 and Fig.10 An operation example of the information processing device 100 according to this embodiment will be described.

[0105] First, the processing performance calculation unit 101 obtains the network structure information 111 and the processing capacity information 112, and uses the obtained network structure information 111 and the processing capacity information 112 to calculate the processing performance of the embedded device 20 when the neural network 10 is installed in the embedded device 20 (step S101).

[0106] The processing performance calculation unit 101 calculates the processing time of each layer based on “the total product and number of operations per layer (OP) / the processing capacity of the device (OP / second)”, and sums up the calculated processing time of each layer to obtain the processing performance of the embedded device 20.

[0107] Next, the requirement achievement determination unit 102 determines whether the processing performance of the embedded device 20 calculated by the processing performance calculation unit 101 satisfies the required processing performance described in the required processing performance information 113 (step S102 ).

[0108] When the processing performance of the embedded device 20 satisfies the required processing performance (step S103 : Yes), the process ends.

[0109] When the processing performance of the embedded device 20 does not satisfy the required processing performance (step S103: No), the reduction layer specifying unit 103 performs bottleneck analysis (step S104), and specifies the reduction layer and the reduction amount of the calculation amount of the reduction layer (step S105).

[0110] Specifically, the reduction layer specifying unit 103 obtains the Figure 4 The example shown in describes the total product, the number of calculations, and the processing time of each layer, and the layer with the largest total product and the largest number of calculations is designated as the reduction layer.

[0111] Furthermore, the reduction tier designation unit 103 outputs information notifying the reduction tier and the reduction amount to the network conversion unit 104 .

[0112] Next, the network conversion unit 104 converts the neural network 10 so that the total number of product sum calculations of the reduction layer specified by the reduction layer specifying unit 103 is reduced by the reduction amount determined by the reduction layer specifying unit 103 (step S106).

[0113] The network conversion unit 104 converts the neural network with reference to the network structure information 111 .

[0114] Furthermore, the network conversion unit 104 notifies the learning unit 105 of the converted neural network 10 .

[0115] Next, the learning unit 105 uses the learning data set 114 to learn the neural network 10 converted by the network conversion unit 104 (step S107 ).

[0116] The learning unit 105 outputs the learning result to the recognition rate determination unit 106 .

[0117] Next, the recognition rate determination unit 106 analyzes the learning result of the learning unit 105 and determines whether the recognition rate of the converted neural network 10 satisfies the required recognition rate described in the required recognition rate information 115 (step S108 ).

[0118] When the recognition rate of the neural network 10 after conversion does not satisfy the required recognition rate, the recognition rate determination unit 106 notifies the reduction layer specifying unit 103 that the recognition rate does not satisfy the required recognition rate.

[0119] On the other hand, when the recognition rate of the neural network 10 after conversion satisfies the required recognition rate, the recognition rate determination unit 106 notifies the processing performance calculation unit 101 that the recognition rate satisfies the required recognition rate.

[0120] If the recognition rate of the converted neural network 10 does not satisfy the required recognition rate (step S108: No), the reduction layer designation unit 103 redesignates the reduction amount (step S109). In redesignating the reduction amount, the reduction layer designation unit 103 relaxes the reduction amount.

[0121] That is, when the recognition rate of the neural network 10 after the reduced amount of calculation is implemented in the embedded device 20 does not satisfy the required recognition rate, the reduction level specifying unit 103 determines the reduced amount of calculation.

[0122] For example, the reduction layer designation unit 103 performs Fig.11 The reductions shown are moderated.

[0123] exist Fig.11In this case, the reduction layer designation unit 103 increases the number of product sum calculations of the L0 layer from 9.89 GOP to 9.895 GOP, thereby alleviating the reduction amount. In this case, the processing performance becomes 1.0005 seconds, which is slightly lower than the required processing performance.

[0124] When the recognition rate of the converted neural network 10 satisfies the required recognition rate (step S108 : Yes), the processing performance calculation unit 101 calculates the processing performance of the embedded device 20 for the converted neural network 10 (step S110 ).

[0125] That is, the processing performance calculation unit 101 calculates the processing performance of the embedded device 20 using the network structure information 111 and the processing capacity information 112 related to the converted neural network 10 .

[0126] Next, the requirement achievement determination unit 102 determines whether the processing performance of the embedded device 20 calculated by the processing performance calculation unit 101 satisfies the required processing performance described in the required processing performance information 113 (step S111 ).

[0127] When the processing performance of the embedded device 20 satisfies the required processing performance (step S112: Yes), the process ends. At this time, the requirement achievement determination unit 102 outputs the lightweight network configuration information 116 to a predetermined output destination.

[0128] When the processing performance of the embedded device 20 does not satisfy the required processing performance (step S112: No), the reduction layer designation unit 103 performs bottleneck analysis (step S113), and again designates the reduction layer and the reduction amount of the calculation amount of the reduction layer (step S114).

[0129] In step S114 , the reduction layer designation unit 103 designates a layer that has not been designated as a reduction layer as an additional reduction layer.

[0130] For example, the reduction layer designation unit 103 designates, as an additional reduction layer, a layer having the largest number of product sum calculations among the layers that have not yet been designated as reduction layers.

[0131] exist Fig.12 In the example, L0 layer has been designated as a reduction layer, and the number of product sum operations of L1 layer is greater than that of L2. Therefore, the reduction layer designation unit 103 designates L1 layer as an additional reduction layer. Fig.12 In the example of , the reduction layer designation unit 103 determines to reduce the number of product sum calculations of the L1 layer to 0.04 GOP (reduction amount: 0.06 GOP). As a result, the processing performance becomes 1 second, which satisfies the required processing performance.

[0132] Furthermore, when all layers have been designated as reduction layers, the reduction layer designation unit 103 designates the layer with the largest amount of calculation after reduction as an additional reduction layer.

[0133] Steps S115 to S118 are the same as steps S106 to S109 , and thus description thereof will be omitted.

[0134] In the above, an example is used in which the number of product sum calculations in the L0 layer is greater than that in the L1 layer and the L2 layer.

[0135] However, depending on the neural network, there may be multiple layers with the same total product and number of operations. In this case, the reduction layer designation unit 103 preferentially designates the later layer as the reduction layer. That is, when there are two or more layers with the largest total product and number of operations, the reduction layer designation unit 103 designates the last layer of the two or more layers with the largest total product and number of operations as the reduction layer. This is because the later the layer, the less likely it is that the recognition rate will be reduced due to the reduction in the amount of operations.

[0136] For example, Fig.13 As shown, when the total product sum calculation number of the L0 layer and the total product sum calculation number of the L1 layer are both 100 GOPs, the reduction layer designation unit 103 designates the subsequent layer, that is, the L1 layer, as the reduction layer.

[0137] In addition, when the difference between the amount of calculation of the layer with the largest amount of calculation and the amount of calculation of the layer with the second largest amount of calculation is less than a threshold, and the layer with the second largest amount of calculation is located at a later stage than the layer with the largest amount of calculation, the reduction layer designation unit 103 may also designate the layer with the second largest amount of calculation as the reduction layer.

[0138] For example, suppose the threshold is 10% of the computational effort of the layer with the largest computational effort. Fig.14 As shown, when the total product sum calculation number of the L0 layer is 100 GOP and the total product sum calculation number of the L1 layer is 95 GOP, the difference in the total product sum calculation number between the L0 layer and the L1 layer is less than 10% of the total product sum calculation number of the L0 layer. Therefore, the reduction layer designation unit 103 designates the subsequent layer, i.e., the L1 layer, as the reduction layer.

[0139] The threshold value is not limited to 10%, and the user of the information processing device 100 can arbitrarily set the threshold value.

[0140] ***Description of the Effects of the Implementation Method***

[0141] As described above, according to the present embodiment, the reduction layer is specified according to the amount of calculation of each layer, so that effective reduction of the amount of calculation can be performed according to the distribution of the amount of calculation in the neural network.

[0142] Furthermore, according to the present embodiment, even if the designer of the neural network does not have knowledge about the embedded device as the installation destination, a neural network that satisfies the required processing performance of the embedded device can be automatically obtained.

[0143] Similarly, according to the present embodiment, even if the person in charge of installing the embedded device has no knowledge related to neural networks, a neural network that satisfies the required processing performance of the embedded device can be automatically obtained.

[0144] ***Description of the hardware structure***

[0145] Finally, a supplementary description of the hardware structure of the information processing device 100 will be given.

[0146] The storage device 902 stores an OS (Operating System).

[0147] Furthermore, at least a part of the OS is executed by the CPU 901 .

[0148] The CPU 901 executes a program that realizes the functions of the processing performance calculation unit 101 , the achievement requirement determination unit 102 , the reduction layer designation unit 103 , the network conversion unit 104 , and the recognition rate determination unit 106 while executing at least a part of the OS.

[0149] The CPU 901 executes the OS, thereby performing task management, memory management, file management, communication control, and the like.

[0150] In addition, at least any one of the information, data, signal values ​​and variable values ​​representing the processing results of the processing performance calculation unit 101, the requirement achievement determination unit 102, the reduction layer specification unit 103, the network conversion unit 104, the learning unit 105 and the recognition rate determination unit 106 is stored in at least any one of the storage device 902, the register and the cache memory.

[0151] In addition, the program that realizes the functions of the processing performance calculation unit 101, the requirement achievement determination unit 102, the reduction layer specification unit 103, the network conversion unit 104, the learning unit 105, and the recognition rate determination unit 106 may be stored in a portable recording medium such as a magnetic disk, a floppy disk, an optical disk, a high-density disk, a Blu-ray (registered trademark) disk, or a DVD. Furthermore, a portable recording medium storing the program that realizes the functions of the processing performance calculation unit 101, the requirement achievement determination unit 102, the reduction layer specification unit 103, the network conversion unit 104, the learning unit 105, and the recognition rate determination unit 106 may be commercially distributed.

[0152] In addition, the "unit" of the processing performance calculation unit 101, the requirement achievement determination unit 102, the reduction layer specification unit 103, the network conversion unit 104, the learning unit 105 and the recognition rate determination unit 106 can also be rewritten as "circuit" or "process" or "step" or "processing".

[0153] In addition, the information processing device 100 may also be implemented by a processing circuit, such as a logic IC (Integrated Circuit), a GA (Gate Array), an ASIC (Application Specific Integrated Circuit), or an FPGA (Field-Programmable Gate Array).

[0154] In addition, in this specification, a processor and a processing circuit are collectively referred to as a "processing circuit".

[0155] That is, a processor and a processing circuit are each specific examples of a “processing circuit”.

[0156] Description of symbols

[0157] 10: neural network; 20: embedded device; 100: information processing device; 101: processing performance calculation unit; 102: requirement achievement determination unit; 103: reduction layer designation unit; 104: network conversion unit; 105: learning unit; 106: recognition rate determination unit; 111: network structure information; 112: processing capacity information; 113: required processing performance information; 114: learning data set; 115: required recognition rate information; 116: lightweight network structure information; 901: CPU; 902: storage device; 903: GPU; 904: communication device; 905: bus.

Claims

1. An information processing device, comprising: a processing performance calculation unit that calculates the processing performance of a device when a neural network having a plurality of layers is installed; a requirement-reaching determination unit for determining whether the processing performance of the device when the neural network is installed meets the required processing performance; a reduction layer specifying unit that specifies a layer whose amount of computation is to be reduced, i.e., a reduction layer, from among the plurality of layers, when the requirement achievement determining unit determines that the processing performance of the device when the neural network is installed does not satisfy the required processing performance, based on the amount of computation of each layer of the neural network, and determines an amount of reduction in the amount of computation of the reduction layer in such a manner that the processing performance of the device when the neural network having the reduced amount of computation is installed satisfies the required processing performance; and A network conversion unit converts the neural network in such a way as to reduce the amount of computation of the reduction layer by the reduction amount.

2. The information processing device according to claim 1, wherein: The reduction layer designation unit designates a layer having the largest amount of calculation as the reduction layer.

3. The information processing device according to claim 2, wherein: When there are two or more layers with the largest amount of computation, the reduction layer designation unit designates the last layer of the two or more layers with the largest amount of computation as the reduction layer.

4. The information processing device according to claim 1, wherein: When the difference between the amount of computation of the layer with the largest amount of computation and the amount of computation of the layer with the second largest amount of computation is smaller than a threshold value and the layer with the second largest amount of computation is located at a later stage than the layer with the largest amount of computation, the reduction layer designation unit designates the layer with the second largest amount of computation as the reduction layer.

5. The information processing device according to claim 1, wherein: The reduction layer specifying unit specifies an additional reduction layer from among the plurality of layers when the processing performance of the device does not satisfy the required processing performance when the neural network having the reduced amount of calculation is installed in the device.

6. The information processing device according to claim 5, wherein: The reduction layer designation unit designates, as the additional reduction layer, a layer having the largest amount of computation among the layers that have not yet been designated as the reduction layers.

7. The information processing device according to claim 5, wherein: When all of the plurality of layers have been designated as the reduction layers, the reduction layer designation unit designates a layer having the largest amount of computation after reduction as the additional reduction layer.

8. The information processing device according to claim 1, wherein: The reduction level specifying unit determines a reduced reduction amount when a recognition rate of the neural network after the reduced computation amount is installed in the device does not satisfy a required recognition rate.

9. An information processing method, wherein: Computer calculations of the processing performance of a device equipped with a neural network having multiple layers, The computer determines whether the processing performance of the device when the neural network is installed meets the required processing performance, When it is determined that the processing performance of the device when the neural network is installed does not meet the required processing performance, the computer specifies a layer in which the amount of computation is reduced, i.e., a reduction layer, from among the plurality of layers based on the amount of computation of each layer of the neural network, and determines the amount of reduction in the amount of computation of the reduction layer in such a manner that the processing performance of the device when the neural network with the reduced amount of computation is installed meets the required processing performance. The computer converts the neural network in such a manner as to reduce the amount of computation of the reduction layer by the reduction amount.

10. A computer-readable recording medium having an information processing program recorded thereon, the information processing program causing a computer to execute the following processing: a processing performance calculation process for calculating the processing performance of a device when a neural network having a plurality of layers is installed; A requirement determination process is performed to determine whether the processing performance of the device when the neural network is installed meets the required processing performance; a reduction layer specifying process of specifying a layer whose amount of computation is to be reduced from among the plurality of layers, i.e., a reduction layer, in accordance with the amount of computation of each layer of the neural network, when it is determined by the requirement attainment determination process that the processing performance of the device when the neural network is installed does not satisfy the required processing performance, and determining the amount of computation to be reduced in the reduction layer in such a manner that the processing performance of the device when the neural network having the reduced amount of computation is installed satisfies the required processing performance; and The network conversion process converts the neural network in such a way as to reduce the amount of computation of the reduction layer by the reduction amount.

Citation Information

Patent Citations

  • Device and method for increasing processing speed of neural network, and application of the same

    JP2018109947A

  • Neural network method and apparatus

    CN107665364A

  • Device and method for improving processing speed of neural network and application thereof

    US20180189650A1

  • Information processing method and information processing apparatus

    US20180365557A1

  • Methods and algorithms of reducing computation for deep neural networks via pruning

    US20190050735A1