Calculation device, recognition device and control device
By deleting filter information based on sensitivity and optimizing calculation direction, the arithmetic unit's utilization is improved, addressing the inefficiencies in in-vehicle computing devices and enhancing CNN operation performance.
Patent Information
- Application Number
- JP2021085284
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2021-05-20
- Publication Date
- 2025-05-14
- Estimated Expiration
- 2041-05-20
AI Technical Summary
Existing in-vehicle computing devices with limited arithmetic performance struggle to fully utilize processing speed due to inconsistent operation units and the number of filters and channels used in convolution operations, leading to constraints in dividing weight parameters and utilizing the arithmetic unit efficiently.
The proposed solution involves deleting filter information based on sensitivity in the convolution calculation and determining the calculation direction, allowing for improved utilization of the arithmetic unit by optimizing the number of operations and resource allocation.
This approach enhances the utilization rate of the arithmetic unit, enabling more efficient execution of CNN operations within limited operational performance, thereby improving processing speed and resource utilization.
Smart Images

Figure 0007675480000003 
Figure 0007675480000004 
Figure 0007675480000005
Abstract
Description
[Technical field]
[0001] The present invention relates to a computing device that performs a calculation based on input data and a computing method thereof. The present invention also relates to a recognition device that uses the computing device to recognize input data and a control device that performs control according to the input data. In particular, the present invention relates to a technology for collecting external information using sensors such as a camera or a LIDAR (Light Detection and Ranging) and detecting the type of an object and its coordinates from the information. [Background technology]
[0002] In recent years, traffic accidents have become a social issue, and there is an increasing demand for safety during vehicle travel. To meet these demands, various autonomous driving and driving assistance technologies have been proposed. Among these technologies, object recognition and behavior prediction methods using CNN (Convolutional Neural Network), a type of DNN (Deep Neural Network), are known to have high recognition performance, and their application to autonomous driving is progressing.
[0003] CNN is a neural network that takes image data, which is information about the outside world, as input and is composed of multiple convolution layers connected in cascade. Here, a convolution layer is a series of operations that consists of product-sum operations and activation function operations, multiplying the pixels in the input data by the corresponding weight parameters, accumulating the results a certain number of times to create output data, and then performing activation function operations to output the results.
[0004] By performing a convolutional layer operation on image data, the type of a specific object in the input image data and the coordinates of the object are output. First, a specific configuration will be described. In the recognition device described in this patent, the first layer performs a product-sum operation on the input image data and the weight parameters of the convolution operation of the first layer to output the convolution operation result. The jth convolutional layer of a plurality of neural networks is called the jth layer, and the jth layer (an integer satisfying 1≦j≦L) outputs the operation result of the jth convolutional layer from the output data of the j-1th layer and the weight parameters of the convolution operation of the jth layer. If the final layer is the Lth layer, the output data of the L-1th layer before the Lth layer and the weight parameters of the convolution operation of the Lth layer are input, and the type of object and the coordinates of the object are output. Each convolutional layer performs a convolution operation using the input data and weight parameters, and then performs an activation function operation and outputs the result.
[0005] Here, when implementing a computationally intensive process such as CNN, which mainly uses product-sum operations, in an in-vehicle ECU (Electronic Control Unit), it is necessary to satisfy a throughput equivalent to the processing speed of the external information acquisition device and a latency equal to or less than the vehicle control cycle. Conventionally, when implementing CNN, it is common to use a graphic processing unit (GPU) or a field programmable logic array (FPGA) designed for edge devices. However, in-vehicle devices have only limited computing performance due to restrictions on the conditions of use, implementation costs, and device size that can be installed, and there is a problem that the required computing performance cannot be met. Therefore, in order to satisfy the required computing performance, it is common to achieve the target processing speed by improving the device itself or by using a computing unit specialized for CNN calculations.
[0006] Next, we will discuss the conventional technology when CNN is implemented in hardware as a computing device. External input is acquired using an external information acquisition device such as a camera or LIDAR. The acquired information is stored in memory. The computing device is composed of a memory, a model storage unit, a parameter storage unit, and multiple convolution computing units, and outputs recognition results such as the type of object and its coordinates.
[0007] Here, the external world information stored in the memory is sent to the convolution operation unit. The model storage unit also stores data that has been previously trained, and saves the trained data in the parameter storage unit. The parameter storage unit selects weight parameters for each layer, the number of channels for each layer, and the number of filters for each layer from the received trained data, and sends them to the convolution operation units from the 1st layer to the Lth layer. The convolution operation unit in the 1st layer receives input data from the memory, the weight parameters for the 1st layer, the number of channels for the 1st layer, and the number of filters for the 1st layer as inputs, and outputs the operation result to the 2nd layer. The convolution operation units are also connected in series. The convolution operation unit in the jth layer, which is the jth layer, receives the output of the convolution operation unit in the j-1th layer, the weight parameters for the jth layer, the number of channels for the jth layer, and the number of filters for the jth layer as inputs, and outputs the operation result to the j+1th layer.
[0008] In addition, the convolution calculation unit performs a convolution calculation based on the input data sent from the memory and the weight parameters, number of channels, and number of filters sent from the parameter storage unit, and then performs an activation function calculation and outputs the calculation result to the next layer.
[0009] However, in the case of in-vehicle computing devices, where the computing performance of the calculator is limited, there is an issue that the computing device cannot be fully utilized due to a mismatch between the calculation unit of the computing device and the number of filters and channels used for the convolution calculation, resulting in a decrease in processing speed.
[0010] Therefore, the method of Patent Document 1 has been proposed as a method and device that can efficiently execute CNN calculations within limited calculation performance in CNN-related calculations. The configuration of Patent Document 1 is shown below. First, the calculation device described in Patent Document 1 is composed of a divider, a calculator, and a generator. Next, the connection relationship of Patent Document 1 is shown.
[0011] The output from the divider is input to the calculator. The calculator outputs to the generator, and the generator receives input from the calculator.
[0012] Next, the operation of the arithmetic device proposed in Patent Document 1 is shown. First, the divider divides the weight parameters of a selected layer of the convolutional neural network in at least one of the depth dimension and the dimension of the number of kernels. Then, an arithmetic parameter array including a plurality of arithmetic parameters is obtained. Next, the arithmetic unit executes the arithmetic of the selected layer using each arithmetic parameter in the arithmetic parameter array, thereby obtaining a partial arithmetic result array. Finally, the generator generates the output of the layer based on the partial arithmetic result array. As a result, the divider divides the weight parameters according to the arithmetic unit of the arithmetic unit, and then transmits the weight parameters to the arithmetic unit. This improves the operating efficiency or utilization rate of the arithmetic unit, and avoids the hardware being limited by the size of the parameters. [Prior art documents] [Patent documents]
[0013] [Patent Document 1] JP 2019-82996 A Summary of the Invention [Problem to be solved by the invention]
[0014] In Patent Document 1, weight parameters input to a computing unit are divided in advance in at least one direction, either in the depth direction of the CNN calculation or in the dimension of the number of kernels, and the divided data is transferred from the divider to the computing unit to perform a convolution calculation. Then, for a computing unit of a limited size based on the result of this calculation, a convolution calculation is performed using weight parameters divided into a size that can be calculated by the computing unit.
[0015] However, in Patent Document 1, a division operation is required in advance to obtain a partial array of weight parameters using a divider outside the calculator, and a generator is required to recombine the partial array operations of the divided weight parameters. Furthermore, in the convolution operation, it is necessary to maintain the consistency of the operation before and after the division. For this reason, there are restrictions on the division method of the weight parameters and the size of the weight parameters after division, which poses a problem that the calculator cannot be fully utilized. [Means for solving the problem]
[0016] In order to solve the above problem, in the present invention, filter information is deleted based on the sensitivity and calculation direction (execution order) in the convolution calculation.
[0017] More specifically, in a computing device that performs CNN computation based on input data, the computing device includes a model storage unit that stores a model used in the CNN computation, a CNN computation unit that performs the CNN computation by executing a convolution computation for each of a plurality of convolution layers using the model, a weight parameter acquisition unit that acquires weight parameters of a convolution filter used in the convolution computation from the model storage unit, a channel information acquisition unit that acquires channel information for each of the plurality of convolution layers from the model stored in the model storage unit, a filter information acquisition unit that acquires filter information for each of the plurality of convolution layers from the model storage unit, a computation allocation unit that assigns a combination of weight information and computation required for the CNN computation and transmits it to the CNN computation unit, and a deletion index determination unit that deletes a part of filter information used in the convolution computation based on the maximum number of computations possible for the CNN computation unit, the execution order of the CNN computation, the channel information for each convolution layer of the CNN computation unit acquired from the channel information acquisition unit, the weight information acquired from the weight parameter acquisition unit, and the filter information for each of the plurality of convolution layers acquired from the filter information acquisition unit. Effect of the Invention
[0018] According to a representative embodiment of the present invention, it is possible to improve the utilization rate of the arithmetic unit. Note that problems, configurations and effects other than those described above will become apparent from the following description of the embodiment. [Brief description of the drawings]
[0019] [Figure 1] FIG. 2 is a functional block diagram of a recognition device which is a computing device in the first embodiment. [Diagram 2] FIG. 11 is an explanatory diagram for explaining deletion channel index information in the first embodiment. [Diagram 3] FIG. 11 is an explanatory diagram for explaining deletion filter index information in the first embodiment. [Figure 4] FIG. 2 is a diagram illustrating an internal configuration of a CNN calculation unit in the first embodiment. [Diagram 5]13 is a diagram illustrating an internal configuration of a deletion index determination unit in the first embodiment. [Figure 6] FIG. 2 is a diagram illustrating an internal configuration of a convolution calculation unit in the first embodiment. [Figure 7] 10 is a flowchart showing a process flow of a deletion filter determination unit in the first embodiment. [Figure 8] 11 is a flowchart showing a process flow of a deletion channel determination unit in the first embodiment. [Figure 9] 11 is a flowchart showing a process flow of a computation allocation unit in the first embodiment. [Figure 10] FIG. 11 is a diagram for explaining an example of a stride and a removable filter in the first embodiment. [Figure 11] FIG. 11 is an explanatory diagram illustrating an example of a deletion filter in the first embodiment; [Figure 12] FIG. 11 is a functional block diagram of a recognition device which is a calculation device in a second embodiment. [Figure 13] FIG. 11 is a diagram illustrating an internal configuration of a deletion index determination unit in the second embodiment. [Figure 14] FIG. 11 is an explanatory diagram showing a state of filter storage when a calculation is performed by a calculation unit in the second embodiment. [Figure 15] FIG. 11 is a functional block diagram of a control device according to a third embodiment. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS EXAMPLES
[0020] 1 is a functional block diagram of a recognition device 1000 according to a first embodiment. The recognition device 1000 is a type of computing device, and is a device that recognizes the situation of the outside world using outside world information from a camera or a LIDAR.
[0021] 1 is connected to an external world information acquisition device 101. The recognition device 1000 is configured with a CNN calculation unit 102, a model storage unit 104, a channel information acquisition unit 105, a weight parameter acquisition unit 106, a filter information acquisition unit 107, a deletion index determination unit 108, and a calculation allocation unit 109. As a result, the recognition device 1000 outputs a recognition result 103.
[0022] Here, each component will be described. First, the external world information acquisition device 101 acquires external world information and transmits the acquired information to the CNN calculation unit 102. This external world information acquisition device 101 can be realized by a sensor such as a camera or LIDAR. Furthermore, the channel information acquisition unit 105 receives model information 110 which is an output from the model storage unit 104. Furthermore, the weight parameter acquisition unit 106 receives the model information 110 which is an output from the model storage unit 104.
[0023] Furthermore, filter information acquisition unit 107 receives model information 110 which is an output from model storage unit 104. Furthermore, deletion index determination unit 108 receives channel information 111 indicating the number of channels in each layer from channel information acquisition unit 105. Furthermore, deletion index determination unit 108 receives weight parameter information 112 which indicates the weight parameters of each layer which is an output from weight parameter acquisition unit 106. Furthermore, deletion index determination unit 108 receives filter information 113 which indicates the number of filters in each layer which is an output from filter information acquisition unit 107.
[0024] The computation allocation unit 109 also receives model information 110 which is the output of the model storage unit 104. Furthermore, the computation allocation unit 109 receives deletion filter index information 114 for each layer which is the output of the deletion index determination unit. Furthermore, the computation allocation unit 109 also receives deletion channel index information 116 for each layer.
[0025] In addition, the CNN calculation unit 102 receives input data 117 which is external world information from the external world information acquisition device 101, a CNN calculation control signal 115 which is output from the calculation allocation unit 109, channel information 111, and filter information 113. Then, the CNN calculation unit 102 outputs the recognition result 103 using these.
[0026] Next, the operation and signal flow of the recognition device 1000 will be described. First, external world information is transmitted from the external world information acquisition device 101 to the CNN calculation unit 102. Next, the CNN calculation unit 102 performs calculation based on the model information 110 of the model storage unit 104, the CNN calculation control signal 115 of the calculation allocation unit 109, the channel information 111, and the filter information 113. Note that in FIG. 1, the arrows indicating the input channel information 111 and filter information 113 are omitted. Then, the CNN calculation unit 102 uses these to calculate and output the recognition result 103.
[0027] Next, channel information acquisition unit 105 extracts the number of channels for each layer from the model information in model storage unit 104, and outputs channel information 111 to deletion index determination unit 108. In addition, weight parameter acquisition unit 106 extracts weight parameters for each layer from model information 110 received from model storage unit 104, and outputs weight parameter information 112 to deletion index determination unit 108.
[0028] Furthermore, filter information acquisition unit 107 acquires the number of filters in each layer from model information 110 in model storage unit 104, and outputs filter information 113 for each layer including this to deletion index determination unit .
[0029] Furthermore, deletion index determination section 108 calculates the indexes of the filters and channels for which operations are to be deleted, from channel information 111, weight parameter information 112 and filter information 113. Then, deletion index determination section 108 outputs deletion filter index information 114 and deletion channel index information 116 including these indexes to computation allocation section 109.
[0030] Furthermore, the computation allocation unit 109 determines the computation order and identifies filters and channels to be deleted from the model based on the deleted filter index information 114, deleted channel index information 116 for each layer, and model information 110. Here, the computation allocation unit 109 only needs to identify at least one of a filter and a channel as the deletion target. In particular, it is preferable to identify a filter as the deletion target. Then, the computation allocation unit 109 transmits a CNN computation control signal 115 including the identified filter and channel to the CNN computation unit 102.
[0031] Here, the deletion channel index information 116 will be described with reference to FIG. 2. The deletion filter index information 114 will be described with reference to FIG. 3. Furthermore, the operation of the deletion index determination unit 108 will be described in detail with reference to FIG. 5. The recognition device of this embodiment can also be realized by a so-called computer. In this case, the functions of each unit are executed by a processing device such as a CPU in accordance with a program. Moreover, this program is stored in a storage medium. Moreover, each unit can also be realized by dedicated hardware such as an FPGA (Field Programmable Gate Array) or a dedicated circuit.
[0032] 2 is an explanatory diagram for explaining the deletion channel index information 116. First, there are L layers of CNN in this embodiment, and each layer is connected in a cascade manner from the 1st layer to the Lth layer. Also, there are channels 251 in each layer, and the channel 251 in one layer is connected to all the channels 251 in the next layer. Here, a different index is given to the channel in each layer, and the deletion channel index information 116 specifies the channel index to be deleted for each layer. Hereinafter, this specified index number will be called the deletion channel index.
[0033] Next, Fig. 3 is an explanatory diagram for explaining the deletion filter index information 114. First, in CNN, filters 351 exist in each of the 1st to Lth layers. In Fig. 3, for a certain jth layer (1 ≤ j ≤ L), a 3 × 3 filter is illustrated. Note that the number of filters is merely an example and is not limited to this.
[0034] In this embodiment, all filters in one layer are assigned the same index number 352. The filter index 352 to be deleted for each layer is specified in this deletion filter index information 114. Hereinafter, this specified index number is referred to as a deletion filter index.
[0035] Fig. 4 is a diagram showing the internal configuration of the CNN calculation unit 102. In Fig. 4, the CNN calculation unit 102 has a memory 201, a parameter storage unit 202, and a convolution calculation unit 203. The CNN calculation unit 102 also handles input data 117, a CNN calculation control signal 115, model information 110, a CNN calculation control signal 204 for one layer, weight parameters 205 for all layers, the number of filters in all layers 206, the number of channels in all layers 207, individual weight parameters 208 for one layer, that is, weight parameters for each layer, the number of individual filters in each layer 209, and the number of individual channels in each layer 210. The CNN calculation unit 102 uses these to output a recognition result 103.
[0036] Next, the connections of the CNN calculation unit 102 shown in Fig. 4 will be described. Input data 117 is input to the memory 201. This input data 117 is the above-mentioned external world information. Furthermore, the parameter storage unit 202 holds the model information 110.
[0037] The first-layer convolution operation unit 203-1 receives the following signals or information as input: An individual CNN operation control signal 204-1, i.e., a first layer CNN operation control signal, is branched from the CNN operation control signal 115. Saved data 211 (input data 117) saved in memory 201 Individual weighting parameters 208-1 branched from the weighting parameters 205 output from the parameter storage unit 202 The number of individual filters 209-1 branched from the number of filters 206 output from the parameter storage unit 202 The number of individual channels 210-1 branched from the number of channels output from the parameter storage unit 202 Thereafter, the following signals or information are input to the second-layer and subsequent convolution operation units, that is, the j-th layer convolution operation unit 203-j. A CNN operation control signal 204-j of an individual layer, i.e., a j-th layer, is branched from the CNN operation control signal 115. Output data 212-j of the convolution operation unit of the j-1th layer Individual weight parameters 208-j branched from the weight parameters 205 output from the parameter storage unit 202 The number of individual filters 209-j branched from the number of filters 206 output from the parameter storage unit 202 The number of individual channels 210-j branched from the number of channels output from the parameter storage unit 202 The following signals or information are input to the L-th convolution calculation unit 203-L in the final layer: An individual CNN operation control signal 204-L branched from the CNN operation control signal 115 The output data 212-L of the L-1th layer convolution operation unit 203-j Individual weight parameters 208-L branched from the weight parameters 205 output from the parameter storage unit 202 The number of individual filters 209-L branched from the number of filters 206 output from the parameter storage unit 202 The number of individual channels 210-L branched from the number of channels output from the parameter storage unit 202 Then, the Lth layer convolution calculation unit 203-L outputs the recognition result 103 obtained by using these.
[0038] Next, the operation of the CNN calculation unit 102 will be described in detail. First, the CNN calculation unit 102 performs a convolution calculation based on the input data 117 and the model information 110, and outputs the recognition result 103. At this time, each convolution calculation unit 203 receives a CNN calculation control signal 204 for each layer branched from the CNN calculation control signal 115, and skips the calculation of some filters and channels. This skipping of calculation is an example of the deletion of the present application. Here, the deletion of the present application also includes passive processing such as omitting or ignoring calculation. Here, the deletion, that is, the skipped filter and channel are the same until the model information 110 is changed due to an external model update. Note that, as described above, the target to be skipped may be either the filter or the channel. As a result, in this embodiment, in other words, in this application, at least one of the filter information and the channel information used for convolution is deleted.
[0039] Moreover, the memory 201 stores the input data 117. Then, the saved data 211, which is the input data 117, is output to the convolution calculation unit 203. Next, the convolution calculation unit 203 performs a convolution calculation based on the input saved data 211 (output data 212) and the individual weighting parameters 208. As a result of this convolution calculation, the number of calculations, the number of output data, and the allocation of data to the calculation units are determined by a CNN calculation control signal 204, which is a signal for controlling the calculation order and data reading, the number of individual filters 209, and the number of individual channels 210. The convolution calculation unit 203 outputs output data 212, which is the calculation result determined in this manner. This is repeated from the first layer to the Lth layer.
[0040] Furthermore, the model information 110 received by the parameter storage unit 202 is divided into weight parameters 205 for each layer, the number of filters 206 for each layer, and the number of channels 207 for each layer, and these are output to the convolution operation unit 203 .
[0041] This concludes the explanation of CNN computation unit 102, and next we will explain deleting index determination unit 108. FIG.
[0042] 5, deletion index determination unit 108 receives as inputs channel information 111, weight parameter information 112, and filter information 113. Based on these inputs, deletion index determination unit 108 outputs deletion channel index information 116 and deletion filter index information 114. To this end, deletion index determination unit 108 has a target processing speed storage unit 301, a maximum number of operations storage unit 302, a calculation speed analysis unit 303, a sensitivity information analysis unit 304, a deletion channel determination unit 305, a deletion filter determination unit 306, a deletion channel index output unit 307, and a deletion filter index output unit.
[0043] 5 shows the connections of the deletion index determination unit 108. The calculation speed analysis unit 303 receives as inputs a target processing speed 310 which is the output of the target processing speed storage unit 301, a maximum number of calculations which can be performed per calculation unit 315 which is the output of the maximum number of calculations storage unit 302, channel information 111, and filter information 113.
[0044] Moreover, the sensitivity information analysis unit 304 receives as input the channel information 111, the weight parameter information 112, and the filter information 113. Moreover, the deletion channel decision unit 305 receives as input the number of deletion channels 311 which is the result of the calculation speed analysis unit 303, and the sensitivity information 313 which is the result of the sensitivity information analysis unit 304.
[0045] Furthermore, the deletion filter determination unit 306 receives as input the number of filters to be deleted 312 output from the calculation speed analysis unit 303 and sensitivity information 316 output from the sensitivity information analysis unit 304. Furthermore, the deletion channel index output unit 307 receives as input priority information 317 output from the deletion channel determination unit 305 and the number of operations 315 output from the maximum number of operations storage unit 302, and outputs deletion channel index information 116 for each layer.
[0046] Furthermore, the deletion filter index output unit 308 receives as input the output 318 of the deletion filter determination unit 306 and the number of operations 315 that is the output of the maximum number of operations storage unit 302, and outputs the deletion filter index information 114 for each layer.
[0047] Next, the operation of the deletion index determination unit 108 will be described with reference to Fig. 5. First, the calculation speed analysis unit 303 estimates the time required for the convolution calculation for each layer and the difference between the target processing speed and the time required for the convolution calculation, based on the input channel information 111, filter information 113, and the maximum number of calculations that can be performed per calculation unit 315.
[0048] Further, the sensitivity information analysis unit 304 estimates sensitivity information for the recognition accuracy of all channels in each layer and sensitivity information for the recognition accuracy of all filters in each layer from the channel information 111, the weight parameter information 112, and the filter information 113.
[0049] Furthermore, the removal channel determination unit 305 receives difference information between the calculation speed of each layer and the target processing speed input from the calculation speed analysis unit 303, sensitivity information 313 to the recognition accuracy for each layer output from the sensitivity information analysis unit 304, and sensitivity information 313 of channels within the layer. Based on these, the removal channel determination unit 305 sets priorities of channels to be removed in order from the channel with the lowest sensitivity to recognition accuracy so that the difference between the target processing speed and the speed becomes 0. Details of the operation of the removal channel determination unit 305 will be described later with reference to FIG. 8.
[0050] Further, the deletion filter determination unit 306 prioritizes filters to be deleted for each layer based on the difference information between the processing speed of each layer and the target processing speed from the calculation speed analysis unit 303 and the filter sensitivity information 316 .
[0051] Furthermore, the deletion channel index output unit 307 receives priority information 317 of channels to be deleted for each layer determined by the deletion channel determination unit 305, and information on the maximum number of operations that can be performed by the computing unit output from the maximum number of operations storage unit 302. Based on these, the deletion channel index output unit 307 sets an index number so that the number of remaining channels is a multiple of the maximum number of operations that can be performed by the computing unit or a divisor other than 1, and outputs this.
[0052] Here, removal channel index output unit 307 performs processing according to the following (Equation 1), where Mc is the maximum number of channels that can be calculated, C is the number of channels in a certain j layer (1≦j≦L), and Cd is the number of channels to be removed from the j layer. In other words, removal channel index output unit 307 sets Cd that satisfies (Equation 1), and sets the calculation order of the Cd channels in ascending order of priority information to removal channel index information 116 for the j layer.
[0053]
number
[0054] Furthermore, based on the priority order of filters to be deleted set by the deletion filter determination unit 306, the deletion filter index output unit 308 sets index information so that the number of remaining filters is a multiple of the maximum number of operations that the calculator can perform or a divisor other than 1. The deletion filter index output unit 308 then outputs this. Note that the specification of the filter priority order by the deletion filter determination unit 306 will be described later with reference to FIG. 7.
[0055] Here, deletion filter index output unit 308 performs processing according to the following (Equation 2), where Mf is the maximum number of filters that can be calculated, F is the number of filters in a certain j-th layer, and Fd is the number of filters to be deleted from the j-th layer. That is, deletion filter index output unit 308 sets Fd that satisfies (Equation 2), and sets deletion filter index information 114 for the j-th layer for the calculation order of the Fd filters in ascending order of priority information.
[0056]
number
[0057] Next, a description will be given of the convolution calculation unit 203. Fig. 6 is a diagram showing the internal configuration of the convolution calculation unit 203 in this embodiment.
[0058] First, the configuration of the convolution calculation unit 203 will be described. Here, the jth layer will be described as an example. The convolution calculation unit 203 receives the CNN calculation control signal 204, the number of individual channels 210, the number of individual filters 209, the output data 212-j of the previous layer, and the individual weighting parameters 208 as inputs. Then, the convolution calculation unit 203 uses these to output the output data 212-j, which is the calculation result. For this purpose, the convolution calculation unit 203 has a calculation control unit 401, an input data temporary storage unit 402, a calculation unit 403, and an output data temporary storage unit 404.
[0059] Next, the connections in the internal configuration of the convolution calculation unit 203 are shown. The calculation control unit 401 receives as inputs the CNN calculation control signal 204, the number of individual channels 210, and the number of individual filters 209. In addition, the input data temporary storage unit 402 receives as inputs the output data 212-j of the previous layer, the individual weighting parameters 208, and a control signal 410 output from the calculation control unit 401 for controlling the output order and output timing of the stored data.
[0060] The calculation unit 403 also receives as input a calculation stop control signal 412 for controlling the order of calculation, the start of calculation, and the stop of calculation, and input data 411 required for the calculation. The output data temporary storage unit 404 also receives as input a calculation result 413 of the calculation unit 403 and a control signal 414 for controlling the output timing of the stored data, and outputs output data 212-j+1 which is the calculation result.
[0061] The operation of the convolution operation unit 203 will be described below. First, the output data 212-j of the previous layer and the individual weight parameters 208 are stored in the input data temporary storage unit 402. Then, the operation control unit 401 outputs the data output timing and the information of the channel and filter for which the output is to be skipped in accordance with the order of the operations to be performed, based on the information of the CNN operation control signal 204, the number of individual channels 210, and the number of individual filters 209. In addition, the operation control unit 401 outputs the operation stop control signal 412 to the operation unit 403, which controls the number of operations, the operation start timing, and the end timing.
[0062] The output data temporary storage unit 404 holds a control signal 414 that controls the timing of outputting the data after the calculation, allowing it to be used in other parts. The calculation unit 403 performs a convolution calculation and activation function processing on the input data, and outputs the calculation result 413 to the output data temporary storage unit 404. The output data temporary storage unit 404 sequentially outputs the output data 212-j+1, which is the calculation result, in accordance with the control signal 414 from the calculation control unit 401. As for the CNN calculation control signal 204, the number of individual channels 210, and the number of individual filters 209, the calculation control unit 401 continues to hold the data unless the model information 110 is updated.
[0063] Next, a detailed description will be given of the operation of the deletion filter determination unit 306. Fig. 7 is a flowchart showing the operation of the deletion filter determination unit 306 for setting the priority of filters to be deleted.
[0064] First, the deletion filter determination unit 306 starts operation when the number of channels to be deleted 311 output from the calculation speed analysis unit 303 and the sensitivity information 313 output from the sensitivity information analysis unit 304 are input (step S1001).
[0065] Next, the deletion filter determination unit 306 extracts n layers (n is an integer that satisfies n>=0) of layers with processing speeds equal to or lower than the target processing speed based on the output of the operation speed analysis unit 303. Then, the deletion filter determination unit 306 arranges the layers in order of decreasing processing speed (step S1002).
[0066] Next, when the layers having a processing speed equal to or slower than the processing speed are arranged, the deletion filter determination unit 306 sets i as a parameter indicating the currently targeted layer, and sets i=n (step S1003).
[0067] Next, the deletion filter determination unit 306 determines whether i=0 (step S1004). As a result, if i=0, the process proceeds to step S1006. If i is not 0, the process proceeds to step S1005.
[0068] Furthermore, the deletion filter determination unit 306 sets the deletion priority for each filter size in ascending order of sensitivity based on the information on sensitivity to recognition accuracy (step S1005). For example, in the case of a filter consisting of 3×3=9 weight parameters, the priority is set in ascending order of the amount of degradation of recognition accuracy when the nine weight parameters are deleted. Note that the number of filters is not limited to this example. Furthermore, the criterion for deletion is not limited to the amount of degradation, and may be any criterion that satisfies a predetermined rule. Furthermore, the amount of degradation of calculation accuracy may be used in addition to the degradation of recognition accuracy.
[0069] Furthermore, the deletion filter determination unit 306 decrements i by 1 and returns to step S1004 (step S1007).
[0070] Finally, the deletion filter determination unit 306 outputs the deletion priority information of the filters calculated individually for each of the n layers (step S1006).
[0071] This concludes the explanation of the operation of setting the priority of filters to be deleted in deletion filter determination section 306.
[0072] Next, a description will be given of the operation of the deletion channel determination section 305. FIG.
[0073] First, the deletion channel decision unit 305 starts operation when the number of deletion channels 311 output from the calculation speed analysis unit 303 and the sensitivity information 313 output from the sensitivity information analysis unit 304 are input (step S2001).
[0074] Next, the deletion channel determination unit 305 extracts n layers (n is an integer satisfying n>=0) whose processing speed is below the target based on the number of deletion channels 311 output from the calculation speed analysis unit, and sorts them in order of slowest processing speed (step S2002).
[0075] Next, the deletion channel determination unit 305 sets i as a parameter representing the layer currently being targeted when layers having a processing speed equal to or slower than the processing speed are arranged, and sets i=n (step S2003).
[0076] Next, the deletion channel determination unit 305 judges whether i=0 or not (step S2004). As a result, if i=0, the process proceeds to step S2006. If i is not 0, the process proceeds to step S2005.
[0077] Furthermore, the deletion channel determination unit 305 calculates the number of channels that need to be deleted based on the sensitivity information 313 that is the output of the sensitivity information analysis unit 304 for the channels of the i layer and the information on the number of deletion channels 311 that is the output of the calculation speed analysis unit 303. It is desirable to calculate the number of channels that need to be deleted until the processing speed of the i layer exceeds the target processing speed. Then, the deletion channel determination unit 305 sets deletion priorities in ascending order of sensitivity to recognition accuracy (step S2005).
[0078] Finally, the deletion channel determination unit 305 outputs the deletion priority information of the channels calculated individually for each of the n layers (step S2006).
[0079] Next, a description will be given of the operation in the computation allocation unit 109. Fig. 9 is a flowchart showing the operation in the computation allocation unit 109. The processing flow will be described below.
[0080] First, the computation allocation unit 109 starts operation when it receives the deletion channel index information 116 and the deletion filter index information 114 from the deletion index determination unit 108 (step S3001). Next, the computation allocation unit 109 stores the input deletion channel index information (step S3002).
[0081] Next, the computation allocation unit 109 stores the deletion filter index information and the deletion priority (step S3002). Next, the computation allocation unit 109 judges whether the filter to be deleted is at both ends as viewed from the stride direction of the deletion filter index information (step S3004). The stride direction and both ends will be described later with reference to FIG. 10. If it is judged that the filter is at both ends, the process proceeds to step S3007. If it is judged that the filter is not at both ends, the process proceeds to step S3005.
[0082] The computation allocation unit 109 also determines whether the stride direction can be changed in terms of implementation (step S3005). For this purpose, it is desirable for the computation allocation unit 109 to make a determination based on implementation constraints. If the result of this determination is that the stride direction can be changed, the process proceeds to step S3009. If the stride direction cannot be changed, the process proceeds to step S3006.
[0083] Next, the computation allocation unit 109 changes the priority of the filter index to be deleted to the next highest priority among both ends (step S3006). The deletion channel index information is output (step S3007).
[0084] Then, the computation allocation unit 109 outputs the deletion filter index information and the stride direction (step S3008).
[0085] Furthermore, if it is determined in step S3006 that this is not possible, the computation allocation unit 109 changes the stride direction by 90 degrees (step S3009), and the process then proceeds to step S3007.
[0086] Next, the stride and the filter that can be removed will be described. Fig. 10 is a diagram for explaining an example of the stride and the filter that can be removed. In the convolution operation, the convolution operation is performed while a convolution filter 803 and a filter 805 move in the stride directions 802 and 804 on the intermediate data 801 in Fig. 10(a) and (b). Here, the filter is 3 x 3, but the size of the filter is not limited to this.
[0087] The two ends of the filter refer to the filters shown in gray in the convolution filter 803 in Fig. 10(a) and the filter 805 in Fig. 10(b). Here, the stride direction 802 of the convolution filter 803 is left and right, so the filters at both ends are left and right. Also, the stride direction 804 of the filter 805 is up and down, so the filters at both ends are up and down.
[0088] Next, FIG. 11 shows an example of deleting a 3×3 filter. Here, an actual example of deleting when the maximum number of operations possible is 8 is shown. The maximum number of operations is 8 as an example here, but is not limited to this and can be 2 n (n≧1). In this way, when the maximum number of operable filters is 8, the number of filters to be deleted is one. Therefore, one filter is deleted from the gray area shown in FIG. 11. Deleted filter 913 is the filter after deletion has been performed with horizontal stride 912. Deleted filter 914 is the filter after deletion has been performed with vertical stride 915. This concludes the explanation of the first embodiment. EXAMPLES
[0089] Next, a description will be given of Example 2. Example 2 is an example in which the recognition device 1000 of Example 1 is applied to recognition of the outside world when a vehicle is traveling. For this reason, in Example 2, the processing speed of the convolution calculation is changed according to the traveling speed and the speed required for the calculation processing, such as when traveling on a road with no speed limit, when traveling on a normal highway, or when traveling in an urban area.
[0090] Here, the processing speed required by the ECU differs between normal highway driving or city driving and speed limit road driving. The width of the road surface and surrounding conditions also differ. Therefore, this embodiment shows an example in which the change in driving speed is observed and the number of calculations in which a single computing unit performs multiple calculations is increased to accommodate the increase in the required processing speed.
[0091] In the second embodiment, when the vehicle speed information is acquired and the average speed V per unit time exceeds the upper limit X of the vehicle speed determined in advance at the time of design, the target processing speed is changed and the number of deleted indexes is increased. This is an embodiment for increasing the number of calculations and improving the processing speed. Note that the same reference numerals are used in the drawings to designate parts in common with the first embodiment, and their explanations are omitted.
[0092] FIG. 12 is a functional block diagram of a recognition device 1000 which is a calculation device in the second embodiment. Here, differences from the first embodiment will be described with reference to FIG. 12. In the recognition device 1000 of the second embodiment, a driving speed acquisition unit 901 is added. In addition, the present embodiment has a deletion index determination unit 903 which is configured differently from the first embodiment. Next, the connection relationship of the recognition device 1000 of the second embodiment will be described. First, the driving speed acquisition unit 901 outputs the driving speed to the deletion index determination unit 903. The deletion index determination unit 108 receives outputs from the driving speed acquisition unit 901, the channel information acquisition unit 105, the weight parameter acquisition unit 106, and the filter information acquisition unit 107.
[0093] Next, a description will be given of the operation of the recognition device 1000 according to the second embodiment, with respect to differences from the first embodiment. First, a traveling speed acquisition unit 901 monitors the traveling speed of the vehicle, and continuously outputs the current traveling speed to the deletion index determination unit .
[0094] The recognition device 1000 of this embodiment can also be realized by a so-called computer. In this case, the functions of each part are executed by a processing device such as a CPU according to a program. The program is stored in a storage medium. Each part can also be realized by dedicated hardware such as an FPGA (Field Programmable Gate Array) or a dedicated circuit.
[0095] 13 is a diagram showing the internal configuration of the deletion index determination unit 903 in the embodiment 2. Here, the difference between the configuration of the deletion index determination unit 903 and the deletion index determination unit 108 will be described with reference to FIG.
[0096] The deletion index determination unit 903 further includes a target processing speed determination unit 902 and a calculation speed analysis unit 905 in addition to the deletion index determination unit 108 of the first embodiment. Here, the target processing speed determination unit 902 receives as an additional input the running speed information output from the running speed acquisition unit 901. The calculation speed analysis unit 905 receives as an additional input the output 916 from the maximum number of operations storage unit 302 and the target processing speed 310 which is the output from the target processing speed determination unit, and outputs the maximum number of operations 906.
[0097] Next, the connection relationship of the deletion index determination unit 903 in the second embodiment will be described. First, travel speed information is input to the target processing speed determination unit 902. Also, the output 916 of the maximum number of operations storage unit 302, the target processing speed 310 which is the output of the target processing speed determination unit 902, the channel information 111, and the filter information 113 are input to the calculation speed analysis unit 905. Then, the calculation speed analysis unit 905 outputs the number of channels to be deleted 311, the number of filters to be deleted 312, and the changed maximum number of operations 906. Also, the changed maximum number of operations 906 is input to the deletion channel index output unit 307 and the deletion filter index output unit 308.
[0098] The operation of the deletion index determination unit 903 in Embodiment 2 will be described below. In the target processing speed determination unit 902, the traveling speed V per unit time is calculated from the input traveling speed information. Here, when V < X is satisfied, the target processing speed determination unit 902 maintains the existing target processing speed G. On the other hand, when V > X is satisfied, the target processing speed determination unit 902 changes the target processing speed G to 2G and transmits the changed target processing speed to the operation speed analysis unit 905.
[0099] Also, when the target processing speed increases in the operation speed analysis unit 905, the target processing speed determination unit 902 changes the maximum number of operations 2 n to 2 n-1 When it decreases, the target processing speed determination unit 902 changes it back to 2 n Next, the target processing speed determination unit 902 obtains the number of deletion filter channels based on the changed maximum number of operations. Then, the target processing speed determination unit 902 transmits the changed maximum number of operations 906 to the deletion channel index output unit 307 and the deletion filter index output unit 308. Since the other operations are the same as those in FIG. 5, they are omitted.
[0100] Next, FIG. 14 is an explanatory diagram of the storage state of filters when performing operations in the arithmetic unit. FIG. 14(a) shows an example before reducing the maximum number of operations, and FIG. 14(b) shows an example after reducing the maximum number of operations. Also, these are configured based on the arithmetic unit 151 and the weight parameter 152.
[0101] Here, the case where the maximum number of operations is 8 (n = 3) and the filter size is 9 will be described with reference to FIG. 14(a). When the maximum number of operations is 8, one index is deleted from the filter, the weight parameter 152 is set in the arithmetic unit 151, and the operation is performed. Also, an example when the traveling speed increases and the maximum number of operations becomes 4 will be described with reference to FIG. 14(b). When the maximum number of operations is 4, 5 are deleted from the filter and set in the arithmetic unit 151. In this case, compared with FIG. 14(a), twice the weight parameter 152 can be put into the arithmetic unit 151. The description of Embodiment 2 ends here. EXAMPLES
[0102] The third embodiment is an embodiment in which the recognition device 1000 of the first and second embodiments is applied to a control device 2000. Fig. 15 is a functional block diagram of the control device 2000 in the third embodiment. This control device 2000 is implemented as, for example, an ECU.
[0103] 3, a control device 2000 includes a recognition device 1000 and a control signal generating unit 2001. Then, the recognition result 103 output from the recognition device 1000 that executes the processing of the first or second embodiment is transmitted to the control signal generating unit 2001. Next, the control signal generating unit 2001 generates a control signal 2002 according to the recognition result, and controls a control target 3000 based on the control signal 2002.
[0104] Here, when the control device 2000 is executed by an ECU, the controlled object 3000 is a vehicle. In this case, automatic driving or driving assistance of the vehicle can be realized based on the processes described in each embodiment.
[0105] Although the description of each embodiment is finished above, various modifications and application examples are assumed for each embodiment. For example, in each embodiment, the recognition device 1000 is used as an example, but a calculation device that performs calculations not limited to recognition is also included in the category of each embodiment.
[0106] In addition, according to each embodiment, when performing calculations using general image data, the utilization rate of the calculator is about 50%, but by using the calculation deletion mechanism according to the present invention, it is expected that the utilization rate of the calculator can be improved to nearly 100%. In other words, the utilization rate can be improved.
[0107] Furthermore, in this embodiment, when performing calculations in the convolutional neural network, some of the filters and channels used in the convolution calculation are deleted. For this reason, deletion index information, which is a number indicating the position or order of the filters and channels that differ for each layer, is provided to the calculation control unit. This allows skipping the reading of some of the input data for the convolution calculation and some of the weight parameters in accordance with the calculation unit of the calculator, and deleting part of the convolution calculation. This allows the calculator to be used efficiently in a device with limited calculation performance. [Explanation of symbols]
[0108] 101 External world information acquisition device 102 CNN calculation unit 104 Model Storage Unit 105 Channel information acquisition unit 106 Weight parameter acquisition unit 107 Filter information acquisition unit 108 Delete index determination part 109 Operation allocation unit 111 Channel Information 112 Weight Parameter Information 113 Filter Information 114 Delete filter index information 115 CNN operation control signal 116 Delete channel index information 117 Input Data 202 Parameter storage section 203 Convolution Calculation Unit 204 Branched CNN operation control signal 208 Individual Weight Parameters 209 individual filters 210 individual channels 211 Saved Data 301 Target Processing Speed Storage 302 Maximum number of operations storage section 303 Computation speed analysis section 304 Sensitivity Information Analysis Section 305 Deleted channel decision unit 306 Removal Filter Decision Unit 401 Calculation and control unit 412 Operation stop control signal 801 Intermediate Data 802 Stride Direction 803 Convolution Filter 901 Travel speed acquisition unit 902 Target processing speed determination unit
Claims
1. In a computing device that performs CNN calculations based on input data, A model storage unit for storing a model used in the CNN calculation; A CNN calculation unit that performs the CNN calculation by executing a convolution calculation for each of a plurality of convolution layers using the model; a weight parameter acquisition unit that acquires weight parameters of a convolution filter used in the convolution operation from the model storage unit; a channel information acquisition unit that acquires channel information for each of the plurality of convolution layers from the model stored in the model storage unit; a filter information acquisition unit that acquires filter information for each of the plurality of convolution layers from the model storage unit; An operation allocation unit that allocates a combination of weight information and operations required for the CNN operation and transmits the combination to the CNN operation unit; a deletion index determination unit that deletes a portion of filter information used for the convolution calculations, based on the maximum number of calculations possible for the CNN calculation unit, an execution order of the CNN calculations, channel information for each convolution layer of the CNN calculation unit acquired from the channel information acquisition unit, weight information acquired from the weight parameter acquisition unit, and filter information for each of the plurality of convolution layers acquired from the filter information acquisition unit.
2. 2. The computing device according to claim 1, The deletion index determination unit is a calculation device that further deletes a part of the channel information used in the convolution calculation.
3. 2. The computing device according to claim 1, The deletion index determination unit is a calculation device that determines filter information to be deleted in response to a change in the weight information.
4. 3. The computing device according to claim 2, The deletion index determination unit is a calculation device that determines channel information to be deleted in response to a change in the weight information.
5. 2. The computing device according to claim 1, The filter indicated by the filter information is composed of a plurality of weight parameters, The deletion index determination unit is a calculation device that determines a weight parameter that satisfies a predetermined rule from among the plurality of parameters.
6. 6. The computing device according to claim 5, The deletion index determination unit is a calculation device that determines weight parameters to be deleted based on an amount of degradation in calculation accuracy as the predetermined rule.
7. 2. The computing device according to claim 1, The deletion index determination unit is a calculation device that determines filter information to be deleted in accordance with a processing speed.
8. 8. The computing device according to claim 7, The deletion index determination unit is a calculation device that determines filter information to be deleted based on sensitivity information indicating sensitivities of the plurality of convolution layers and the size of the filter.
9. The computing device according to any one of claims 1 to 8, As the input data, external environment information acquired from an external environment acquisition device is used, A recognition device that recognizes an external situation using the external world information.
10. The recognition device according to claim 9, A control device characterized in that the result of the calculation in the recognition device is output as a control signal for an object in accordance with the recognized external situation.
Citation Information
Patent Citations
Vehicle electronic control apparatus
JP2018190045A
Method and apparatus for performing operations in convolutional neural network, and non-temporary storage medium
JP2019082996A
Neural network weight reducing device, neural network weight reducing method, and program
JP2020190996A
Neural network hardware accelerator system with zero-skipping and hierarchical structured pruning methods
US20200401895A1
Multi-task processing system and multi-task processing method
WO2005096634A1