Arithmetic processing unit
By designing an operation processing device that includes memory for saving accumulated results, the problem that convolutional neural network cannot be calculated at one time due to too much filter coefficients or input feature image data in ASIC is solved, and more efficient image recognition processing is achieved.
Patent Information
- Application Number
- CN201880096920.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2018-10-12
- Publication Date
- 2025-07-01
- Estimated Expiration
- 2039-02-15
AI Technical Summary
When using deep learning of convolutional neural networks for image recognition, too much filter coefficient or too much input feature map data leads to the problem of not being able to perform calculations at once, especially when large-capacity memory cannot be installed in ASICs.
An operation processing device is designed, including a data storage memory management unit, a filter coefficient storage memory management unit, an external memory, an operation unit and a controller. By setting the memory for storing the accumulation result in the calculation unit, the accumulated intermediate results can be temporarily saved, and the filter coefficients and input feature quantity map data can be updated if necessary to ensure the continuity of the calculation.
It effectively avoids the problem of inability to calculate at one time due to too much filter coefficients or input feature map data, and improves the efficiency and accuracy of operation processing, especially in ASIC environments with limited resources.
Smart Images

Figure CN112639838B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to an arithmetic processing device, and more particularly, to a circuit structure of an arithmetic processing device that performs deep learning using a convolutional neural network. Background Art
[0002] Conventionally, there has been an arithmetic processing device that performs arithmetic operations using a neural network obtained by hierarchically connecting a plurality of processing layers. In particular, in an arithmetic processing device for image recognition, deep learning using a convolutional neural network (Convolutional Neural Network, hereinafter referred to as CNN) has been widely performed.
[0003] Figure 18 FIG. is a diagram showing a process of image recognition based on deep learning using CNN. In image recognition based on deep learning using CNN, by sequentially performing processes in a plurality of processing layers of CNN on the input image data (pixel data), final arithmetic result data for recognizing an object included in the image is obtained.
[0004] The processing layers of CNN are roughly classified into a Convolution layer (convolution layer) that performs convolution processing including convolution operation processing, non-linear processing, downsampling processing (pooling processing), etc., and a Full Connect layer (fully connected layer) that performs a fully connected (Full Connect) process of multiplying all inputs (pixel data) by filter coefficients and accumulating them. However, there are also convolutional neural networks without a fully connected (Full Connect) layer.
[0005] Image recognition based on deep learning using CNN is performed as follows. First, a combination of a convolution operation process (Convolution process) that extracts a certain region from the image data and applies a plurality of filters with different filter coefficients (filter coefficients) to generate a feature map (Feature Map, FM) and a downsampling process (pooling process) that downsamples a part of the feature map is used as one processing layer, and the above combination of processes is performed multiple times (in a plurality of processing layers). These processes are the processes of the Convolution layer.
[0006] In addition to maxpooling that extracts the maximum value of the nearby 4pix and downsizes it to 1 / 2×1 / 2, there are also variations such as average pooling that calculates the average value of the nearby 4pix (without extraction) in the pooling process.
[0007] Figure 19This is a diagram showing the process of convolution processing. First, the input image data is separately subjected to filtering processes with different filter coefficients. By summing them all up, data corresponding to one pixel can be obtained. Nonlinear transformation and downsampling processing (pooling processing) are performed on the generated data, and the above-mentioned processing is carried out for all pixels of the image data, thereby generating an output feature map (oFM) for one plane. By repeating the above operations multiple times, oFMs for multiple planes are generated. In an actual circuit, all of the above are processed in a pipeline.
[0008] Furthermore, by using the above output feature map (oFM) as an input feature map (iFM) and further performing filtering processes with different filter coefficients, the above convolution processing is repeated. In this way, multiple convolution processes are carried out to obtain an output feature map (oFM).
[0009] When performing convolution processing and reducing the FM to a certain extent, the image data is read as a one-dimensional data string. Multiple (in multiple processing layers) full connection processes are carried out. In this full connection process, each data in the one-dimensional data string is multiplied by a different coefficient and then summed up. These processes are the processes of the full connection layer.
[0010] Moreover, after the full connection process, as the final operation result, that is, the object estimation result, the probability of detecting the object contained in the image (the probability of object detection) is output. In Figure 18 this example, as the final operation result data, the probability of detecting a dog is 0.01 (1%), the probability of detecting a cat is 0.04 (4%), the probability of detecting a small boat is 0.94 (94%), and the probability of detecting a bird is 0.02 (2%).
[0011] In this way, high recognition rates can be achieved for image recognition based on deep learning using CNN. However, in order to increase the types of detected objects and improve the object detection accuracy, the network needs to be enlarged. As a result, the data storage buffer and the filter coefficient storage buffer inevitably become large-capacity, but large-capacity memories cannot be mounted in an ASIC (Application Specific Integrated Circuit).
[0012] In addition, in deep learning for image recognition processing, the relationship between the FM (Feature Map) sizes and the number of FMs (the number of faces of the FM) in the (K-1)-th layer and the K-th layer is mostly as shown in the following formula, so it is difficult to optimize when determining the memory size of the circuit.
[0013] FM size [K] = 1 / 4 × FM size [K-1]
[0014] Number of FMs [K] = 2 × number of FMs [K-1]
[0015] For example, when considering the memory size of a circuit corresponding to Yolo_v2, which is a variant of CNN, if it is determined only based on the maximum values of the FM size and the number of FMs, about 1 GB is required. In fact, since the number of FMs and the FM size have an inverse proportional relationship, a memory of about 3 MB is sufficient computationally. However, for an ASIC mounted on a battery-powered mobile device, there is a need to minimize power consumption and chip cost as much as possible, and efforts are required to minimize the memory as much as possible.
[0016] Due to such problems, CNNs are usually installed through software processing using high-performance PCs or GPUs (Graphics Processing Units). However, in order to achieve high-speed processing, it is necessary to construct the heavy processing parts in hardware. An example of such hardware installation is described in Patent Document 1.
[0017] In Patent Document 1, an arithmetic processing device is disclosed that realizes the efficiency of arithmetic processing by respectively mounting arithmetic blocks and a plurality of memories in a plurality of arithmetic processing units. The arithmetic block and the buffer paired with it perform convolution arithmetic processing in parallel via a relay unit, and accumulate data is transmitted and received between the arithmetic units. As a result, even if the input network is increased, the input for activation processing can be generated at once.
[0018] Prior Art Documents
[0019] Patent Documents
[0020] Patent Document 1: Japanese Unexamined Patent Application Publication No. 2017-151604 Summary of the Invention
[0021] Problems to be Solved by the Invention
[0022] The structure of Patent Document 1 is an asymmetric structure with an up-and-down relationship (directional). Since all the arithmetic blocks are cascade-connected, the accumulated intermediate result has to pass through all the arithmetic blocks. Therefore, if a large network is to be handled, the accumulated intermediate result must pass through the relay section and the redundant data holding section multiple times, forming a long cascade connection path, which takes processing time. In addition, when a huge network is subdivided, since the same data or filter coefficients are read in (re-read) from the DRAM (external memory) multiple times, the access amount to the DRAM may increase. However, Patent Document 1 does not describe nor consider a specific control method for avoiding such a possibility.
[0023] In view of the above circumstances, an object of the present invention is to provide an arithmetic processing device that can avoid the problem that calculation cannot be performed at once when there are too many filter coefficients to fit into the WBUF or when there are too many iFMs to fit into the IBUF.
[0024] Means for Solving the Problem
[0025] A first aspect of the present invention is an arithmetic processing device for performing deep learning of convolutional processing and fully connected processing, characterized in that the arithmetic processing device includes: a data storage memory management unit having a data storage memory for storing input feature map data and a data storage memory control circuit for managing and controlling the data storage memory; a filter coefficient storage memory management unit having a filter coefficient storage memory for storing filter coefficients and a filter coefficient storage memory control circuit for managing and controlling the filter coefficient storage memory; an external memory for storing the input feature map data and output feature map data; a data input unit for obtaining the input feature map data from the external memory; a filter coefficient input unit for obtaining the filter coefficients from the external memory; an arithmetic unit having a structure of N parallel inputs and M parallel outputs for obtaining the input feature map data from the data storage memory and obtaining the filter coefficients from the filter coefficient storage memory, and performing filtering processing, accumulation processing, non-linear operation processing, and pooling processing, where N and M are integers greater than or equal to 1; a data output unit for concatenating the M parallel data output from the arithmetic unit and outputting the output feature map data to the external memory; an accumulation result storage memory management unit having an accumulation result storage memory, an accumulation result storage memory storage unit, and an accumulation result storage memory readout unit, where the accumulation result storage memory temporarily records the intermediate results of the accumulation processing in units of each pixel of the input feature map, the accumulation result storage memory storage unit receives valid data to generate an address and writes the valid data into the accumulation result storage memory, and the accumulation result storage memory readout unit reads out specified data from the accumulation result storage memory; and a controller for controlling within the arithmetic processing device, and the arithmetic unit includes: a filter arithmetic unit for performing filtering processing in N parallel; a first adder for accumulating all the operation results of the filter arithmetic unit; a second adder for accumulating the result of the accumulation processing of the first adder at a subsequent stage; and a flip-flop for holding the result of the accumulation processing of the second adder.And an arithmetic control unit that controls the arithmetic unit. The arithmetic control unit performs control in the following manner: during the filtering process and the accumulation process for a specific pixel used to calculate the output feature amount map, when it is impossible to save all the input feature amount map data required for the filtering process and the accumulation process into the data storage memory, or when it is impossible to save all the filter coefficients required for the filtering process and the accumulation process into the filter coefficient storage memory, the intermediate result is temporarily saved into the accumulation result storage memory, and the processing of other pixels is performed. After all the intermediate results of the accumulation process for all pixels are saved in the accumulation result storage memory, it returns to the first pixel, reads out the value saved in the accumulation result storage memory and uses it as the initial value of the accumulation process, and continues to execute the accumulation process.;
[0026] The arithmetic control unit may also perform control in the following manner: when the filtering process and the accumulation process that can be executed with all the filter coefficients saved in the filter coefficient storage memory are completed, the intermediate result is temporarily saved into the accumulation result storage memory, and after updating the filter coefficients saved in the filter coefficient storage memory, the accumulation process is continued.
[0027] The arithmetic control unit may also perform control in the following manner: when the filtering process and the accumulation process that can be executed with all the input feature amount map data that can be input are completed, the intermediate result is temporarily saved into the accumulation result storage memory, and after updating the input feature amount map data saved in the data storage memory, the accumulation process is continued.
[0028] Alternatively, the accumulation result storage memory management unit may include: an accumulation result storage memory reading unit that reads out the accumulation intermediate result from the accumulation result storage memory and writes it to the external memory; and an accumulation result storage memory saving unit that reads in the accumulation intermediate result from the external memory and saves it to the accumulation result storage memory. The arithmetic control unit performs control in the following manner: during the filtering process and the accumulation process for a specific pixel used to calculate the output feature amount map, when the intermediate result is written from the accumulation result storage memory to the external memory, and the input feature amount map data saved in the data storage memory or the filter coefficients saved in the filter coefficient storage memory are updated and the accumulation process is continued, the accumulation intermediate result written to the external memory is read into the accumulation result storage memory from the external memory and the accumulation process is continued.
[0029] Advantages of the Invention
[0030] The arithmetic processing device according to various embodiments of the present invention can temporarily store the accumulated intermediate results in units of pixels of the iFM size, thus avoiding problems such as being unable to perform a single calculation because all iFM data does not completely enter the IBUF or the filter coefficients do not completely enter the WBUF. BRIEF DESCRIPTION OF THE DRAWINGS
[0031] Figure 1 It is a schematic diagram of obtaining an output feature map (oFM) from an input feature map (iFM) through convolution processing.
[0032] Figure 2 It is a schematic diagram showing a situation where the WBUF (memory for storing filter coefficients) for storing filter coefficients is insufficient in convolution processing.
[0033] Figure 3 It is a schematic diagram showing the operation when the filter coefficients are updated once in the middle in the convolution processing in the arithmetic processing device according to the first embodiment of the present invention.
[0034] Figure 4 It is a block diagram showing the overall structure of the arithmetic processing device according to the first embodiment of the present invention.
[0035] Figure 5 It is a block diagram showing the structure of the SBUF management unit in the arithmetic processing device according to the first embodiment of the present invention.
[0036] Figure 6 It is a diagram showing the structure of the arithmetic unit of the arithmetic processing device according to the first embodiment of the present invention.
[0037] Figure 7A It is a flowchart showing the control process performed by the arithmetic control unit in the arithmetic processing device according to the first embodiment of the present invention.
[0038] Figure 7B It shows Figure 7A The flowchart of the filter coefficient update control in step S2.
[0039] Figure 8 It is a schematic diagram of dividing and inputting iFM data into the arithmetic unit in the second embodiment of the present invention.
[0040] Figure 9 It is a schematic diagram showing the operation when the iFM data is updated n1 times in the middle in the convolution processing in the arithmetic processing device according to the second embodiment of the present invention.
[0041] Figure 10AIt is a flowchart showing the control performed by the arithmetic control unit in the arithmetic processing unit according to the second embodiment of the present invention.
[0042] Figure 10B It shows Figure 10A The flowchart of the iFM data update control in step S22.
[0043] Figure 11 It is a schematic diagram showing the mid-course update of iFM data and filter coefficients in the arithmetic processing unit according to the third embodiment of the present invention.
[0044] Figure 12A It is a flowchart showing the control performed by the arithmetic control unit in the arithmetic processing unit according to the third embodiment of the present invention.
[0045] Figure 12B It shows Figure 12A The flowchart of the iFM data update control in step S42 and the filter coefficient update control in step S44.
[0046] Figure 13 It is a diagram showing the schematic diagram of the convolution process when two SBUF are prepared for each oFM in the case where the number m of oFM for generating one output channel is 2.
[0047] Figure 14 It is a diagram showing the schematic diagram of the convolution process in the arithmetic processing unit according to the fourth embodiment of the present invention.
[0048] Figure 15 It is a block diagram showing the overall structure of the arithmetic processing unit according to the fourth embodiment of the present invention.
[0049] Figure 16 It is a block diagram showing the structure of the SBUF management unit in the arithmetic processing unit according to the fourth embodiment of the present invention.
[0050] Figure 17A It is a flowchart showing the control performed by the arithmetic control unit in the arithmetic processing unit according to the fourth embodiment of the present invention.
[0051] Figure 17B It shows Figure 17A The flowchart of the iFM data update control in step S72.
[0052] Figure 17C It shows Figure 17A The flowchart of the filter coefficient update control in step S76.
[0053] Figure 17D It shows Figure 17AFlowchart of the SBUF update control process in step S74.
[0054] Figure 17E is a diagram showing Figure 17A Flowchart of the SBUF backoff control process in step S82.
[0055] Figure 18 is a diagram showing the process flow of image recognition based on deep learning using CNN.
[0056] Figure 19 is a diagram showing the process flow of the prior art convolution process. Detailed implementation mode
[0057] The embodiments of the present invention will be described with reference to the accompanying drawings. First, the background of the structure adopting the embodiments of the present invention will be described.
[0058] Figure 1 is a schematic diagram of obtaining an output feature map (oFM) from an input feature map (iFM) through convolution processing. By performing filtering processing, accumulation, non-linear transformation, pooling (shrinking), etc. on the iFM, the oFM is obtained. As the information required to calculate one pixel of the oFM, the information (iFM data and filter coefficients) of all pixels near the coordinates of the iFM corresponding to the output (one pixel of the oFM) is required.
[0059] Figure 2 is a schematic diagram showing the situation where the WBUF (memory for storing filter coefficients) for storing filter coefficients is insufficient in convolution processing. In Figure 2 In the example, based on the information (iFM data and filter coefficients) of 9 pixels near the coordinates (X, Y) of 6 iFMs, the data of one pixel (oFM data) at the coordinates (X, Y) of the oFM is calculated. At this time, each iFM data read from the IBUF (memory for storing data) is multiplied by the filter coefficient read from the WBUF (memory for storing filter coefficients) and accumulated.
[0060] As Figure 2 shown, when the size of the WBUF is small, the filter coefficients corresponding to all iFM data cannot be stored in the WBUF. In Figure 2In the example, WBUF can only store the filter coefficients corresponding to 3 pieces of iFM data. In this case, the first half of the 3 pieces of iFM data are each multiplied by the corresponding filter coefficient and accumulated, and the result (accumulation result) is temporarily stored (step 1). Next, the filter coefficients stored in WBUF are updated (step 2), and the second half of the 3 pieces of iFM are each multiplied by the corresponding filter coefficient and further accumulated (step 3). Then, the accumulation result of (step 1) is added to the accumulation result of (step 3). After that, by performing non-linear processing and pooling processing, 1 pixel data (oFM data) at the coordinates (X, Y) of oFM is obtained.
[0061] In this case, when calculating the pixel data (oFM data) at the next coordinate of oFM, since the filter coefficients stored in WBUF have been updated, WBUF needs to read the filter coefficients from DRAM again. Since such re-reading of the filter coefficients is performed for each pixel number, the bandwidth of DRAM is consumed, and power is also wasted.
[0062] (First Embodiment)
[0063] Next, the first embodiment of the present invention will be described with reference to the accompanying drawings. Figure 3 It is a schematic diagram showing the operation in the case where the filter coefficients are updated once in the middle in the convolution process in this embodiment. In the convolution process, all the input iFM data are multiplied by different filter coefficients, and all these values are accumulated to calculate 1 pixel data (oFM data) of oFM.
[0064] If the number of iFM (number of sheets) = N, the number of oFM (number of sheets) = M, and the filter kernel size is 3×3 (=9), then the total number of elements of the filter coefficients is 9×N×M. N and M vary depending on the network, but sometimes become extremely large sizes exceeding tens of millions. In such a case, since it is impossible to set a huge WBUF that can store all the filter coefficients, it is necessary to update the data stored in WBUF in the middle. However, when the size of WBUF is a small capacity that cannot even form 1 pixel data (oFM data) of oFM (specifically, less than 9N), it is necessary to re-read the filter coefficients in units of oFM pixels, so the efficiency is very poor.
[0065] Therefore, in the present embodiment, an SRAM (hereinafter referred to as SBUF (accumulation result storage memory)) having a capacity equal to (or larger than) the size of one iFM is prepared. Then, all accumulations that can be performed with the filter coefficients stored in the WBUF are carried out, and the intermediate results (accumulation results) are written (stored) in the SBUF (accumulation result storage memory) in pixel units. In Figure 3 the example of Figure 3 , the corresponding filter coefficients are multiplied by the first half of the three iFM data respectively and accumulated, and the intermediate results are stored in the SBUF (accumulation result storage memory). Then, when the filter coefficients stored in the WBUF are updated and the subsequent accumulations (accumulations of the second half of the three) are started, the values taken out from the SBUF are used as the accumulation initial values, and the corresponding filter coefficients are multiplied by the second half of the three iFM data respectively and accumulated. Then, by performing non-linear processing and pooling processing on the accumulation results, one pixel data of the oFM (oFM data) is obtained.
[0066] Figure 4 FIG. 6 is a block diagram showing the overall structure of the arithmetic processing device according to the present embodiment. The arithmetic processing device 1 includes a controller 2, a data input unit 3, a filter coefficient input unit 4, an IBUF (data storage memory) management unit 5, a WBUF (filter coefficient storage memory) management unit 6, an arithmetic unit (arithmetic block) 7, a data output unit 8, and an SBUF (accumulation result storage memory) management unit 11. The data input unit 3, the filter coefficient input unit 4, and the data output unit 8 are connected to a DRAM (external memory) 9 via a bus 10. The arithmetic processing device 1 generates an output feature map (oFM) based on the input feature map (iFM).
[0067] The IBUF management unit 5 includes a memory for storing input feature map (iFM) data (data storage memory, IBUF) and a management / control circuit for the data storage memory (data storage memory control circuit). The IBUF is composed of a plurality of SRAMs respectively.
[0068] The IBUF management unit 5 counts the number of valid data in the input data (iFM data), transforms it into coordinates, and further transforms it into an IBUF address (address in the IBUF), stores the data in the data storage memory, and retrieves the iFM data from the IBUF by a specified method.
[0069] The WBUF management unit 6 includes a memory for storing filter coefficients (filter coefficient storage memory, WBUF) and a management / control circuit for the filter coefficient storage memory (filter coefficient storage memory control circuit). The WBUF management unit 6 refers to the state of the IBUF management unit 5 and retrieves the filter coefficients corresponding to the data retrieved from the IBUF management unit 5 from the WBUF.
[0070] The DRAM 9 stores the iFM data, oFM data, and filter coefficients. The data input unit 3 obtains the input feature map (iFM) from the DRAM 9 by a prescribed method and transfers it to the IBUF (memory for data storage) management unit 5. The data output unit 8 writes the output feature map (oFM) data into the DRAM 9 by a prescribed method. Specifically, the data output unit 8 concatenates the M-parallel data output from the arithmetic unit 7 and outputs it to the DRAM 9. The filter coefficient input unit 4 obtains the filter coefficients from the DRAM 9 by a prescribed method and transfers them to the WBUF (memory for storing filter coefficients) management unit 6.
[0071] Figure 5 It is a block diagram showing the structure of the SBUF management unit 11. The SBUF management unit 11 includes an SBUF (memory for storing accumulation results) storage unit 111, an SBUF (memory for storing accumulation results) 112, and an SBUF (memory for storing accumulation results) readout unit 113. The SBUF 112 is a buffer for temporarily storing the accumulated intermediate results for each pixel unit (pixel unit) of the iFM. The SBUF readout unit 113 reads out the desired data (accumulation result) from the SBUF 112. When the SBUF storage unit 111 receives valid data (accumulation result), it generates an address and writes the valid data into the SBUF 112.
[0072] The arithmetic unit 7 obtains data from the IBUF (memory for data storage) management unit 5 and obtains filter coefficients from the WBUF (memory for storing filter coefficients) management unit 6. In addition, the arithmetic unit 7 obtains the data (accumulation result) read out from the SBUF 112 by the SBUF readout unit 113, and performs data processing such as filtering / accumulation / nonlinear operation / pooling processing. The data (accumulation result) after the arithmetic unit 7 has performed data processing is stored in the SBUF 112 by the SBUF storage unit 111. The controller 2 controls the entire circuit.
[0073] In the CNN, the processing of the required number of layers is repeatedly executed in multiple processing layers. Then, the arithmetic processing device 1 outputs the subject estimation result as the final output data, and processes this final output data using a processor (which can also be a circuit), thereby obtaining the subject estimation result.
[0074] Figure 6 It is a diagram showing the structure of the arithmetic unit 7 of the arithmetic processing device according to the present embodiment. The number of input channels of the arithmetic unit 7 is N (N is an integer of 1 or more), that is, the input data (iFM data) is N-dimensional, and the N-dimensional input data is processed in parallel (input N-parallel).
[0075] The number of output channels of the arithmetic unit 7 is M (M is an integer of 1 or more), that is, the output data is M-dimensional, and the M-dimensional input data is output in parallel (M-parallel output). As Figure 6 shown, in one layer, for each channel (ich_0 to ich_N-1), input iFM data (d_0 to d_N-1) and filter coefficients (k_0 to k_N-1), 1 oFM data is output. This process is performed in parallel for M layers, and M oFM data och_0 to och_M-1 are output.
[0076] In this way, the arithmetic unit 7 has a structure in which the number of input channels is set to N, the number of output channels is set to M, and the degree of parallelism becomes N×M. The sizes of the number of input channels N and the number of output channels M can be set (changed) according to the size of the CNN, so it is necessary to appropriately set them in consideration of processing performance and circuit scale.
[0077] The arithmetic unit 7 has an arithmetic control unit 71 that controls each part within the arithmetic unit. In addition, the arithmetic unit 7 has a filter arithmetic unit 72, a first adder 73, a second adder 74, an FF (flip-flop) 75, a non-linear transformation unit 76, and a pooling processing unit 77 for each layer. There are exactly the same circuits for each plane, and there are M such layers.
[0078] The arithmetic control unit 71 requests the previous stage of the arithmetic unit 7 to input specified data to the filter arithmetic unit 72. The filter arithmetic unit 72 is configured to be able to simultaneously execute a multiplier and an adder in N-parallel internally. The filter arithmetic unit 72 performs filtering processing on the input data and outputs the results of the filtering processing in N-parallel.
[0079] The first adder 73 adds up all the results of the filtering processing in the filter arithmetic unit 72 that are executed and output in N-parallel. That is, the first adder 73 can be called an accumulator in the spatial direction. The second adder 74 accumulates the operation results of the first adder 73 input in a time-division manner. That is, the second adder 74 can be called an accumulator in the time direction.
[0080] In the present embodiment, there are two cases for the second adder 74 to start processing with an initial value set to zero and to start processing with the value stored in the SBUF (accumulation result storage memory) 112 as the initial value. That is, in Figure 6 the shown switch box 78, the input of the initial value of the second adder 74 is switched between zero and the value obtained from the SBUF management unit 11 (accumulation intermediate result).
[0081] The controller 2 makes this switch based on the current stage of the ongoing accumulation. Specifically, in each operation (stage), an instruction such as the write destination of the operation result is sent from the controller 2 to the operation control unit 71, and when the operation ends, the operation end is notified to the controller 2. At this time, the controller 2 makes a judgment based on the current stage of the ongoing accumulation and issues an instruction to switch the input of the initial value of the second adder 74.
[0082] The operation control unit 71 performs all accumulations that can be executed with the filter coefficients stored in the WBUF through the second adder 74 and the FF 75, and writes (stores) the intermediate result (accumulation intermediate result) thereof to the SBUF (accumulation result storage memory) 112 in pixel units. An FF 75 for holding the result of the accumulation is provided at the subsequent stage of the second adder 74.
[0083] The operation control unit 71 controls in the following manner: during the filter processing / accumulation operation processing of the data (oFM data) of a specific pixel for calculating oFM, the intermediate result is temporarily stored in the SBUF 112, and the processing of another pixel of oFM is performed. Then, the operation control unit 71 controls in the following manner: after the accumulation intermediate results for all pixels are stored in the SBUF 112, it returns to the first pixel, reads out the value stored in the SBUF 112 and uses it as the initial value of the accumulation processing, and continues to execute the accumulation processing.
[0084] In the present embodiment, the timing of storing the accumulation intermediate result in the SBUF 112 is set to when the filter / accumulation processing that can be executed with all the filter coefficients stored in the WBUF is completed, and it is controlled in the following manner: after the filter coefficients stored in the WBUF are updated, the processing is continued.
[0085] The non-linear transformation unit 76 performs non-linear operation processing on the accumulation result in the second adder 74 and the FF 75 based on an activation function or the like. There is no particular limitation on the specific implementation method. For example, non-linear operation processing is performed by piecewise approximation.
[0086] The pooling processing unit 77 performs pooling processing such as selecting and outputting the maximum value (MaxPooling) or calculating the average value (Average Pooling) from the multiple data input from the non-linear transformation unit 76. In addition, the processing in the non-linear transformation unit 76 and the pooling processing unit 77 can be omitted by the operation control unit 71.
[0087] With such a structure, in the arithmetic unit 7, the sizes of the number of input channels N and the number of output channels M can be set (changed) according to the size of the CNN. Therefore, it is necessary to appropriately set them in consideration of processing performance and circuit scale. In addition, since it is N-parallel processing without an up-down relationship, the accumulation is competitive, and no long path like a cascaded connection is generated, and the latency is short.
[0088] Figure 7A FIG. is a flowchart showing the control flow performed by the arithmetic control unit in the arithmetic processing device of the present embodiment. When starting the convolution process, first, it enters the "iFM number loop 1" (step S1). Then, the filter coefficients stored in the WBUF are updated (step S2). Next, it enters the "iFM number loop 2" (step S3).
[0089] Next, it enters the "arithmetic unit execution loop" (step S4). Then, "coefficient save determination" is performed (step S5). In the "coefficient save determination", it is determined whether the filter coefficients stored in the WBUF are the desired filter coefficients. When the result of the "coefficient save determination" is OK, it enters the "data save determination" (step S6). When the result of the "coefficient save determination" is not OK, it waits until the result of the "coefficient save determination" becomes OK.
[0090] In the "data save determination" in step S6, it is determined whether the iFM data stored in the IBUF is the desired data. When the result of the "data save determination" is OK, it enters the "arithmetic unit execution" (step S7). When the result of the "data save determination" is not OK, it waits until the result of the "data save determination" becomes OK.
[0091] In the "arithmetic unit execution" in step S7, the arithmetic unit performs filtering / accumulation processing. When the filtering / accumulation processing that can be performed with all the filter coefficients stored in the WBUF ends, the process ends. If not, it returns to steps S1, S3, and S4, and the processing is repeated.
[0092] If the number of iFM data is set to n1×n2×N, the number of times of the "iFM number loop 1" (step S1) = n1, and the number of times of the "iFM number loop 2" (step S3) = n2, then the accumulation of the second adder 74 is n2 times, and the number of times of temporarily writing as an intermediate result in the SBUF 112 is n1 times.
[0093] Figure 7B is shown Figure 7AFlowchart of the filtering coefficient update control process in step S2. First, in step S11, the filtering coefficient is read into WBUF. Then, in step S12, the number of updates of the filtering coefficient is counted. When the filtering coefficient update is the first time, it enters step S13, and the initial accumulation value is set to zero. When the filtering coefficient update is not the first time, it enters step S14, and the initial value of the accumulation calculation is set to the value stored in SBUF.
[0094] Next, in step S15, the number of updates of the filtering coefficient is counted. When the filtering coefficient update is the last time, it enters step S16, and the output destination of the data (accumulation result) is set to the non-linear transformation unit. When the filtering coefficient update is not the last time, it enters step S17, and the output destination of the data (accumulation result) is set to SBUF.
[0095] In addition, in the filtering coefficient update control, the initial accumulation value (step S13 or S14) and the output destination of the data (accumulation result) (step S16 or S17) are passed as status information to the operation control unit of the operation unit, and the switches of each unit are controlled according to the status in the operation control unit.
[0096] (Second Embodiment)
[0097] The first embodiment of the present invention addresses the case of a large number of filtering coefficients (the case of a small WBUF), but the same problem also occurs when there is too much iFM data instead of filtering coefficients. That is, consider the case where only a part of the iFM data is stored in IBUF. At this time, if the iFM data stored in IBUF is updated midway to calculate the data (oFM data) of 1 pixel (1 pixel) of oFM, then it is necessary to re-read the iFM data to calculate the data (oFM data) of the next pixel of oFM.
[0098] In addition, the iFM data required to process 1 pixel of oFM is only the nearby information of the same pixel. However, even if only the local area is stored in IBUF, in the case where the network is enlarged and thousands of iFM data are required, or when IBUF is reduced to the limit in order to reduce the scale, the data buffer (IBUF) is insufficient, and it is inevitable to split and read the iFM data.
[0099] Therefore, in the second embodiment of the present invention, it is possible to address the case of too much iFM data (the case of a small IBUF). In addition, the setting of SBUF (memory for storing the accumulation result) is the same as that in the first embodiment. Figure 8 It is a schematic diagram of splitting the iFM data and inputting it into the operation unit in this embodiment.
[0100] First, the iFM data is stored in the data buffers (IBUF_0 to IBUF_N-1) of the n2×N planes. The arithmetic unit performs accumulation n2 times based on the second adder 74 (accumulator in the time direction), and writes the intermediate result (accumulation intermediate result) to the SBUF (memory for storing the accumulation result) 112. If the intermediate results have been written for all pixels, the next iFM data of the n2×N planes is read in, and the accumulation intermediate result is taken out from the SBUF 112 as the initial value, and the accumulation operation is continued. By repeating the above operation n1 times, processing of n×N (=n1×n2×N) planes can be performed.
[0101] Figure 9 It is a schematic diagram showing the operation in the case where the iFM data is updated n1 times in the middle during the convolution processing in the present embodiment. First, each data of the first iFM group (iFM_0) is multiplied by the filter coefficient and accumulated, and the intermediate result (accumulation intermediate result) is written to the SBUF (memory for storing the accumulation result) 112. Then, all the calculations that can be performed using the first iFM group (iFM_0) are carried out.
[0102] Next, the second iFM group (iFM_1) is read into the IBUF. Then, the accumulation intermediate result is taken out from the SBUF 112 as the initial value, each data of the second iFM group (iFM_1) is multiplied by the filter coefficient and accumulated, and the intermediate result (accumulation intermediate result) is written to the SBUF (memory for storing the accumulation result) 112. Then, all the calculations that can be performed using the second iFM group (iFM_1) are carried out.
[0103] The same operation is repeated until the n1th iFM group (iFM_n1), and pooling processing such as non-linear processing / shrinking processing is performed on the obtained accumulation result to obtain the data (oFM data) of 1 pixel of the oFM. In this way, performing all the calculations up to the point where it is possible is the same as in the first embodiment.
[0104] The structure of the present embodiment is the same as the structure of the first embodiment shown in Figures 4 - 6 Therefore, the description is omitted. As a difference from the first embodiment, the second adder 74 performs all the accumulations that can be executed on the iFM data stored in the IBUF, and writes (stores) the intermediate result (accumulation intermediate result) to the SBUF (memory for storing the accumulation result) 112 in pixel units.
[0105] In addition, in the present embodiment, the timing for storing the accumulated intermediate result in SBUF 112 is set to when all the filtering / accumulation processes that can be performed on the inputtable iFM data are completed, and the control is such that the process continues to be implemented after the iFM data is updated.
[0106] Figure 10A FIG. is a flowchart showing the control performed by the arithmetic control unit in the arithmetic processing device of the present embodiment. When starting the convolution process, first, it enters "iFM number loop 1" (step S21). Then, the iFM data stored in IBUF is updated (step S22). Next, it enters "iFM number loop 2" (step S23).
[0107] Next, it enters "arithmetic unit execution loop" (step S24). Then, "coefficient storage determination" is performed (step S25). In the "coefficient storage determination", it is determined whether the filter coefficient stored in WBUF is the desired filter coefficient. If the result of the "coefficient storage determination" is OK, it enters "data storage determination" (step S26). If the result of the "coefficient storage determination" is not OK, it waits until the result of the "coefficient storage determination" becomes OK.
[0108] In the "data storage determination" of step S26, it is determined whether the iFM data stored in IBUF is the desired data. If the result of the "data storage determination" is OK, it enters "arithmetic unit execution" (step S27). If the result of the "data storage determination" is not OK, it waits until the result of the "data storage determination" becomes OK.
[0109] In the "arithmetic unit execution" of step S27, the arithmetic unit performs filtering / accumulation processing. When the filtering / accumulation processing that can be performed on all the iFM data stored in IBUF is completed, the process ends. If not, it returns to steps S21, S23, and S24, and the processing is repeated.
[0110] Figure 10B FIG. shows Figure 10A a flowchart of the iFM data update control in step S22. First, in step S31, the iFM data is read into IBUF. Then, in step S32, the update count of the iFM data is counted. When the iFM data is updated for the first time, it enters step S33, and the accumulation initial value is set to zero. When the iFM data is not updated for the first time, it enters step S34, and the accumulation initial value is set as the value stored in SBUF.
[0111] Next, in step S35, the number of updates of the iFM data is counted. When the iFM data is updated for the last time, step S36 is entered, and the output destination of the data (accumulation result) is set to the non-linear transformation unit. When the iFM data is not updated for the last time, step S37 is entered, and the output destination of the data (accumulation result) is set to SBUF.
[0112] In addition, in the iFM data update control, the accumulation initial value (step S33 or S34) and the output destination of the data (accumulation result) (step S36 or S37) are passed as status information to the operation control unit of the operation unit, and each unit switch is controlled according to the status in the operation control unit.
[0113] (Third Embodiment)
[0114] In the first embodiment, it is the case where all filter coefficients cannot be stored in WBUF, and in the second embodiment, it is the case where all iFM data cannot be stored in IBUF, but there is also a case where both occur simultaneously. That is, as the third embodiment, the case where all filter coefficients cannot be stored in WBUF and all iFM data cannot be stored in IBUF is described.
[0115] Figure 11 It is a schematic diagram of updating the iFM data and filter coefficients midway in this embodiment. Figure 11 It is an example where the number of iFM groups n1 = 2 and the filter coefficients are updated once.
[0116] First, each data of the initial iFM group (iFM_0) is multiplied by the filter coefficient and accumulated, and the intermediate result (accumulation intermediate result) is written into SBUF (memory for storing accumulation result) 112.
[0117] Next, the filter coefficient group stored in WBUF is updated. Then, the accumulation intermediate result is taken out from SBUF 112 as the initial value, each data of the iFM group (iFM_0) is multiplied by the filter coefficient and accumulated, and the intermediate result (accumulation intermediate result) is written into SBUF 112. In this way, all calculations that can be performed using the initial iFM group (iFM_0) are carried out.
[0118] Next, the iFM group stored in IBUF is updated (the second iFM group (iFM_1) is read into IBUF), and the filter coefficient group stored in WBUF is updated. Then, the accumulation intermediate result is taken out from SBUF112 as the initial value, each data of the second iFM group (iFM_1) is multiplied by the filter coefficient and accumulated, and the intermediate result (accumulation intermediate result) is written into SBUF (memory for storing accumulation result) 112.
[0119] Next, update the filter coefficients stored in WBUF. Then, take out the accumulated intermediate result as the initial value from SBUF112, multiply each data of the second iFM group (iFM_1) by the filter coefficients and accumulate them, and write the intermediate result (accumulated intermediate result) to SBUF (memory for storing the accumulated result) 112. In this way, perform all the calculations that can be performed using the second iFM group (iFM_1).
[0120] By performing pooling processing such as non-linear processing / downscaling processing on the accumulated result obtained in this way, data (oFM data) of one pixel of oFM is obtained. In this way, performing all the calculations up to the point where it is possible is the same as in the first embodiment and the second embodiment.
[0121] In this way, also in the present embodiment, it is possible to cope with the situation where both WBUF and IBUF are insufficient.
[0122] Figure 12A It is a flowchart showing the control performed by the operation control unit in the operation processing device of the present embodiment. Figure 12A It shows the case where the update frequency of the filter coefficient group is higher than the update frequency of the iFM data. The one with the higher update frequency becomes the inner loop.
[0123] When starting the convolution processing, first enter the "iFM number loop 1" (step S41). Then, update the iFM data stored in IBUF (step S42). Next, enter the "iFM number loop 2" (step S43). Then, update the filter coefficients stored in WBUF (step S44). Next, enter the "iFM number loop 3" (step S45).
[0124] Next, enter the "operation unit execution loop" (step S46). Then, perform "coefficient save determination" (step S47). In the "coefficient save determination", it is determined whether the filter coefficients stored in WBUF are the desired filter coefficients. If the result of the "coefficient save determination" is OK, enter the "data save determination" (step S48). If the result of the "coefficient save determination" is not OK, wait until the result of the "coefficient save determination" becomes OK.
[0125] In the "data save determination" of step S48, it is determined whether the iFM data stored in IBUF is the desired data. If the result of the "data save determination" is OK, enter the "operation unit execution" (step S49). If the result of the "data save determination" is not OK, wait until the result of the "data save determination" becomes OK.
[0126] In the "Execution by the arithmetic unit" of step S49, the arithmetic unit performs filtering / accumulation processing. When the filtering / accumulation processing that can be performed on all the iFM data stored in the IBUF is completed, the process ends. If not, the process returns to steps S41, S43, and S46, and the processing is repeated.
[0127] Figure 12B is a flowchart showing Figure 12A the flow of the iFM data update control in step S42 and the filter coefficient update control in step S44.
[0128] First, the update control of the iFM data, which is the outer loop, is performed. In step S51, the iFM data is read into the IBUF. Then, in step S52, the number of updates of the iFM data is counted. If the iFM data update is the first time, the process proceeds to step S53, and the value Si1 is set to zero. If the iFM data update is not the first time, the process proceeds to step S54, and the value Si1 is set to the value stored in the SBUF.
[0129] Then, in step S55, the number of updates of the iFM data is counted. If the iFM data update is the last time, the process proceeds to step S56, and Od1 is set to the non-linear transformation unit. If the iFM data update is not the last time, the process proceeds to step S57, and Od1 is set to the SBUF.
[0130] Next, the update control of the filter coefficient, which is the inner loop, is performed. In step S61, the filter coefficient is read into the WBUF. Then, in step S62, the number of updates of the filter coefficient is counted. If the filter coefficient update is the first time, the process proceeds to step S63, and the accumulation initial value is set to the value Si1. If the filter coefficient update is not the first time, the process proceeds to step S64, and the accumulation initial value is set to the value stored in the SBUF.
[0131] Then, in step S65, the number of updates of the filter coefficient is counted. If the filter coefficient update is the last time, the process proceeds to step S66, and the output destination of the data (accumulation result) is set to Od1. If the filter coefficient update is not the last time, the process proceeds to step S67, and the output destination of the data (accumulation result) is set to the SBUF.
[0132] In addition, in the iFM data update control and the filter coefficient control, the value Si1 (step S53 or S54), Od1 (step S56 or S57), the accumulation initial value (step S63 or S64), and the output destination of the data (accumulation result) (step S66 or S67) are passed as status information to the arithmetic control unit of the arithmetic unit, and the arithmetic control unit controls the switches of each unit according to its status.
[0133] In the above control flow, the loop count is set to n, which is divided into n = n1 × n2 × n3. Among them, the number of times of "iFM number loop 1" (step S41) = n1, the number of times of "iFM number loop 2" (step S43) = n2, and the number of times of "iFM number loop 3" (step S45) = n3. At this time, the accumulation of the second adder 74 is n3 times, and the number of times of temporarily writing as an intermediate result in SBUF is n1 × n2 times.
[0134] Thus, in the first to third embodiments, the following method is shown: a structure that can achieve high-speed processing corresponding to a moving image and can change the filter size of the CNN, thereby enabling easy handling of either convolution processing or full connect processing in a structure where, in an input N-parallel / output M-parallel circuit, even when the iFM number > N and the oFM number > M, specific control can be performed, and further, it can handle the case where the iFM number or the number of parameters is large due to the degree of input segmentation required as N and M increase. That is, even when the CNN network expands, it can be handled.
[0135] (Fourth Embodiment)
[0136] In the case of outputting multiple oFMs from one output channel, consider the situation where the number of oFMs needs to exceed the number of planes of the output parallelism M. In Figure 11 the shown processing, both the filter coefficient and the iFM are updated during this processing to generate one oFM data. And in this processing, if the number of oFMs that must be generated by one output channel is m (m > 1), then consider repeatedly performing the Figure 11 shown processing m times for processing.
[0137] In this method, the IBUF is rewritten sequentially, so it is necessary to reread all the iFMs m times. Therefore, the DRAM access amount increases, and the desired performance cannot be obtained. Therefore, if multiple SBUFs are prepared for each oFM, the SBUF can store all the accumulated results of m planes and can prevent rereading, but the circuit scale increases.
[0138] As such an example, Figure 13 is a diagram showing a schematic diagram of convolution processing in the case where the number of oFMs m = 2 that must be generated by one output channel, and two SBUFs are prepared for each oFM. Since two oFM data (oFM 0 and oFM 1) are generated, in order to prevent rereading, it is necessary to have a first SBUF that stores the accumulated result of oFM0 and a second SBUF that stores the accumulated result of oFM 1.
[0139] First, for the oFM 0 data, each data of the initial iFM group (n1 = 0) is multiplied by a filtering coefficient and accumulated, and the accumulated intermediate result is stored in the 1st SBUF. Then, after updating the filtering coefficient stored in the WBUF, the value in the 1st SBUF is used as the initial value for accumulation, and the accumulated intermediate result is stored in the 1st SBUF.
[0140] Next, for the oFM 1 data, after updating the filtering coefficient stored in the WBUF, each data of the initial iFM group (n1 = 0) is multiplied by a filtering coefficient and accumulated, and the accumulated intermediate result is saved to the 2nd SBUF. Then, after updating the filtering coefficient stored in the WBUF, the value in the 2nd SBUF is used as the initial value for accumulation, and the accumulated intermediate result is stored in the 2nd SBUF.
[0141] Next, the second iFM group (n1 = 1) is read into the IBUF. Then, for the oFM 0 data, using the value in the 1st SBUF as the initial value, each data of the second iFM group (n1 = 1) is multiplied by a filtering coefficient and accumulated, and the accumulated intermediate result is saved to the 1st SBUF. Then, after updating the filtering coefficient stored in the WBUF, the value in the 1st SBUF is used as the initial value for accumulation, and the accumulated intermediate result is stored in the 1st SBUF.
[0142] Next, for the oFM 1 data, after updating the filtering coefficient stored in the WBUF, using the value in the 2nd SBUF as the initial value, each data of the second iFM group (n1 = 1) is multiplied by a filtering coefficient and accumulated, and the accumulated intermediate result is saved to the 2nd SBUF. Then, after updating the filtering coefficient stored in the WBUF, the value in the 2nd SBUF is used as the initial value for accumulation, and the accumulated intermediate result is stored in the 2nd SBUF.
[0143] By performing pooling processing such as non - linear processing / downscaling processing on the accumulated results obtained in this way (the values finally stored in the 1st and 2nd SBUF), two oFM data are obtained.
[0144] In this way, when the number of oFM planes needs to exceed the number of planes of the output parallelism M, in order to prevent re - reading, the SBUF needs to be set to the number of SBUF for the number of planes of the oFM output by 1 output channel, resulting in an increase in SRAM and an increase in the circuit scale.
[0145] Therefore, as the fourth embodiment, a method for coping without increasing the scale even when the number of oFM increases will be described. Figure 14 It is a diagram showing a schematic diagram of convolution processing in the arithmetic processing device of this embodiment.
[0146] In this embodiment, an SBUF having a capacity equal to (or larger than) the size of one iFM (one iFM) is prepared in the same manner as in the first to third embodiments. That is, the SBUF is sized to be able to store the accumulated intermediate result of all the pixel amounts on one side of the iFM.
[0147] In this embodiment, the accumulated intermediate result generated in the middle of the process of processing one oFM amount is temporarily written to the DRAM. This process is performed for the m-side amount. When updating the iFM and continuing the accumulation, the output accumulated intermediate result is read from the DRAM and the process is continued.
[0148] Use Figure 14 The process flow of this embodiment will be described. Similar to Figure 13 the same, Figure 14 Fig. shows a schematic diagram of the convolution process when generating data for two oFMs (oFM 0 and oFM 1).
[0149] First, for the oFM 0 data, each data in the first iFM group (n1 = 0) is multiplied by the filter coefficient and accumulated, and the accumulated intermediate result is stored in the SBUF. Then, after updating the filter coefficient stored in the WBUF, the value in the SBUF is used as the initial value for accumulation, and the accumulated intermediate result is stored in the SBUF. The accumulated intermediate result stored in the SBUF is sequentially transmitted to the DRAM as the intermediate result of the oFM 0 data.
[0150] Next, for the oFM 1 data, after updating the filter coefficient stored in the WBUF, each data in the first iFM group (n1 = 0) is multiplied by the filter coefficient and accumulated, and the accumulated intermediate result is stored in the SBUF. Then, after updating the filter coefficient stored in the WBUF, the value in the SBUF is used as the initial value for accumulation, and the accumulated intermediate result is stored in the SBUF. The accumulated intermediate result stored in the SBUF is sequentially transmitted to the DRAM as the intermediate result of the oFM 1 data.
[0151] Next, the second iFM group (n1 = 1) is read into the IBUF. Then, for the oFM 0 data, the intermediate result of the oFM 0 data stored in the DRAM is stored in the SBUF as the initial value. Next, using the value in the SBUF as the initial value, each data in the second iFM group (n1 = 1) is multiplied by the filter coefficient and accumulated, and the accumulated intermediate result is stored in the SBUF. Then, after updating the filter coefficient stored in the WBUF, the value in the SBUF is used as the initial value for accumulation, and the accumulated intermediate result is stored in the SBUF. By performing pooling processing such as non-linear processing / downsampling processing on the accumulated result obtained in this way, the data of oFM 0 is obtained.
[0152] Next, for the oFM 1 data, after updating the filter coefficients stored in WBUF, the intermediate result of the oFM 1 data stored in DRAM is stored in SBUF as the initial value. Next, using the value in SBUF as the initial value, each data of the second iFM group (n1 = 1) is multiplied by the filter coefficient and accumulated, and the accumulated intermediate result is stored in SBUF. Then, after updating the filter coefficients stored in WBUF, the value in SBUF is used as the initial value for accumulation, and the accumulated intermediate result is stored in the second SBUF. By performing pooling processing such as non-linear processing / shrinking processing on the accumulated result obtained in this way, the data of oFM 1 is obtained.
[0153] In this way, the data obtained from DRAM is temporarily stored in SBUF. Thus, it becomes the same state as the previous case where the initial value enters SBUF, and thus the previous processing can be started. Non-linear processing etc. is also performed at the end of the processing before outputting to DRAM.
[0154] This embodiment has the disadvantage that the processing speed is reduced due to outputting the accumulated intermediate result to DRAM. However, the processing of this embodiment can be handled with little increase in the circuit, so if a certain degree of degradation in performance can be allowed, it can handle the latest networks.
[0155] Next, the structure for performing the processing of this embodiment will be described. Figure 15 It is a block diagram showing the overall structure of the arithmetic processing device of this embodiment. Figure 15 The arithmetic processing device 20 shown is Figure 1 different from the arithmetic processing device 1 of the first embodiment shown in that it is the structure of the SBUF (memory for storing accumulated results) management unit.
[0156] Figure 16 It is a block diagram showing the structure of the SBUF management unit 21 of this embodiment. The SBUF management unit 21 has an SBUF control unit 210, a first SBUF storage unit 211, a second SBUF storage unit 212, an SBUF 112, a first SBUF readout unit 213, and a second SBUF readout unit 214.
[0157] The SBUF 112 is a buffer for temporarily storing the accumulated intermediate results in units of each pixel of the iFM (pixel unit). The first SBUF storage unit 211 and the first SBUF readout unit 213 are I / Fs for reading and writing values to and from DRAM.
[0158] When the 1st SBUF storage unit 211 receives data (intermediate result) from the DRAM 9 via the data input unit 3, it generates an address and writes it into the SBUF 112. When the 2nd SBUF storage unit 212 receives valid data (accumulated intermediate result) from the arithmetic unit 7, it generates an address and writes it into the SBUF 112.
[0159] The 1st SBUF readout unit 213 reads out the desired data (intermediate result) from the SBUF 112 and writes the data into the DRAM 9 via the data output unit 8. The 2nd SBUF readout unit 214 reads out the desired data (accumulated intermediate result) from the SBUF 112 and outputs the data as the initial value for accumulation to the arithmetic unit 7.
[0160] The structure of the arithmetic unit 7 is the same as that of the arithmetic unit in the 1st embodiment shown in Figure 6 and thus the description thereof is omitted. The arithmetic unit 7 acquires data from the IBUF (data storage memory) management unit 5 and acquires filter coefficients from the WBUF (filter coefficient storage memory) management unit 6. In addition, the arithmetic unit 7 acquires the data (accumulated intermediate result) read out from the SBUF 112 by the 2nd SBUF readout unit 214 and performs data processing such as filtering / accumulation / non-linear operation / pooling processing. The data (accumulated intermediate result) after the arithmetic unit 7 has performed data processing is stored in the SBUF 112 by the 2nd SBUF storage unit 212.
[0161] The SBUF control unit 210 controls the loading of the initial value (accumulated intermediate result) from the DRAM to the SBUF and the writing of the intermediate result from the SBUF to the DRAM. In the loading of the initial value from the DRAM to the SBUF, as described above, the 1st SBUF storage unit 211 receives data (initial value) from the DRAM 9 via the data input unit 3, generates an address and writes it into the SBUF 112.
[0162] Specifically, when inputting from the DRAM, when rtrig (read trigger) is input from the upper controller 2, the SBUF control unit 210 acquires data from the DRAM 9 and takes the data into the SBUF 112. After the taking-in is completed, the SBUF control unit 210 sends a rend (read end) signal to the upper controller 2 and waits for the next operation.
[0163] In the writing of the result from the SBUF to the DRAM, as described above, the first SBUF reading unit 213 reads the desired data (intermediate result) from the SBUF 112 and writes the data to the DRAM 9 via the data output unit 8. Specifically, when outputting to the DRAM, when the wtrig (write trigger) signal is output from the SBUF control unit 210 to the upper controller 2, all the data in the SBUF is output to the data output unit 8. After that, the SBUF control unit 210 sends the rend (read end) signal to the upper controller 2 and waits for the next operation.
[0164] In addition, the SBUF control unit 210 controls the first SBUF storage unit 211, the second SBUF storage unit 212, the first SBUF reading unit 213, and the second SBUF reading unit 214. Specifically, the SBUF control unit 210 outputs a trig (trigger) signal when giving an instruction and receives an end (end) signal when the process ends.
[0165] The data input unit 3 loads the accumulated intermediate result (intermediate result) from the DRAM 9 according to the request from the SBUF management unit 21. The data output unit 8 writes the accumulated intermediate result (intermediate result) to the DRAM 9 according to the request from the SBUF management unit 21.
[0166] With such a structure, it is possible to handle the situation where both the input and output become a huge FM.
[0167] Figure 17A It is a flowchart showing the control performed by the operation control unit in the arithmetic processing device of this embodiment.
[0168] When starting the convolution processing, first enter the "iFM number loop 1" (step S71). Then, update the iFM data stored in the IBUF (step S72). Next, enter the "oFM number loop" (step S73). Then, update the data stored in the SBUF (step S74). Next, enter the "iFM number loop 2" (step S75). Then, update the filter coefficient stored in the WBUF (step S76). Next, enter the "iFM number loop 3" (step S77).
[0169] Next, enter the "operation unit execution loop" (step S78). Then, perform "coefficient save determination" (step S79). In the "coefficient save determination", it is determined whether the filter coefficient stored in the WBUF is the desired filter coefficient. If the result of the "coefficient save determination" is OK, enter the "data save determination" (step S80). If the result of the "coefficient save determination" is not OK, standby until the result of the "coefficient save determination" becomes OK.
[0170] In the "data save determination" of step S80, it is determined whether the iFM data stored in IBUF is the desired data. When the result of the "data save determination" is OK, it proceeds to "execution of the arithmetic unit" (step S81). When the result of the "data save determination" is not OK, it waits until the result of the "data save determination" becomes OK.
[0171] In the "execution of the arithmetic unit" of step S81, the arithmetic unit performs filtering / accumulation processing. When the filtering / accumulation processing that can be performed with all the iFM data stored in IBUF is completed, it advances to "SBUF backup" (step S82). Otherwise, it returns to steps S75, S77, and S78, and the processing is repeated.
[0172] In the "SBUF backup" of step S82, the data stored in SBUF is backed up to DRAM. Then, it returns to steps S71 and S73, and the processing is repeated. When all the operations are completed, the process ends.
[0173] Figure 17B It shows Figure 17A The flowchart of the iFM data update control in step S72. First, in step S91, the iFM data is read into IBUF. Then, in step S92, the update count of the iFM data is counted. When the iFM data is updated for the first time, it proceeds to step S93 and sets the value Si1 to zero. When the iFM data is not updated for the first time, it proceeds to step S94 and sets the value Si1 as the value stored in SBUF.
[0174] Then, in step S95, the update count of the iFM data is counted. When the iFM data is updated for the last time, it proceeds to step S96 and sets Od1 as the non-linear transformation unit. When the iFM data is not updated for the last time, it proceeds to step S97 and sets Od1 to SBUF.
[0175] Figure 17C It shows Figure 17A The flowchart of the filter coefficient update control in step S76. First, in step S101, the filter coefficient is read into WBUF. Then, in step S102, the update count of the filter coefficient is counted. When the filter coefficient is updated for the first time, it proceeds to step S103 and sets the accumulation initial value to the value Si1. When the filter coefficient is not updated for the first time, it proceeds to step S104 and sets the accumulation initial value as the value stored in SBUF.
[0176] Then, in step S105, the number of updates of the filter coefficient is counted. When the filter coefficient is updated for the last time, step S106 is entered, and the output destination of the data (accumulation result) is set to Od1. When the filter coefficient is not updated for the last time, step S107 is entered, and the output destination of the data (accumulation result) is set to SBUF.
[0177] In addition, in Figure 17B the iFM data update control and Figure 17C the filter coefficient control, the values Si1 (step S93 or S94), Od1 (step S96 or S97), the accumulation initial value (step S103 or S104), and the output destination of the data (accumulation result) (step S106 or S107) are passed as status information to the operation control unit of the operation unit, and each unit switch is controlled according to its status in the operation control unit.
[0178] Figure 17D is a flowchart showing Figure 17A the process of SBUF update control in step S74 of
[0179] Figure 17E In step S111, the number of times of the iFM loop 1 is determined. When the iFM loop 1 is the first time, no processing is performed (ended). When the iFM loop 1 is not the first time, step S112 is entered, and the SBUF value is read from the DRAM. Figure 17A is a flowchart showing
[0180] the process of SBUF save control in step S82 of
[0181] Figure 17A In step S121, the number of times of the iFM loop 1 is determined. When the iFM loop 1 is the last time, no processing is performed (ended). When the iFM loop 1 is not the last time, step S122 is entered, and the SBUF value is written into the DRAM.
[0182] In the above control process, the number of loops is set to n, and it is divided into n = n1 × n2 × n3. Among them, the number of times of "iFM loop 1" (step S71) = n1, the number of times of "iFM loop 2" (step S75) = n2, and the number of times of "iFM loop 3" (step S77) = n3. At this time, the accumulation by the second adder 74 is n3 times, the number of times of temporarily writing as an intermediate result in the SBUF is n2 times, and the number of times of writing the intermediate result into the DRAM is n1 times. The control process of
[0182] As described above, one embodiment of the present invention has been described. However, the technical scope of the present invention is not limited to the above embodiment, and within the scope not departing from the gist of the present invention, the combination of components can be changed, or various changes can be made to each component or deletion can be performed.
[0183] Each component is a component for explaining the functions and processes of the respective components. The functions and processes of multiple components can also be implemented by one structure (circuit) simultaneously.
[0184] Each component, separately or as a whole, can be implemented by a computer composed of one or more processors, logic circuits, memories, input / output interfaces, and computer-readable recording media, etc. In this case, a program for implementing the functions of each component or the whole can be recorded in the recording medium, and the computer system reads in the recorded program and executes it, thereby implementing the above various functions and processes.
[0185] In this case, for example, the processor is at least one of a CPU, a DSP (Digital Signal Processor), and a GPU (Graphics Processing Unit). For example, the logic circuit is at least one of an ASIC (Application Specific Integrated Circuit) and an FPGA (Field-Programmable Gate Array).
[0186] In addition, the "computer system" mentioned here may also include hardware such as an OS or peripheral devices. In addition, when using a WWW system, the "computer system" also includes a homepage providing environment (or display environment). In addition, the "computer-readable recording medium" refers to writable non-volatile memories such as floppy disks, optical disks, ROMs, flash memories, removable media such as CD-ROMs, and storage devices such as hard disks built into the computer system.
[0187] Moreover, the "computer-readable recording medium" is a recording medium that holds a program for a certain period of time, such as a volatile memory (e.g., DRAM (Dynamic Random Access Memory)) inside a computer system as a server or a client when the program is transmitted via a network such as the Internet or a communication line such as a telephone line.
[0188] In addition, the above program can also be transmitted from a computer system that stores the program in a storage device or the like to other computer systems via a transmission medium or by a transmission wave in the transmission medium. Here, the "transmission medium" for transmitting the program is a medium having a function of transmitting information, such as a network (communication network) such as the Internet or a communication line (communication wire) such as a telephone line. In addition, the above program may also be a program for implementing a part of the above functions. Further, it may also be a so-called differential file (differential program) that implements the above functions by combining with a program already recorded in the computer system.
[0189] Industrial Applicability
[0190] The present invention can be widely applied to an arithmetic processing device that performs deep learning using a convolutional neural network.
[0191] Reference Numeral Explanation
[0192] 1, 20: Arithmetic processing device; 2: Controller; 3: Data input unit; 4: Filter coefficient input unit; 5: IBUF management unit (memory management unit for data storage); 6: WBUF management unit (memory management unit for filter coefficient storage); 7: Arithmetic unit; 8: Data output unit; 9: DRAM (external memory); 10: Bus; 11, 21: SBUF management unit (memory management unit for accumulated result storage); 71: Arithmetic control unit; 72: Filter arithmetic unit; 73: First adder; 74: Second adder; 75: FF (flip-flop); 76: Nonlinear transformation unit; 77: Pooling processing unit; 111: SBUF storage unit (memory storage unit for accumulated result storage); 112: SBUF (memory for accumulated result storage); 113: SBUF readout unit (memory readout unit for accumulated result storage); 210: SBUF control unit (memory control unit for accumulated result storage); 211: First SBUF storage unit (memory storage unit for accumulated result storage); 212: Second SBUF storage unit (memory storage unit for accumulated result storage); 213: First SBUF readout unit (memory readout unit for accumulated result storage); 214: Second SBUF readout unit (memory readout unit for accumulated result storage).
Claims
1. An operation processing device for performing deep learning of convolution processing and fully connected processing, characterized in that: The operation processing device has: A data storage memory management unit having a data storage memory for storing input feature map data and a data storage memory control circuit for managing and controlling the data storage memory; A filter coefficient storage memory management unit having a filter coefficient storage memory for storing filter coefficients and a filter coefficient storage memory control circuit for managing and controlling the filter coefficient storage memory; An external memory for storing the input feature map data and the output feature map data; A data input unit for obtaining the input feature map data from the external memory; A filter coefficient input unit for obtaining the filter coefficients from the external memory; An operation unit that obtains the input feature map data from the data storage memory in an input N-parallel and output M-parallel structure, and obtains the filter coefficients from the filter coefficient storage memory, and performs filtering processing, accumulation processing, non-linear operation processing, and pooling processing, where N and M are integers greater than or equal to 1; A data output unit that concatenates the M-parallel data output from the operation unit and outputs it as output feature map data to the external memory; An accumulation result storage memory management unit having an accumulation result storage memory, an accumulation result storage memory storage unit, and an accumulation result storage memory readout unit. The accumulation result storage memory temporarily records the intermediate results of the accumulation processing in units of each pixel of the input feature map. The accumulation result storage memory storage unit receives valid data to generate an address and writes the valid data into the accumulation result storage memory. The accumulation result storage memory readout unit reads out specified data from the accumulation result storage memory; and A controller for controlling the overall operation within the operation processing device, The operation unit has: A filter operation unit that performs filtering processing in N parallel; A first adder that accumulates all the operation results of the filter operation unit; A second adder that accumulates the result of the accumulation processing of the first adder at a later stage; A flip-flop that holds the result of the accumulation processing of the second adder; and An operation control unit that controls each unit within the operation unit, The operation control unit controls in the following manner: during the filtering process and the accumulation process for specific pixels used to calculate the output feature amount map, when all the input feature amount map data required for the filtering process and the accumulation process cannot be stored in the data storage memory, or when all the filter coefficients required for the filtering process and the accumulation process cannot be stored in the filter coefficient storage memory, the intermediate result is temporarily stored in the accumulation result storage memory to process other pixels. After the intermediate results of the accumulation process for all pixels are stored in the accumulation result storage memory, it returns to the initial pixel, reads out the value stored in the accumulation result storage memory as the initial value of the accumulation process, and continues to execute the accumulation process; Among them, the operation control unit also controls in the following manner: when the filtering process and the accumulation process that can be executed with all the filter coefficients stored in the filter coefficient storage memory are completed, the intermediate result is temporarily stored in the accumulation result storage memory, and after updating the filter coefficients stored in the filter coefficient storage memory, the accumulation process is continued.
2. The operation processing device according to claim 1, wherein, The operation control unit controls in the following manner: when the filtering process and the accumulation process that can be executed with all the input feature amount map data that can be input are completed, the intermediate result is temporarily stored in the accumulation result storage memory, and after updating the input feature amount map data stored in the data storage memory, the accumulation process is continued.
3. The operation processing device according to any one of claims 1 to 2, wherein, The accumulation result storage memory management unit has: An accumulation result storage memory reading unit that reads the accumulation intermediate result from the accumulation result storage memory and writes it to the external memory; And An accumulation result storage memory storage unit that reads the accumulation intermediate result from the external memory and stores it in the accumulation result storage memory, The operation control unit controls in the following manner: during the filtering process and the accumulation process for specific pixels used to calculate the output feature amount map, when the intermediate result is written from the accumulation result storage memory to the external memory, and the input feature amount map data stored in the data storage memory or the filter coefficients stored in the filter coefficient storage memory are updated and the accumulation process is continued, the accumulation intermediate result written to the external memory is read from the external memory into the accumulation result storage memory and the accumulation process is continued.
Citation Information
Patent Citations
Arithmetic processing unit
JP2017151604A