On-chip high-capacity buffer applied to deep learning acceleration core
By introducing a hierarchical design of on-chip large-capacity buffer L2_buffer into the deep learning acceleration core, using two-dimensional DMA transmission method to process off-core memory data, the problem of low computing efficiency caused by untimely data supply is solved, and efficient data supply and computational efficiency are improved.
Patent Information
- Application Number
- CN202510435160.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-08
- Publication Date
- 2025-07-22
AI Technical Summary
The problem of insufficient data supply in existing deep learning acceleration cores that lead to reduced computing efficiency is mainly due to the fact that the first-level buffer consumes a lot of time and resources during data handling.
The on-chip large-capacity buffer L2_buffer, which adopts a hierarchical design, uses the instruction decoding module, data interaction interface module, address processing module, data processing module and exception judgment module to process the data from the off-core memory using two-dimensional DMA transmission method, and stores it in the L2_buffer according to four data placement methods to prepare data for the first-level cache L1_buffer in advance.
The data calculation efficiency is improved, and the problem of data exchange consumes a lot of time and resources due to small storage space is avoided, so as to ensure efficient data supply of computing components, and to improve address calculation efficiency.
Smart Images

Figure CN120353754A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to deep learning hardware acceleration, and specifically to a large-capacity on-chip cache applied to a deep learning acceleration core. Background Art
[0002] The main functions of a deep learning acceleration core are to implement inference calculations of a deep neural network model, input / output processing, and have a certain data calculation ability in the form of vectors and matrices. To match its high computing power, the capacity of the memory should be able to cache a large amount of data.
[0003] Currently, in the industry, for data processing and transfer of a deep learning acceleration core, only a first-level cache is relied on. However, in the field of deep learning, the storage format of data is generally in the form of four-dimensional tensors, and there are generally two data placement modes for these data in off-chip memory - NCHW and NHWC. The first-level cache needs to store the data in these two data placement modes into the internal memory according to the requirements of the arithmetic unit in different placement ways, and this process will consume a large amount of time and resources. While the first-level cache transfers data from off-chip, the arithmetic unit will also perform read and write operations on the first-level cache, making the time consumed in data movement much greater than the time consumed in data calculation, thus resulting in a decrease in data calculation efficiency due to the inability to supply data efficiently. Summary of the Invention
[0004] (1) Technical Problems to be Solved
[0005] In view of the above-mentioned drawbacks of the prior art, the present invention provides a large-capacity on-chip cache applied to a deep learning acceleration core, which can effectively overcome the defect that the data calculation efficiency is reduced due to the inability to supply data efficiently in the prior art.
[0006] (2) Technical Solutions
[0007] To achieve the above object, the present invention is realized through the following technical solutions:
[0008] A large-capacity on-chip cache applied to a deep learning acceleration core includes an on-chip large-capacity cache L2_buffer with a hierarchical design located between the first-level cache L1_buffer and off-chip memory. The L2_buffer includes an instruction decoding module, a data interaction interface module, an address processing module, a data processing module, an exception judgment module, and a storage module;
[0009] When the L2_buffer works, the instruction control unit sends L2_buffer read / write instructions to the secondary decoding buffer queue, and the external interface depend / inform unit in the instruction decoding module controls the instruction execution process;
[0010] If the received instruction is an illegal instruction, an exception is reported and the L2_buffer read / write instruction sent by the instruction control unit is received again;
[0011] If the received instruction is a legal instruction and there is no dependency, the instruction is executed sequentially and a read / write request to the DMA is initiated;
[0012] If the received instruction is a legal instruction and there is a dependency, after receiving the begin_inform signal returned by the external interface depend / inform unit, the next operation is continued and a read / write request to the DMA is initiated;
[0013] Among them, if the above read / write process is completed normally, the idle state is entered; otherwise, an exception is reported and then the idle state is entered.
[0014] Preferably, the functions of the instruction decoding module include:
[0015] The L2_buffer receives the instruction signal sent by the instruction control unit, determines whether it is an L2_buffer read / write instruction according to the lower 5 bits, puts the L2_buffer read instruction into the load_fifo, and puts the L2_buffer write instruction into the store_fifo;
[0016] Among them, both the load_fifo and the store_fifo are synchronous fifos with a data width of 138 bits and a depth of 16.
[0017] Preferably, the state machine of the instruction decoding module includes:
[0018] When the fifo is not empty and all the instructions in the fifo have not been executed, the current state changes from the idle state to the waiting state;
[0019] When in the waiting state, the instruction is read out from the fifo, and whether there is a dependency is judged according to the higher 5 bits. When there is no dependency or the request response signal returned by the external interface depend / inform unit is received, the current state changes from the waiting state to the decoding state;
[0020] When the number of decoded feature maps is not 0, the current state changes from the decoding state to the execution state;
[0021] Among them, when the current instruction enters the execution state, the next instruction sends the dependency to the external interface depend / inform unit. After the current instruction is executed, when the instruction decoding module receives the request response signal returned by the external interface depend / inform unit, the next instruction enters the decoding state;
[0022] The write instruction completion flag indicates that all data has been written, and the end flag bit flag is pulled high.
[0023] The read instruction completion flag is judged according to the resp returned by the DMA. When the first bit of the resp is 1, it indicates that the read instruction is completed.
[0024] Preferably, the functions of the data interaction interface module include:
[0025] The L2_buffer performs read and write operations on the off-chip memory through the read and write ports of the custom on-chip bus protocol master. A two-dimensional DMA transfer method is adopted between the L2_buffer and the DMA. The control word consists of a start field, the start address transferred to the target unit, the data transfer length in the x direction, the number of 1D DMA transfers in the y direction, the start address deviation between two 1D DMA transfers, and a read / write control field, and the control word is sent to the DMA.
[0026] The L2_buffer is responsible for processing the read and write access requests of the L1_buffer. The access between the L2_buffer and the L1_buffer supports multi-address simultaneous access based on the control segment. The custom bus protocol maintains the control segment and the data segment in the form of data packets and performs physical transmission based on the custom bus port. A two-dimensional transfer method is also adopted between the L2_buffer and the L1_buffer. The L2_buffer, as the custom bus slave interface, receives the control word signal of the L1_buffer, parses the control word signal, and performs read and write operations on the storage units inside the L2_buffer.
[0027] Among them, the start field is generated by notifying the DMA to establish a connection. The data transfer length in the x direction is the one-dimensional length, and the number of 1D DMA transfers in the y direction is the two-dimensional length.
[0028] A. When the data layout mode of the off-chip memory is NCHW, the one-dimensional length is the number of times the current feature map width transfer is completed. Whether the feature map width can divide the data volume determines the one-dimensional length, following the ceiling principle. The calculation formula for the one-dimensional length is as follows:
[0029] 1) The feature map width can divide the data volume:
[0030] One-dimensional length = feature map width / data volume;
[0031] 2) The feature map width cannot divide the data volume:
[0032] One-dimensional length = feature map width / data volume + 1;
[0033] In the above two one-dimensional length calculation formulas, the data volume is the number of data that can be stored in one row of a bank in the L2_buffer, which is related to the data type being transmitted. When the bank bit width is 128 bits, for an 8-bit data type, the corresponding data volume is 16; for a 16-bit data type, the corresponding data volume is 8; and for a 32-bit data type, the corresponding data volume is 4.
[0034] The two-dimensional length is the number of data transmission completion times of the one-dimensional length, and the address step size is equal to the feature map width. The calculation formula for the two-dimensional length is as follows:
[0035] Two-dimensional length = number of feature maps * number of feature map channels * feature map height;
[0036] B. When the data placement mode in the off-chip memory is NHWC, whether the number of feature map channels can be divided evenly by the data volume determines the one-dimensional length. Following the ceiling principle, the calculation formula for the one-dimensional length is as follows:
[0037] 1) The number of feature map channels can be divided evenly by the data volume:
[0038] One-dimensional length = number of feature map channels / data volume;
[0039] 2) The number of feature map channels cannot be divided evenly by the data volume:
[0040] One-dimensional length = number of feature map channels / data volume + 1;
[0041] In the above two one-dimensional length calculation formulas, the data volume is the number of data that can be stored in one row of a bank in the L2_buffer, which is related to the data type being transmitted. When the bank bit width is 128 bits, for an 8-bit data type, the corresponding data volume is 16; for a 16-bit data type, the corresponding data volume is 8; and for a 32-bit data type, the corresponding data volume is 4.
[0042] The two-dimensional length is the number of data transmission completion times of the one-dimensional length, and the address step size is equal to the number of feature map channels. The calculation formula for the two-dimensional length is as follows:
[0043] Two-dimensional length = number of feature maps * feature map height * feature map width.
[0044] Preferably, the state machine of the data interaction interface module includes:
[0045] The start state is the idle state. When the control word is assembled, the decoding is valid and it enters the waiting state;
[0046] When in the waiting state, the control word is sent to the DMA. When writing to the DMA, the write enable signal for writing the L2_buffer to the DMA is pulled high. When the wready returned by the DMA is 1, it indicates that the DMA has received the control word, and the L2_buffer sends the corresponding length of data to the DMA. When the first bit of the resp returned by the DMA is 1, it indicates that the current DMA transfer has ended; when reading from the DMA, the read request signal for reading the L2_buffer from the DMA is pulled high. When the 0th bit of the resp returned by the DMA is 0, it indicates that the DMA has received the control word and sends the corresponding length of data to the L2_buffer. When the 0th bit of the resp returned by the DMA is 1, it indicates that the current DMA transfer has ended.
[0047] Preferably, the address calculation of the address processing module adopts an incremental calculation method. According to the different data placement modes of the off-chip memory and the different data placement methods in the L2_buffer, the corresponding address calculation formulas are used:
[0048] A. When the data placement mode of the off-chip memory is NCHW and the data placement method in the L2_buffer is row bank priority mode, the address calculation formula is as follows:
[0049] bank_num = h % BANK_NUM;
[0050] bank_addr = start_addr + (n * C + c) * (m_ceil(H, BANK_NUM)) * (m_ceil(W, data_num)) + (h / BANK_NUM) * (m_ceil(W, data_num)) + w / data_num;
[0051] B. When the data placement mode of the off-chip memory is NCHW and the data placement method in the L2_buffer is column bank priority mode, the address calculation formula is as follows:
[0052] bank_num = w % BANK_NUM;
[0053] bank_addr = start_addr + (n * C + c) * (m_ceil(W, BANK_NUM)) * (m_ceil(H, data_num)) + (w / BANK_NUM) * (m_ceil(H, data_num)) + h / data_num;
[0054] C. When the data placement mode of the off-chip memory is NCHW and the data placement method in the L2_buffer is row priority mode, the address calculation formula is as follows:
[0055] bank_num = ((n * C * H + c * H + h) * (m_ceil(W, data_num)) + w / data_num) % BANK_NUM;
[0056] bank_addr = start_addr + ((n * C * H + c * H + h) * (m_ceil(W, data_num)) + w / data_num) / BANK_NUM;
[0057] D. When the data layout mode in the off-chip memory is NHWC and the data layout mode in the L2_buffer is bank-channel priority followed by column priority, the address calculation formula is as follows:
[0058] bank_num = w / (m_ceil(W, BANK_NUM));
[0059] bank_addr = start_addr + n * (m_ceil(W, BANK_NUM)) * H * (m_ceil(C, data_num)) + (w % (m_ceil(W, BANK_NUM))) * H * (m_ceil(C, data_num)) + h * (m_ceil(C, data_num)) + c / data_num;
[0060] In the above four address calculation formulas, m_ceil(a, b) is the m_ceil function, which is defined as judging whether the result after the modulo operation of a and b is 0. If it is 0, the quotient value is taken; otherwise, the quotient value plus 1 is taken. start_addr is the starting address, BANK_NUM is the number of banks in the L2_buffer, data_num is the number of data that can be stored in one row of a bank in the L2_buffer, which is related to the data type to be transmitted. N, C, H, and W are the number of feature maps, the number of channels of the feature maps, the height of the feature maps, and the width of the feature maps respectively. n, c, h, and w are the per-clock-cycle counts of the number of feature maps, the number of channels of the feature maps, the height of the feature maps, and the width of the feature maps respectively. bank_num is the current bank number, and the value range of bank_num is 0 to 15. The [11:10] bits in bank_addr are used to represent the block number, 00: block0; 01: block1; 10: block2; 11: block3. The [9:0] bits in bank_addr are used to represent the depth address of the bank.
[0061] Preferably, the data stored in the off-chip memory is a four-dimensional tensor, and the data layout modes of the four-dimensional tensor in the off-chip memory include NCHW and NHWC;
[0062] The data placement methods in the L2_buffer include row bank - first mode, column bank - first mode, column - first mode after bank - channel first, and row - first mode:
[0063] In the row bank - first mode, the data of the feature map height is stored in one bank. Since the bank bit - width is large, multiple data can be stored inside the depth of one bank. Therefore, the data of each feature map height is folded by P bytes and placed in one bank. The value of P is related to the bank bit - width and the data type. When supporting three data types of 8bit, 16bit, and 32bit, when the data bit - width is 16bit and the bank bit - width is 128bit, 8 data can be stored inside the depth of one bank. For this case, in the row bank - first mode, the first 8 data of the feature map height are stored in one bank first, then the second batch of 8 data of the feature map height are stored, and then the third batch of 8 data of the feature map height are stored until all the data of the feature map height are stored. In the adjacent next bank, the data of the next feature map height are stored in the same way as the previous bank. After the feature maps of one channel are placed, the feature maps of the next channel start to be placed from the starting boundary of bank0;
[0064] In the column bank - first mode, the data of each feature map width is folded by P bytes and placed in one bank. After the feature maps of one channel are placed, the feature maps of the next channel start to be placed from the starting boundary of bank0;
[0065] In the column - first mode after bank - channel first, let the feature map width be FC and the number of banks be R. Then each bank can store at most column data, and only the first banks store column data, and the next bank stores the remaining column data. Each batch of data starts to be placed from the starting boundary of bank0. is the ceiling function, is the floor function;
[0066] In the row - first mode, the height dimension of the four - dimensional tensor is arranged first, followed by the width dimension, then the channel dimension, and finally the batch dimension. If the data of the feature map height does not end at the bank boundary, 0 is filled to the bank boundary. The data of the next feature map height starts to be placed from the starting boundary of the next bank. After the feature maps of one channel are placed, the feature maps of the next channel start to be placed from the starting boundary of bank0;
[0067] Among them, when the data layout mode of the off-chip memory is NCHW, the data layout mode in the L2_buffer is row bank priority mode or column bank priority mode or row priority mode; when the data layout mode of the off-chip memory is NHWC, the data layout mode in the L2_buffer is column priority mode after bank channel priority.
[0068] Preferably, the functions of the data processing module include:
[0069] Padding zeros at the bank boundary before the data is written into the L2_buffer:
[0070] A. When the data layout mode of the off-chip memory is NCHW:
[0071] 1) When the width of the feature map can divide the data volume evenly, all the data transmitted in one dimension is valid and no padding is required;
[0072] 2) When the width of the feature map cannot divide the data volume evenly, padding operation needs to be performed on the last group of data transmitted in one dimension. The invalid data volume is equal to the total data volume transmitted in one dimension minus the width of the feature map, and then the bit positions to be cleared are confirmed according to the current data type;
[0073] B. When the data layout mode of the off-chip memory is NHWC:
[0074] 1) When the number of channels of the feature map can divide the data volume evenly, all the data transmitted in one dimension is valid and no padding is required;
[0075] 2) When the number of channels of the feature map cannot divide the data volume evenly, padding operation needs to be performed on the last group of data transmitted in one dimension. The invalid data volume is equal to the total data volume transmitted in one dimension minus the number of channels of the feature map, and then the bit positions to be cleared are confirmed according to the current data type;
[0076] Write the data that does not need padding and has been cleared into the L2_buffer according to the calculated address.
[0077] Preferably, the functions of the exception judgment module include:
[0078] There are two types of abnormal situations. One is that the address exceeds the range, an address out-of-bounds exception; the other is that two operations simultaneously hit the same bank;
[0079] For the bank of the single-port SRAM, only one read operation or one write operation is supported per clock cycle. The read and write operations cannot be performed simultaneously. It is necessary to compare the addresses of each clock cycle of the read and write operations. If the addresses are the same, an exception occurs. The exception judgment module sends an interrupt to the CPU, and the CPU controls a soft reset. The instructions in the L2_buffer will not be executed further, and the scene is preserved to facilitate the later drive to read and write the register to locate the cause.
[0080] Preferably, the size of the storage module is 1MB, which is divided into 4 blocks in total. Each block is divided into 16 banks. Each bank is a single-port SRAM with a width of 128 bits and a depth of 1024. Each bank can read or write data of 16 bytes of the same depth within one clock cycle.
[0081] (III) Beneficial effects
[0082] Compared with the prior art, an on-chip large-capacity cache provided by the present invention has the following beneficial effects:
[0083] 1) By adopting the two-dimensional DMA transmission mode, the data transmission bandwidth of the custom bus is maximally utilized. After data processing and address calculation of the data in the off-chip memory placed in two data placement modes, the data is stored in the L2_buffer according to four data placement methods, preparing data in advance for the first-level cache L1_buffer, ensuring the efficient supply of data to the computing components, greatly improving the data calculation efficiency, and effectively solving the problem that a large amount of time and resources are consumed by data exchange due to the small storage space of the first-level cache L1_buffer;
[0084] 2) Since the data transmission rate of the L2_buffer matches the data transmission rate of the DMA, during the write operation, as long as the DMA data is valid, it can be immediately written into the L2_buffer; during the read operation, as long as the DMA is idle and the write operation is allowed, and at the same time the L2_buffer can continuously transmit data outward;
[0085] 3) Avoid using multipliers and dividers to calculate addresses, effectively improving the address calculation efficiency. BRIEF DESCRIPTION OF THE DRAWINGS
[0086] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0087] Figure 1Schematic diagram of the hardware structure of the on-chip large-capacity cache L2_buffer in the present invention;
[0088] Figure 2 Schematic diagram of the working process of the on-chip large-capacity cache L2_buffer in the present invention;
[0089] Figure 3 State transition diagram of the instruction decoding module in the present invention;
[0090] Figure 4 Interface diagram of the instruction decoding module in the present invention;
[0091] Figure 5 State transition diagram when the off-chip memory is read by the data interaction interface module in the present invention;
[0092] Figure 6 State transition diagram when the off-chip memory is written by the data interaction interface module in the present invention;
[0093] Figure 7 Schematic diagram of the data layout mode of NCHW for the off-chip memory in the present invention;
[0094] Figure 8 Schematic diagram of the data layout mode of NHWC for the off-chip memory in the present invention;
[0095] Figure 9 Schematic diagram of the data layout in the L2_buffer with the row-by-bank priority mode in the present invention;
[0096] Figure 10 Schematic diagram of the data layout in the L2_buffer with the column-by-bank priority mode in the present invention;
[0097] Figure 11 Schematic diagram of the data layout in the L2_buffer with the bank-channel priority followed by column priority mode in the present invention;
[0098] Figure 12 Schematic diagram of the data layout in the L2_buffer with the row priority mode in the present invention;
[0099] Figure 13 Schematic diagram of the working process of the exception judgment module in the present invention;
[0100] Figure 14 Schematic diagram of the working process when the on-chip large-capacity cache L2_buffer reads the off-chip memory in the present invention;
[0101] Figure 15This is a schematic diagram of the working process when the on-chip large-capacity cache L2_buffer in the present invention performs a write operation on the off-core memory. Detailed implementation manners
[0102] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Apparently, the described embodiments are some but not all of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0103] An on-chip large-capacity cache applied to a deep learning acceleration core, such as Figure 1 shown, includes an on-chip large-capacity cache L2_buffer with a hierarchical design located between the first-level cache L1_buffer and the off-core memory. L2_buffer includes an instruction decoding module, a data interaction interface module, an address processing module, a data processing module, an exception judgment module, and a storage module;
[0104] As Figure 2 shown, when L2_buffer works, the instruction control unit sends L2_buffer read / write instructions to the secondary decoding buffer queue, and the external interface depend / inform unit in the instruction decoding module controls the instruction execution process;
[0105] If the received instruction is an illegal instruction, an exception is reported, and the L2_buffer read / write instruction sent by the instruction control unit is received again;
[0106] If the received instruction is a legal instruction and there is no dependency relationship, the instruction is executed sequentially, and a read / write request to the DMA is initiated;
[0107] If the received instruction is a legal instruction and there is a dependency relationship, after waiting for the begin_inform signal returned by the external interface depend / inform unit, the next operation is continued, and a read / write request to the DMA is initiated;
[0108] Among them, if the above read / write process is completed normally, it enters the idle state, otherwise it enters the idle state after reporting an exception.
[0109] ① As Figure 4 shown, the functions of the instruction decoding module include:
[0110] The L2_buffer receives the instruction signal sent by the instruction control unit, determines whether it is an L2_buffer read / write instruction according to the lower 5 bits, puts the L2_buffer read instruction into the load_fifo, and puts the L2_buffer write instruction into the store_fifo;
[0111] Among them, both the load_fifo and the store_fifo are synchronous fifos with a data width of 138 bits and a depth of 16.
[0112] As Figure 3 shown, the state machine of the instruction decoding module includes:
[0113] When the fifo is not empty and the instructions in the fifo have not all been executed, the current state enters the waiting state from the idle state;
[0114] When in the waiting state, the instructions are read out from the fifo, and it is judged whether there is a dependency relationship according to the higher 5 bits. When there is no dependency relationship or the request response signal returned by the external interface depend / inform unit is received, the current state enters the decoding state from the waiting state;
[0115] When the number of feature maps obtained by decoding is not 0, the current state enters the execution state from the decoding state;
[0116] Among them, when the current instruction enters the execution state, the next instruction sends the dependency relationship to the external interface depend / inform unit. After the current instruction is executed, when the instruction decoding module receives the request response signal returned by the external interface depend / inform unit, the next instruction enters the decoding state;
[0117] The write instruction completion flag is that all data is written and the end flag bit flag is pulled high;
[0118] The read instruction completion flag is judged according to the resp returned by the DMA. When the first bit of the resp is 1, it means the read instruction is completed.
[0119] ② The functions of the data interaction interface module include:
[0120] The L2_buffer performs read and write operations on the off-chip memory through the read and write ports of the custom on-chip bus protocol master. The L2_buffer and the DMA adopt a two-dimensional DMA transmission method. The control word consists of a start field, the start address transferred to the target unit, the data length transmitted in the x direction, the number of 1D DMA transmissions in the y direction, the start address deviation between two 1D DMA transmissions, and a read / write control field, and the control word is sent to the DMA;
[0121] The L2_buffer is responsible for processing the read and write access requests of the L1_buffer. The access between the L2_buffer and the L1_buffer supports multi-address simultaneous access based on the control segment. A custom bus protocol maintains the control segment and data segment in the form of data packets and performs physical transmission based on the custom bus port. A two-dimensional transmission method is also adopted between the L2_buffer and the L1_buffer. The L2_buffer, as a custom bus slave interface, receives the control word signal of the L1_buffer, parses the control word signal, and performs read and write operations on the storage units inside the L2_buffer.
[0122] Among them, the start field is generated by notifying the DMA to establish a connection. The data transmission length in the x direction is the one-dimensional length, and the number of 1D DMA transmissions in the y direction is the two-dimensional length.
[0123] A. When the data layout mode of the off-chip memory is NCHW (as shown in Figure 7 ), the one-dimensional length is the number of times the current feature map width transmission is completed. Whether the feature map width can divide the data volume determines the one-dimensional length, following the ceiling principle. The calculation formula for the one-dimensional length is as follows:
[0124] 1) When the feature map width can divide the data volume:
[0125] One-dimensional length = feature map width / data volume;
[0126] 2) When the feature map width cannot divide the data volume:
[0127] One-dimensional length = feature map width / data volume + 1;
[0128] In the above two calculation formulas for the one-dimensional length, the data volume is the number of data that can be stored in one row of a bank in the L2_buffer, which is related to the data type being transmitted. When the bank bit width is 128bit, the data volume corresponding to the 8bit data type is 16, the data volume corresponding to the 16bit data type is 8, and the data volume corresponding to the 32bit data type is 4.
[0129] The two-dimensional length is the number of times the one-dimensional length data transmission is completed, and the address step size is equal to the feature map width. The calculation formula for the two-dimensional length is as follows:
[0130] Two-dimensional length = number of feature maps * number of feature map channels * feature map height;
[0131] B. When the data layout mode of the off-chip memory is NHWC (as shown in Figure 8 ), whether the number of feature map channels can divide the data volume determines the one-dimensional length, following the ceiling principle. The calculation formula for the one-dimensional length is as follows:
[0132] 1) The number of channels of the feature map can divide the data volume:
[0133] One-dimensional length = number of channels of the feature map / data volume;
[0134] 2) The number of channels of the feature map cannot divide the data volume:
[0135] One-dimensional length = number of channels of the feature map / data volume + 1;
[0136] In the above two one-dimensional length calculation formulas, the data volume is the number of data that can be stored in one row of a bank in L2_buffer, which is related to the data type being transmitted. When the bank bit width is 128bit, the data volume corresponding to the 8bit data type is 16, the data volume corresponding to the 16bit data type is 8, and the data volume corresponding to the 32bit data type is 4;
[0137] The two-dimensional length is the number of data transmission completion times of the one-dimensional length, the address step size is equal to the number of channels of the feature map, and the calculation formula of the two-dimensional length is as follows:
[0138] Two-dimensional length = number of feature maps * feature map height * feature map width.
[0139] The state machine of the data interaction interface module includes:
[0140] The start state is the idle state. When the control word is spliced well, decoding is valid and it enters the waiting state;
[0141] When in the waiting state, the control word is sent to the DMA. As Figure 6 shown, when writing to the DMA operation, the write enable signal of the L2_buffer write DMA is pulled high. When the wready returned by the DMA is 1, it means the DMA has received the control word, and L2_buffer sends the corresponding length of data to the DMA. When the first bit of the resp returned by the DMA is 1, it means the current DMA transmission ends; As Figure 5 shown, when reading from the DMA operation, the read request signal of the L2_buffer read DMA is pulled high. When the 0th bit of the resp returned by the DMA is 0, it means the DMA has received the control word and sends the corresponding length of data to L2_buffer. When the 0th bit of the resp returned by the DMA is 1, it means the current DMA transmission ends.
[0142] ③ The address calculation of the address processing module adopts an incremental calculation method. According to the different data placement modes of the off-chip memory and the different data placement methods in L2_buffer, the corresponding address calculation formula is adopted:
[0143] A. When the data placement mode of the off-chip memory is NCHW (as Figure 7as shown), and when the data arrangement pattern in L2_buffer is row bank - priority mode (such as Figure 9 as shown), the address calculation formula is as follows:
[0144] bank_num = h % BANK_NUM;
[0145] bank_addr = start_addr+(n * C + c)*(m_ceil(H,BANK_NUM))*(m_ceil(W,data_num))+(h / BANK_NUM)*(m_ceil(W,data_num))+w / data_num;
[0146] B. When the data arrangement pattern of the off - core memory is NCHW (such as Figure 7 as shown), and the data arrangement pattern in L2_buffer is column bank - priority mode (such as Figure 10 as shown), the address calculation formula is as follows:
[0147] bank_num = w % BANK_NUM;
[0148] bank_addr = start_addr+(n * C + c)*(m_ceil(W,BANK_NUM))*(m_ceil(H,data_num))+(w / BANK_NUM)*(m_ceil(H,data_num))+h / data_num;
[0149] C. When the data arrangement pattern of the off - core memory is NCHW (such as Figure 7 as shown), and the data arrangement pattern in L2_buffer is row - priority mode (such as Figure 12 as shown), the address calculation formula is as follows:
[0150] bank_num = ((n * C * H + c * H + h)*(m_ceil(W,data_num))+w / data_num) % BANK_NUM;
[0151] bank_addr = start_addr+((n * C * H + c * H + h)*(m_ceil(W,data_num))+w / d ata_num) / BANK_NUM;
[0152] D. When the data arrangement pattern of the off - core memory is NHWC (such as Figure 8 as shown), and the data arrangement pattern in L2_buffer is column - priority mode after bank - channel - priority (such as Figure 11 as shown), the address calculation formula is as follows:
[0153] bank_num = w / (m_ceil(W, BANK_NUM));
[0154] bank_addr = start_addr + n * (m_ceil(W, BANK_NUM)) * H * (m_ceil(C, data_num)) + (w % (m_ceil(W, BANK_NUM))) * H * (m_ceil(C, data_num)) + h * (m_ceil(C, data_num)) + c / data_num;
[0155] In the above four address calculation formulas, m_ceil(a, b) is the m_ceil function, which is defined as judging whether the result of the modulo operation on a and b is 0. If it is 0, the quotient value is taken; otherwise, the quotient value plus 1 is taken. start_addr is the starting address, BANK_NUM is the number of banks in the L2_buffer, data_num is the number of data that can be stored in a row of a bank in the L2_buffer, which is related to the data type to be transmitted. N, C, H, and W are the number of feature maps, the number of channels of the feature map, the height of the feature map, and the width of the feature map respectively. n, c, h, and w are the per-clock-cycle counts of the number of feature maps, the number of channels of the feature map, the height of the feature map, and the width of the feature map respectively. bank_num is the current bank number, and the value range of bank_num is 0 to 15. The [11:10] bits in bank_addr are used to represent the block number, 00: block0; 01: block1; 10: block2; 11: block3. The [9:0] bits in bank_addr are used to represent the depth address of the bank.
[0156] In the technical solution of this application, multiplication is divided into two forms. One is that the multiplicand is a power of 2, and the other is that the multiplicand is an unsigned number that is not fixed. When the multiplicand is a power of 2, in the hardware language, the multiplication result is obtained by shifting the multiplier to the left by the corresponding exponent. When the multiplicand is an unsigned number that is not fixed, the multiplication calculation is implemented by accumulating the counter. First, according to the characteristics of the m_ceil function, the bank_addr calculation formula in each mode is subdivided into two forms: divisible and non-divisible using the multiplication distribution law, and the bank_addr calculation formula is split. Then, according to the characteristics of the address increment in each mode, an internal register is used to temporarily store the accumulated result of the previous cycle, and the accumulated result of the counter in each clock cycle is calculated in an incremental manner, which is the current multiplication result.
[0157] The modulo operation calculates the remainder of the division of two numbers. In hardware language, when the dividend is a power of 2, the modulo result is the [exponent - 1:0] bits of the divisor. In the technical solution of this application, BANK_NUM is 16, which is the 4th power of 2, and the exponent is 4. Then the result of h % BANK_NUM is the value corresponding to the [3:0] bits of h.
[0158] In the division operation in hardware language, when the dividend is a power of 2, the division result is the value obtained by shifting the divisor to the right by the exponent. In the technical solution of this application, when data_num is 16, the corresponding exponent is 4, then the result of w / data_num is the value of w >> 4; when data_num is 32, the corresponding exponent is 5, then the result of w / data_num is the value of w >> 5.
[0159] According to the above hardware calculation method, calculate the multiplication results corresponding to the changes of the counters n, c, h, and w in each clock cycle, add them to the modulo results and division results to obtain the bank step address when the counter changes in each clock cycle. Finally, add it to the starting address start_addr to obtain the bank depth address of the current clock cycle, and at the same time, the bank number of the current clock cycle can also be calculated to perform point-to-point read and write operations on the L2_buffer.
[0160] In the technical solution of this application, the data stored in the off-chip memory is a four-dimensional tensor, and the data layout modes of the four-dimensional tensor in the off-chip memory include NCHW and NHWC;
[0161] The data layout methods in the L2_buffer include row bank - first mode, column bank - first mode, column - first mode after bank - channel - first mode, and row - first mode:
[0162] The row bank - first mode, such as Figure 9As shown in the figure, the data of the height of the feature map is stored in one bank. Since the bank bit width is relatively large and multiple data can be stored inside one bank depth, the data of each feature map height is folded into P bytes and placed in one bank. The value of P is related to the bank bit width and data type. When supporting three data types of 8bit, 16bit, and 32bit, when the data bit width is 16bit and the bank bit width is 128bit, 8 data can be stored inside one bank depth. For this case, in the row-by-bank priority mode, the first 8 data of the feature map height are stored in one bank first, then the second batch of 8 data of the feature map height are stored, and then the third batch of 8 data of the feature map height are stored until all the data of the feature map height are stored. In the next adjacent bank, the data of the next feature map height are stored in the same way as the previous bank. After the feature maps of one channel are placed, the feature maps of the next channel start to be placed from the starting boundary of bank0;
[0163] Column-by-bank priority mode, as Figure 10 shown in the figure, the data of each feature map width is folded into P bytes and placed in one bank. After the feature maps of one channel are placed, the feature maps of the next channel start to be placed from the starting boundary of bank0;
[0164] After bank-channel priority and then column priority mode, as Figure 11 shown in the figure, denote the feature map width as FC and the number of banks as R. Then each bank can store at most column data, and only the first banks store column data, and the next bank stores the remaining column data. Each batch of data starts to be placed from the starting boundary of bank0, is the ceiling function, is the floor function;
[0165] Row priority mode, as Figure 12 shown in the figure, the height dimension of the four-dimensional tensor is arranged first, followed by the width dimension, then the channel dimension, and finally the batch dimension. If the data of the feature map height does not end at the bank boundary, 0 is filled to the bank boundary. The data of the next feature map height starts to be placed from the starting boundary of the next bank. After the feature maps of one channel are placed, the feature maps of the next channel start to be placed from the starting boundary of bank0;
[0166] Among them, when the data layout mode of the off-chip memory is NCHW, the data layout in the L2_buffer is the row bank-priority mode or the column bank-priority mode or the row-priority mode; when the data layout mode of the off-chip memory is NHWC, the data layout in the L2_buffer is the column-priority mode after bank-channel priority.
[0167] ④ The functions of the data processing module include:
[0168] Zero-padding at the bank boundary is performed before the data is written into the L2_buffer:
[0169] A. When the data layout mode of the off-chip memory is NCHW (as Figure 7 shown):
[0170] 1) When the width of the feature map can divide the data volume evenly, all the data transmitted in one dimension is valid and no zero-padding is required;
[0171] 2) When the width of the feature map cannot divide the data volume evenly, zero-padding operation needs to be performed on the last group of data transmitted in one dimension. The invalid data volume is equal to the total data volume transmitted in one dimension minus the width of the feature map, and then the bit positions to be cleared are confirmed according to the current data type;
[0172] B. When the data layout mode of the off-chip memory is NHWC (as Figure 8 shown):
[0173] 1) When the number of channels of the feature map can divide the data volume evenly, all the data transmitted in one dimension is valid and no zero-padding is required;
[0174] 2) When the number of channels of the feature map cannot divide the data volume evenly, zero-padding operation needs to be performed on the last group of data transmitted in one dimension. The invalid data volume is equal to the total data volume transmitted in one dimension minus the number of channels of the feature map, and then the bit positions to be cleared are confirmed according to the current data type;
[0175] The data that does not require zero-padding and has been cleared is written into the L2_buffer according to the calculated address.
[0176] ⑤ As Figure 13 shown, the functions of the exception judgment module include:
[0177] There are two types of exception situations. One is that the address exceeds the range, an address out-of-bounds exception; the other is that two operations simultaneously hit the same bank;
[0178] For the bank of the single-port SRAM, only one read operation or one write operation is supported per clock cycle. The read and write operations cannot be performed simultaneously. It is necessary to compare the addresses of each clock cycle of the read and write operations. If the addresses are the same, an exception occurs. The exception judgment module sends an interrupt to the CPU, and the CPU controls a soft reset. The instructions in the L2_buffer do not execute further, and the scene is preserved to facilitate later driving the read and write registers to locate the cause.
[0179] ⑥ The size of the storage module is 1MB, which is divided into 4 blocks in total. Each block is divided into 16 banks. Each bank is a single-port SRAM with a width of 128 bits and a depth of 1024. Each bank can read or write 16 bytes of data with the same depth within one clock cycle. The working processes of the on-chip large-capacity buffer L2_buffer during read and write operations on off-chip memories are respectively as Figure 14 、 Figure 15 shown.
[0180] The above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements will not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. An on-chip large-capacity cache applied to a deep learning acceleration core, characterized in that: It includes an on-chip large-capacity buffer L2_buffer with a hierarchical design located between the first-level cache L1_buffer and off-chip memory. The L2_buffer includes an instruction decoding module, a data interaction interface module, an address processing module, a data processing module, an exception judgment module, and a storage module; When the L2_buffer works, the instruction control unit sends L2_buffer read / write instructions to the secondary decoding buffer queue, and the external interface depend / inform unit in the instruction decoding module controls the instruction execution process; If the received instruction is an illegal instruction, an exception is reported and the L2_buffer read / write instruction sent by the instruction control unit is received again; If the received instruction is a legal instruction and there is no dependency relationship, the instruction is executed sequentially and a read / write request to the DMA is initiated; If the received instruction is a legal instruction and there is a dependency relationship, after waiting for the begin_inform signal returned by the external interface depend / inform unit, the next operation is continued and a read / write request to the DMA is initiated; Among them, if the above read / write process is completed normally, it enters the idle state, otherwise it enters the idle state after reporting an exception.
2. The on-chip large-capacity cache applied to the deep learning acceleration core according to claim 1, wherein: The functions of the instruction decoding module include: The L2_buffer receives the instruction signal sent by the instruction control unit, judges whether it is an L2_buffer read / write instruction according to the lower 5 bits, puts the L2_buffer read instruction into the load_fifo, and puts the L2_buffer write instruction into the store_fifo; Among them, both the load_fifo and the store_fifo are synchronous fifos with a data width of 138 bits and a depth of 16.
3. The on-chip large-capacity cache applied to the deep learning acceleration core according to claim 2, wherein: The state machine of the instruction decoding module includes: When the fifo is not empty and the instructions in the fifo have not all been executed, the current state enters the waiting state from the idle state; When in the waiting state, the instruction is read out from the fifo, and it is judged whether there is a dependency relationship according to the higher 5 bits. When there is no dependency relationship or the request response signal returned by the external interface depend / inform unit is received, the current state enters the decoding state from the waiting state; When the number of feature maps obtained by decoding is not 0, the current state enters the execution state from the decoding state; Among them, when the current instruction enters the execution state, the next instruction sends the dependency relationship to the external interface depend / inform unit. After the current instruction is executed, when the instruction decoding module receives the request response signal returned by the external interface depend / inform unit, the next instruction enters the decoding state; The write instruction completion flag is that all data is written and the end flag bit flag is pulled high; The read instruction completion flag is judged according to the resp returned by the DMA. When the first bit of the resp is 1, it means the read instruction is completed.
4. The on-chip large-capacity cache applied to the deep learning acceleration core according to claim 1, wherein: The functions of the data interaction interface module include: The L2_buffer performs read and write operations on off-chip memory through the read and write ports of the custom on-chip bus protocol master. A two-dimensional DMA transfer mode is adopted between the L2_buffer and the DMA. The control word consists of a start field, the start address transferred to the target unit, the data transfer length in the x direction, the number of 1D DMA transfer times in the y direction, the start address deviation between two 1D DMA transfers, and a read / write control field, and the control word is sent to the DMA; The L2_buffer is responsible for processing the read and write access requests of the L1_buffer. The access between the L2_buffer and the L1_buffer supports multi-address simultaneous access based on the control segment. The custom bus protocol maintains the control segment and the data segment in the form of data packets and performs physical transmission based on the custom bus port. A two-dimensional transfer mode is also adopted between the L2_buffer and the L1_buffer. The L2_buffer, as the custom bus slave interface, receives the control word signal of the L1_buffer, parses the control word signal, and performs read and write operations on the storage units inside the L2_buffer; Among them, the start field is generated by notifying the DMA to establish a connection. The data transfer length in the x direction is the one-dimensional length, and the number of 1D DMA transfer times in the y direction is the two-dimensional length; A. When the data layout mode of the off-chip memory is NCHW, the one-dimensional length is the number of times the current feature map width transfer is completed. Whether the feature map width can divide the data volume determines the one-dimensional length, following the principle of rounding up. The calculation formula for the one-dimensional length is as follows: 1) The feature map width can divide the data volume: One-dimensional length = feature map width / data volume; 2) The feature map width cannot divide the data volume: One-dimensional length = feature map width / data volume + 1; In the above two calculation formulas for the one-dimensional length, the data volume is the number of data that can be stored in a row of a bank in the L2_buffer, which is related to the data type to be transferred. When the bank bit width is 128bit, the data volume corresponding to the 8bit data type is 16, the data volume corresponding to the 16bit data type is 8, and the data volume corresponding to the 32bit data type is 4; The two-dimensional length is the number of times the one-dimensional length data transfer is completed. The address step size is equal to the feature map width. The calculation formula for the two-dimensional length is as follows: Two-dimensional length = number of feature maps * number of feature map channels * feature map height; B. When the data layout mode of the off-chip memory is NHWC, whether the number of feature map channels can divide the data volume determines the one-dimensional length, following the principle of rounding up. The calculation formula for the one-dimensional length is as follows: 1) The number of feature map channels can divide the data volume: One-dimensional length = number of feature map channels / data volume; 2) The number of feature map channels cannot divide the data volume: One-dimensional length = number of feature map channels / data volume + 1; In the above two one-dimensional length calculation formulas, the data volume is the number of data that can be stored in one row of a bank in L2_buffer, which is related to the data type being transmitted. When the bank bit width is 128bit, for an 8bit data type, the corresponding data volume is 16; for a 16bit data type, the corresponding data volume is 8; for a 32bit data type, the corresponding data volume is 4. The two-dimensional length is the number of data transmission completion times of the one-dimensional length, and the address step size is equal to the number of feature map channels. The calculation formula for the two-dimensional length is as follows: Two-dimensional length = number of feature maps * feature map height * feature map width.
5. The on-chip large-capacity cache applied to the deep learning acceleration core according to claim 4, wherein: The state machine of the data interaction interface module includes: The start state is the idle state. When the control word is spliced, the decoding is valid and it enters the waiting state; When in the waiting state, the control word is sent to the DMA. When writing to the DMA, the write enable signal of the L2_buffer write DMA is pulled high. When wready returned by the DMA is 1, it means the DMA has received the control word, and L2_buffer sends the corresponding length of data to the DMA. When the first bit of resp returned by the DMA is 1, it means the current DMA transmission ends; When reading from the DMA, the read request signal of the L2_buffer read DMA is pulled high. When the 0th bit of resp returned by the DMA is 0, it means the DMA has received the control word and sends the corresponding length of data to L2_buffer. When the 0th bit of resp returned by the DMA is 1, it means the current DMA transmission ends.
6. The on-chip large-capacity cache applied to the deep learning acceleration core according to claim 1, wherein: The address calculation of the address processing module adopts an incremental calculation method. According to the different data placement modes of the off-chip memory and the different data placement methods in L2_buffer, the corresponding address calculation formulas are used: A. When the data placement mode of the off-chip memory is NCHW and the data placement method in L2_buffer is row bank priority mode, the address calculation formula is as follows: bank_num = h % BANK_NUM; bank_addr = start_addr + (n * C + c) * (m_ceil(H, BANK_NUM)) * (m_ceil(W, data_num)) + (h / BANK_NUM) * (m_ceil(W, data_num)) + w / data_num; B. When the data placement mode of the off-chip memory is NCHW and the data placement method in L2_buffer is column bank priority mode, the address calculation formula is as follows: bank_num = w % BANK_NUM; bank_addr = start_addr + (n * C + c) * (m_ceil(W, BANK_NUM)) * (m_ceil(H, data_num)) + (w / BANK_NUM) * (m_ceil(H, data_num)) + h / data_num; C. When the data layout mode of the off-chip memory is NCHW and the data layout mode in the L2_buffer is row-major mode, the address calculation formula is as follows: bank_num = ((n * C * H + c * H + h) * (m_ceil(W, data_num)) + w / data_num) % BANK_NUM; bank_addr = start_addr + ((n * C * H + c * H + h) * (m_ceil(W, data_num)) + w / data_num) / BANK_NUM; D. When the data layout mode of the off-chip memory is NHWC and the data layout mode in the L2_buffer is column-major mode after bank-channel priority, the address calculation formula is as follows: bank_num = w / (m_ceil(W, BANK_NUM)); bank_addr = start_addr + n * (m_ceil(W, BANK_NUM)) * H * (m_ceil(C, data_num)) + (w % (m_ceil(W, BANK_NUM))) * H * (m_ceil(C, data_num)) + h * (m_ceil(C, data_num)) + c / data_num; In the above four address calculation formulas, m_ceil(a, b) is the m_ceil function, which is defined as judging whether the result of the modulo operation of a and b is 0. If it is 0, the quotient value is taken; otherwise, the quotient value plus 1 is taken. start_addr is the starting address, BANK_NUM is the number of banks in the L2_buffer, data_num is the number of data that can be stored in one row of a bank in the L2_buffer, which is related to the data type to be transmitted. N, C, H, and W are the number of feature maps, the number of feature map channels, the height of the feature map, and the width of the feature map respectively. n, c, h, and w are the per-clock-cycle counts of the number of feature maps, the number of feature map channels, the height of the feature map, and the width of the feature map respectively. bank_num is the current bank number, and the value range of bank_num is 0 to 15. The [11:10] bits in bank_addr are used to represent the block number: 00: block0; 01: block1; 10: block2; 11: block3. The [9:0] bits in bank_addr are used to represent the depth address of the bank.
7. The on-chip large-capacity cache applied to the deep learning acceleration core according to claim 6, wherein: The data stored in the off-chip memory is a four-dimensional tensor, and the data layout modes of the four-dimensional tensor in the off-chip memory include NCHW and NHWC; The data layout modes in the L2_buffer include row-bank priority mode, column-bank priority mode, column-major mode after bank-channel priority, and row-major mode: In the row bank - first mode, the data of the feature map height is stored within one bank. Since the bank has a large bit - width and can store multiple data within one bank depth, the data of each feature map height is folded by P bytes and placed in one bank. The value of P is related to the bank bit - width and data type. In the case of supporting three data types: 8bit, 16bit, and 32bit, when the data bit - width is 16bit and the bank bit - width is 128bit, 8 data can be stored within one bank depth. For this case, in the row bank - first mode, the first 8 data of the feature map height are stored in one bank first, then the second batch of 8 data of this feature map height are stored, and then the third batch of 8 data of this feature map height are stored until all the data of this feature map height are stored. In the next adjacent bank, the data of the next feature map height are stored in the same way as the previous bank. After the feature maps of one channel are placed, the feature maps of the next channel start to be placed from the starting boundary of bank0; In the column bank - first mode, the data of each feature map width is folded by P bytes and placed in one bank. After the feature maps of one channel are placed, the feature maps of the next channel start to be placed from the starting boundary of bank0; According to the bank channel priority and then column priority mode, if the width of the feature map is FC and the number of banks is R, then each bank can store at most column data, and only the first banks store column data, and the next bank stores the remaining column data. Each batch of data starts to be placed at the starting boundary of bank0. is the ceiling function, is the floor function; In the row - first mode, the height dimension of the four - dimensional tensor is arranged first, followed by the width dimension, then the channel dimension, and finally the batch dimension. If the data of the feature map height does not end at the bank boundary, 0s are filled to the bank boundary. The data of the next feature map height starts to be placed from the starting boundary of the next bank. After the feature maps of one channel are placed, the feature maps of the next channel start to be placed from the starting boundary of bank0; Among them, when the data placement mode of the off - chip memory is NCHW, the data placement mode in the L2_buffer is row bank - first mode or column bank - first mode or row - first mode; when the data placement mode of the off - chip memory is NHWC, the data placement mode in the L2_buffer is bank - channel - first and then column - first mode.
8. The on-chip large-capacity cache applied to the deep learning acceleration core according to claim 1, characterized in that: The functions of the data processing module include: Zero padding at the bank boundary before data is written into the L2_buffer: A. When the data placement mode of the off - chip memory is NCHW: 1) When the feature map width can be divided evenly by the data volume, all the data in one - dimensional transmission are valid and no zero padding is required; 2) When the feature map width cannot be divided evenly by the data volume, zero padding operation needs to be performed on the last group of data in one - dimensional transmission. The invalid data volume is equal to the total data volume in one - dimensional transmission minus the feature map width, and then the bit positions to be cleared are confirmed according to the current data type; B. When the data placement mode of the off - chip memory is NHWC: 1) When the number of feature map channels can be divided evenly by the data volume, all the data in one - dimensional transmission are valid and no zero padding is required; 2) When the number of channels of the feature map cannot divide the data volume evenly, it is necessary to perform zero-padding on the last group of data transmitted in one dimension. The invalid data volume is equal to the total data volume transmitted in one dimension minus the number of channels of the feature map, and then the bit positions to be cleared are determined according to the current data type; Write the data that does not need to be zero-padded and cleared as described above to the L2_buffer according to the calculated address.
9. The on-chip large-capacity cache applied to the deep learning acceleration core according to claim 1, wherein: The functions of the abnormal judgment module include: There are two types of abnormal situations. One is that the address exceeds the range, an address out-of-bounds exception; the other is that two operations simultaneously hit the same bank; For the bank of the single-port SRAM, only one read operation or one write operation is supported per clock cycle, and the read and write operations cannot be performed simultaneously. It is necessary to compare the addresses of each clock cycle of the read and write operations. If the addresses are the same, an exception occurs. The abnormal judgment module sends an interrupt to the CPU, and the CPU controls a soft reset. The instructions in the L2_buffer will not be executed further, and the scene is saved to facilitate later driving the read and write registers to locate the cause.
10. The on-chip large-capacity cache applied to the deep learning acceleration core according to any one of claims 1 to 9, wherein: The size of the storage module is 1MB, which is divided into 4 blocks in total. Each block is divided into 16 banks. Each bank is a single-port SRAM with a width of 128 bits and a depth of 1024. Each bank can read or write 16 bytes of data with the same depth within one clock cycle.