Reading organization optimization method for accelerating convolution calculation and reducing fragment data

By building block scanning and intra-blocking coefficient organization models, the data access sequence and address mapping of convolutional calculations in deep neural networks are optimized, and the problem of fragmentation of DRAM access is solved, achieving efficient data multiplexing and improvement of DRAM access efficiency.

CN120179585AInactive Publication Date: 2025-06-20HANGZHOU DIANZI UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510302011.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-14
Publication Date
2025-06-20
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

In multi-layer convolutional calculation of deep neural networks, due to the limited capacity of on-chip SRAM, data is frequently loaded from off-chip DRAM to on-chip SRAM, resulting in a large amount of data handling operations and DRAM access delay and energy consumption.

Method used

By constructing a block scanning and intra-blocking coefficient organization model, three block scanning working modes and block coefficient access and address mapping models are proposed to optimize the block division and access order of data, reduce the fragmentation of DRAM access, and improve data multiplexing rate.

Benefits of technology

By optimizing data organization and address mapping, the fragmentation of DRAM access is reduced, the data multiplexing rate is improved, the amount of fragmented data read is reduced by 47%, and the efficiency of DRAM access is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120179585A_ABST
    Figure CN120179585A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of deep learning, and discloses an organization optimization method for accelerating convolution calculation and reducing fragment data reading, which comprises the following steps of: 1, constructing a block scanning and in-block coefficient organization model; 2, block scanning: providing three block scanning working modes according to the dimension of each layer and the block size configuration condition; and step 3, constructing a block coefficient group access and address mapping model. According to parameter configuration of different layers of a network model, a method for optimizing different block data access sequences and address mapping is provided, convolution calculation is accelerated, fragment data reading is reduced, simulation results show that the fragment data reading amount can be reduced by 47%, on-chip DRAM access and on-chip data multiplexing are optimized, and the method is suitable for large-scale popularization and application. The method is of great significance to efficient deployment and computational mapping of a network model on a specific convolution accelerator hardware platform.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of deep learning, and particularly relates to an optimization method for accelerating convolution calculation and reducing fragmented data reading and organization. Background Art

[0002] With the rapid development of deep learning technology, deep neural networks (DNNs) have been widely used in various application fields, such as computer vision, natural language processing, speech recognition, etc.

[0003] In actual deployment, DNN models usually need to process a large number of input feature map coefficients, weight coefficients, and output feature map coefficients. These data are usually stored in a large-capacity off-chip DRAM, while the on-chip SRAM capacity of the computing core is limited and cannot load all the data at once. Therefore, in actual multi-layer convolution calculations, some input feature map coefficients and weight coefficients need to be loaded from off-chip DRAM into the on-chip SRAM cache near the computing core. After the computing core completes part of the convolution calculation, the output feature map coefficients are stored back into the DRAM, and then the computing core continues to perform convolution calculations on some other input feature coefficients.

[0004] As Figure 1 shown, DRAM is organized as channel, rank, chip, bank, row, and column in a top-down manner. When accessing DRAM, one Rank responds to the access at the same time. A rank consists of a series of chips in parallel. The chips in a rank can access in parallel to form an access word. In a chip, the access request locates a bank, and the location includes row and column addresses. A row address activates a row in the bank, and the data in the activated row appears in the row buffer. When the data to be accessed is not in the already activated row, a row activation operation is required, and a precharge operation needs to be triggered before activation. This precharge + activation requires a certain delay and energy consumption.

[0005] Logical addresses are used when a program accesses data in memory. The specific physical address needs to be converted by the Memory Management Unit (MMU), which converts the logical address into the actual physical address. The fields that determine the actual physical address include: the channel index value (concurrent), the rank index (concurrent), the bank index (concurrent), the row address index (15 bits, a total of 32K rows), and the column address within the row (13 bits, a total of 8K bytes). If each memory die has a bit width of 8 bits, 8 dies are connected in parallel to form a RANK (64 bits). Here, taking a 4G-byte DRAM as an example: 8K bytes per row * 32K rows * 2 ranks * 8 banks = 4096MB = 4GB.

[0006] During the DRAM read and write processes, the latency and energy consumption of DRAM access are related to the DRAM state: The energy consumption for a single DRAM access includes: standby (STBY) and the consumption during read (RD), write (WR), activation (ACT), and precharge (PRE). For a row cache hit (the accessed data is in the currently activated row), it requires the consumption during read (RD), write (WR), and precharge (PRE), and the data can be continuously accessed in a burst manner through DMA. For a row cache conflict operation, the currently activated row needs to be closed and another row needs to be opened, which requires the consumption of precharge (PRE), activation (ACT), and standby energy. A row miss indicates that there is no currently opened row, which requires the energy for activation (ACT) and standby (STBY).

[0007] Due to the large data volume and diversity of DNN models, a large number of complex data transfer operations will occur between the computing core, on-chip cache, and off-chip DRAM. The access to off-chip DRAM is one of the most energy-consuming operations in DNN accelerators, which is also the "memory wall" problem in intelligent computing.

[0008] Two problems need to be solved: data organization and address mapping. A coefficient matrix with dimensions of W * H * D is divided into multiple blocks with dimensions of Tw * Th * Td. The scanning order in which these coefficients are converted into a one-dimensional logical address space has a significant impact on the efficient continuous reading of data within a block. Additionally, how the one-dimensional logical addresses of this data are converted into physical addresses in DRAM has an important impact on the concurrent burst access efficiency of DRAM data.

[0009] The principle of optimizing the coefficient scanning order is to minimize the problem of data access fragmentation caused by discontinuous data chunking. Since different chunking strategies result in different Tw*Th*Td configurations, it is necessary to select an appropriate scanning order according to the magnitudes of the three chunking coefficients. On the other hand, optimizing DRAM data concurrent burst access is also a factor to be considered. Series data from the same tile (which will be accessed and reused in a short time) should be stored in the same row of the same bank as much as possible, so as to increase the hit rate of the row buffer. To further reduce row buffer conflicts and increase throughput, it is necessary to make the best use of chip and bank-level parallel access. Chip-level access parallelism means that data of the same tile should be stored concurrently in different chips as much as possible, so that concurrent access to the chunked data can be achieved. Similarly, data of adjacent unified chunks should be stored in adjacent banks as much as possible to achieve bank-level concurrent access. DRAM data access and address mapping need to minimize the probability of row cache conflicts and row cache misses, make use of the concurrent characteristics at the bank and chip levels, and the advantages of in-row data burst mode access, so as to improve the energy consumption efficiency of DRAM data and reduce unnecessary delays in DRAM data access.

[0010] To optimize data access between DRAM and SRAM, existing technologies usually reduce the overhead of data transfer by optimizing the batch concurrency intensity, chunking strategy, loop order strategy, etc. of DRAM. Selecting an appropriate tile size can maximize the data reuse efficiency in on-chip SRAM. According to the size of the feature map and the dimension size of the weight coefficient of each convolutional layer, the most appropriate input feature coefficient ifmaps, output feature coefficient ofmaps, and weight coefficient weight can be determined, and logically, the minimum amount of on-chip DRAM access data can be achieved. However, this minimum is only in a logical sense. Whether the minimum amount of data transfer between DRAM and SRAM can be truly achieved depends on the data organization and address mapping scheme of various coefficients in DRAM. Different address mapping schemes will result in some unnecessary rank, bank, and row switching operations, and only a very limited amount of data can be continuously read in burst mode each time. Reading a chunk of data may require multiple bank or row address switches, which will lead to delays and unnecessary power consumption caused by some operations such as row switching, row activation, and precharge. Unoptimized data organization and address mapping may result in the theoretically various amounts of chunked DRAM read data, which are only a part of the actual DRAM read data.

[0011] Existing work usually only has a single data organization mode and does not consider the impact of different coefficient Tw*Th*Td configurations on the efficiency of continuous data reading. When constructing the bandwidth model, some work considers the bandwidth consumption in the ideal logical sense and does not consider the actual DRAM structure, resulting in the inconsistency between the actually consumed DRAM access bandwidth and the logically accessed bandwidth. Therefore, how to further optimize the data organization and address mapping strategy, reduce the fragmentation of DRAM access, and improve the data reuse rate under the limited on-chip cache and computing resources is an important challenge in the current DNN accelerator design. Summary of the Invention

[0012] The purpose of the present invention is to provide an optimized method for accelerating convolution calculation and reducing fragmented data reading organization to solve the above technical problems.

[0013] To solve the above technical problems, the specific technical solution of an optimized method for accelerating convolution calculation and reducing fragmented data reading organization of the present invention is as follows:

[0014] An optimized method for accelerating convolution calculation and reducing fragmented data reading organization includes the following steps:

[0015] Step 1: Construct a block scanning and in-block coefficient organization model;

[0016] Step 2: Block scanning: According to the configuration of each layer dimension and block size, three block scanning working modes are proposed;

[0017] Step 3: Construct a block coefficient group access and address mapping model.

[0018] Further, the said Step 1 includes the following steps:

[0019] The input feature map ifmaps with size W*H*I is divided into ceil(W / Tw)*ceil(H / Th)*ceil(I / Ti)

[0020] blocks with size T w *T h *T i ; the output feature map ofmaps with size N*M*J is divided into ceil(M / Tm)*

[0021] ceil(N / Tn)*ceil(J / Tj) blocks with size T m *T n *T j ; the four-dimensional weight coefficient matrix with size I*J*P*Q is merged into a three-dimensional matrix of I*J*(P*Q), and thus the three-dimensional coefficient matrix is divided into ceil(I / Ti)*ceil(J / Tj) blocks with size Ti*Tj*(P*Q);

[0022] The access of Ifmaps and ofmps coefficient blocks in a cyclic manner is abstracted as a process of three-dimensional traversal of internal tiles in a three-dimensional matrix. The weight coefficient matrix is understood as a two-dimensional traversal loop of internal blocks in a three-dimensional matrix, which is a special case of the three-dimensional traversal loop.

[0023] Furthermore, in step 2, according to the dimension of each layer and the configuration of the block size, three block scanning working modes are proposed. The blocks first complete a raster scan of a block plane in a certain raster order, and then continue with the raster scan of the next block plane in one direction until the last block of the last block plane, thus completing the raster scan of the entire three-dimensional matrix.

[0024] Furthermore, the three block scanning working modes in step 2 are as follows: The first one: In the three-dimensional coefficient matrix, the block scanning traverses and scans the block planes from front to back - the front-back mode wh , and the blocks in the block plane are raster scanned. Starting from the first block, raster scanning is performed until the last block of the first block plane, and then the first block of the second block plane is started, and so on until the last block of the last block plane; The second one: The block scanning traverses and scans the block planes from left to right - the left-right mode dh , and the blocks in the block plane are raster scanned. Starting from the first block, raster scanning is performed until the last block of the first block plane, and then the first block of the second block plane is started, and so on until the last block of the last block plane; Similarly, the third one: The block scanning traverses and scans the block planes from top to bottom - the top-down mode wd , and the blocks in the block plane are raster scanned. Starting from the first block, raster scanning is performed until the last block of the first block plane, and then the first block of the second block plane is started, and so on until the last block of the last block plane.

[0025] Furthermore, in the case of the three access modes in step 2, three coefficient scanning orders are respectively adopted for the coefficients inside each block. In the front-back mode, the coefficients in a coefficient plane of the block are raster scanned in the row-column (h*w) order. In the left-right mode, the coefficients in a coefficient plane are raster scanned in the d*h direction; In the top-down mode, the coefficients in a coefficient plane are raster scanned in the d*w direction;

[0026] Different types of coefficients adopt different scanning orders. The general principle is to minimize the switching of block planes, that is, to minimize the number of block plane switches as much as possible.

[0027] Further, in step 2, the block scanning mode mode is determined according to three coefficient dimension parameters and block configuration. Considering the fragmentation minimization angle based on block plane switching, the appropriate scanning mode is selected according to the number of blocks in three block scanning modes. The three numbers of blocks are ceil(W / T w )*ceil(H / T h )、ceil(H / T h )*ceil(D / T d ) and ceil(W / T w )*ceil(D / T d ). Calculate the maximum number of blocks maxTileNum in these three cases. If maxTileNum is equal to ceil(W / T w )*ceil(H / T h ), select the front - back mode mode wh ; if maxTileNum is equal to ceil(H / T h )*ceil(D / T d ), select the left - right mode mode dh ; if maxTileNum is equal to ceil(W / Tw)*ceil(D / Td), select the up - down mode mode wd .

[0028] Further, step 3 includes the following steps:

[0029] Assume that the block parameter variables are (w, h, d), and the block size is T w *T h *T d . The values of Tw*Th*Td and (w, h, d) in three modes are: input feature coefficient w*h*i, output feature coefficient n*m*j, weight coefficient i*1*j; the block index numbers are w0, h0, d0, and the coefficient sequence numbers within the block are w1, h1, d1; the input feature coefficient address access function is BWA(addr0, mode, <T w ,T h ,T i , <W, H, I>); the output feature coefficient address access function is BWA(addr0, mode, <T n ,T m ,T j , <N, M, J>); the weight coefficient address access function is BWA(addr0, mode,

[0030] <T i ,1,T j , <I, P*Q, J>).

[0031] Address access and data volume calculation:

[0032] Front and back mode wh , left and right mode hd , up and down mode mode hd Three modes, according to the three mode data scanning strategies, respectively in the d / h / w, w / d / h and h / w / d direction of the loop, traversal ceil (D / T d )*

[0033] ceil(H / T h )*ceil(W / T w )、ceil(W / T w )*ceil(D / T d )*ceil(H / T h )、ceil(H / T h )*ceil(W / T w )

[0034] *ceil(D / T d ) blocks. In the loop, the GetTileDim function is used to calculate the actual block size; when the last block plane in the scanning direction is encountered, the GetTileDim function is used to fine-tune the actual size of each block T w 1 ,T h 1 ,T d 1 ;

[0035] For a given address addr, the number of coefficients to be accessed is l, the single coefficient data width is DW, and the data bus width is BW bytes, then the amount of data to access l coefficients is as follows:

[0036] In the three modes of the algorithm, based on the starting offset address addr0, the address offset of different scanning block planes is calculated to obtain the block plane address addr1, and then the starting address addr2 of each line is calculated based on addr1, and a line of data is read by ReadLine; in the three modes, T is read from the logical address addr2 respectively. d 1 、T h 1 and T w 1 The above are the three mode selection traversal and address mapping processes.

[0037] The method for accelerating convolution calculation and reducing fragmented data reading organization optimization of the present invention has the following advantages:

[0038] In the scenario of deep learning intensive computing provided by the present invention, the data volumes of input / output feature coefficients and weight coefficients are so huge that they cannot all be loaded into the on-chip SRAM and are stored in the off-chip SDRAM. When organizing the data, three-dimensional block partitioning needs to be performed on the three types of coefficient data. Only partial block data resides in the on-chip SRAM at the same time, and the calculation core of the MAC array in the CNN computing engine directly accesses it at high speed. A large number of small blocks are partitioned from the three types of data, the scanning order of the small blocks in the three-dimensional direction, and the data organization method are embodied as the mapping method between the data logical address and the direct physical address of the SDRAM, improving the data access efficiency between the SDRAM and the SRAM.

[0039] The block scanning and in-block coefficient organization model, as well as the block coefficient access and address mapping model proposed by the present invention, can be applied to the compilation optimization process of the compiler tool chain for various inference chip architectures. According to the parameter configurations of different layers of the network model, an optimization method for different block data access orders and address mappings is given to accelerate the convolution calculation and reduce the fragmented data reading. The simulation results show that the fragmented data reading volume can be reduced by 47%, realizing the optimization of on-chip DRAM access and on-chip data reuse, which is of great significance for the efficient deployment and calculation mapping of the network model on a specific convolutional accelerator hardware platform. BRIEF DESCRIPTION OF THE DRAWINGS

[0040] Figure 1 It is a schematic diagram of the DRAM structure;

[0041] Figure 2 It is a schematic diagram of the partitioning of the input / input feature coefficients and weight coefficients of the present invention;

[0042] Figure 3 It is a schematic diagram of the input / input feature coefficients and weight coefficients represented by a three-dimensional matrix of the present invention;

[0043] Figure 4 It is a schematic diagram of the block scanning of the present invention;

[0044] Figure 5 It is a schematic diagram of the partitioning of the input / input feature coefficients and weight coefficients of the present invention; DETAILED DESCRIPTION OF THE EMBODIMENTS

[0045] In order to better understand the purpose, structure and function of the present invention, the following further describes in detail an organization optimization method for accelerating convolution calculation and reducing fragmented data reading of the present invention with reference to the accompanying drawings.

[0046] The present invention mainly aims to solve the following problems:

[0047] 1. The problem of optimizing the organization of data with coefficient blocks of different sizes. After determining the coefficient block size and loop order through the bandwidth optimization algorithm, the original complete three-dimensional matrices W*H*D of ifmaps / ofmaps / weight are divided into ceil(W / Tw)*ceil(H / Th)*ceil(D / Td) coefficient blocks of size Tw*Th*Td. What scanning order should these blocks be traversed and accessed, and how should the coefficients within the block be accessed in sequence to achieve the minimum row buffer conflict?

[0048] 2. The problem of mapping logical addresses to physical addresses for bandwidth optimization. How to achieve efficient mapping between logical addresses and physical addresses under a given scanning order, and determine the initialization logical address and the function for updating the loop access address to maximize the utilization of multi-bank concurrency and in-row consecutive data burst reading as much as possible.

[0049] An organization optimization method for accelerating convolution calculation and reducing fragmented data reading according to the present invention includes the following steps:

[0050] Step 1: Construct a block scanning and in-block coefficient organization model;

[0051] As Figure 2 shown, the input feature map (ifmaps) of size W*H*I is divided into ceil(W / Tw)*ceil(H / Th)*ceil(I / Ti) blocks of size T w *T h *T i ; the output feature map (ofmaps) of size N*M*J is divided into ceil(M / Tm)*ceil(N / Tn)*ceil(J / Tj) blocks of size T m *T n *T j ; considering that the convolution kernel P*Q is generally small, the four-dimensional weight coefficient matrix of size I*J*P*Q can be combined into a three-dimensional matrix of I*J*(P*Q), and this three-dimensional coefficient matrix is divided into ceil(I / Ti)*ceil(J / Tj) blocks of size Ti*Tj*(P*Q), as Figure 3 shown.

[0052] The loop access of ifmaps and ofmps coefficient blocks can be abstracted as a three-dimensional traversal loop process of internal blocks (tiles) in a three-dimensional matrix, and the weight coefficient matrix can be understood as a two-dimensional traversal loop of internal blocks in a three-dimensional matrix, which is a special case of the three-dimensional traversal loop.

[0053] Step 2: Block scanning: According to the dimension and block size configuration of each layer, three block scanning working modes are proposed;

[0054] The input and output feature coefficients and weight coefficients of different convolutional layers have different dimensions. The optimized block configuration T w *T h *T i 、T m *T n *T j and Ti*Tj*(P*Q) are each different. Different configurations result in different amounts of data transferred in a single data transfer from off-chip DRAM to on-chip SRAM and different numbers of loop transfers. From the perspective of minimizing external DRAM access fragmentation, the present invention proposes three block scanning working modes according to the dimensions of each layer and the block size configuration, as Figure 4 shown. The blocks first complete a raster scan of a block plane in a certain raster order, and then continue with the raster scan of the next block plane in one direction until the last block of the last block plane, thus completing the raster scan of the entire three-dimensional matrix.

[0055] The three block scanning working modes are: First: As shown in (a) of Figure 4 , in the three-dimensional coefficient matrix, the block scan traverses the block planes from front to back (front-back mode, mode wh ), and the blocks inside the block plane are raster scanned, starting from the first block (sequence number 1) and raster scanning to the last block (sequence number 2) of the first block plane, then starting from the first block (sequence number 3) of the second block plane, and so on until the last block of the last block plane. Second: Similarly, as shown in (b) of Figure 4 , the block scan traverses the block planes from left to right (left-right mode, mode dh ), and the blocks inside the block plane are raster scanned, starting from the first block (sequence number 1) and raster scanning to the last block (sequence number 2) of the first block plane, then starting from the first block (sequence number 3) of the second block plane, and so on until the last block of the last block plane. Third: Similarly, as shown in (c) of Figure 4 , the block scan traverses the block planes from top to bottom (top-down mode, mode wd ), and the blocks inside the block plane are raster scanned, starting from the first block (sequence number 1) and raster scanning to the last block (sequence number 2) of the first block plane, then starting from the first block (sequence number 3) of the second block plane, and so on until the last block of the last block plane.

[0056] Correspondingly, in the case of the three access modes, the coefficients inside each block are respectively as shown in Figure 5For the three coefficient scanning orders shown in (a), (b), and (c), in the front-back mode, within a block, a coefficient plane is raster scanned row by row (h*w). Similarly, in the left-right mode, the coefficients in a coefficient plane are raster scanned in the d*h direction; Figure 5 In (c) of Figure 5 , in the up-down mode, the coefficients in a coefficient plane are raster scanned in the d*w direction.

[0057] Which scanning order should a specific coefficient matrix adopt? This depends on the matrix size and the block size, and these parameters are different for different layers. Therefore, different layers should adopt different scanning orders. The present invention supports different types of coefficients to adopt different scanning orders. The general principle is to minimize the switching of block planes, that is, to minimize the number of block plane switches as much as possible. This can make the continuous blocks in a block plane as many as possible. The blocks in a block plane are usually stored continuously, so there is a greater probability of reducing row switching to maximize the utilization efficiency of DRAM parallel access.

[0058] This method determines the block scanning mode mode according to three coefficient dimension parameters and block configuration, as shown in Algorithm 1. Based on the consideration of minimizing the fragmentation of block plane switching, this method selects the appropriate scanning mode according to the number of blocks in three block scanning mode cases. The three numbers of blocks are ceil(W / T w )*ceil(H / T h ), ceil(H / T h )*ceil(D / T d ), and ceil(W / T w )*ceil(D / T d ). Calculate the maximum number of blocks maxTileNum in these three cases. If maxTileNum is equal to ceil(W / T w )*ceil(H / T h ), select the front-back mode mode wh ; if maxTileNum is equal to ceil(H / T h )*ceil(D / T d ), select the left-right mode mode dh ; if maxTileNum is equal to ceil(W / Tw)*ceil(D / Td), select the up-down mode mode wd .

[0059] Here, the algorithm is described using unified W*H*D and T w *T h *T d to correspond to three coefficient three-dimensional sizes W*H*I, N*M*J, I*(P*Q)*J, and at the same time correspond to the block size T w *T h*T i 、T n *T m *T j and Ti*(P*Q)*Tj, the corresponding relationship is shown in the following table.

[0060] Coefficient type W*H*D <![CDATA[T w *T h *T d > w*h*d Input feature coefficient W*H*I <![CDATA[T w *T h *T i > w*h*i Output feature coefficient N*M*J <![CDATA[T n *T m *T j > n*m*j Weight coefficient I*(P*Q)*J <![CDATA[T i *(P*Q)*T j > i*1*j

[0061] Algorithm 1: Block Scanning Mode Strategy Input: The block sizes T of the input, output feature coefficients, and weight coefficients w *T h *T i 、T m *T n *T j and T i *T j *(P*Q)

[0062] The three-dimensional sizes of the three coefficients W*H*I, N*M*J, I*J*(P*Q)

[0063] Output: Block scanning mode mode

[0064]

[0065]

[0066] Step 3: Construct a block coefficient group access and address mapping model;

[0067] The block numbers are as Figure 3 shown. Assume that the block parameter variable is (w, h, d) and the block size is T w *T h *T d In the case of the three modes, the values of T w *T h *T d and (w, h, d) are shown in the rightmost column of Table 1. The block index numbers are w0, h0, d0, and the coefficient index numbers within the block are w1, h1, d1.

[0068] The input feature coefficient address access function is BWA(addr0, mode, <T w , T h , T i , <W, H, I>); The output feature coefficient address access function is BWA(addr0, mode, <T n , T m , T j , <N, M, J>); The weight coefficient address access function is BWA(addr0, mode, <T i , 1, T j, <I, P*Q, J>).

[0069] Algorithm 2: Address Access and Data Volume Calculation

[0070] Function BWA(addr0, mode, <T w , T h , T d , <W, H, D>)

[0071] Input: T w , T h , T d ; W, H, D, scanning mode mode, starting offset address of feature coefficients

[0072] Output: TotalDataAccess (data access volume)

[0073]

[0074]

[0075] Divide into front - back mode mode wh , left - right mode mode hd , up - down mode mode hd These three modes. Respectively according to the data scanning strategies of the three modes, loop in the d / h / w, w / d / h, and h / w / d directions respectively, and traverse ceil(D / T d ) * ceil(H / T h ) * ceil(W / T w ), ceil(W / T w ) * ceil(D / T d ) * ceil(H / T h ), ceil(H / T h ) * ceil(W / T w ) * ceil(D / T d ). In the loop, the function GetTileDim is used to calculate the actual tile size. Generally, the tile size is T w , T h , T d is determined. However, when encountering the last tile plane in the scanning direction, there may be insufficient remaining data to fill a tile, and there will be a few actual tile sizes smaller than T w , T h , T d . The function GetTileDim is used to fine - tune the actual size T of each tile w 1 , T h 1 , Td 1 For a given address addr, access coefficient number l, single coefficient data bit width (number of bytes) DW, and data bus bit width BW bytes, the amount of l coefficient data accessed is as follows:

[0076] In the three modes of the algorithm, based on the starting offset address addr0, the address offsets of different scan block planes are calculated to obtain the block plane address addr1, and then the starting address addr2 of each row is calculated based on addr1, and a row of data is read by ReadLine. In the three modes, starting from the logical address addr2, T d 1 , T h 1 and T w 1 coefficients are read. The above is the process of traversal and address mapping for the three-mode selection.

[0077] Function GetTileDim(w, h, d, <T w , T h , T d >, <W, H, D>)

[0078] w’ ← (w - 1) * T w

[0079] h’ ← (h - 1) * T h

[0080] d’ ← (d - 1) * T d

[0081] T w 1 ← min(T w , W - w’)

[0082] T h 1 ← min(T h , H - h’)

[0083] T d 1 ← min(T d , D - d’)

[0084] return <T w 1 , T h 1 , T d 1 >

[0085] It will be understood that the present invention is described by way of some embodiments, and those skilled in the art will be aware that various changes or equivalent substitutions can be made to these features and embodiments without departing from the spirit and scope of the present invention. Additionally, under the teaching of the present invention, these features and embodiments can be modified to adapt to specific circumstances and materials without departing from the spirit and scope of the present invention. Therefore, the present invention is not limited by the specific embodiments disclosed herein, and all embodiments falling within the scope of the claims of the present application belong to the scope protected by the present invention.

Claims

1. A method for accelerating convolution calculation and reducing fragmented data reading organization optimization, characterized in that: The steps include: Step 1: Construct a block scanning and coefficient organization model within the block; Step 2: Block scanning: According to the configuration of each layer dimension and block size, three block scanning working modes are proposed; Step 3: Construct the block coefficient group access and address mapping model.

2. The method for optimizing the organization of accelerated convolution calculation and reduced fragmented data reading according to claim 1, characterized in that: The step 1 comprises the following steps: The input feature map ifmaps of size W*H*I is divided into ceil(W / Tw)*ceil(H / Th)*ceil(I / Ti) of size T w *T h *T i The output feature map ofmaps of size N*M*J is divided into ceil(M / Tm)*ceil(N / Tn)*ceil(J / Tj) blocks of size T m *T n *T j The four-dimensional weight coefficient matrix of size I*J*P*Q is merged into a three-dimensional matrix of size I*J*(P*Q), so that the three-dimensional coefficient matrix is ​​divided into ceil(I / Ti)*ceil(J / Tj) blocks of size Ti*Tj*(P*Q); The block-by-block iterative access of ifmaps and ofmps coefficients is abstracted as a three-dimensional traversal loop process of tiles within a three-dimensional matrix. The weight coefficient matrix is ​​understood as a two-dimensional traversal loop of tiles within a three-dimensional matrix, which is a special case of a three-dimensional traversal loop.

3. The method for optimizing the organization of accelerated convolution calculation and reduced fragmented data reading according to claim 1, characterized in that: The step 2 proposes three block scanning working modes according to the configuration of each layer dimension and block size. The blocks first complete a block plane scan according to a certain raster order, and then continue the next block plane scan in one direction until the last block of the last block plane, thereby completing the raster scan of the entire three-dimensional matrix.

4. The method for optimizing the organization of accelerated convolution calculation and reduced fragmented data reading according to claim 3, characterized in that: The three block scanning working modes in step 2 are: the first one: in the three-dimensional coefficient matrix, the block scanning is performed according to the block plane traversal scanning from front to back - front-back mode, mode wh , the blocks in the block plane are scanned according to the raster scan, starting from the first block to the last block of the first block plane, and then starting from the first block of the second block plane, and so on, until the last block of the last block plane; the second type: block scanning is scanned according to the traversal of the block plane from left to right - left-right mode left-right, mode dh , the blocks in the block plane are raster scanned, starting from the first block to the last block of the first block plane, and then starting from the first block of the second block plane, and so on, until the last block of the last block plane; similarly, the third type: block scanning is scanned from top to bottom in the block plane traversal - top-down mode, mode wd The blocks in the block plane are raster scanned, starting from the first block to the last block of the first block plane, and then starting from the first block of the second block plane, and so on, until the last block of the last block plane.

5. The method for optimizing the organization of accelerated convolution calculation and reduced fragmented data reading according to claim 3, characterized in that: In the three access modes of step 2, three coefficient scanning orders are respectively adopted for the coefficients in each block. In the front-to-back mode, a coefficient plane in the block is raster scanned according to the row and column (h*w); in the left-right mode, the coefficients in a coefficient plane are raster scanned according to the d*h direction; in the top-to-bottom mode, the coefficients in a coefficient plane are raster scanned according to the d*w direction; Different types of coefficients use different scanning orders. The general principle is to minimize the block plane switching, that is, to make the number of block plane switching as small as possible.

6. The method for optimizing the organization of accelerated convolution calculation and reduced fragmented data reading according to claim 3, characterized in that: The step 2 determines the block scanning mode mode according to the three coefficient dimension parameters and the block configuration, and selects the appropriate scanning mode according to the number of blocks in the three block scanning modes based on the perspective of minimizing the fragmentation of the block plane switching. The three block numbers are ceil(W / T w )*ceil(H / T h )、ceil(H / T h )*ceil(D / T d ) and ceil(W / T w )*ceil(D / T d ), calculate the maximum number of tiles maxTileNum in these three cases. If maxTileNum is equal to ceil(W / T w )*ceil(H / T h ), select the front and back mode mode wh ; If maxTileNum is equal to ceil(H / T h )*ceil(D / T d ), select left and right mode dh ; If maxTileNum is equal to ceil(W / Tw)*ceil(D / Td), select up and down mode mode wd .

7. The method for optimizing the organization of accelerated convolution calculation and reduced fragmented data reading according to claim 1, characterized in that: The step 3 comprises the following steps: Assume that the block parameter variables are (w, h, d) and the block size is T w *T h *T d , the values ​​of Tw*Th*Td and (w,h,d) in the three modes are: input feature coefficient w*h*i, output feature coefficient n*m*j, weight coefficient i*1*j; the block index number is w0,h0,d0, and the coefficient number index within the block is w1,h1,d1; the input feature coefficient address access function is BWA(addr0,mode, <T w ,T h ,T i >,<W,H,I> ); the output characteristic coefficient address access function is BWA(addr0, mode, <T n ,T m ,T j ,<N,M,J> ); the weight coefficient address access function is BWA(addr0, mode, <T i ,1,T j ,<I,P*Q,J> ). Address access and data volume calculation: Front and back mode wh , left and right mode hd , up and down mode mode hd Three modes, according to the three mode data scanning strategies, respectively in the d / h / w, w / d / h and h / w / d direction of the loop, traversal ceil (D / T d )*ceil(H / T h )*ceil(W / T w )、ceil(W / T w )*ceil(D / T d )*ceil(H / T h )、ceil(H / T h )*ceil(W / T w )*ceil(D / T d ) blocks; inside the loop, the GetTileDim function is used to calculate the actual block size; when encountering the last block plane in the scanning direction, the GetTileDim function is used to fine-tune the actual size of each block T w 1 ,T h 1 ,T d 1 ; For a given address addr, the number of coefficients to be accessed is l, the single coefficient data width is DW, and the data bus width is BW bytes, then the amount of data to access l coefficients is as follows: (addr,l,DW,BW)=[ceil((addr+l*DW) / BW)-ceil(addr / BW)]*BW; In the three modes of the algorithm, based on the starting offset address addr0, the address offset of different scanning block planes is calculated to obtain the block plane address addr1, and then the starting address addr2 of each line is calculated based on addr1, and a line of data is read by ReadLine; in the three modes, T is read from the logical address addr2 respectively. d 1 , T h 1 and T w 1 The above are the three mode selection traversal and address mapping processes.