A Computation Scheduling Mapping Method for a CNN Network Model with High Computational and Memory Access Efficiency
By optimizing on-chip SRAM storage, MAC computing unit configuration and flow scheduling, the joint optimization problems of computing mapping, storage mapping and soft flow scheduling are solved, and the computing and memory access efficiency of the CNN network model is improved.
Patent Information
- Application Number
- CN202111586693.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-12-23
- Publication Date
- 2025-07-04
- Estimated Expiration
- 2041-12-23
AI Technical Summary
The existing methods lack joint optimization for the three dimensions of computing mapping, storage mapping and soft flow scheduling, resulting in low computing efficiency and data reuse efficiency, and the inability to make full use of the computing resources of the hardware platform.
By determining the on-chip SRAM storage configuration, on-chip concurrent MAC computing unit configuration and flow scheduling optimization scheme, combined with the structural characteristics of the neural network model, the computing, storage and memory access bandwidth are optimized to achieve multi-objective optimization mapping methods.
The utilization rate of the computing unit is improved, data waiting time and access bandwidth consumption of external memory are reduced, and computing and data multiplexing efficiency is improved.
Smart Images

Figure CN114330653B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of deep learning, and particularly relates to a calculation scheduling mapping method for a CNN network model with high computing and memory access efficiency. Background Art
[0002] TVM, an OctoML company, has launched the deep learning automatic code generation toolset TVM. AWS and a research team from the University of Washington have launched the end-to-end compiler NNVM based on the TVM stack. Tencent Youtu has designed and launched the cross-platform deep learning inference framework TNN based on NCNN. The AI LAB Open Intelligence Laboratory has launched the Tengine inference framework that supports multiple platforms. Megvii has established a full-process and one-stop artificial intelligence algorithm platform MegEngine from algorithm research and development to deployment and application. Google has open-sourced the MLIR architecture that is closely integrated with TensorFlow, supporting representation formats and compiler utility libraries. NVIDIA has launched the high-performance deep learning Inference optimizer TensorRT for GPUs. Huawei has launched its self-developed AI computing framework MindSpore for its self-developed NPU processor, providing a unified API for the entire scene, and providing end-to-end capabilities for model development, model operation, and model deployment of the entire-scene AI. In these works, some have adopted the methods of automatic network model mapping and code production.
[0003] Papers such as Optimizing the Convolution Operation to Accelerate Deep Neural Networks on FPGA and Optimizing Accelerator on FPGA for Deep Convolutional Neural Networks have explored heuristic optimization methods on the FPGA platform. Optimizing Memory Efficiency for Deep Convolutional Neural Networks on GPUs has explored storage and computing optimization methods on the GPU platform. Dynamic Look-up Table Method for Optimizing the Training of Deep Neural Networks on Many-Core Architecture has explored optimization and computing methods on the multi-core processor architecture.
[0004] Currently, the existing methods lack an automatic structure mapping and code generation method for joint optimization in three dimensions: computational mapping, storage mapping, and software pipelining scheduling.
[0005] (1) Different neural network model structures have differences, consisting of different numbers of network layers. The convolutional kernel sizes configured for each layer may be different, and the input and output channel numbers configured for each layer are also different. High-efficiency convolutional computing requires designing optimized computational mapping, storage mapping schemes, and software pipelining optimization scheduling strategies according to the characteristics of the target platform system architecture.
[0006] (2) The on-chip memory (SRAM) sizes of different target hardware platforms are configured differently. How to reasonably allocate storage space for input feature coefficients (Tix * Tiy * Tif), convolutional kernel coefficients (Tkx * Tky * Tif * Tof), and output feature coefficients (Tox * Toy * Tof), thereby determining the frequency of data interaction between external memory and on-chip memory (cache), directly affects data reuse efficiency, and overall affects the operating efficiency of the computing unit, the throughput effect of the external memory, and the access bandwidth consumption.
[0007] (3) The on-chip computing resources of different target hardware platforms are configured differently, and the concurrency intensity that can simultaneously implement MAC calculations (multiplication and addition calculations) may vary. How to map dense multiple-loop MAC calculations to a computing unit array (Pkx, Pky, Pif, Pix, Piy, Pof) with limited concurrency has an important impact on making full use of the efficiency of the on-chip MAC computing array.
[0008] (4) Different computing configurations (Pkx, Pky, Pif, Pix, Piy, Pof) and cache configuration combinations (Tix * Tiy * Tif) + (Tkx * Tky * Tif * Tof) + (Tox * Toy * Tof) need to be jointly optimized to maximize computing and data reuse efficiency, and achieve the optimal comprehensive performance under the constraints of target parameters such as minimum computing waiting, maximum on-chip data reuse efficiency, and minimum off-chip data on-chip cache update bandwidth consumption. Summary of the Invention
[0009] The purpose of the present invention is to provide a computational scheduling mapping method for a computationally and memory-access efficient CNN network model to solve the technical problem that the existing methods lack an automatic structure mapping and code generation method for joint optimization in three dimensions of computational mapping, storage mapping, and software pipelining scheduling.
[0010] To solve the above technical problem, the specific technical solution of a computational scheduling mapping method for a computationally and memory-access efficient CNN network model of the present invention is as follows:
[0011] A computational scheduling mapping method for a computationally and memory-access efficient CNN network model includes the following steps:
[0012] Step 1: Determine the storage mapping scheme according to the on-chip SRAM storage configuration;
[0013] Step 2: Determine the computing mapping scheme according to the configuration of on-chip concurrent MAC computing units;
[0014] Step 3: Determine the pipelining scheduling optimization scheme according to the network model, storage, and computing mapping scheme.
[0015] Furthermore, the specific steps of the said Step 1 are as follows:
[0016] Typical convolutional neural network multiple loops contain six loop indices, namely: output feature index, xy coordinates of output feature coefficients, input feature channel index, xy offsets within the convolution kernel, and S is the sliding step length; assume the input feature parameter dimension is Nif*Nix*Niy , and the convolutional weight coefficient dimension is Nif*Nof*Nkx*Nky , and the output feature parameter dimension is Nof*Nox*Noy;
[0017] The hardware core component of typical convolutional computing is the dense MAC computing unit array, directly connected to which is the register array to achieve data movement and reuse; there is also an on-chip cache tightly coupled with the MAC computing array for caching input feature coefficients and weight coefficients;
[0018] Assume the on-chip cache dimension for storing input feature coefficients is Tif*Tix*Tiy , the on-chip cache dimension for storing weight coefficients is Tif*Tof*Tkx*Tky , and the on-chip cache for storing output feature coefficients is Tof*Tox*Toy ; thus, the input, output feature coefficient, and weight coefficient cache storage unit Tm is Tm = Tif*Tix*Tiy + Tif* Tof*Tkx*Tky + Tof*Tox*Toy ; assume the internal cache capacity of the chip is Ttot, then the parameter Tif,Tix,Tiy,Tof,Tkx,Tky,Tox,Toy is selected to satisfy the following several factors:
[0019] Tm ≤ Ttot The frequency of data exchange finter between the on-chip cache and the external memory is as small as possible; Minimize the data waiting time of the MAC computing unit;
[0020] finter = ( Nif*Nix*Niy ) / ( Tif*Tix*Tiy ) + ( Nif*Nkx*Nky ) / ( Tif*Tkx*Tky )
[0021] + ( Nof*Nox*Noy ) / ( Tof*Tox*Toy ).
[0022] Furthermore, the said Step 2 includes the following specific steps:
[0023] After determining the framework of the storage mapping scheme, the computing mapping scheme is determined according to the on-chip MAC computing resources of the chip. Assume that the number of concurrently available MAC computing units is Pm = Pif * Pix * Piy * Pkx * Pky * Pof, where one MAC computing activation can cover the input feature coefficients in the dimension of Pif * Pix * Piy, cover the weight coefficients in the dimension of Pif * Pof * Pkx * Pky, and cover the output feature coefficients in the dimension of Pof * Pox * Poy.
[0024] Assume that the number of MAC computing units available inside the chip is Ptot. Then, the selection of the parameters Pif, Pix, Piy, Pof, Pkx, Pky, Pox, Poy needs to satisfy the following factors:
[0025] Pm ≤ Ptot The frequency fintra of data exchange between the on-chip cache and the register array should be as small as possible. Minimize the data waiting sorting of the MAC computing units. Maximize the external software pipelining efficiency.
[0026] fintra = (Tif * Tix * Tiy) / (Pif * Pix * Piy) + (Tif * Tkx * Tky) / (Pif * Pkx * Pky)
[0027] + (Tof * Tox * Toy) / (Pof * Pox * Poy).
[0028] Further, in step 3, according to the characteristics of the target computing platform and the network structure of the neural network to be mapped, appropriate parameter combinations are designed. The parameters include computing mapping related parameters Pif, Pix, Piy, Pof, Pkx, Pky, Pox, Poy, storage mapping related parameters Tif, Tix, Tiy, Tof, Tkx, Tky, Tox, Toy, batch processing intensity batch, loop unrolling combination strategy batch, if, of, ix / iy, kx / ky, ox / oy, convolution window sliding strategy scan_mode in the feature map, ping-pong strategy pingpong, and loop unrolling and vectorization strategy vec / unroll.
[0029] A calculation and memory access efficient CNN network model calculation scheduling mapping method of the present invention has the following advantages: A calculation and memory access efficient CNN network model calculation scheduling mapping method of the present invention includes the computing power of a unit MAC computing unit, the concurrency intensity, the on-chip cache granularity, the cache size, and combines the characteristics of the algorithm network structure to implement an optimized mapping for each network layer, and proposes a network structure mapping implementation method for multi-objective optimization of calculation, storage, and memory access bandwidth. BRIEF DESCRIPTION OF THE DRAWINGS
[0030] Figure 1 It is a schematic diagram of a convolution calculation loop;
[0031] Figure 2 It is a schematic diagram of a typical convolution MAC calculation array structure;
[0032] Figure 3 It is a schematic diagram of the input, output, and weight coefficient storage mapping of the present invention;
[0033] Figure 4 It is a schematic diagram of the MAC calculation mapping of the present invention;
[0034] Figure 5 It is a schematic diagram of the multi-dimensional parameter selection optimization of the present invention. DETAILED DESCRIPTION OF THE INVENTION
[0035] In order to better understand the purpose, structure, and function of the present invention, the following further describes in detail a calculation and memory access efficient CNN network model calculation scheduling mapping method of the present invention with reference to the accompanying drawings.
[0036] A calculation and memory access efficient CNN network model calculation scheduling mapping method of the present invention includes the following steps:
[0037] (1) Determine the storage mapping scheme according to the on-chip SRAM storage configuration
[0038] As Figure 1 shown is a schematic diagram of a typical multiple loop of a convolutional neural network, including six loop indices which are: output feature index, xy coordinates of the output feature coefficient, input feature channel index, xy offsets within the convolution kernel, and S is the sliding step length. Here, it is assumed that the input feature parameter dimension is Nif*Nix*Niy , the convolution weight coefficient dimension is Nif*Nof*Nkx*Nky , and the output feature parameter dimension is Nof*Nox*Noy.
[0039] Figure 2Figure 0 shows a schematic diagram of a typical hardware architecture for convolution calculations. The core component is an array of dense MAC (multiplication and addition operation) calculation units, which is directly connected to a register array to achieve data movement and reuse. There is also an on-chip cache (usually SRAM) tightly coupled to the MAC calculation array for caching input feature coefficients and weight coefficients.
[0040] During the six-fold loop calculation process, it is necessary to repeatedly transfer the input feature coefficients in(ni; S*x+kx, S*y+ky) and the weight coefficients weight(ni,no;kx,ky) to the on-chip storage unit tightly coupled to the calculation unit. As Figure 3 shown, assume that the on-chip cache dimension for storing input feature coefficients is Tif*Tix*Tiy , and the on-chip cache dimension for storing weight coefficients is Tif* Tof*Tkx*Tky , and the on-chip cache for storing output feature coefficients is Tof*Tox*Toy . Thus, the input, output feature coefficient, and weight coefficient cache storage unit Tm is Tm = Tif*Tix*Tiy + Tif* Tof*Tkx*Tky + Tof*Tox*Toy . Assume the internal cache capacity of the chip is Ttot. Then, when selecting these parameters ( Tif,Tix,Tiy,Tof,Tkx,Tky,Tox,Toy ), the following factors need to be considered:
[0041] Tm ≤ Ttot The frequency of data exchange finter between the on-chip cache and the external memory should be as small as possible; Minimize the data waiting time of the MAC calculation unit.
[0042] Nif*Nix*Niy ) / ( Tif*Tix*Tiy ) + ( Nif*Nkx*Nky ) / ( Tif*Tkx*Tky )
[0043] + ( Nof*Nox*Noy ) / ( Tof*Tox*Toy )
[0044] (2) Determine the calculation mapping scheme according to the configuration of on-chip concurrent MAC calculation units
[0045] After determining the framework of the storage mapping scheme, the calculation mapping scheme is determined according to the on-chip MAC calculation resources of the chip. As Figure 4 shown, assume that the number of concurrent MAC calculation units is Pm = Pif * Pix * Piy * Pkx * Pky * Pof, where one MAC calculation activation can cover the input feature coefficients in the dimension of Pif * Pix * Piy, cover the weight coefficients in the dimension of Pif * Pof * Pkx * Pky, and cover the output feature coefficients in the dimension of Pof * Pox * Poy.
[0046] Assume that the number of MAC computing units available inside the chip is Ptot. Then, the selection of these parameters (Pif, Pix, Piy, Pof, Pkx, Pky, Pox, Poy) needs to consider the following factors:
[0047] Pm ≤ Ptot The frequency fintra of data exchange between the on-chip cache and the register array should be as small as possible; Minimize the data waiting time of the MAC computing units; Maximize the external software pipelining efficiency.
[0048] fintra = (Tif * Tix * Tiy) / (Pif * Pix * Piy) + (Tif * Tkx * Tky) / (Pif * Pkx * Pky)
[0049] + (Tof * Tox * Toy) / (Pof * Pox * Poy)
[0050] (3) Determine the pipelining scheduling optimization scheme according to the network model, storage, and computing mapping scheme
[0051] As described above, the computing mapping-related parameters (Pif, Pix, Piy, Pof, Pkx, Pky, Pox, Poy) and the storage mapping-related parameters (Tif, Tix, Tiy, Tof, Tkx, Tky, Tox, Toy) have an important impact on the computing efficiency of the neural network model on the target platform; in addition, the batch processing intensity batch, the loop unrolling combination strategy (batch, if, of, ix / iy, kx / ky, ox / oy), the sliding strategy scan_mode of the convolution window within the feature map, whether to use the ping-pong strategy pingpong for the on-chip cache, and the loop unrolling and vectorization strategy vec / unroll these control parameters have an important impact on the external software pipelining scheduling efficiency. The combination of the above parameters constitutes the parameter space for the implementation of the network result mapping, as Figure 5 shown.
[0052] According to the characteristics of the target computing platform (such as Ptot and Ttot), as well as the network structure of the neural network to be mapped, design an appropriate parameter combination, as Figure 5 shown. In fact, it is necessary to select the optimal path in the Figure 5 combination paths, similar to the Viterbi optimization algorithm.
[0053] It is understood that the present invention is described by way of some embodiments. Those skilled in the art will appreciate that various changes or equivalent substitutions can be made to these features and embodiments without departing from the spirit and scope of the present invention. Additionally, under the teachings of the present invention, these features and embodiments can be modified to adapt to specific circumstances and materials without departing from the spirit and scope of the present invention. Therefore, the present invention is not limited by the specific embodiments disclosed herein, and all embodiments falling within the scope of the claims of the present application belong to the scope protected by the present invention.
Claims
1. A calculation scheduling mapping method for a CNN network model with high computational and memory access efficiency, characterized in that It includes the following steps: Step 1: Determine the storage mapping scheme according to the on-chip SRAM storage configuration; A typical convolutional neural network multiple loop contains six loop indices, namely: output feature index, xy coordinates of the output feature coefficient, input feature channel index, xy offsets within the convolutional kernel, and S is the sliding step length; assume the input feature parameter dimension is Nif*Nix*Niy, the convolutional weight coefficient dimension is Nif*Nof*Nkx*Nky, and the output feature parameter dimension is Nof*Nox*Noy; The typical hardware core component for convolutional calculation is an array of dense MAC calculation units, directly connected to which is a register array to achieve data movement and reuse; there is also an on-chip cache tightly coupled with the MAC calculation array for caching input feature coefficients and weight coefficients; assume the on-chip cache dimension for storing input feature coefficients is Tif*Tix*Tiy, the on-chip cache dimension for storing weight coefficients is Tif*Tof*Tkx*Tky, and the on-chip cache for storing output feature coefficients is Tof*Tox*Toy; thus, the input, output feature coefficient, and weight coefficient cache storage unit Tm is Tm = Tif*Tix*Tiy + Tif*Tof*Tkx*Tky + Tof*Tox*Toy; assume the internal cache capacity of the chip is Ttot, then the selection of parameters Tif, Tix, Tiy, Tof, Tkx, Tky, Tox, Toy needs to meet the following factors: ① Tm ≤ Ttot ② The frequency of data exchange finter between the on-chip cache and the external memory should be as small as possible; ③ Minimize the data waiting rearrangement of the MAC calculation unit; finter = (Nif*Nix*Niy) / (Tif*Tix*Tiy) + (Nif*Nkx*Nky) / (Tif*Tkx*Tky) + (Nof*Nox*Noy) / (Tof*Tox*Toy); Step 2: Determine the calculation mapping scheme according to the on-chip concurrently available MAC calculation unit configuration; After determining the framework of the storage mapping scheme, determine the calculation mapping scheme according to the on-chip MAC calculation resources of the chip; assume the number of concurrently available MAC calculation units is Pm = Pif*Pix*Piy*Pkx*Pky*Pof, where one MAC calculation activation can cover the input feature coefficients of the Pif*Pix*Piy dimension, cover the weight coefficients of the Pif*Pof*Pkx*Pky dimension, and cover the output feature coefficients of the Pof*Pox*Poy dimension; Assume the number of MAC calculation units that can be obtained inside the chip is Ptot, then the selection of parameters Pif, Pix, Piy, Pof, Pkx, Pky, Pox, Poy needs to meet the following factors: ① Pm ≤ Ptot ② The frequency of data exchange fintra between the on-chip cache and the register array should be as small as possible; ③ Minimize the data waiting rearrangement of the MAC calculation unit; ④ Maximize the external software pipelining efficiency; fintra = (Tif * Tix * Tiy) / (Pif * Pix * Piy) + (Tif * Tkx * Tky) / (Pif * Pkx * Pky) + (Tof * Tox * Toy) / (Pof * Pox * Poy); Step 3: Determine the optimized stream scheduling scheme according to the network model, storage, and computing mapping scheme.
2. The calculation and memory access efficient CNN network model calculation scheduling mapping method according to claim 1, characterized in that In Step 3, appropriate parameter combinations are designed according to the characteristics of the target computing platform and the network structure of the neural network to be mapped. The parameters include computing mapping-related parameters Pif, Pix, Piy, Pof, Pkx, Pky, Pox, Poy, storage mapping-related parameters Tif, Tix, Tiy, Tof, Tkx, Tky, Tox, Toy, batch processing intensity batch, loop unrolling combination strategy batch, if, of, ix / iy, kx / ky, ox / oy, convolution window sliding strategy scan_mode within the feature map, ping-pong strategy pingpong, and loop unrolling and vectorization strategy vec / unroll.
Citation Information
Patent Citations
Convolutional neural network hardware accelerator for solidifying full network layer on reconfigurable platform
CN112116084A
KR20200063958A