Bandwidth optimization method supporting deep neural network acceleration cross-layer memory access
By designing a single-layer and cross-layer block optimization bandwidth model of deep neural networks and optimizing data access mode, the cross-layer memory access bandwidth optimization problem of deep neural networks on the CNN accelerator platform is solved, efficient computing and memory access efficiency is achieved, and energy consumption is reduced.
Patent Information
- Application Number
- CN202510216697.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-26
- Publication Date
- 2025-06-10
- Estimated Expiration
- 2045-02-26
AI Technical Summary
The prior art is difficult to effectively optimize the cross-layer memory access bandwidth of deep neural networks on the CNN accelerator platform, resulting in inefficient data handling between computing cores, on-chip caches and off-chip DRAMs, increasing energy consumption.
By designing a single-layer block optimization and cyclic sequence optimization bandwidth model, combining cross-layer joint block optimization and cyclic sequence optimization, the data access mode of input feature maps, output feature maps and weights is optimized, and the data blocking strategy and cyclic expansion control strategy with minimal DRAM access are realized.
It realizes the minimization of DRAM access bandwidth, optimizes computing and memory access efficiency, reduces energy consumption, and supports optimized code generation and computing mapping of computing and memory access efficiency.
Smart Images

Figure CN120123274A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of neural network optimization, and in particular, to a method for optimizing the cross-layer memory access bandwidth to support the acceleration of deep neural networks. Background Art
[0002] When various deep neural network models are actually deployed and applied, a large number of input feature map coefficients and weight coefficients are stored in a large-capacity DRAM memory. The computing core of the CNN accelerator has limited computing resources, and at the same time, the capacity of the on-chip cache SRAM is also limited. When performing actual multi-layer convolution calculations, the input feature map coefficients and weight coefficients are loaded from off-chip DRAM into the on-chip SRAM cache near the computing core. The computing core MAC array usually can only perform partial convolution calculations. The output feature map coefficients generated by the calculations need to be stored in the DRAM, and then the computing core continues to perform convolution calculations on some other input feature coefficients. The large amount of data volume and diversity in the CNN network layer will cause a large amount of complex data transfer between the computing core, the on-chip cache, and the off-chip DRAM. The off-chip memory access based on DRAM is the most energy-consuming operation in the deep neural network DNN accelerator, which is the intelligent computing storage wall problem.
[0003] To achieve the storage access optimization of the network model on a specific CNN acceleration engine platform, it is necessary to solve problems including the block method of CNN convolution weight coefficients and feature coefficients, the computing mapping of the core computing unit, and the external pipeline scheduling, that is, the loop order strategy, as well as the data storage mapping update strategy, so as to realize the code generation and implementation mapping for optimizing the computing and memory access efficiency. The solution space dimension of this optimization problem is very large. The variables involved include the batchsize concurrency intensity (Tz), the width and height of the feature map participated in the calculation by the computing core (Tn*Tm), the input feature dimension of the on-chip cache (Ti), the output feature dimension of the on-chip cache (Tj), the number of cross-layer optimization layers, and the loop loop priority strategy category. The loop loop priority strategy category also shows the following data reuse modes: input coefficient reuse iro / output coefficient reuse oro / weight coefficient reuse wro. These control variables respectively take a certain number of candidate values discretely, constituting a very large possible solution space. If there is no analytical function expression of the SRAM cache consumption and off-chip memory access bandwidth model under different parameter combinations, it is very difficult to find the optimal solution. The on-chip SRAM cache is an inherent configuration of the computing core and is a constraint condition that needs to be considered during the solution. If it is possible to construct an analytical mathematical expression model of the on-chip SRAM consumption and off-chip memory access bandwidth model under different parameter combinations, then it is possible to solve offline for the minimization of the memory access bandwidth that satisfies the SRAM cache constraint.
[0004] Currently, existing work usually only focuses on optimizing single-layer bandwidth access and does not consider that the output feature data of the previous layer is already in the cache and can be reused by the input feature coefficients of the subsequent layer. This cross-layer data reuse is very effective for the efficient deployment of network layers with a relatively high proportion of feature coefficients. Some work does not consider input batch concurrency, which can effectively reuse weight coefficients and is very effective for the WRO mode that prioritizes weight coefficients. There is also some work that organizes data in the data access paradigm required for core computing when constructing a bandwidth access function model. During this process, there will be some overlapping data in the input and output feature map coefficients, and the repeated loading of this overlapping data reduces the accuracy of the model and the coefficient optimization efficiency. In addition, some work considers the bandwidth consumption in the ideal logical sense when constructing a bandwidth model and does not consider the actual DRAM structure, resulting in the problem that the actual DRAM access bandwidth consumption is inconsistent with the logical access bandwidth. Most existing work does not consider this problem.
[0005] Simultaneously considering the optimal data flow strategy of batch concurrency intensity optimization, block strategy optimization, loop order strategy optimization, and cross-layer data reuse optimization, and achieving global optimization of on-chip DRAM access and on-chip data reuse under the constraints of limited on-chip cache and computing core computing resources is an issue that needs further research and solution, and is of great significance for the efficient deployment and computing mapping of network models on specific convolutional accelerator hardware platforms. Summary of the Invention
[0006] In order to solve the above technical problems existing in the prior art, the present invention proposes a method for optimizing cross-layer memory access bandwidth to support deep neural network acceleration, and its specific technical solution is as follows: A method for optimizing cross-layer memory access bandwidth to support deep neural network acceleration, including: Step 1, design a bandwidth model for single-layer block optimization and loop order optimization; Step 2, single-layer independent block optimization and loop order optimization; Step 3, cross-layer joint block optimization and loop order optimization; Step 4, composite block optimization and loop order optimization.
[0007] Furthermore, the model design in the above Step 1 specifically includes: According to the characteristics of the data flow of deep convolution calculation, three data access modes are set, including the input feature map coefficient reuse IRO mode, the output feature map coefficient reuse ORO mode, and the weight coefficient reuse WRO mode; According to the data flow characteristics of the three data access modes, the single access data volume and loop access times of the input feature map, output feature map, and weights in each mode are set.
[0008] Further, the amount of data accessed once for the input feature map, output feature map, and weights in each mode specifically includes: input feature map coefficients: Bif1, Bif2, Bif3, output feature map coefficients: Bof1, Bof2, Bof3, and weight coefficients: Bwgt1, Bwgt2, Bwgt3; the number of loop accesses for the input feature map, output feature map, and weights in each mode specifically includes: input feature map coefficient access times: Cif1, Cif2, Cif3, output feature map coefficient access times: Cof1, Cof2, Cof3, and weight coefficient access times: Cwgt1, Cwgt2, Cwgt3.
[0009] Further, step 2 specifically includes: Step 2.1: Divide the feature data of single-layer convolution calculation into blocks to obtain block parameters; Step 2.2: Traverse all possible block combinations, calculate the memory bandwidth and storage occupancy under three data access modes respectively, and compare the storage threshold and the minimum memory bandwidth to determine the optimal block combination and data access mode.
[0010] Further, step 2.1 specifically includes: using step sizes of 1, stepM, stepN, stepI, and stepJ to divide Z, M, N, I, and J into blocks respectively to obtain block parameters Tz, Th, Tw, Tm, Tn, Ti, and Tj; where Z is the maximum batch number, I is the number of input feature map channels, J*M*N is the feature map coefficient dimension, Th = S*(Tm - 1) + P, Tw = S*(Tn - 1) + Q, S is the convolution step size, I*J*P*Q is the convolution kernel size, take Tp = P, Tq = Q; Tz is the batch loop block size, Tm*Tn and Th*Tw are the horizontal and vertical block sizes of the data feature map coefficients, Ti is the input feature map channel block size, and Tj is the output feature map channel block size.
[0011] Further, in step 2.2, the calculation of the memory bandwidth and storage occupancy is specifically as follows: Bandwidth of the input feature map ifmaps: BW if = B if *C if *size(ifmaps); Bandwidth of the output feature map ofmaps: BW of = B of *C of *size(ofmaps); Bandwidth of the convolution kernel weight weight: BW wgt = B wgt *Cwgt *size(weight); Storage occupancy: SRAM tot = B if *size(ifmaps)+B wgt *size(weight)+B of *size(ofmaps); Total bandwidth: BW tot = BW if +BW of +BWwgt; Wherein, B if , B of , B wgt respectively represent the data sizes of the input feature map, output feature map, and convolution kernel, C if , C of , C wgt respectively represent the loop access times of the input feature map, output feature map, and convolution kernel, and size represents the size; Compare the storage threshold and the minimum bandwidth, and determine the optimal block combination and data access mode, specifically: Set the storage threshold SRAM_TH and the minimum bandwidth minBW. In SRAM tot <=SRAM_TH and BW tot2 <=minBW conditions, that is, if the storage occupancy of the current mode does not exceed the storage threshold SRAM_TH and the total bandwidth is smaller, then update the minimum bandwidth and the data access mode, that is, finally obtain the optimal block combination and mode to minimize the total bandwidth and meet the storage limit.
[0012] Further, step 3 specifically includes: Step 3.1, traverse the block combinations of the convolutional calculation data of adjacent two layers; Step 3.2, cyclically traverse and calculate the bandwidth and storage occupancy of adjacent two layers of different block combinations according to the ORO mode, set the storage constraint, cyclically traverse and check whether the calculated storage occupancy meets the storage constraint according to the IRO mode, and select the block combination with the minimum total bandwidth.
[0013] Further, step 3.1 specifically includes: For the u-th layer and the (u + 1)-th layer, that is, adjacent two layers, first block the feature data of the convolutional calculation of the u-th layer to obtain the block combination of the u-th layer; then reuse the block parameters of the u-th layer to obtain the block combination of the (u + 1)-th layer, where Tw(u + 1)=Tn(u), Th(u + 1)=Tm(u), that is: the width and height of the feature map of the convolutional calculation of the u-th layer are reused as the weight parameters of the feature map of the convolutional calculation of the (u + 1)-th layer.
[0014] Furthermore, step 4 specifically includes: Step 4.1: First, optimize each layer separately, and then jointly optimize adjacent layers; Step 4.2: Dynamically select the optimal block strategy by comparing the results of separate optimization and joint optimization.
[0015] Furthermore, step 4.1 is specifically as follows: Use the single-layer independent block optimization and loop order optimization to optimize each layer separately, and then use the cross-layer joint block optimization and loop order optimization to jointly optimize adjacent layers.
[0016] The present invention can implement a data block strategy, a loop unrolling control strategy, and a data storage mapping update strategy with the minimum DRAM access, and realize a data access and data flow scheme with maximized data reuse at a fine granularity, which can effectively support the automatic code generation for optimizing computing and memory access efficiency and the computing mapping optimization. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] Figure 1 is a schematic diagram of the code design of the IRO / ORO / WRO three data access modes in the embodiment of the present invention; Figure 2 is a schematic diagram of the block relationship in the embodiment of the present invention; Figure 3 is a flowchart of the code implementation of the single-layer independent block optimization and loop order optimization in the embodiment of the present invention; Figure 4 is a flowchart of the code implementation of the cross-layer joint block optimization and loop order optimization in the embodiment of the present invention; Figure 5 is a flowchart of the code implementation of the composite block optimization and loop order optimization in the embodiment of the present invention; Figure 6 is a comparison result diagram between the separate optimization and the joint optimization of the VGG network layer in the embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0018] In order to make the objectives, technical solutions, and technical effects of the present invention more clear and understandable, the present invention will be further described in detail below with reference to the accompanying drawings of the specification and embodiments.
[0019] This embodiment provides a method for optimizing the cross-layer memory access bandwidth to support the acceleration of deep neural networks, which specifically includes the following content: Step 1: Design a bandwidth model for single-layer block optimization and loop order optimization.
[0020] According to the characteristics of the data flow of deep convolutional calculations, determine three data access modes: input feature coefficient reuse IRO, output feature coefficient reuse ORO, and weight coefficient reuse WRO, asFigure 1 as shown
[0021] According to the data stream characteristics of IRO, ORO, and WRO, determine the amount of data accessed per single access of the input feature map, output feature map, and weights under the three data access modes, specifically including the input feature map coefficients: Bif1, Bif2, Bif3, the output feature map coefficients: Bof1, Bof2, Bof3, and the weight coefficients: Bwgt1, Bwgt2, Bwgt3; the number of loop accesses of the input feature map, output feature map, and weights under the three data access modes, specifically including the input feature map coefficient access times: Cif1, Cif2, Cif3, the output feature map coefficient access times: Cof1, Cof2, Cof3, and the weight coefficient access times: Cwgt1, Cwgt2, Cwgt3. Among them, set the batch loop block size as Tz, the horizontal and vertical block sizes of the data feature map coefficients as Tm*Tn and Th*Tw, the input feature map channel block size as Ti, and the output feature map channel block size as Tj, as shown in Table 1 below.
[0022]
[0023] Table 1: Amount of data accessed per single access and number of loop accesses under the three IRO / ORO / WRO modes.
[0024] Step 2, single-layer independent block optimization and loop order optimization.
[0025] The goal of single-layer convolution calculation block and loop order optimization is to determine the block parameters Tz, Tm, Tn, Th, Tw, Ti, Tj. Assume the maximum batch number is Z, the number of input feature map channels is I, the dimension of the input feature map coefficients is J*M*N, and the convolution kernel size is I*J*P*Q. Since the current mainstream convolution kernel P*Q is relatively small, generally taking 3*3, then Tp = P and Tq = Q can be taken. The relationship between the block parameters Tz, Th, Tw, Tm, Tn, Ti, Tj and Z, M, N, I, J is as Figure 2 shown. Among them, Th = S*(Tm - 1) + P, Tw = S*(Tn - 1) + Q, and S is the convolution step.
[0026] In Table 1, under different modes, the step sizes of 1, stepM, stepN, stepI, and stepJ are used to perform block partitioning on Z, M, N, I, and J respectively, obtaining block parameters Tz, Th, Tw, Tm, Tn, Ti, and Tj. Under different values of Tz, Th, Tw, Tm, Tn, Ti, and Tj, calculate the data storage sizes Bif / Bof / Bwgt in the computing core and on-chip cache, as well as the number of DRAM loop reads Cif / Cof / Cwgt. Assume that the total on-chip SRAM cache size is SRAM_TH, and calculate the required SRAM size and the data volume read from the DRAM external memory according to the data flow characteristics of each combination. Then compare the bandwidth consumption under the values of Tz, Th, Tw, Tm, Tn, Ti, and Tj that satisfy the condition that the total SRAM consumption does not exceed SRAM_TH, and determine the minimum bandwidth case minBW as the optimal block combination, and input different loop order modes mode. The specific implementation process above can be referred to as Figure 3 the algorithm function 1 shown as follows.
[0027] Step 3: Cross-layer joint block optimization and loop order optimization.
[0028] There is a possibility of coefficient reuse between the output and input features of adjacent two layers. For example, the coefficients of the output feature map of the u-th layer are the coefficients of the input feature map of the (u + 1)-th layer. If the adjacent two layers are optimized separately for block partitioning, after the coefficients of the u-th layer feature map are calculated, a DRAM write operation will be initiated to store these output feature map coefficients. Then, when the next layer performs convolution calculation, a DRAM read operation will be initiated again to read these feature coefficients, and there are redundant write and read operations for the feature map coefficients here.
[0029] If the u-th layer is optimized to the ORO mode alone, when the output feature map coefficients are stored in the on-chip SRAM, continue to initiate the convolution calculation of the (u + 1)-th layer. Assume that the (u + 1)-th layer takes the IRO mode in the loop, which can maximize the reuse of the feature map coefficients cached on the on-chip SRAM in the previous layer. This combination of ORO + IRO for adjacent two layers can maximize the reuse between the output and input feature map coefficients.
[0030] Theoretically, data reuse can be achieved between more layers for the input and output feature map coefficients, but it is necessary to consider the reuse of weight coefficients and the trade-off of batch concurrency factors. Since the on-chip SRAM cache is limited, generally it is difficult to achieve cross-layer data reuse for more than two layers simultaneously. Therefore, the method of this embodiment only considers data reuse between adjacent two layers.
[0031] The combination of adjacent two layers of ORO + IRO means that the weight coefficients of adjacent two layers need to be cached. Since the parameters of different network models are different, not all such combinations of adjacent layers have better memory access bandwidth consumption. First, it is necessary to ensure that the overall SRAM consumption does not exceed the available on-chip SRAM budget, and at the same time, it is also necessary to meet the constraint of maximizing data reuse. This optimization selection process can refer to, for example, Figure 4 the algorithm function 2 shown in it. cross_check is used to identify whether the cross-layer combination optimization has better data reuse performance than the individual optimization.
[0032] Compound block optimization and loop order optimization.
[0033] A deep neural network usually contains many layers, and there are different parameters M, N, I, J in each layer. Different combinations of parameters in different adjacent layers will lead to different block and loop optimization strategies. Some adjacent layers may be optimized by individual block, and the results obtained by optimizing the loop order IRO / ORO / WRO for each layer may be better; for some adjacent layers, the combination optimization of ORO + IRO may have better performance. It is necessary to input the parameters of different layers and compare and select the combination optimization / individual optimization for every two adjacent layers, so as to determine the optimization selection of the entire convolutional layer sequence. The implementation strategy can refer to, for example, Figure 5 the process description of algorithm 3 shown in it.
[0034] Adopting this technical solution can make full use of the characteristics of different layers to achieve the best cross-layer block and loop order optimization. Taking the VGG network as an example, the comparison between the optimization results of individual layers and the compound optimization results is given. As Figure 6 shown, in the case of 512K SRAM, the comparison results between the individual optimization (Bandwidth no_cross) and the joint optimization (BANDWIDTH) of the VGG network layers show that there is about 21.19% bandwidth savings. For each layer, the parameters M, N, I, J are shown first, and then the optimized selection (Tz, Tm, Tn, Ti, Tj) and the loop order mode (1, 2, 3 correspond to IRO, ORO, WRO respectively) are shown. It can be seen from the figure that by adopting the compound optimization method, in some ORI + IRO loop order optimization modes, the overall memory access cost is significantly smaller than that of individual optimization of each layer, although each layer also adopts its own optimal loop order optimization mode and the optimal (Tz, Tm, Tn, Ti, Tj) combination.
[0035] In summary, the present invention's content targets the structural differences of different neural network models, which are composed of different numbers of network layers. The convolutional kernel sizes of each layer may be different, and the input and output channel numbers of each layer are also different. According to the different parameters of each layer, mathematical models for cache consumption and external memory access bandwidth consumption in different combination modes are determined, and it is analyzed that bandwidth and memory consumption are important bases for the optimal block and loop order strategy selection. It also considers the possibility of reusing the output and input feature coefficients between adjacent layers. Whether to optimize the ORO+IRO combination of two adjacent layers or to optimize the selection of the separate block loop order for the two layers to minimize the overall bandwidth consumption? Different network model parameters lead to very different results. While satisfying the overall SRAM consumption not exceeding the on-chip available SRAM budget and also meeting the constraint of maximizing data reuse, a cross-layer combination optimization or separate single-layer optimization strategy is determined to achieve the optimal selection of block and loop order between multiple layers.
[0036] As described above, the above are only the preferred implementation cases of the present invention and do not impose any form of limitation on the present invention. Although the implementation process of the present invention has been described in detail above, for those familiar with the art, they can still modify the technical solutions recorded in the foregoing examples or make equivalent replacements for some of the technical features. Any modifications, equivalent replacements, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.
Claims
1. A method for optimizing cross-layer memory access bandwidth to support deep neural network acceleration, characterized in that: include: Step 1: Design a single-layer block optimization and loop sequence optimization bandwidth model; Step 2: Single-layer independent block optimization and loop sequence optimization; Step 3: Cross-layer joint block optimization and loop order optimization; Step 4: Composite block optimization and loop sequence optimization.
2. The optimization method according to claim 1, characterized in that: The model design of step 1 specifically includes: According to the characteristics of deep convolution calculation data flow, three data access modes are set, including input feature map coefficient reuse IRO mode, output feature map coefficient reuse ORO mode and weight coefficient reuse WRO mode; According to the data flow characteristics of the three data access modes, the single access data volume and the number of loop accesses of the input feature map, output feature map and weight in each mode are set.
3. The optimization method according to claim 2, characterized in that: The single access data volume of the input feature map, output feature map and weight in each mode specifically includes: input feature map coefficients: Bif1, Bif2, Bif3, output feature map coefficients: Bof1, Bof2, Bof3, and weight coefficients: Bwgt1, Bwgt2, Bwgt3; the number of cyclic accesses of the input feature map, output feature map and weight in each mode specifically includes: input feature map coefficient access times: Cif1, Cif2, Cif3, output feature map coefficient access times: Cof1, Cof2, Cof3, and weight coefficient access times: Cwgt1, Cwgt2, Cwgt3.
4. The optimization method according to claim 2, characterized in that: The step 2 specifically includes: Step 2.1, divide the feature data calculated by the single-layer convolution into blocks and obtain the block parameters; Step 2.2: Traverse all possible block combinations, calculate the memory bandwidth and storage occupancy under the three data access modes respectively, and compare the storage threshold and the minimum memory bandwidth to determine the optimal block combination and data access mode.
5. The optimization method according to claim 4, characterized in that: The step 2.1 specifically includes: using the step sizes of 1, stepM, stepN, stepI, and stepJ to divide Z, M, N, I, and J into blocks respectively to obtain block parameters Tz, Th, Tw, Tm, Tn, Ti, and Tj; wherein Z is the maximum batch number, I is the number of input feature map channels, J*M*N is the feature map coefficient dimension, Th=S*(Tm-1)+P, Tw=S*(Tn-1)+Q, S is the convolution step size, I*J*P*Q is the convolution kernel size, and Tp=P, Tq=Q is taken; Tz is the batch cycle block size, Tm*Tn and Th*Tw are the horizontal and vertical block sizes of the data feature map coefficients, Ti is the input feature map channel block size, and Tj is the output feature map channel block size.
6. The optimization method according to claim 5, characterized in that: In step 2.2, the calculation of the memory bandwidth and storage occupancy is specifically as follows: Input feature map ifmaps bandwidth: BW if = B if *C if *size(ifmaps); Bandwidth of output feature maps ofmaps: BW of = B of *C of *size(ofmaps); Bandwidth of convolution kernel weight: BW wgt = B wgt *C wgt *size(weight); Storage space: SRAM tot = B if *size(ifmaps)+B wgt *size(weight)+B of *size(ofmaps); Total bandwidth: BW tot = BW if +BW of +BWwgt; Among them, B if , B of , B wgt Respectively represent the data size of the input feature map, output feature map and convolution kernel, C if , C of , C wgt Respectively represent the number of cyclic visits to the input feature map, output feature map and convolution kernel, and size represents the size; The comparison of the storage threshold and the minimum bandwidth to determine the optimal block combination and data access mode is specifically as follows: Assume storage threshold SRAM_TH, minimum bandwidth minBW, in SRAM tot <=SRAM_TH and BW tot2 <=minBW, that is, if the storage occupancy of the current mode does not exceed the storage threshold SRAM_TH and the total bandwidth is smaller, the minimum bandwidth and data access mode are updated, that is, the optimal block combination and mode are finally obtained, so that the total bandwidth is minimized and the storage limit is met.
7. The optimization method according to claim 6, characterized in that: The step 3 specifically includes: Step 3.1, traverse the block combination of two adjacent layers of convolution calculation data; Step 3.2: Calculate the bandwidth and storage occupancy of two adjacent layers of different block combinations in a loop according to the ORO mode, set storage constraints, check whether the calculated storage occupancy meets the storage constraints in a loop according to the IRO mode, and select the block combination with the smallest total bandwidth.
8. The optimization method according to claim 7, characterized in that: The step 3.1 specifically includes: For the u-th layer and the u+1-th layer, that is, the two adjacent layers, the feature data of the u-th layer convolution calculation is first divided into blocks to obtain the u-th layer block combination; then the u-th layer block parameters are reused to obtain the u+1-th layer block combination, where Tw(u+1)=Tn(u), Th(u+1)=Tm(u), that is: the width and height of the feature map calculated by the u-th layer convolution are reused as the weight parameters of the feature map calculated by the u+1-th layer convolution.
9. The optimization method according to claim 8, characterized in that: The step 4 specifically includes: Step 4.1: Optimize each layer individually first, and then jointly optimize adjacent layers; Step 4.2: Dynamically select the optimal block partitioning strategy by comparing the results of individual optimization and joint optimization.
10. The optimization method according to claim 9, characterized in that: The step 4.1 is specifically as follows: each layer is optimized separately by using the single-layer independent block optimization and loop sequence optimization, and then adjacent layers are jointly optimized by using the cross-layer joint block optimization and loop sequence optimization.
Citation Information
Patent Citations
Implementing method of deep learning multilayer neural network
CN106295799A
Neural network processor incorporating inter-device connectivity
CN110476174A
Convolutional neural network accelerator and working method thereof
CN113312285A
Neural network accelerator and neural network acceleration method and device
CN117273094A
Deep neural network distributed training method based on model structure automatic analysis
CN117808081A