Hardware Accelerator Data Flow Design Optimization Method Based on Neural Network Backpropagation

Through the method based on neural network backpropagation, the data flow and DNN layer of the DNN accelerator are configured and encoded, and the data flow code is optimized by backpropagation. The problem of insufficient data flow design in the existing technology is solved, efficient and automated data flow optimization is achieved, and the efficiency and versatility of DNN execution is improved.

CN116384474BActive Publication Date: 2025-06-13SHANGHAI ARTIFICIAL INTELLIGENCE INNOVATION CENT
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310369513.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-04-07
Publication Date
2025-06-13
Estimated Expiration
2043-04-07

AI Technical Summary

Technical Problem

The existing DNN accelerator data flow is insufficient, resulting in the generalization and efficiency of DNN execution being limited, and the manual workload is large and inefficient when designing data flows.

Method used

Using a hardware accelerator data flow design optimization method based on neural network backpropagation, the hardware accelerator data flow and its accompanying DNN layer configuration are represented by encoding. The trained neural network predictor is used to search the optimized data flow code through backpropagation to achieve automated optimization of the data flow code.

Benefits of technology

Through automated data flow code optimization methods, the workload of manual design data flow is significantly reduced, the efficiency and versatility of DNN execution is improved, and the best data flow that meets the optimization goals can be quickly searched in the vast design space.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116384474B_ABST
    Figure CN116384474B_ABST
Patent Text Reader

Abstract

The present invention relates to a method for optimizing the data flow design of a hardware accelerator based on neural network backpropagation, belonging to the optimization design of hardware accelerators in neural network applications, and is used to reduce the manual workload when designing the data flow. The technical solution is specifically as follows: The data flow of the hardware accelerator and its accompanying DNN layer configuration are represented by encoding to obtain the initial data flow code; the weights of the trained neural network predictor are fixed, the initial data flow code is input into the trained neural network predictor, and according to the set target performance indicators, differentiable backpropagation is performed to search for the optimized data flow code; the neural network predictor is established using a neural network model, which can evaluate the performance of all sampled hardware accelerator data flows and their accompanying DNN layer configurations, and forms a dataset encoding representation with a unified input and a performance metric output.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to neural network applications and hardware accelerator optimization design, and particularly to a method for optimizing the data flow design of a hardware accelerator based on neural network backpropagation. Background Art

[0002] Deep neural networks (DNNs) have achieved remarkable breakthroughs in many fields, such as vision and language, autonomous driving, and bioscience. However, the exponentially growing model sizes usually increase the latency and energy consumption of DNN applications. Compared with general-purpose hardware processors, DNN accelerators can achieve higher efficiency and lower energy when executing DNNs. This is achieved by designing a more suitable microarchitecture and optimizing the hardware mapping strategy (referred to as data flow) of DNNs, including the order of executing DNN layer computations and how these computations are mapped to hardware resources (e.g., processing elements and memories).

[0003] Designing a data flow to achieve optimal on-chip performance and efficiency is a fundamental but cumbersome and challenging task. However, existing DNN accelerators are usually too specialized and lack data flow design, which hinders the generalization and efficiency of DNN execution. Summary of the Invention

[0004] In order to reduce the manual workload when designing a data flow and to be able to design a better data flow for DNN applications, the object of the present invention is to propose a method for optimizing the data flow design of a hardware accelerator based on neural network backpropagation.

[0005] In order to achieve the above object of the invention, the technical solution of the present invention is as follows.

[0006] In a first aspect, a method for optimizing the data flow design of a hardware accelerator based on neural network backpropagation, the method comprising the following steps:

[0007] Represent the hardware accelerator data flow and its accompanying DNN layer configuration using an encoding to obtain an initial data flow code;

[0008] Fix the weights of the trained neural network predictor, input the initial data flow code into the trained neural network predictor, and perform differentiable backpropagation according to the set target performance metrics to search for an optimized data flow code;

[0009] The neural network predictor is established using a neural network model, and is used to evaluate the performance of all sampled hardware accelerator data flows and their accompanying DNN layer configurations, and forms a dataset encoding representation with a unified input and a performance metric output.

[0010] In the above technical solution, the DNN layer configuration encodes and represents seven dimensions, and the seven dimensions that need to be encoded and represented are the input row / column (Y / X), filter row / column (R / S), output / input channel (K / C), and an additional dimension (T) that describes the DNN layer type.

[0011] In the above technical solution, the encoding format of the hardware accelerator data stream is (M, N, 2), with a total of M×N×2 dimensions;

[0012] M represents the number of memory levels of the set hardware accelerator, N represents the number of DNN layer configuration dimensions, and "2" means that the encoding on each part of the DNN accelerator has two aspects. One represents the number indexing the corresponding dimension, and the other is the accompanying number corresponding to the indexed dimension;

[0013] The indexed dimensions include the input row / column (Y / X), filter row / column (R / S), and output / input channel (K / C).

[0014] In the above technical solution, the first dimension of the encoding of the hardware accelerator data stream is the parallel dimension selected to be expanded and perform parallel computing, and the remaining dimensions are the dimensions reordered according to the computing order; the accompanying number of the parallel dimension specifies the number of processing units (PE, Processing Element) to be allocated to perform parallel computing, and the accompanying numbers of the remaining six dimensions determine their partition sizes.

[0015] In the above technical solution, the neural network model adopted by the neural network predictor is MobileNet-V2, ResNet101, or Vision Transformer (ViT).

[0016] In the above technical solution, for each DNN layer, 2000 data streams are randomly sampled.

[0017] In the above technical solution, the training method of the predictor includes supervised, semi-supervised, and unsupervised learning methods.

[0018] In the above technical solution, the target performance metrics include the latency, energy, and power of single-layer / multi-layer.

[0019] In a second aspect, the present invention proposes a hardware accelerator data stream design optimization device based on neural network backpropagation, including a memory and a processor, and a computer program capable of being loaded and executed by the processor is stored on the memory and can execute any of the above methods.

[0020] In a third aspect, the present invention proposes a computer-readable storage medium storing a computer program capable of being loaded and executed by the processor to execute any of the above methods.

[0021] The technical effects of the present invention are as follows:

[0022] (1) Different from previous work, the present invention uses an encoding scheme to convert data streams and DNN layer configurations into a unified encoded representation, thereby implementing a data stream code propagation (DCP), and constructing a huge encoded space for data stream design. By utilizing this encoding scheme, a benchmark is constructed using the input of the unified code and the output of the corresponding evaluation metrics (e.g., latency, energy, power, etc.), so as to accelerate the acquisition of data stream search that meets the optimization objectives.

[0023] (2) The present invention is the first technical solution to utilize differentiable backpropagation in data stream optimization. By backpropagating the gradients of the neural network predictor, the data stream code can be effectively updated in a vast design space to achieve the desired optimization objectives.

[0024] (3) A large number of experiments have also verified the generality and data efficiency of data stream code propagation, and data streams can be customized for many optimization objectives (e.g., latency / energy / statistical verification force of single / multiple layers) in various DNNs with visible and invisible hardware configurations. Description of the Drawings

[0025] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.

[0026] Figure 1 、 one Schematic diagram of the DCP encoding strategy in a specific embodiment;

[0027] Figure 2 、 one Schematic overview of the DCP (Data stream code propagation) method in a specific embodiment;

[0028] Figure 3 、 one Schematic diagram of the performance of the DCP neural network predictor in a specific embodiment;

[0029] Figure 4 、 one DCP performance and search time consumption in a specific embodiment. Specific Embodiments

[0030] When designing a DNN accelerator, several components should be considered. These components also constitute the design space of the DNN accelerator, including data flow, hardware resources, specific circuit design, etc. To design an efficient dedicated DNN accelerator for target applications, many works have been proposed to explore the DNN accelerator design space. For example, Apollo is a framework that optimizes DNN accelerators based on features extracted from the knowledge of integrated circuit (IC) design.

[0031] Since data flow is an important part of the DNN accelerator design space, the method of the present invention can be broadly classified as a method for exploring the DNN accelerator design space. The size of the data flow design space of a single layer of DNN network is 10 36 , and existing conventional search techniques for data flow design, such as exhaustive search, reinforcement learning, and genetic algorithms, are difficult to handle the huge data flow design space. To address this challenge, most design choices are to narrow the design space, but this will lead to sub-optimal designs.

[0032] To efficiently search for the best data flow in the vast design space, the present invention proposes a data flow code propagation method to effectively search for the optimal data flow of the DNN accelerator in a differentiable manner (the method process is as Figure 1 shown). To achieve this goal, first, the data flow parameters and DNN layer configurations are converted into a unified coding representation, and a coding space for data flow design is constructed. By using this coding scheme, a benchmark is constructed using the unified code input and the output of the corresponding evaluation metric. Then, given the data flow and DNN layer configuration, a neural network predictor is trained to project the unified code into the corresponding evaluation metric. Finally, under the set goal, the best data flow under the DNN layer configuration is searched by backpropagating the gradient.

[0033] Compared with existing data flow design techniques, data flow code propagation has the following key points. First, data flow code propagation (DCP) is data-efficient because once the neural network predictor is trained, it can search for the best data flow of various DNNs within seconds. However, previous works needed to search for the data flow of each DNN through a time-consuming simulation process. Second, it can optimize the data flow optimization at the entire model level by accumulating the gradients of all layers. Finally, data flow code propagation (DCP) can be easily extended to multi-objective data flow optimization, thus generating data flow configurations that can better balance multiple hardware metrics.

[0034] A specific implementation manner is given below in conjunction with the accompanying drawings.

[0035] S0 1. Represent the hardware accelerator data stream and its accompanying DNN layer configuration in an encoded form to obtain the initial data stream code.

[0036] By encoding the data stream and the accompanying DNN layer configuration into an encoded representation, data stream optimization can be effectively performed. The encoded representations of the data stream and the DNN model can be constructed in different ways.

[0037] Figure 1 An example of a detailed encoding scheme is shown.

[0038] In the shown encoding scheme, the configuration of the DNN layer is represented by seven dimensions, which are the input row / column (Y / X), filter row / column (R / S), output / input channels (K / C), and an additional dimension (T) that describes the type of the DNN layer. The output row / column can be inferred from the input row / column and the filter row / column. Exemplarily, the DNN layer is encoded as a seven-dimensional code, and its order is exemplified as K, C, Y, X, R, S, T, recorded as numerical values 1, 2, 3, 4, 5, 6, 7. The encoding summary is as follows in Table 1:

[0039] Table 1

[0040]

[0041] For the encoded representation of the data stream, it is exemplarily encoded as a (3, 7, 2)-dimensional code. The "3" therein refers to the three memory levels of the accelerator. For each memory level, there will be a (7, 2)-dimensional code to describe the data stream of a cluster in the corresponding memory level. The DNN accelerator may have an L2 cluster and several L1 / L0 clusters, which contain a memory buffer and corresponding sub-clusters. The "7" in the (7, 2)-dimensional code of each memory level represents seven dimensions, and the "2" contains a number indexing the corresponding dimension and an accompanying number.

[0042] In Figure 1 the DNN layer is exemplarily encoded as:

[0043] {(K, 192), (C, 192), (Y, 14), (X, 14), (R, 1), (S, 1), (T, 2)}

[0044] Dimension mapping:

[0045] {192, 192, 14, 14, 1, 1, 2}

[0046] The DNN accelerator may have an L2 cluster and several L1 / L0 clusters, which contain a memory buffer and corresponding sub-clusters. In Figure 1 the DNN accelerator has one L2 cluster, one L1 cluster, and one L0 cluster.

[0047] An exemplary encoding of the L2 cluster is as follows:

[0048] {(K, 168), (R, 1), (S, 1), (K, 192), (C, 96), (X, 14), (Y, 14)}

[0049] An exemplary encoding of the L1 cluster is as follows:

[0050] {(K, 24), (Y, 2), (S, 1), (K, 48), (C, 24), (R, 1), (X, 2)}

[0051] An exemplary encoding of the L0 cluster is as follows:

[0052] {(C, 1), (Y, 1), (K, 2), (R, 1), (C, 1), (S, 1), (X, 1)}

[0053] It should be noted that the first dimension refers to the parallel dimension selected from K, C, Y, X, R, S to be expanded and perform parallel computing, and the remaining six dimensions are obtained by reordering the dimensions of K, C, Y, X, R, S in the calculation order. The accompanying number of the parallel dimension specifies how many PEs will be allocated to perform parallel computing, while the accompanying numbers of the remaining six dimensions determine their partition sizes. The size of the memory buffer in each sub-cluster can be calculated through the partition size, because the buffer stores the tensors of the input, output, and weights. In addition, the sizes of these tensors are linear combinations of the partition sizes.

[0054] The dimension mappings corresponding to the exemplary encodings of the L2 cluster, L1 cluster, and L0 cluster are respectively:

[0055] {(1, 168), (5, 1), (6, 1), (1, 192), (2, 96), (4, 14), (3, 14)}

[0056] {(1, 24), (3, 2), (6, 1), (1, 48), (2, 24), (5, 1), (4, 2)}

[0057] {(2, 1), (3, 1), (1, 2), (5, 1), (2, 1), (6, 1), (4, 1)}

[0058] After the dimension mapping, the first dimension of each encoding is the index of the corresponding value of K, C, Y, X, R, S, and the specific values vary according to the settings.

[0059] Taking the dimension mapping of the above L1 cluster as an example and referring to Table 1 above, it can be seen that for (1, 24), "1" is the parallel dimension for performing parallel computing, which is the "output channel", "24" refers to the number of PEs allocated in a cluster at the corresponding memory level, and "3", "6", "1", "2", "5", "4" are the calculation orders after the "output channel", which are "input row", "filter column", "output channel", "input channel", "filter row", "input column", and their accompanying numbers "2", "1", "48", "24", "1", "2" represent the partition sizes of the corresponding dimensions, so that the memory buffer size of the L1 cluster can be obtained.

[0060] S02. Fix the weights of the trained neural network predictor, input the initial data flow code into the trained neural network predictor, and perform differentiable backpropagation according to the set target performance metrics to search for the optimized data flow code.

[0061] The role of the neural network predictor is to evaluate the performance of all sampled hardware accelerator data flows and their accompanying DNN layer configurations, and form a dataset encoding representation with a unified input and a performance metric output. That is: the neural network predictor is a cost model that takes the DNN layer configuration and data flow encoding as inputs and the performance matrix as the output. See Figure 2 The schematic diagram of the left part of the block diagram.

[0062] The neural network predictor is established using a neural network model. The available neural network models include MobileNet-V2, ResNet101, or Vision Transformer (ViT), etc. Traditional machine learning methods can also be used to train the predictor, such as SVM, Softmax, etc.

[0063] The neural network predictor constructs a benchmark by evaluating the evaluation metrics of each pair of code representations of the data flow and the DNN layer, so that the predictor can quickly search for the optimal data flow of the DNN accelerator. This makes the neural network predictor different from the previous data flow design methods, and the data flow code propagation does not require a time-consuming performance evaluation process for each model during the search.

[0064] However, the above encoding scheme introduces a huge search space, making it difficult to enumerate all DNN layer dimensions and their corresponding data flows. Therefore, for each DNN layer, a trained predictor is obtained by controlling the dataset size. In one embodiment, randomly sampling 2000 data flows can satisfy the training of a good predictor. The training methods of the predictor include supervised, semi-supervised, and unsupervised learning methods. The target performance metrics include latency, energy, power, etc. of single-layer / multi-layer. Figure 2An example is given. When training or predicting, the performance metrics component is [R latency , R energy , R power , and the performance metrics component in the middle of the iteration is [P latency , P energy , P power . Data flow code propagation aims to train predictors using the benchmarks introduced above to optimize the data flow of all models.

[0065] After training the neural network predictor, data flow optimization becomes extremely fast by backpropagating along the direction of the target metric gradient. As shown in the right part of the block diagram Figure 2 , the predictor fixes the weights, inputs the initial data flow code, and performs backpropagation on the data flow code according to the performance metrics that set one or more targets. After a set number of iterations, the optimized data flow code is returned. In Figure 2 , the set target performance metrics are represented as a target matrix (TargetMetrics), and the performance metrics obtained during the iteration are represented as a predicted matrix (Predicted Metrics). By calculating the iteration loss (Loss) between the target matrix and the predicted matrix, the optimized data flow code is obtained through backpropagation, thereby obtaining the optimized data flow.

[0066] In one embodiment, the performance of the neural network predictor used in this method is verified. As shown in Figure 3 , the predictor can accurately predict various performance metrics for a given data flow and DNN layer configuration. Based on three DNN models with different topologies and complexities to compare and verify the effectiveness of this method, the selected DNN models are MobileNet-V2, ResNet101, and Vision Transformer (ViT). The other methods used for comparison and verification are: Random Search (Random), which randomly samples the design points and retains the best solution; Genetic Algorithm (GA), including pureGA and GAMMA; Particle Swarm Optimization (PSO); Passive Portfolio; (1+1) Evolutionary Strategy (OnePlusOne); Covariance Matrix Adaptation ES (CMA); Differential Evolution (DE); and Test-Based Population Size Adaptation (TBPSA). The relevant search performance and search time consumption are shown in Figure 4 . From the data in this figure, it can be seen that the method of the present invention has lower latency and lower energy consumption.

[0067] Through the description of the above embodiments, those skilled in the art can clearly understand that the method of the present disclosure can be implemented by means of software plus necessary general hardware. Of course, it can also be implemented by dedicated hardware including application specific integrated circuits, dedicated CPUs, dedicated memories, dedicated components, etc. Generally, functions accomplished by computer programs can be easily implemented by corresponding hardware, and the specific hardware structures for implementing the same function can also be diverse, such as analog circuits, digital circuits or dedicated circuits, etc. However, in more cases for the present disclosure, implementation by software programs is a better embodiment.

[0068] Although the embodiments of the present invention have been described above in conjunction with the accompanying drawings, the present invention is not limited to the above specific embodiments and application fields. The above specific embodiments are merely illustrative and guiding, rather than restrictive. Those of ordinary skill in the art can also make many forms under the inspiration of this specification and without departing from the scope protected by the claims of the present invention, and all of these fall within the scope of protection of the present invention.

Claims

1. A method for optimizing the data flow design of a hardware accelerator based on neural network backpropagation, characterized in that, the method comprises the following steps: Represent the hardware accelerator data flow and its accompanying DNN layer configuration using a unified encoding to obtain the initial data flow code; where: The initial data flow code consists of DNN layer configuration encoding and data flow encoding; The DNN layer configuration dimensions include an index dimension and an additional dimension describing the DNN layer type. The index dimension includes output channels, input channels, input rows, output columns, filter rows, and filter columns; The length of each data flow encoding is the same as the length of the DNN layer configuration encoding. The first dimension of each data flow encoding is the parallel dimension selected from the index dimension to be expanded for parallel computing, and the remaining dimensions of each data flow encoding are the dimensions reordered according to the calculation order; the accompanying number of the parallel dimension specifies the number of processing units to be allocated to perform parallel computing, and the accompanying numbers of the remaining dimensions are the sizes of their partitions; The data stream encoding is a total of M×N×2 dimensions, indicating the number of memory levels of the set hardware accelerator, indicating the configuration dimension number of each DNN layer. The 2 indicates that the data stream encoding on each level of memory of the hardware accelerator has two aspects. One represents the number corresponding to the index dimension, and the other is the accompanying number corresponding to the index dimension; Fix the weights of the trained neural network predictor, input the initial data flow code into the trained neural network predictor, and perform differentiable backpropagation according to the set target performance metrics to search for the optimized data flow code; The neural network predictor is established using a neural network model, which is used to evaluate the performance of all sampled hardware accelerator data flows and their accompanying DNN layer configurations, and forms a dataset encoding representation with a unified input and a performance metric output.

2. The method according to claim 1, characterized in that, The neural network model adopted by the neural network predictor is MobileNet-V2, ResNet101 or Vision Transformer.

3. The method according to claim 1, characterized in that, For each DNN layer, 2000 data flows are randomly sampled.

4. The method according to claim 1, characterized in that, The training method of the predictor includes supervised, semi-supervised and unsupervised learning methods.

5. The method according to claim 1, characterized in that, The target performance metrics include single-layer / multi-layer latency, energy, and power.

6. A device for optimizing the data flow design of a hardware accelerator based on neural network backpropagation, characterized in that: It includes a memory and a processor, and a computer program capable of being loaded and executed by the processor as any one of the methods in claims 1 to 5 is stored on the memory.

7. A computer-readable storage medium, characterized in that: A computer program capable of being loaded and executed by the processor as any one of the methods in claims 1 to 5 is stored.

Citation Information

Patent Citations

  • Hardware acceleration method, system and application of convolutional neural network convolutional layer

    CN115238863A

  • Multi-target neural architecture search method based on semi-supervised performance predictor

    CN115620046A